A log audit and analysis method and system

By building a variety of log analysis models on the cloud computing center, the real-time and historical log data are automated, and the problems of low efficiency and poor accuracy of log audit analysis in the existing technology are solved, and efficient and accurate log audit analysis is achieved, which is suitable for large-scale data scenarios.

CN118939701BActive Publication Date: 2025-06-17INSPUR QILU SOFTWARE IND +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410950988.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2025-06-17
Estimated Expiration
2044-07-16

AI Technical Summary

Technical Problem

The existing log audit analysis technology has problems such as low efficiency, poor accuracy, high cost and low practicality, especially in large-scale data scenarios, which are difficult to meet security needs.

Method used

The cloud computing center is adopted to automatically process and analyze real-time and historical log data by constructing a log audit analysis model built with data mapping and transformation models, log feature initial classification model, log attribute sub-classification model, and deep learning algorithm.

Benefits of technology

It improves the efficiency and accuracy of log audit analysis, reduces computing resources and time costs, and enhances practicality and security analysis capabilities in large-scale data scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118939701B_ABST
    Figure CN118939701B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of log audit and analysis, and discloses a log audit and analysis method and system. The method includes the following steps: based on a cloud computing center, constructing a data mapping and transformation model, a log feature primary classification model, a log attribute secondary classification model, and several log audit and analysis models; performing data mapping and transformation on real-time log big data; performing primary classification of log features on several standard real-time log data; performing secondary classification of log attributes on each real-time log feature primary classification cluster; performing log audit and analysis model matching; performing log audit and analysis on each real-time log attribute secondary classification cluster; and visually displaying the standard real-time log data, the real-time log audit and analysis results, and the real-time alarm signals. The present invention solves the problems of low efficiency, poor accuracy, high cost, and low practicability existing in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of log audit analysis, and particularly relates to a log audit analysis method and system. Background Art

[0002] With the continuous development of information technology, network security issues have become increasingly prominent. As an important security means, log audit analysis can discover potential security threats by analyzing the logs of systems, networks, and applications. However, the existing log audit analysis technologies generally have problems such as low efficiency and high false alarm rates, and cannot meet the growing security needs; moreover, in the scenario of processing and analyzing a large amount of data, the existing log audit analysis technologies have high computing resource costs and time costs, and cannot meet the log audit analysis under a large amount of data, with low practicability. Summary of the Invention

[0003] In order to solve the problems of low efficiency, poor accuracy, high cost, and low practicability existing in the prior art, the purpose of the present invention is to provide a log audit analysis method and system.

[0004] The technical solution adopted by the present invention is as follows:

[0005] A log audit analysis method, characterized by comprising the following steps:

[0006] Based on a cloud computing center, according to historical log big data, construct a data mapping and conversion model, a log feature primary classification model, a log attribute secondary classification model, and several log audit analysis models;

[0007] Receive real-time log big data uploaded by a log data acquisition device, and use the data mapping and conversion model to perform data mapping and conversion on the real-time log big data to obtain several homogeneous standard real-time log data;

[0008] Use the log feature primary classification model to perform primary classification of log features on several standard real-time log data to obtain several real-time log feature primary classification clusters;

[0009] Use the log attribute secondary classification model to perform secondary classification of log attributes on each real-time log feature primary classification cluster to obtain several real-time log attribute secondary classification clusters;

[0010] According to the real-time log feature primary classification and real-time log attribute secondary classification of each real-time log attribute secondary classification cluster, perform matching in several log audit analysis models to obtain a target log audit analysis model;

[0011] Use the target log audit analysis model to perform log audit analysis on each real-time log attribute secondary classification cluster to obtain the real-time log audit analysis result of each standard real-time log data;

[0012] Based on the real-time log audit analysis results, generate corresponding real-time alarm signals, and visually display the standard real-time log data, real-time log audit analysis results, and real-time alarm signals.

[0013] Furthermore, based on the cloud computing center, according to the historical log big data, construct a data mapping and transformation model, a log feature primary classification model, a log attribute secondary classification model, and several log audit analysis models, including the following steps:

[0014] Based on the cloud computing center, collect historical log big data from different data sources, and preprocess the historical log big data to obtain several heterogeneous preprocessed historical log data;

[0015] According to the data structure differences between the system data structure of the cloud computing center and the preprocessed historical log data, construct a unified data model;

[0016] According to the unified data model and several preprocessed historical log data, use the reinforcement learning algorithm to construct a data mapping and transformation model, and generate several homogeneous standard historical log data;

[0017] According to several standard historical log data, use the clustering algorithm to perform primary classification of log features, construct a log feature primary classification model, and generate several historical log feature primary classification clusters corresponding to the primary classification of log features;

[0018] According to each historical log feature primary classification cluster, use the clustering algorithm to perform secondary classification of log attributes, construct a log attribute secondary classification model, and generate several historical log attribute secondary classification clusters corresponding to the secondary classification of log attributes;

[0019] Perform data dimensionality reduction and matrix transformation on the historical log attribute secondary classification clusters to obtain several dimensionality-reduced historical log attribute secondary classification cluster matrices;

[0020] According to several dimensionality-reduced historical log attribute secondary classification cluster matrices, use the deep learning algorithm to perform log audit analysis, and construct a log audit analysis model corresponding to each historical log attribute secondary classification in the primary classification of historical log features.

[0021] Furthermore, according to several preprocessed historical log data and historical data mapping and transformation strategies, use the DQN algorithm to construct a data mapping and transformation model, and generate several homogeneous standard historical log data.

[0022] Furthermore, according to several standard historical log data, use the AP clustering algorithm to perform primary classification of log features, construct a log feature primary classification model, and generate several historical log feature primary classification clusters corresponding to the primary classification of log features.

[0023] Furthermore, for each initially classified cluster of historical log features, the FCM clustering algorithm is used to perform secondary classification of log attributes, construct a secondary classification model of log attributes, and generate several clusters of secondary classification of historical log attributes corresponding to the secondary classification of historical log attributes.

[0024] Furthermore, based on several matrices of clusters of secondary classification of historical log attributes after dimensionality reduction, the N-GAN-MLP algorithm is used to perform log audit analysis, and a log audit analysis model corresponding to each secondary classification of historical log attributes in the initial classification of historical log features is constructed, where N is the dimension of concern.

[0025] Furthermore, real-time log big data uploaded by a log data collection device is received, and the data mapping and transformation model is used to perform data mapping and transformation on the real-time log big data to obtain several homogeneous standard real-time log data, including the following steps:

[0026] Receive real-time log big data from different data sources uploaded by the log data collection device, preprocess the real-time log big data to obtain several heterogeneous preprocessed real-time log data;

[0027] According to the unified data model and the preprocessed real-time log data, use the data mapping and transformation model to generate corresponding real-time data mapping and transformation strategies;

[0028] According to the real-time data mapping and transformation strategy, perform data mapping and transformation on the preprocessed real-time log data to obtain standard real-time log data;

[0029] Traverse all the preprocessed real-time log data to obtain several homogeneous standard real-time log data.

[0030] Furthermore, use the initial classification model of log features to perform initial classification of log features on several standard real-time log data to obtain several clusters of initial classification of real-time log features, including the following steps:

[0031] Input several standard real-time log data into the initial classification model of log features to initialize the AP clustering algorithm, obtain M initial first clustering centers, as well as the initial attraction information and initial belonging information of each standard real-time log data, where M is the total number of first clustering centers;

[0032] Introduce an iterative attenuation coefficient to update the initial attraction information and initial belonging information of each standard real-time log data to obtain updated attraction information and updated belonging information;

[0033] Update the M initial first clustering centers to obtain M updated first clustering centers, and based on the M updated first clustering centers, according to the updated attraction information and updated membership information, conduct an initial classification of log characteristics for all standard real-time log data to obtain M real-time log characteristic initial classification clusters;

[0034] Repeat the above steps. If the number of iterations exceeds the iteration number threshold or the first clustering centers do not change, output the M final first clustering centers and the final real-time log characteristic initial classification clusters, and use the log characteristic initial classification corresponding to the first clustering centers as the real-time log characteristic initial classification.

[0035] Furthermore, use the log attribute secondary classification model to conduct a secondary classification of log attributes for each real-time log characteristic initial classification cluster to obtain several real-time log attribute secondary classification clusters, including the following steps:

[0036] Input the real-time log characteristic initial classification cluster into the log attribute secondary classification model to initialize the FCM clustering algorithm, obtaining the fuzzy factor and H initial second clustering centers, as well as the fuzzy membership degrees of all standard real-time log data in the real-time log characteristic initial classification cluster, where H is the total number of second clustering centers;

[0037] According to the fuzzy membership degrees, update the H initial second clustering centers to obtain the corresponding H updated second clustering centers;

[0038] Use the Lagrange multiplier method to calculate the merging function to obtain the merging function value and the merging function change value. If the merging function value is greater than the function threshold or the merging function change value is greater than the change value threshold, continue to update the second clustering centers. Otherwise, use the current second clustering centers as the final second clustering centers;

[0039] Repeat the above steps to obtain H final second clustering centers, and divide each standard real-time log data in the real-time log characteristic initial classification cluster into the final second clustering center with the closest Euclidean distance to obtain several real-time log attribute secondary classification clusters.

[0040] A log audit analysis system for implementing the log audit analysis method. The system includes a cloud computing center and several log data collection devices, and the cloud computing center is respectively communicatively connected with the several log data collection devices;

[0041] The cloud computing center is provided with a model construction unit, a data mapping and conversion unit, a log characteristic initial classification unit, a log attribute secondary classification unit, a model matching unit, a log audit analysis unit, and a visualization display unit that are connected in sequence.

[0042] The beneficial effects of the present invention are:

[0043] The present invention discloses a log audit analysis method and system. The data mapping and conversion model constructed by the reinforcement learning algorithm can automatically map and convert the log data from different data sources, convert heterogeneous data into homogeneous data, and improve the efficiency of log audit analysis. The log characteristic primary classification model and log attribute secondary classification model constructed by the machine learning algorithm can perform primary classification on a large volume of real-time log big data according to log characteristics, perform secondary classification according to log attributes, and uniformly process and analyze log data with the same characteristics and attributes, thereby improving the efficiency of log audit analysis, being more suitable for large-volume data scenarios, improving practicality, and reducing cost investment. The log audit analysis model constructed by the deep learning algorithm can mine the deep data characteristics of log data, perform accurate and efficient log audit analysis, and reduce the error rate of log audit analysis.

[0044] Other beneficial effects of the present invention will be further described in the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a flowchart of the log audit analysis method in the present invention.

[0046] Figure 2 It is a structural block diagram of the log audit analysis system in the present invention. DETAILED DESCRIPTION

[0047] The present invention will be further explained below in conjunction with the accompanying drawings and specific embodiments.

[0048] Embodiment 1:

[0049] like Figure 1 As shown, this embodiment provides a log audit analysis method, which is characterized by comprising the following steps:

[0050] S1: Based on the cloud computing center and historical log big data, a data mapping and conversion model, a log feature initial classification model, a log attribute secondary classification model, and several log audit analysis models are constructed, including the following steps:

[0051] S1-1: Based on the cloud computing center, historical log big data from different data sources is collected and preprocessed to obtain a number of heterogeneous preprocessed historical log data;

[0052] S1-2: Construct a unified data model based on the data structure differences between the system data structure of the cloud computing center and the pre-processed historical log data;

[0053] S1-3: According to the unified data model and a number of preprocessed historical log data, use the Deep Q Network (DQN) algorithm to construct a data mapping and transformation model, and generate a number of isomorphic standard historical log data, including the following steps:

[0054] S1-2-1: Define the simulation environment of the Deep Q Network (DQN) algorithm according to the data mapping and transformation problem, and construct the deep Q network, agent, and experience storage pool of the DQN algorithm

[0055] S1-2-2: Define the action space of the DQN algorithm according to the preset data mapping and transformation strategy of the unified data model, and define the state space of the DQN algorithm according to the preprocessed historical log data;

[0056] S1-2-3: Define the reward function of the DQN algorithm according to the action space and state space;

[0057] S1-2-4: Based on the simulation environment, state space, action space, reward function, deep Q network, agent, and experience storage pool, construct an initial data mapping and transformation model;

[0058] The neural network structure is a deep Q network, which includes two neural networks with the same structure but different parameters: one is used to predict the Q value (online Q network), and the other is used to generate the target Q value (target Q network). The parameters of the target Q network are periodically copied from the online Q network but not updated frequently, which helps to improve the stability of the learning process;

[0059] S1-2-5: Optimize and train the initial data mapping and transformation model according to the preset data mapping and transformation strategy of the unified data model and a number of preprocessed historical log data, and store the generated data mapping and transformation experience in the experience storage pool to obtain the final data mapping and transformation model;

[0060] S1-4: According to a number of standard historical log data, use the Affinity-Propagation (AP) clustering algorithm to conduct an initial classification of log characteristics, construct an initial classification model of log characteristics, and generate a number of historical log characteristic initial classification clusters corresponding to the initial classification of historical log characteristics;

[0061] S1-5: According to each historical log characteristic initial classification cluster, use the Fuzzy C-mean (FCM) clustering algorithm to conduct a secondary classification of log attributes, construct a secondary classification model of log attributes, and generate a number of historical log attribute secondary classification clusters corresponding to the secondary classification of historical log attributes;

[0062] S1-6: Use the Principal Component Analysis (PCA) method to perform data dimensionality reduction on the historical log attribute sub-classification clusters, and perform matrix transformation to obtain several historical log attribute sub-classification cluster matrices after dimensionality reduction;

[0063] S1-7: According to several historical log attribute sub-classification cluster matrices after dimensionality reduction, use the N-Generative Adversarial Network (GAN)-Multilayer Perceptron (MLP) algorithm to perform log audit analysis, and construct a log audit analysis model corresponding to each historical log attribute sub-classification in the initial classification of historical log characteristics, where N is the dimension of concern;

[0064] S2: Receive the real-time log big data uploaded by the log data collection device, and use the data mapping and transformation model to perform data mapping and transformation on the real-time log big data to obtain several homogeneous standard real-time log data, including the following steps:

[0065] S2-1: Receive the real-time log big data from different data sources uploaded by the log data collection device, and preprocess the real-time log big data to obtain several preprocessed real-time log data in heterogeneous form;

[0066] S2-2: According to the unified data model and the preprocessed real-time log data, use the data mapping and transformation model to generate the corresponding real-time data mapping and transformation strategy, including the following steps:

[0067] S2-2-1: According to the preprocessed real-time log data, update the state space of the data mapping and transformation model to obtain the updated state space S' = [s'1,..., s' i" ,..., s' I , where s' i" is the updated i"th state value, and i" is the state indicator;

[0068] S2-2-2: According to the preset data mapping and transformation strategy of the unified data model, update the action space of the data mapping and transformation model A' = [a'1,..., a' j" ,..., a' I , where a' j" is the updated j"th action value, and j" is the action indicator;

[0069] S2-2-3: The updated state space S' = [s'1,..., s' i' ,..., s' IAs the input of the data mapping and transformation model, based on the deep Q-network, use an agency to obtain the updated action space A' = [a'1,..., a' j' ,..., a' I ; predict the Q-values of possible actions in

[0070] S2-2-4: Use the reward function to obtain the reward values of possible actions, and update the predicted Q-values of possible actions according to the reward values to obtain the updated Q-values of possible actions;

[0071] The formula is:

[0072] Q(s' p' , a' p' ) = (1 - α)·Q(s p' , a p' ) + α·(R(s p' , a p' , s' p' ) + γ·Q max (s p' , a p' ))

[0073] In the formula, Q(s' p' , a' p' ) is the updated Q-value corresponding to the updated state value s' p' and the updated action value a' p' ; Q(s p' , a p' ) is the predicted Q-value corresponding to the state value s p' and the action value a p' ; α is the learning rate; Q max (s p' , a p' ) is the highest predicted Q-value;

[0074] S2-2-5: Repeat the above steps until the iteration number threshold is reached, and use the greedy strategy to take the possible action corresponding to the highest updated Q-value as the execution action;

[0075] S2-2-6: Optimize the preset data mapping and transformation strategy according to the execution action to obtain the real-time data mapping and transformation strategy;

[0076] S2-3: According to the real-time data mapping and transformation strategy, perform data mapping and transformation on the preprocessed real-time log data to obtain the standard real-time log data;

[0077] S2-4: Traverse all preprocessed real-time log data to obtain several isomorphic standard real-time log data;

[0078] S3: Use the initial classification model of log features to perform initial classification of log features on a number of standard real-time log data, and obtain a number of initial classification clusters of real-time log features, including the following steps:

[0079] S3-1: Input a number of standard real-time log data into the initial classification model of log features, perform initialization of the AP clustering algorithm, and obtain M initial first cluster centers, as well as the initial attraction information and initial belonging information of each standard real-time log data, where M is the total number of first cluster centers;

[0080] The formula for the attraction information is:

[0081]

[0082] In the formula, r t+1 (i,k) is the degree to which the k-th standard real-time log data is suitable as the first cluster center of the i-th standard real-time log data at the (t + 1)-th iteration, that is, the attraction information of the k-th standard real-time log data to the i-th standard real-time log data; a t (i,j) is the degree to which the i-th standard real-time log data selects the j-th standard real-time log data as its first cluster center at the t-th iteration; r t (i,j) is the degree to which the standard real-time log data j is suitable as the first cluster center of the i-th standard real-time log data at the t-th iteration; s(i,k) is the similarity of the k-th standard real-time log data as the first cluster center of the i-th standard real-time log data; i, j, and k are all standard real-time log data indicators; t is the iteration number indicator;

[0083] The formula for the belonging information is:

[0084]

[0085] In the formula, r t+1 (k,k) is the degree to which the k-th standard real-time log data is suitable as the first cluster center as a whole at the (t + 1)-th iteration; Σ j≠i,k max{r t+1 (j,k),0} is the degree to which the k-th standard real-time log data is suitable as the first cluster center other than the standard real-time log data i at the (t + 1)-th iteration; a t+1 (i,k) is the degree to which the i-th standard real-time log data selects the k-th standard real-time log data as its first cluster center at the (t + 1)-th iteration, that is, the belonging information of the k-th standard real-time log data to the i-th standard real-time log data;

[0086] S3-2: Introduce an iterative attenuation coefficient to update the initial attraction information and initial belonging information of each standard real-time log data, and obtain the updated attraction information and updated belonging information;

[0087] The iterative update formula is as follows:

[0088] r' t+1 r'(i,k) = λ * r(i,k) t +(1 - λ) * r(i,k) t+1 r(i,k)

[0089] a' t+1 a'(i,k) = λ * a(i,k) t +(1 - λ) * a(i,k) t+1 a(i,k)

[0090] Wherein, r' t+1 (i,k), a' t+1 (i,k) are the updated attraction information and updated attribution information of the k-th standard real-time log data to the i-th standard real-time log data at the (t + 1)-th iteration; λ is the iterative attenuation coefficient; r t (i,k), r t+1 (i,k) are the attraction information of the k-th standard real-time log data to the i-th standard real-time log data at the t-th and (t + 1)-th iterations; a t (i,k), a t+1 (i,k) are the attribution information of the k-th standard real-time log data to the i-th standard real-time log data at the t-th and (t + 1)-th iterations; λ is the iterative attenuation coefficient;

[0091] S3-3: Update the M initial first clustering centers to obtain M updated first clustering centers, and based on the M updated first clustering centers, classify the initial log characteristics of all standard real-time log data according to the updated attraction information and updated attribution information to obtain M initial clusters of real-time log characteristics;

[0092] The judgment formula for the clustering center is:

[0093] k = argmax{a(i,k) + r(i,k)}

[0094] Wherein, both i and k are standard real-time log data indicators; if i = k, the i-th standard real-time log data is the first clustering center of the k-th standard real-time log data; if i ≠ k, the k-th standard real-time log data is the first clustering center of the i-th standard real-time log data;

[0095] S3-4: Repeat the above steps. If the number of iterations exceeds the iteration threshold or the first clustering center does not change, output the M final first clustering centers and the final initial clusters of real-time log characteristics, and use the initial log characteristic classification corresponding to the first clustering center as the initial classification of real-time log characteristics;

[0096] The initial classification of real-time log features includes security logs, network device logs, system logs, operation logs, performance logs, traffic logs, and application logs;

[0097] S4: Use the secondary classification model of log attributes to perform secondary classification of log attributes for each initial classification cluster of real-time log features, obtaining several secondary classification clusters of real-time log attributes, including the following steps:

[0098] S4-1: Input the initial classification cluster of real-time log features into the secondary classification model of log attributes, perform initialization of the FCM clustering algorithm, obtain the fuzzy factor and H initial second cluster centers, as well as the fuzzy membership degrees of all standard real-time log data in the initial classification cluster of real-time log features, where H is the total number of second cluster centers;

[0099] The formula is:

[0100] d ij =||x i -z j' || 2

[0101] In the formula, d ij is the Euclidean distance between the i-th standard real-time log data and the j'-th second cluster center; x i is the i-th standard real-time log data; z j' is the j'-th second cluster center; i is the standard real-time log data indicator; j' is the second cluster center indicator;

[0102] S4-2: Update the H initial second cluster centers according to the fuzzy membership degrees to obtain the corresponding H updated second cluster centers;

[0103] The formula for updating the fuzzy membership degree is:

[0104]

[0105] In the formula, x i is the i-th standard real-time log data; i is the standard real-time log data indicator; j' and k' are both second cluster center indicators; c is the total number of cluster centers; is the distance from the i-th standard real-time log data to the j'-th and k'-th second cluster centers; u ij' is the updated fuzzy membership degree of the i-th standard real-time log data belonging to the j'-th second cluster center;

[0106] The formula for updating the cluster center is:

[0107]

[0108] In the formula, z' j'is the second clustering center updated for the j'th time; m is the fuzzy factor; i is the standard real-time log data indicator; n' is the total number of data; j' is the second clustering center indicator; x i is the i'th standard real-time log data; u ij' is the fuzzy membership degree of the i'th standard real-time log data belonging to the j'th second clustering center;

[0109] S4-3: Use the Lagrange multiplier method to calculate the merging function, obtain the merging function value and the merging function change value. If the merging function value is greater than the function threshold, or the merging function change value is greater than the change value threshold, then continue to update the second clustering center. Otherwise, take the current second clustering center as the final second clustering center;

[0110] The formula is:

[0111]

[0112] In the formula, J t 、J t-1 are the merging function values of the Lagrange multiplier method for the t'th and (t - 1)'th iterations; ΔJ t is the corresponding change value; λ' i is the i'th characteristic parameter; t is the iteration number indicator; m is the fuzzy factor; i is the standard real-time log data indicator; n' is the total number of data; c is the total number of second clustering centers;

[0113] S4-4: Repeat the above steps to obtain H final second clustering centers, and use the log attribute sub-classifications corresponding to the final second clustering centers as the real-time log attribute sub-classifications. Divide each standard real-time log data in the real-time log characteristic initial classification cluster into the final second clustering center with the closest Euclidean distance to obtain several real-time log attribute sub-classification clusters;

[0114] The real-time log attribute sub-classifications include firewall logs, intrusion detection logs, and security event management logs belonging to security logs, router logs and switch logs belonging to network device logs, database system logs and operating system logs belonging to system logs, operator logs and audit logs belonging to operation logs, performance counter logs and performance monitoring logs belonging to performance logs, access traffic logs, operation traffic logs, and network traffic logs belonging to traffic logs, and chat program logs, office program logs, and management program logs belonging to application program logs;

[0115] S5: According to the real-time log characteristic initial classification and real-time log attribute sub-classification of each real-time log attribute sub-classification cluster, perform matching in several log audit analysis models to obtain the target log audit analysis model;

[0116] S6: Use the target log audit analysis model to perform log audit analysis on each real-time log attribute sub-classification cluster, and obtain the real-time log audit analysis results of each standard real-time log data, including the following steps:

[0117] S6-1: Perform data dimensionality reduction and matrix transformation on the real-time log attribute sub-classification cluster to obtain the dimensionality-reduced real-time log attribute sub-classification cluster matrix;

[0118] S6-2: Use the target log audit analysis model to extract N real-time dimensionality features from the dimensionality-reduced real-time log attribute sub-classification cluster matrix;

[0119] The concerned dimensions include the log time attribute dimension, the log source attribute dimension, the log operation attribute dimension, the log structure attribute dimension, and the log resource attribute dimension. Through each concerned dimension, the comprehensiveness of log audit analysis is improved, the log data information can be better characterized, and potential security threats, performance bottlenecks or other problems can be discovered;

[0120] S6-3: Perform feature fusion on the N real-time dimensionality features to obtain real-time fusion features, and perform log audit analysis based on the real-time fusion features to obtain the corresponding real-time log audit analysis results;

[0121] S6-4: Traverse all real-time log attribute sub-classification clusters, and extend the real-time log audit analysis results to all standard real-time log data in the real-time log attribute sub-classification clusters to obtain the real-time log audit analysis results of each standard real-time log data;

[0122] S7: Generate corresponding real-time alarm signals according to the real-time log audit analysis results, and visually display the standard real-time log data, the real-time log audit analysis results, and the real-time alarm signals.

[0123] Embodiment 2:

[0124] As Figure 2 shown, this embodiment provides a log audit analysis system for implementing the log audit analysis method. The system includes a cloud computing center and several log data collection devices, and the cloud computing center is respectively communicatively connected with the several log data collection devices;

[0125] The cloud computing center is provided with a model construction unit, a data mapping and transformation unit, a log feature primary classification unit, a log attribute sub-classification unit, a model matching unit, a log audit analysis unit, and a visual display unit that are connected in sequence;

[0126] The model construction unit is used to construct a data mapping and transformation model, a log feature primary classification model, a log attribute sub-classification model, and several log audit analysis models according to the historical log big data;

[0127] A data mapping and conversion unit is used to receive the real-time log big data uploaded by the log data collection device, and use the data mapping and conversion model to perform data mapping and conversion on the real-time log big data to obtain a plurality of isomorphic standard real-time log data;

[0128] A log feature preliminary classification unit, used to perform log feature preliminary classification on a number of standard real-time log data using a log feature preliminary classification model, and obtain a number of real-time log feature preliminary classification clusters;

[0129] The log attribute sub-classification unit is used to perform log attribute sub-classification on each real-time log feature initial classification cluster using the log attribute sub-classification model to obtain a plurality of real-time log attribute sub-classification clusters;

[0130] A model matching unit, used to match the real-time log characteristic initial classification and the real-time log attribute sub-classification of each real-time log attribute sub-classification cluster among several log audit analysis models to obtain a target log audit analysis model;

[0131] The log audit analysis unit is used to use the target log audit analysis model to perform log audit analysis on each real-time log attribute sub-classification cluster to obtain a real-time log audit analysis result for each standard real-time log data;

[0132] The visualization display unit is used to generate corresponding real-time alarm signals according to the real-time log audit analysis results, and to visualize the standard real-time log data, the real-time log audit analysis results and the real-time alarm signals.

[0133] The present invention discloses a log audit analysis method and system. The data mapping and conversion model constructed by the reinforcement learning algorithm can automatically map and convert the log data from different data sources, convert heterogeneous data into homogeneous data, and improve the efficiency of log audit analysis. The log characteristic primary classification model and log attribute secondary classification model constructed by the machine learning algorithm can perform primary classification on a large volume of real-time log big data according to log characteristics, perform secondary classification according to log attributes, and uniformly process and analyze log data with the same characteristics and attributes, thereby improving the efficiency of log audit analysis, being more suitable for large-volume data scenarios, improving practicality, and reducing cost investment. The log audit analysis model constructed by the deep learning algorithm can mine the deep data characteristics of log data, perform accurate and efficient log audit analysis, and reduce the error rate of log audit analysis.

[0134] The present invention is not limited to the above optional embodiments, and anyone can obtain other various forms of products under the inspiration of the present invention. The above specific embodiments should not be construed as limiting the protection scope of the present invention. The protection scope of the present invention shall be defined by the claims, and the specification can be used to interpret the claims.

Claims

1. A log audit analysis method, characterized in that: The steps include: Based on the cloud computing center and historical log big data, we build data mapping and conversion models, log feature primary classification models, log attribute secondary classification models, and several log audit analysis models, including the following steps: Based on the cloud computing center, historical log big data from different data sources is collected and preprocessed to obtain a number of heterogeneous preprocessed historical log data; According to the data structure differences between the system data structure of the cloud computing center and the pre-processed historical log data, a unified data model is constructed; Based on the unified data model and several pre-processed historical log data, the DQN algorithm is used to build a data mapping and conversion model, and generate several isomorphic standard historical log data; Based on several standard historical log data, a clustering algorithm is used to perform preliminary classification of log characteristics, a log characteristic preliminary classification model is constructed, and several historical log characteristic preliminary classification clusters corresponding to the historical log characteristic preliminary classifications are generated; Based on the initial classification cluster of each historical log feature, a clustering algorithm is used to perform log attribute sub-classification, a log attribute sub-classification model is constructed, and several historical log attribute sub-classification clusters corresponding to the historical log attribute sub-classifications are generated; Performing data dimension reduction and matrix conversion on the historical log attribute sub-classification clusters to obtain several dimension-reduced historical log attribute sub-classification cluster matrices; Based on several historical log attribute sub-classification cluster matrices after dimensionality reduction, the N-GAN-MLP algorithm is used to perform log audit analysis and build a log audit analysis model corresponding to each historical log attribute sub-classification in the initial classification of historical log characteristics, where N is the focus dimension. Receive the real-time log big data uploaded by the log data collection device, use the data mapping and conversion model to perform data mapping and conversion on the real-time log big data, and obtain a number of isomorphic standard real-time log data; Use the log feature preliminary classification model to perform preliminary classification of log features on several standard real-time log data to obtain several real-time log feature preliminary classification clusters; Using the log attribute sub-classification model, perform log attribute sub-classification on each real-time log feature initial classification cluster to obtain a number of real-time log attribute sub-classification clusters; According to the real-time log characteristic initial classification and the real-time log attribute sub-classification of each real-time log attribute sub-classification cluster, matching is performed in several log audit analysis models to obtain a target log audit analysis model; Use the target log audit analysis model to perform log audit analysis on each real-time log attribute sub-classification cluster to obtain the real-time log audit analysis results of each standard real-time log data, including the following steps: Performing data dimension reduction and matrix conversion on the real-time log attribute sub-classification clusters to obtain the real-time log attribute sub-classification cluster matrix after dimension reduction; Use the target log audit analysis model to extract N real-time dimension features of the real-time log attribute sub-classification cluster matrix after dimensionality reduction; The dimensions of concern include log time attribute dimension, log source attribute dimension, log operation attribute dimension, log structure attribute dimension, and log resource attribute dimension; Perform feature fusion on N real-time dimension features to obtain real-time fusion features, and perform log audit analysis based on the real-time fusion features to obtain corresponding real-time log audit analysis results; Traverse all real-time log attribute sub-classification clusters, and extend the real-time log audit analysis results to all standard real-time log data in the real-time log attribute sub-classification clusters to obtain the real-time log audit analysis results for each standard real-time log data; According to the real-time log audit analysis results, the corresponding real-time alarm signal is generated, and the standard real-time log data, real-time log audit analysis results and real-time alarm signal are visualized.

2. A log audit analysis method according to claim 1, characterized in that: According to several preprocessed historical log data and historical data mapping and conversion strategies, the DQN algorithm is used to build a data mapping and conversion model, and generate several isomorphic standard historical log data.

3. A log audit analysis method according to claim 1, characterized in that: According to several standard historical log data, the AP clustering algorithm is used to perform initial classification of log characteristics, build a log characteristic initial classification model, and generate several historical log characteristic initial classification clusters corresponding to the historical log characteristic initial classifications.

4. A log audit analysis method according to claim 1, characterized in that: According to the initial classification cluster of each historical log feature, the FCM clustering algorithm is used to perform log attribute sub-classification, build a log attribute sub-classification model, and generate several historical log attribute sub-classification clusters corresponding to the historical log attribute sub-classifications.

5. A log audit analysis method according to claim 2, characterized in that: Receiving the real-time log big data uploaded by the log data collection device, using the data mapping and conversion model, performing data mapping and conversion on the real-time log big data, and obtaining a plurality of isomorphic standard real-time log data, including the following steps: Receiving real-time log big data from different data sources uploaded by a log data collection device, preprocessing the real-time log big data, and obtaining a plurality of heterogeneous preprocessed real-time log data; According to the unified data model and pre-processed real-time log data, use the data mapping and conversion model to generate the corresponding real-time data mapping and conversion strategy; According to the real-time data mapping and conversion strategy, the pre-processed real-time log data is mapped and converted to obtain standard real-time log data; Traverse all preprocessed real-time log data to obtain several isomorphic standard real-time log data.

6. A log audit analysis method according to claim 3, characterized in that: Using the log feature preliminary classification model, the log feature preliminary classification is performed on several standard real-time log data to obtain several real-time log feature preliminary classification clusters, including the following steps: Input a number of standard real-time log data into the log characteristic initial classification model, initialize the AP clustering algorithm, and obtain M initial first cluster centers, as well as the initial attraction information and initial attribution information of each standard real-time log data, where M is the total number of first cluster centers; Introducing an iterative attenuation coefficient, updating the initial attraction information and initial attribution information of each standard real-time log data, and obtaining updated attraction information and updated attribution information; The M initial first cluster centers are updated to obtain M updated first cluster centers, and based on the M updated first cluster centers, all standard real-time log data are preliminarily classified according to log characteristics according to updated attraction information and updated attribution information to obtain M real-time log characteristic preliminarily classified clusters; Repeat the above steps. If the number of iterations exceeds the iteration number threshold or the first cluster center does not change, output M final first cluster centers and the final real-time log feature initial classification clusters, and use the log feature initial classification corresponding to the first cluster center as the real-time log feature initial classification.

7. A log audit analysis method according to claim 4, characterized in that: Using the log attribute sub-classification model, each real-time log feature initial classification cluster is sub-classified by log attribute to obtain several real-time log attribute sub-classification clusters, including the following steps: Input the real-time log feature initial classification cluster into the log attribute secondary classification model, initialize the FCM clustering algorithm, obtain the fuzzy factor and H initial second cluster centers, and the fuzzy membership of all standard real-time log data in the real-time log feature initial classification cluster, where H is the total number of second cluster centers; According to the fuzzy membership, the H initial second cluster centers are updated to obtain the corresponding H updated second cluster centers; Use the Lagrange multiplier method to calculate the merge function, and obtain the merge function value and the merge function change value. If the merge function value is greater than the function threshold, or the merge function change value is greater than the change value threshold, then continue to update the second cluster center, otherwise, use the current second cluster center as the final second cluster center; Repeat the above steps to obtain H final second cluster centers, divide each standard real-time log data in the real-time log feature initial classification cluster to the final second cluster center with the closest Euclidean distance, and obtain several real-time log attribute sub-classification clusters.

8. A log audit analysis system, used to implement the log audit analysis method according to any one of claims 1 to 7, characterized in that: The system includes a cloud computing center and several log data collection devices, and the cloud computing center is respectively connected to the several log data collection devices for communication; The cloud computing center is provided with a model building unit, a data mapping and conversion unit, a log characteristic primary classification unit, a log attribute secondary classification unit, a model matching unit, a log audit analysis unit and a visualization display unit which are connected in sequence.

Citation Information

Patent Citations

  • A cloud user behavior audit system and method based on cloud log analysis

    CN109471846A