Data security log analysis system and method
By using distributed log collection, precise data filtering and standardized processing, machine learning anomaly detection, comprehensive risk assessment, graph algorithm correlation analysis, Bayesian network threat intelligence fusion, data visualization, flexible alarm and automatic response, online learning optimization and other technical means in the data security log analysis system, the technical limitations of the existing system in terms of acquisition, preprocessing, detection, evaluation, visualization, alarm and response are solved, and more efficient, accurate and flexible data security analysis is achieved.
Patent Information
- Application Number
- CN202510143526.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing data security log analysis system has many challenges and technical limitations in log collection, preprocessing, abnormality detection, risk assessment, visualization, alarm and automatic response, and it is difficult to deal with complex cyber attacks and intelligent threats.
The distributed log acquisition tool and multi-threaded parallel acquisition mechanism are adopted to ensure rapid acquisition in a high-concurrency environment; log preprocessing is performed through precise data filtering and standardized algorithms; abnormal detection is performed using machine learning algorithms such as DBSCAN and hidden Markov models; risk assessment is comprehensively considered to the severity, frequency of occurrence and data sensitivity of the anomalies; correlation analysis and threat intelligence fusion are performed using graph algorithms and Bayesian networks; data visualization library is used to generate an intuitive visual interface; alarm priority and method are adjusted according to risk levels, and automatic response is achieved; and analysis models are continuously optimized through online learning algorithms.
It realizes rapid and complete collection of logs in a high-concurrency environment, removes noise data, and accurately recognizes complex abnormal behaviors, provides accurate risk assessment, intuitive visual display, flexible alarms and automatic responses, timely adapts to changes in the network environment, and improves the efficiency and accuracy of data security analysis.
Smart Images

Figure CN120162776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security log analysis, and particularly to a data security log analysis system and method. Background Art
[0002] In today's digital age, various organizations and enterprises generate a vast amount of data, which is stored, transmitted, and processed in a network environment and faces increasingly severe security threats. To ensure data security, data security logs have become an important means of monitoring and analyzing potential security threats. However, existing data security log analysis faces many challenges and technical limitations.
[0003] First of all, in terms of log collection, traditional methods can usually only obtain logs from limited data sources and are difficult to cope with large-scale distributed environments. Logs generated by different network devices, servers, and applications have different formats, which brings great difficulties to the unified collection of logs. Moreover, existing collection tools may experience collection delays or data loss in a high-concurrency environment, resulting in incomplete log information and affecting subsequent security analysis. At the same time, there are also deficiencies in the media and strategies for storing logs, lacking sufficient redundancy and efficient storage methods, unable to ensure the integrity and timeliness of log data, and prone to performance bottlenecks when storing a large amount of log data, affecting data availability.
[0004] In the log preprocessing stage, the cleaning and standardization of log data are relatively rough. A large amount of noise data will interfere with subsequent analysis work, but existing processing methods often cannot accurately remove this noise and are difficult to convert logs in different formats into a unified structure, resulting in data inconsistency problems. Different log sources may use different timestamp formats, log level identifiers, and event description methods, which pose obstacles to the further processing and analysis of data and make it difficult for subsequent analysis work to be based on accurate and consistent data.
[0005] For anomaly detection, traditional rule-based methods are difficult to adapt to complex and changing network attack forms. Simple rule matching cannot detect complex anomaly behavior patterns. As network attack means become increasingly intelligent and concealed, relying solely on static rules cannot meet the identification requirements for abnormal behaviors and cannot effectively detect according to constantly changing behavior patterns. Moreover, existing anomaly detection methods usually do not consider the correlation between data and lack effective identification capabilities for distributed and complex attack forms, and may miss some potential security threats.
[0006] The risk assessment process is usually relatively simple and static. It assesses abnormal behaviors based on a single dimension only, without comprehensively considering multiple factors such as the severity, occurrence frequency of abnormalities, and the sensitivity of the data involved. This leads to inaccurate risk assessment results and fails to provide effective decision-making basis for security personnel.
[0007] In terms of visualization, existing tools usually can only provide simple charts and cannot intuitively and comprehensively display complex security postures. They cannot perform reasonable visual presentations according to the data volume and dimensions, making it impossible for security personnel to quickly understand and grasp the occurrence trends, risk distributions, and threat types of abnormal events, thus limiting the overall understanding and judgment of the security posture.
[0008] The alarm and automatic response mechanisms are also not perfect enough. Alarm information is often not detailed and accurate enough, and it cannot adjust the priority and method of alarms according to the risk level. The notification for high-risk events is not timely or sufficient. At the same time, the ability of automatic response is limited. There is a lack of automatic blocking and processing mechanisms for high-risk abnormal events, or the processing means are single, and it cannot perform flexible automatic responses according to different risk situations. As a result, effective response measures cannot be taken in a timely manner when security events occur, increasing security risks and potential losses.
[0009] In addition, existing systems lack the ability of continuous learning. They cannot dynamically update and optimize the analysis model according to new log data and analysis results, and cannot adapt to the constantly changing network environment and emerging security threats. Over time, their analysis capabilities will gradually lag behind, making it difficult to ensure the long-term effectiveness of data security.
[0010] In summary, in order to cope with the increasingly complex data security challenges, a more comprehensive, intelligent, and efficient data security log analysis system and method are needed to overcome the various deficiencies of the existing technologies and improve the overall level of data security protection. Summary of the Invention
[0011] The data security log analysis system and method proposed by the present invention are used to solve the problems mentioned in the above existing technologies.
[0012] To achieve the above object, the present invention adopts the following technical solutions: The data security log analysis method includes:
[0013] Log collection step: Collect data security logs. Use a distributed log collection tool to calculate the collection time T according to the formula T = N / (n*r), where N is the total amount of logs, n is the number of threads, and r is the collection rate of each thread, ensuring that log information can be collected quickly and completely under high concurrency. Store the collected logs in a high-speed storage medium, and determine the actual storage capacity S according to the formula S = S0*(1 + k), where S0 is the original storage capacity and k is the redundancy coefficient, ensuring the integrity and timeliness of log data.
[0014] Log preprocessing steps: Clean and standardize the collected log data, remove noise data, and use a data filtering algorithm to calculate according to the formula p = N n / N t where N n is the number of noise data, and N t is the total number of data. Filter out the part where the proportion of noise data exceeds p. Unify the timestamps, log levels, and event description information formats of the logs, convert logs from different sources into a unified data structure, use a data mapping table to map data in different formats to a standard structure, calculate the hash value of the log data through a hash algorithm, and assign a unique hash value H to each log according to the formula H = hash(D).
[0015] Anomaly detection steps: Apply machine learning algorithms. For the clustering algorithm, use the density-based clustering algorithm DBSCAN, and judge whether data points p and q belong to the same cluster according to the formula d(p, q) < ∈, where ∈ is the distance threshold, to discover dense and sparse regions of data points, and then find out abnormal behavior patterns; for the anomaly detection algorithm, use a statistics-based anomaly detection algorithm, calculate the Z-score according to the formula Z = (x - μ) / σ, where x is the data point, μ is the mean, and σ is the standard deviation, and determine the points with Z-scores exceeding the set threshold as anomalies, and identify abnormal behavior patterns, including but not limited to abnormal logins, data tampering, and illegal access. Establish a normal behavior model based on historical data, and use a hidden Markov model to calculate the probability of the state sequence X, where S i is the hidden state, and compare it with the current data to judge whether it is abnormal.
[0016] Risk assessment steps: Conduct a risk assessment on the detected abnormal behaviors. According to the severity S of the anomaly, the occurrence frequency f , and the data sensitivity factor d involved, use the risk assessment formula R = s * f * d to assign a risk level to each abnormal event. The value range of R is from 0 to 100, and the specific rule is that R > 70 is a high risk, 30 < R ≤ 70 is a medium risk, and R ≤ 30 is a low risk, providing a basis for subsequent processing.
[0017] Furthermore, it also includes:
[0018] Association analysis steps: Apply graph algorithms. Use the depth-first search algorithm DFS to traverse the association graph between log events, calculate the vertex set V of the association graph according to the formula V = V0 + E, where V0 is the initial vertex set and E is the edge set, and find multi-step attacks or distributed attacks. Through IP addresses, user identities, and time association information, calculate the association degree C according to the association degree formula where w i is the weight of different association information, and c iIt is the quantization value of associated information.
[0019] Threat intelligence fusion steps: Integrate data from external threat intelligence sources, including open-source threat intelligence libraries and intelligence from professional security agencies. Adopt data fusion algorithms, and fuse threat intelligence I1, I2, I3 from different sources according to the formula F = α×I1 + β×I2 + γ×I3, where α, β, γ are the weights of different intelligence sources. Combine external threat intelligence with the abnormal information detected internally, and update the threat probability through the Bayesian network P(A|B) = P(B|A)P(A) / P(B) to improve the recognition ability of unknown threats and enhance the accuracy and foresight of analysis.
[0020] Visualization steps: Display the analysis results in a visual way, including generating charts and dashboards. Adopt data visualization libraries, including D3.js or ECharts, and select different visualization types according to the data volume V d and dimension D. When V d > V t h and D > D th use 3D charts, otherwise use 2D charts to intuitively present the occurrence trend, risk distribution, and threat type information of abnormal events. Use the color mapping formula C = f(R) to map the risk level R to the color C to facilitate security personnel to quickly understand and master the security situation.
[0021] Alarm steps: When detecting high-risk abnormal events, send alarm information through multiple channels, including emails, text messages, and instant messaging tools. For email alarms, determine the email priority according to the email priority formula P m = R / 100, where R is the risk level; for text message alarms, determine the text message content length L according to the text message length formula L s = k×R to determine the text message content length L s , where k is the proportionality coefficient; for instant messaging tool alarms, determine the urgency U of the information according to the information urgency formula U = 1 - exp(-R / 100), and notify security personnel. The alarm information includes the detailed information of abnormal events, including time, location, event type, and risk level.
[0022] Automatic response steps: For some predefined high-risk abnormal events, the response mechanism can be automatically triggered. Adopt the policy decision tree algorithm to make a decision D according to the formula D = f(R,A), where R is the risk level and A is the abnormal type. The specific rule is to automatically block the IP address of the attack source, determine the blocking duration B according to the formula B = f(R,t), where t is the time parameter, pause the relevant services or accounts, and minimize losses and impacts as much as possible under the premise of ensuring security.
[0023] Continuous learning steps: Continuously update and optimize the model of the machine learning algorithm based on new log data and analysis results. Use online learning algorithms to update the model parameters w according to the gradient descent formula. Here, α is the learning rate, and is the gradient of the loss function, so as to adapt to the changing network environment and security threats and ensure that the analysis ability of the system keeps pace with the times.
[0024] Log collection module: Equipped with distributed log collection tools, it can collect data security logs from multiple data sources, support multiple log formats, and adopt a multi-threaded parallel collection mechanism. The number of threads n is determined according to the formula n = C / r, where C is the number of concurrent connections and r is the processing capacity of a single thread. Store the collected logs in a high-speed storage medium, which uses a hybrid storage of SSD and HDD. Determine the storage performance P according to the storage performance formula P = k1 * k2 + P SSD *P HDD Here, k1 and k2 are the performance weights of SSD and HDD, and it has the characteristics of high throughput and low latency.
[0025] Log preprocessing module: Clean and standardize the collected log data. Use regular expression matching and replacement algorithms to remove noise data, and calculate the number of filtered logs N according to the formula Nf = N - N n Here, N f is the number of noise data. It can process logs with different structures and formats, output a unified data structure, and use pattern matching algorithms to convert data in different formats into a standard structure to improve the efficiency of subsequent analysis. n
[0026] Anomaly detection module: Use machine learning algorithms to perform anomaly detection on log data and establish a normal behavior model. Adopt a neural network model to calculate the output y according to the neuron activation function Here, w i is the weight, xx is the input, and b is the bias, which has high accuracy and low false alarm rate.
[0027] Association analysis module: Perform association analysis on log data from different data sources to find multi-step attacks or distributed attacks. Use a graph database to store association information, and store the association graph according to the graph storage formula G = (V, E), where V is the vertex set and E is the edge set. It has a powerful ability to mine association rules. Use the shortest path algorithm SP = shortestPath(G, s, t) to find the shortest path from the source node s to the target node t in the association graph and mine deep-level association relationships.
[0028] Control Center: Integrates functions of risk assessment, visualization, alerting, automatic response, and continuous learning. Conducts risk assessment according to the risk assessment formula R = s × f × d, generates a visualization interface using a visualization tool library, and sends alerts through the alert scheduling algorithm A s = f(R, T), where T is a time parameter, triggers an automatic response, makes response decisions automatically according to the decision tree algorithm, updates the machine learning model using the online learning algorithm, and realizes the comprehensive management and intelligent analysis of data security logs.
[0029] Compared with the existing technologies, the beneficial effects of the present invention are as follows:
[0030] In terms of log collection, through a distributed log collection tool and a multi-threaded parallel collection mechanism, it can quickly and completely collect logs from multiple data sources in a high-concurrency environment, and uses a redundant storage strategy to ensure the integrity and timeliness of log data, improving the efficiency and reliability of data collection. In the log preprocessing link, through precise data filtering and standardization algorithms, noise is effectively removed, and logs in different formats are converted into a unified structure, laying a good foundation for subsequent analysis. Anomaly detection uses advanced machine learning algorithms to accurately identify complex anomaly behavior patterns, including multi-step attacks and zero-day attacks, improving the accuracy and comprehensiveness of anomaly detection. Risk assessment calculates by integrating multiple factors, assigns a reasonable risk level to abnormal events, and provides a more accurate basis for security decisions. Visualization presentation is optimized according to the data characteristics, and through rich visualization effects, security personnel can intuitively grasp the security situation, which helps for quick response. The alerting and automatic response functions can be flexibly adjusted according to the risk level, improving the effectiveness of alerts and the pertinence of automatic responses, and reducing the impact of security incidents. The continuous learning ability enables the system to keep up with the times and continuously optimize the analysis model, ensuring that the system always maintains a high level of security analysis ability, providing all-round and multi-level effective guarantees for data security. Brief Description of the Drawings
[0031] Figure 1 It is a schematic block diagram of a data security log analysis method proposed by the present invention;
[0032] Figure 2 It is a schematic block diagram of a data security log analysis system proposed by the present invention. Detailed Embodiments
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0035] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined. In addition, the terms "mounted", "connected", and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below with reference to the drawings.
[0036] Refer to Figure 1-2 : A data security log analysis system, comprising:
[0037] Log collection step: Collect data security logs from multiple data sources, and the data sources include but are not limited to network devices, servers, and applications. Use a distributed log collection tool, which adopts a multi-thread parallel collection mechanism, and calculates the collection time T according to the formula T = N / (n * r), where N is the total amount of logs, n is the number of threads, and r is the collection rate of each thread, to ensure that log information can be collected quickly and completely under high concurrency. Support multiple log formats, store the collected logs in a high-speed storage medium, and adopt a redundant storage strategy when storing. Determine the actual storage capacity S according to the formula S = S0 * (1 + k), where S0 is the original storage capacity and k is the redundancy coefficient, to ensure the integrity and timeliness of log data.
[0038] Log preprocessing step: Clean and standardize the collected log data, remove noise data, and use a data filtering algorithm. The specific rule is a threshold-based filtering algorithm, and it will be calculated according to the formula p = N n / N t Calculate, N nis the number of noise data, N t The part where the proportion of noise data in the total data quantity exceeds p is filtered out. It includes the timestamp, log level, and event description information format of unified logs, converts logs from different sources into a unified data structure, uses a data mapping table to map data in different formats into a standard structure, calculates the hash value of log data through a hash algorithm, and assigns a unique hash value H to each log according to the formula H = hash(D) for subsequent analysis.
[0039] Anomaly detection steps: Apply machine learning algorithms, including clustering algorithms and anomaly detection algorithms. For the clustering algorithm, use the density-based clustering algorithm DBSCAN, and judge whether data points p and q belong to the same cluster according to the formula d(p, q) < ∈, where ∈ is the distance threshold, to discover the dense and sparse regions of data points and then find out the abnormal behavior patterns; for the anomaly detection algorithm, use the statistics-based anomaly detection algorithm, calculate the Z-score according to the formula Z = (x - μ) / σ, where x is the data point, μ is the mean value, and σ is the standard deviation, and determine the points with Z-scores exceeding the set threshold as anomalies, and identify abnormal behavior patterns, including but not limited to abnormal logins, data tampering, and illegal access. Establish a normal behavior model based on historical data and use the hidden Markov model Calculate the probability of the state sequence X, S i is the hidden state, and compare it with the current data to judge whether it is abnormal.
[0040] Risk assessment steps: Conduct a risk assessment on the detected abnormal behaviors. According to the factors of the severity s, occurrence frequency f, and data sensitivity d involved in the anomaly, use the risk assessment formula R = s × f × d to assign a risk level to each abnormal event. The value range of R is from 0 to 100, R > 70 is high risk, 30 < R ≤ 70 is medium risk, and R ≤ 30 is low risk, providing a basis for subsequent processing.
[0041] In the present invention, a data security log analysis system further includes the following steps:
[0042] Association analysis steps: Conduct an association analysis on the log data from different data sources to find the correlation between multiple log events. Apply graph algorithms, including the depth-first search algorithm DFS to traverse the association graph between log events, and calculate the vertex set V of the association graph according to the formula V = V0 + E, where V0 is the initial vertex set and E is the edge set, to find multi-step attacks or distributed attacks. Through IP addresses, user identities, and time association information, calculate the association degree C according to the association degree formula calculate the association degree C, w i is the weight of different association information, c i is the quantization value of the association information, and discover the internal connection between abnormal behaviors.
[0043] Threat intelligence fusion steps: Integrate data from external threat intelligence sources, including open-source threat intelligence libraries and intelligence from professional security agencies. Adopt data fusion algorithms to fuse threat intelligence from different sources I1, I2, I3 according to the formula F = α×I1 + β×I2 + γ×I3, where α, β, γ are the weights of different intelligence sources. Combine external threat intelligence with the abnormal information detected internally, and update the threat probability through the Bayesian network P(A|B) = P(B|A)P(A) / P(B) to improve the recognition ability of unknown threats and enhance the accuracy and foresight of analysis.
[0044] Visualization steps: Display the analysis results in a visual way, including generating charts and dashboards. Adopt a data visualization library, and the specific rules are D3.js or ECharts. Select different visualization types according to the data volume V d and the dimension D. When V d >V th and D > D th use 3D charts, otherwise use 2D charts to intuitively present the occurrence trend, risk distribution, and threat type information of abnormal events. Use the color mapping formula C = f(R) to map the risk level R to the color C to facilitate security personnel to quickly understand and master the security situation.
[0045] Alarm steps: When detecting high-risk abnormal events, send alarm information through multiple channels, including emails, text messages, and instant messaging tools. For email alarms, determine the email priority according to the email priority formula P m = R / 100, where R is the risk level; for text message alarms, determine the text message content length L according to the text message length formula L s = k×R to determine the text message content length L s , where k is the proportionality coefficient; for instant messaging tool alarms, determine the urgency U of the information according to the information urgency formula U = 1 - exp(-R / 100), and notify security personnel. The alarm information includes the detailed information of abnormal events, including time, location, event type, and risk level.
[0046] Automatic response steps: For some predefined high-risk abnormal events, the response mechanism can be automatically triggered. Adopt the policy decision tree algorithm to make a decision D according to the formula D = f(R,A), where R is the risk level and A is the abnormal type. The specific rule is to automatically block the IP address of the attack source, and determine the blocking duration B according to the formula B = f(R,t), where t is the time parameter, pause the relevant services or accounts, and minimize losses and impacts as much as possible under the premise of ensuring security.
[0047] Continuous learning steps: Continuously update and optimize the model of the machine learning algorithm according to new log data and analysis results. Use the online learning algorithm according to the gradient descent formula Update the model parameters w, where α is the learning rate, is the gradient of the loss function to adapt to the changing network environment and security threats, ensuring that the system's analysis capabilities keep pace with the times.
[0048] Log collection module: Equipped with a distributed log collection tool, it can collect data security logs from multiple data sources, support multiple log formats, and adopt a multi-threaded parallel collection mechanism. The number of threads n is determined according to the formula n = C / r, where C is the number of concurrent connections and r is the processing capacity of a single thread; the collected logs are stored in a high-speed storage medium, and the storage medium uses a hybrid storage of SSD and HDD. According to the storage performance formula P = k1*k2 + P SSD *P HDD to determine the storage performance P, where k1 and k2 are the performance weights of SSD and HDD, with the characteristics of high throughput and low latency.
[0049] Log preprocessing module: Cleans and standardizes the collected log data, uses a regular expression matching and replacement algorithm to remove noise data, and calculates the number of filtered logs N according to the formula N f = N - N n where N f is the number of noise data. It can process logs with different structures and formats, output a unified data structure, and use a pattern matching algorithm to convert data in different formats into a standard structure to improve the efficiency of subsequent analysis. n
[0050] Anomaly detection module: Applies machine learning algorithms to detect anomalies in log data and establish a normal behavior model. Adopts a neural network model, calculates the output y according to the neuron activation function where w i is the weight, x i is the input, and b is the bias, with high accuracy and low false alarm rate.
[0051] The present invention also discloses a method for a data security log analysis system, including:
[0052] Association analysis module: Performs association analysis on log data from different data sources to find multi-step attacks or distributed attacks. Uses a graph database to store association information, stores the association graph according to the graph storage formula G = (V, E), where V is the vertex set and E is the edge set, with a powerful association rule mining ability. Uses the shortest path algorithm SP = shortestPath(G, s, t) to find the shortest path from the source node s to the target node t in the association graph and mine deep-level association relationships.
[0053] Control Center: Integrates functions of risk assessment, visualization, alerting, automatic response, and continuous learning. Conducts risk assessment according to the risk assessment formula R = s × f × d, generates a visualization interface using the visualization tool library, sends alerts through the alert scheduling algorithm A s = f(R, T), where T is a time parameter, triggers an automatic response, makes response decisions automatically according to the decision tree algorithm, and updates the machine learning model using the online learning algorithm to achieve comprehensive management and intelligent analysis of data security logs.
[0054] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A data security log analysis method, characterized in that: include: Log collection steps: Collect data security logs, use distributed log collection tools, calculate the collection time T according to the formula T = N / (n*r), where N is the total number of logs, n is the number of threads, and r is the collection rate of each thread. Store the collected logs in high-speed storage media, and determine the actual storage capacity S according to the formula S = S0*(1+k), where S0 is the original storage capacity and k is the redundancy coefficient; Log preprocessing step: clean and standardize the collected log data, remove noise data, and use data filtering algorithm according to the formula p = N n / N t Calculation, N n is the number of noise data, N t is the total amount of data. Filter out the part of noise data that accounts for more than p. Unify the timestamp, log level, and event description information format of the log. Convert logs from different sources into a unified data structure. Use a data mapping table to map data in different formats to a standard structure. Calculate the hash value of the log data through a hash algorithm. According to the formula H=hash(D), assign a unique hash value H to each log. Anomaly detection steps: Use machine learning algorithms. For clustering algorithms, the density-based clustering algorithm DBSCAN is used. According to the formula d(p, q) < ∈, it is determined whether data points p and q belong to the same cluster. ∈ is the distance threshold, which is used to find dense and sparse areas of data points and find abnormal behavior patterns. For the anomaly detection algorithm, a statistical anomaly detection algorithm is used to calculate the Z score according to the formula Z = (x-μ) / σ, where x is the data point, μ is the mean, and σ is the standard deviation. Points with Z scores exceeding the set threshold are judged as abnormal, and abnormal behavior patterns are identified, including abnormal logins, data tampering, and illegal access; a normal behavior model is established based on historical data, using a hidden Markov model Calculate the probability of state sequence X, S i It is a hidden state, which is compared with the current data to determine whether it is abnormal; Risk assessment steps: Conduct risk assessment on the detected abnormal behavior. According to the severity S of the abnormality, the frequency f of occurrence, and the sensitivity d of the data involved, use the risk assessment formula R = s*f*d to assign a risk level to each abnormal event. The value range of R is 0 to 100. The specific rules are R>70 for high risk, 30<R≤70 for medium risk, and R≤30 for low risk.
2. The data security log analysis method according to claim 1, characterized in that: Also includes: Correlation analysis steps: Use graph algorithms, specifically the priority search algorithm DFS to traverse the correlation graph between log events, and calculate the vertex set V of the correlation graph according to the formula V=V0+E, where V0 is the initial vertex set and E is the edge set. Find multi-step attacks or distributed attacks, and use IP addresses, user identities, and time correlation information according to the correlation formula Calculate the correlation C, w i is the weight of different associated information, c i It is the quantitative value of the association information.
3. The data security log analysis method according to claim 1, characterized in that: Also includes: Threat intelligence fusion steps: Integrate data from external threat intelligence sources, including open source threat intelligence libraries and intelligence from professional security agencies. Use data fusion algorithms to fuse threat intelligence from different sources according to the formula F = α × I1 + β × I2 + γ × I3. I1, I2, I3, α, β, γ are the weights of different intelligence sources. Combine external threat intelligence with internally detected abnormal information, and update the threat probability through the Bayesian network P(A|B) = P(B|A)P(A) / P(B).
4. The data security log analysis method according to claim 1, characterized in that: Also includes: Visualization step: Display the analysis results in a visual way, including generating charts and dashboards; Use data visualization libraries, including D3.js or ECharts, depending on the amount of data V d and dimension D to select different visualization types. d >V th And D>D th Use 3D charts when possible, otherwise use 2D charts to intuitively present the occurrence trend, risk distribution, and threat type information of abnormal events. Use the color mapping formula C = f(R) to map the risk level R to color C.
5. The data security log analysis method according to claim 1, characterized in that: Also includes: Alarm steps: When a high-risk abnormal event is detected, alarm information is sent through multiple channels, including email, SMS, and instant messaging tools; For email alerts, according to the email priority formula P m =R / 100 determines the priority of the email, where R is the risk level; For SMS alerts, according to the SMS length formula L s = k*R determines the length of the SMS content L s , k is the proportionality coefficient; for instant messaging tool alarms, the urgency U of the information is determined according to the information urgency formula U=1-exp(-R / 100), and the security personnel are notified. The alarm information contains detailed information of the abnormal event, including time, location, event type, and risk level.
6. The data security log analysis method according to claim 1, characterized in that: Also includes: Automatic response steps: For some predefined high-risk abnormal events, the response mechanism is automatically triggered; the policy decision tree algorithm is used to make decision D according to the formula D=f(R,A), where R is the risk level and A is the abnormality type. The specific rule is to automatically block the IP address of the attack source, determine the blocking time B according to the formula B=f(R,t), where t is the time parameter, and suspend related services or accounts.
7. The data security log analysis method according to claim 1, characterized in that: Also includes: Continuous learning step: Continuously update and optimize the machine learning algorithm model based on new log data and analysis results; Using online learning algorithm, according to the gradient descent formula Update the model parameters w, α is the learning rate, is the gradient of the loss function.
8. A system for implementing the data security log analysis method according to any one of claims 1 to 7, characterized in that: include: Log collection module: equipped with distributed log collection tools, collects data security logs from multiple data sources, supports multiple log formats, and adopts a multi-threaded parallel collection mechanism. The number of threads n is determined according to the formula n = C / r, where C is the number of concurrent connections and r is the processing capacity of a single thread; The collected logs are stored in high-speed storage media, which uses SSD and HDD hybrid storage. According to the storage performance formula P = k1*k2+P SSD *P HDD Determine the storage performance P, where k1 and k2 are the performance weights of SSD and HDD; Log preprocessing module: cleans and standardizes the collected log data, uses regular expression matching and replacement algorithms to remove noise data, and uses the formula N f =NN n Calculate the number of filtered logs N f , N n It is the amount of noise data,processing logs of different structures and formats,outputting unified data structures,using pattern matching algorithms to convert data of different formats into standard structures,improving the efficiency of subsequent analysis; Anomaly detection module: Use machine learning algorithms to detect anomalies in log data, establish normal behavior models, and use neural network models based on neuron activation functions. Calculate the output y, w i is the weight, x i is the input and b is the bias.
9. The data security log analysis system according to claim 8, characterized in that: Also includes: Association analysis module: performs association analysis on log data from different data sources to identify multi-step attacks or distributed attacks, uses graph database to store association information, stores association graph according to graph storage formula G=(V, E), where V is the vertex set and E is the edge set. It has the ability to mine association rules, and uses the shortest path algorithm SP=shortestPath(G, s, t) to find the shortest path from source node s to target node t in the association graph, and mines deep association relationships.
10. The data security log analysis system according to claim 8, characterized in that: Control center: Integrates risk assessment, visualization, alarm, automatic response and continuous learning functions, performs risk assessment according to the risk assessment formula R = s × f × d, uses the visualization tool library to generate a visualization interface, and uses the alarm scheduling algorithm A s =f(R, T) sends an alarm, T is the time parameter, triggers an automatic response, automatically makes a response decision based on the decision tree algorithm, and uses the online learning algorithm to update the machine learning model.
Citation Information
Cited By
Data management optimization method and system
CN120915485A
Post-payment intelligent payment method and system based on pre-authorization supervision
CN121365971A