Threat analysis method and system based on big data

The method improves threat analysis by synchronizing user behavior with system logs using HMM models to enhance threat identification speed and accuracy, enabling timely and precise responses to complex threats through dynamic adjustment.

CN120316433AInactive Publication Date: 2025-07-15SHANDONG WANHE BIG DATA CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510386515.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120316433A_ABST
    Figure CN120316433A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of threat analysis, in particular to a threat analysis method and system based on big data, and the method comprises the following steps: based on collected user behavior data, carrying out time labeling processing, synchronizing user activity data and system logs, extracting time labels and behavior indexes, and utilizing an HMM (Hidden Markov Model) to obtain a threat analysis result; and evaluating a state transition probability according to the transition frequency and the time interval to obtain a state anomaly index. According to the method, the state transition probability is evaluated by applying time tagging processing and the hidden Markov model, so that the recognition speed and accuracy of potential threats are greatly improved, user behaviors and system log data are synchronized, the real-time analysis capability of the data is enhanced, early threats can be recognized in time, and the safety of the early threats is improved. The threat source is accurately positioned, the response accuracy is further improved through matching of the threat source and a known threat mode, fine analysis of an abnormal conversion sequence is converted into targeted safety adjustment, and real-time performance and target performance of safety response are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of threat analysis, and particularly to a threat analysis method and system based on big data. Background Art

[0002] Threat analysis technology mainly focuses on identifying, evaluating, and responding to various security threats to protect information systems and networks from potential malicious attacks. The goal of threat analysis is to provide early warnings of potential security threats by predicting and identifying attack patterns, and to take measures to prevent or mitigate their impacts. The core lies in parsing and analyzing network data, user behavior, and application operations in order to detect and respond to various security incidents in a timely manner.

[0003] Among them, the threat analysis method of big data involves applying big data technology to enhance traditional threat analysis methods. By processing and analyzing large-scale data sets, implicit patterns and associations can be mined, thereby effectively predicting and identifying potential security threats. It has a wide range of uses, from network security, fraud detection to internal threat monitoring and other fields. The processing power of big data can be used to improve the speed and accuracy of threat detection, so as to quickly respond to immediate threats and effectively improve the security protection level.

[0004] The prior art lacks effective data synchronization and real-time analysis mechanisms, resulting in insufficient response efficiency for early warnings and the inability to implement security protection measures in a timely manner. Especially when identifying and responding to complex threats such as APT attacks, conventional analysis fails to provide sufficient sensitivity. In addition, the prior art mostly adopts static security policies in threat comparison and risk assessment, lacking the necessary dynamic adjustment ability, and appears inflexible in the face of changing attack strategies. Such a static and lagging security response mechanism cannot detect and effectively prevent new or mutated security threats in a timely manner. Summary of the Invention

[0005] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose a threat analysis method and system based on big data.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A threat analysis method based on big data, including the following steps,

[0007] S1: Based on the collected user behavior data, perform time tagging processing, synchronize user activity data with system logs, extract time tags and behavior metrics, and use the HMM model to evaluate the state transition probability according to the conversion frequency and time interval to obtain the state anomaly index;

[0008] S2: Based on the status anomaly index, filter out the abnormal conversion sequences, perform time series analysis on the filtered sequences, identify the data patterns that differ from the regular patterns, compare them with the known threat patterns, identify potential APT behaviors, and generate threat identification marks;

[0009] S3: Conduct a risk assessment on the threat identification marks, calculate the risk scores, set up a response mechanism according to the risk levels, implement a security adjustment mechanism for high-risk behaviors, update the security configuration, optimize the threat response process, and generate security adjustment records;

[0010] S4: Based on the security adjustment records, record the threat detection results and the response status to threats in real time, conduct continuous behavior analysis, adjust the monitoring parameters and thresholds, match new threat patterns, and generate dynamic threat adaptation configurations.

[0011] The improvement of the present invention is that the extraction steps of the time tags and behavior indicators are specifically as follows:

[0012] S111: Based on the collected user behavior data, extract the timestamps of event records and log entries, verify the consistency of the timestamps, and obtain a synchronized data set;

[0013] S112: Based on the synchronized data set, perform time tagging processing, using the formula:

[0014]

[0015] Calculate the time tag T label , and at the same time extract the behavior indicators from the log entries to obtain a set of behavior indicators with time tags. Among them, T event represents the original timestamp of the event, δT sync represents the time deviation obtained from log analysis, and γ is a time normalization coefficient used to adjust the granularity of the time tag.

[0016] The improvement of the present invention is that the steps for obtaining the status anomaly index are specifically as follows:

[0017] S121: Use the HMM model to conduct state sequence analysis, estimate each state by learning and identifying the statistical characteristics of behavior patterns, and obtain the state transition probabilities;

[0018] S122: According to the state transition probabilities, evaluate the frequency and time interval of transitions between multiple states, using the formula:

[0019]

[0020] Calculate the status anomaly index IS abnormal , where, PS ij represents the probability of transitioning from state i to state j, PSnormal is the expected normal state transition probability, WS ij is the weight for the transition from state i to j.

[0021] The improvement of the present invention is that the step of identifying the differential data pattern is specifically as follows:

[0022] S211: Based on the state anomaly index, analyze the abnormal transition sequence, identify the behavioral characteristics of each data point, and obtain a preliminary set of behavioral characteristics;

[0023] S212: Based on the preliminary set of behavioral characteristics, use the formula:

[0024]

[0025] Judge which behaviors are different from the normal pattern to obtain a set of differential data patterns, where DF cluster is the data variance measure, xf i is the value of the behavioral characteristic of the data point, μf is the average value of the behavioral characteristics of the data point, wf i is the weight based on the frequency of behavior occurrence, n Q is the total number of data points;

[0026] S213: Based on the set of differential data patterns, perform matching analysis with known normal patterns to determine the data patterns with differences.

[0027] The improvement of the present invention is that the step of obtaining the threat recognition identifier is specifically as follows:

[0028] S221: Screen the potential APT behaviors, compare them with the currently known threat behavior patterns, determine the potential threat of the behaviors, and generate a list of potential APT behaviors;

[0029] S222: Analyze the list of potential APT behaviors, determine the similarity, and identify the current security threat to obtain the threat recognition identifier.

[0030] The improvement of the present invention is that the step of calculating the risk score is specifically as follows:

[0031] S311: Conduct a risk assessment on the threat recognition identifier, extract the key features of each identifier, including the attack type, frequency, and potential impact, to obtain a set of key features;

[0032] S312: Based on the set of key features, use the formula:

[0033] GR score = G a × G f + G b × G p

[0034] Calculate the risk score GR for each threat score , where G f represents the occurrence frequency, and G p represents the potential impact. G a is the weight coefficient associated with the occurrence frequency G f and G b is the weight coefficient associated with the potential impact G p .

[0035] The improvement of the present invention is that the step of obtaining the security adjustment record is specifically as follows:

[0036] S321: Set a response mechanism according to the risk level, apply the security adjustment mechanism to the identified high-risk behaviors, and obtain a preliminary mechanism adjustment record;

[0037] S322: Based on the preliminary mechanism adjustment record, update the security configuration, and optimize the threat response process according to the adjustment effect and the feedback information of continuous monitoring to obtain the security adjustment record.

[0038] The improvement of the present invention is that the step of obtaining the dynamic threat adaptation configuration is specifically as follows:

[0039] S411: Based on the security adjustment record, extract key behavior patterns and abnormal features from the real-time recorded threat detection results and response statuses, analyze the potential threat change trend, and obtain a preliminary parameter set for adaptive adjustment by quantifying the frequency, scope, and impact of threat behaviors;

[0040] S412: Evaluate the preliminary parameter set for adaptive adjustment, and use the formula

[0041]

[0042] to calculate the optimal threshold θ for dynamic adjustment opt , and obtain the optimized dynamic adjustment parameters. Among them, θ i represents the current value of the monitoring parameter, μ θ is the parameter mean, is the parameter variance, ∈ D is the adjustment coefficient, wd i is the weight coefficient, and n R is the total number of monitoring parameters;

[0043] S413: Based on the optimized dynamic adjustment parameters, match new threat patterns by analyzing the behavior characteristics of new threats and the existing adjustment records, and update the monitoring and response rules to obtain the dynamic threat adaptation configuration.

[0044] A threat analysis system based on big data, the system includes:

[0045] Based on the collected user behavior data, the behavior data synchronization module synchronizes user activity data and system logs, extracts time tags and behavior metrics, uses the HMM model to calculate the conversion frequency and time interval, evaluates the state transition probability, and obtains the state anomaly index;

[0046] Based on the state anomaly index, the anomaly recognition and analysis module filters out abnormal conversion sequences, performs time series analysis on the filtered sequences, identifies data patterns that differ from the normal pattern, compares with known threat patterns, identifies potential APT behaviors, and generates threat recognition identifiers;

[0047] Using the threat recognition identifier, the risk response configuration module evaluates the associated risks and calculates the risk scores, formulates and implements a security adjustment mechanism according to the risk level, updates the security configuration, optimizes the threat response process, and generates a security adjustment record;

[0048] Based on the security adjustment record, the dynamic monitoring and adjustment module records the threat detection results and the response status to threats in real time, continuously analyzes the behavior patterns, adjusts the monitoring parameters and thresholds, matches new threat patterns, and obtains a dynamic threat adaptation configuration;

[0049] Using the dynamic threat adaptation configuration, the performance optimization module analyzes the current threat recognition performance, evaluates the performance bottlenecks and security weaknesses, and adjusts the response configuration and monitoring parameters again to obtain an optimized result of threat recognition performance.

[0050] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0051] In the present invention, by applying time tagging processing and the Hidden Markov Model to evaluate the state transition probability, the recognition speed and accuracy of potential threats are greatly improved. By synchronizing user behavior and system log data, the real-time data analysis ability is strengthened, enabling early threats to be identified in a timely manner. Precise positioning of the threat source further improves the accuracy of response through matching with known threat patterns. The refined analysis of abnormal conversion sequences is transformed into targeted security adjustments, achieving the real-time and targeted nature of security responses. By dynamically adjusting monitoring parameters to adapt to new threat patterns, the overall threat protection level and adaptability are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a flowchart of the threat analysis method based on big data proposed by the present invention;

[0053] Figure 2 It is a flowchart for extracting time tags and behavior metrics in the present invention;

[0054] Figure 3 It is a flowchart for obtaining the state anomaly index in the present invention;

[0055] Figure 4 It is a flow chart for identifying different data patterns in the present invention;

[0056] Figure 5 It is a flow chart for obtaining threat identification marks in the present invention;

[0057] Figure 6 It is a flow chart for calculating risk scores in the present invention;

[0058] Figure 7 It is a flow chart for obtaining security adjustment records in the present invention;

[0059] Figure 8 It is a flow chart for obtaining dynamic threat adaptation configurations in the present invention. Detailed implementation manners

[0060] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0061] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more unless otherwise specifically defined.

[0062] Embodiment

[0063] Please refer to Figure 1 , the present invention provides a technical solution: a threat analysis method based on big data, including the following steps:

[0064] S1: Based on the collected user behavior data, perform time tagging processing, synchronize the user activity data and system logs, extract the time tags and behavior metrics, and use the HMM model to evaluate the state transition probability according to the transition frequency and time interval to obtain the state anomaly index;

[0065] S2: Based on the state anomaly index, filter out the abnormal transition sequences, perform time series analysis on the filtered sequences, identify the data patterns that are different from the conventional patterns, compare them with the known threat patterns, mark the potential APT behaviors, and generate threat identification marks;

[0066] S3: Conduct a risk assessment on the threat recognition identifier, calculate the risk score, set up a response mechanism according to the risk level, implement a security adjustment mechanism for high-risk behaviors, update the security configuration, optimize the threat response process, and generate a security adjustment record. The security adjustment mechanism specifically refers to: based on the risk assessment results, update security policies and rules, such as stricter access control, enhanced authentication processes, or improved data encryption measures; adjust the configuration to plug security loopholes or restrict potential attack paths, for example, modify firewall rules, update the signature library of the intrusion detection system (IDS), or adjust the network isolation and monitoring settings; deploy or update security software according to the changes in current threats, such as virus protection, malware detection tools, and other security defense measures; implement automated responses to automatically identify and interrupt suspicious activities in the future, thereby reducing the dependence on manual intervention; and adjust the alert sensitivity and response threshold based on the new risk assessment.

[0067] S4: Based on the security adjustment record, record the threat detection results and the response status to threats in real time, conduct continuous behavior analysis, adjust the monitoring parameters and thresholds, match new threat patterns, and generate a dynamic threat adaptation configuration.

[0068] The status anomaly index includes a probability threshold, a behavior deviation level, and an anomaly detection frequency. The threat recognition identifier includes a behavior anomaly pattern, an APT threat type, and a pattern matching degree. The security adjustment record includes adjustment measures, a configuration update timestamp, and security level change information. The dynamic threat adaptation configuration specifically refers to response process adjustment information, monitoring threshold update results, and parameter optimization indicators.

[0069] Please refer to Figure 2 , and the extraction steps of the time tag and behavior metrics are specifically as follows:

[0070] S111: Based on the collected user behavior data, extract the timestamps of event records and log entries, verify the consistency of the timestamps, and obtain a synchronized data set.

[0071] Extract the timestamps of event occurrences from the user behavior data, and at the same time extract the log entries related to the events from the system log records. By comparing the timestamps with the time fields in the log entries one by one, use the time precision and matching rules to verify the two sets of data. By analyzing the format of the timestamps and parsing the time information in the log entries, screen and eliminate the unmatched data entries to ensure that the extracted timestamps and log entries can correspond one by one and the time consistency meets the specified requirements. During the verification process, convert the timestamps to the standard time format for matching, and at the same time parse the content fields in the log entries to confirm the relevance of the events. Use additional time tolerance rules to judge the unclear associated entries. After screening and verification, sort out the synchronized list of event timestamps and log entries to obtain a synchronized data set.

[0072] S112: Perform time tagging based on the synchronized dataset using the formula:

[0073]

[0074] Calculate the time tag T label , and at the same time extract the behavior metrics from the log entries to obtain a set of behavior metrics with time tags. Here, T event represents the original timestamp of the event, δT sync represents the time deviation obtained from log analysis, and γ is the time normalization coefficient used to adjust the granularity of the time tag;

[0075] A specific event timestamp T event = 1622563200 (UNIX timestamp, representing a specific time point), the time deviation δT obtained from the system log sync = 300 seconds, and the time granularity adjustment coefficient γ = 60. The process of calculating the time tag is as follows:

[0076]

[0077] The result shows that the generated time tag is 27042750, which is an adjusted time unit, more suitable for use in time series analysis, can directly reflect the specific time point when the event occurred, and is convenient for subsequent behavior pattern analysis and behavior prediction.

[0078] Please refer to Figure 3 , and the specific steps for obtaining the status anomaly index are as follows:

[0079] S121: Use the HMM model for state sequence analysis. By learning and identifying the statistical characteristics of behavior patterns, estimate each state to obtain the state transition probability;

[0080] Use the HMM model for state sequence analysis. Extract state sequence information from the dataset of time tags and behavior metrics. Serialize each behavior data point according to its category label and time dimension. The serialization process includes adjusting the time interval and reordering the behavior categories. Subsequently, calculate the occurrence probability of each state by performing a ratio operation on the number of occurrences of a specific state and the total number of behavior data points. The calculation of the probability value also requires introducing a time weight parameter, which is set based on the distribution of behavior occurrence times to reflect the time characteristics of behavior patterns. Finally, calculate the probabilities for all states and summarize them into a state transition probability matrix to complete the estimation of each state.

[0081] S122: According to the state transition probability, evaluate the frequency and time interval of transitions between multiple states using the formula:

[0082]

[0083] Calculate the status anomaly index IS abnormal , where PS ij represents the probability of transitioning from state i to state j, and PS normal is the expected normal state transition probability, and WS ij is the weight of the transition from state i to j;

[0084] By considering the difference between the actual state transition probability and the normal probability, as well as the importance of each transition, the severity of the abnormal state can be identified and quantified in a fine-grained manner. Currently, there are three states A, B, and C, where the state transition probability PS AB = 0.1, PS BC = 0.2, the normal state transition probability PS normal = 0.15, the transition weight WS AB = 2, and WS BC = 3. The process of calculating the anomaly index is as follows:

[0085] IS abnormal = |0.1 - 0.15|×2 + |0.2 - 0.15|×3

[0086] IS abnormal = 0.05×2 + 0.05×3

[0087] IS abnormal = 0.1 + 0.15

[0088] IS abnormal = 0.25

[0089] This result indicates that the obtained status anomaly index is 0.25. The value represents that there is a relatively small abnormal fluctuation in the currently observed state transition compared to the normal expected state transition. This index can be used in further monitoring and alarm systems to promptly identify and respond to potential abnormal behaviors.

[0090] Please refer to Figure 4 , and the specific steps for identifying the data pattern of the difference are as follows:

[0091] S211: Based on the status anomaly index, analyze the abnormal transition sequence, identify the behavioral characteristics of each data point, and obtain a preliminary set of behavioral characteristics;

[0092] Based on the state anomaly index, analyze the abnormal conversion sequence. By segmenting the abnormal conversion sequence according to the time dimension, extract the set of behavioral feature values in each time period. The feature values include the frequency of event occurrence, duration, time interval, and sorting pattern. Standardize and normalize the behavioral features of each time period to eliminate the interference of different data dimensions on the analysis results. Then perform calculations on each behavioral feature value, including the extraction of mean, variance, and peak value, so as to generate a complete preliminary set of behavioral features. This set contains the behavioral feature points in each time period and is used for subsequent analysis and pattern recognition.

[0093] S212: Based on the preliminary set of behavioral features, use the formula:

[0094]

[0095] Judge which behaviors are different from the normal pattern to obtain the set of differential data patterns. Among them, DF cluster is the data variance measure, xf i is the behavioral feature value of the data point, μf is the average value of the behavioral feature value of the data point, wf i is the weight based on the frequency of behavior occurrence, n Q is the total number of data points;

[0096] The behavioral feature values of three data points are [10, 20, 15], the average value is 15, the weights are [1.0, 0.5, 1.5] respectively, and the total number of data points is 3. The calculation process is as follows:

[0097] DF cluster =(10 - 15) 2 ×1.0+(20 - 15) 2 ×0.5+(15 - 15) 2 ×1.5

[0098] DF cluster =25×1.0 + 25×0.5 + 0×1.5

[0099] DF cluster =25 + 12.5 + 0 = 37.5

[0100] This result shows that the calculated data variance measure is 37.5, which reflects the degree of difference in the behavioral features of the data points. A higher value indicates a significant behavioral difference, which helps to further identify the data patterns that are significantly different from the normal pattern.

[0101] S213: Based on the set of differential data patterns, perform a matching analysis with the known normal patterns to determine the data patterns with differences;

[0102] Based on the set of data patterns with differences, matching is completed by calculating the similarity between each pattern and specific parameters of the regular pattern. Grouping is performed on each pattern in the set of difference patterns, corresponding the patterns with similar behavioral characteristics in the regular pattern to it, and separately marking the patterns that cannot be matched. The distance value of each group of data is further accurately calculated using the similarity metric formula, and the calculated distance value is compared with the set matching threshold to identify the data patterns that are significantly different from the regular pattern.

[0103] Please refer to Figure 5 , and the steps for obtaining the threat recognition identifier are specifically as follows:

[0104] S221: Screen for potential APT behaviors, compare them with the currently known threat behavior patterns, determine the potential threat of the behaviors, and generate a list of potential APT behaviors;

[0105] Screen for potential APT behaviors, compare them with the currently known threat behavior patterns, analyze the eigenvalue distribution of all abnormal behavior data, construct a behavior feature vector by extracting key characteristics of the behavior such as duration, trigger frequency, and resource types involved, and use vector matching technology to compare with the standard behavior pattern vectors in the known threat pattern library one by one, calculate the matching degree of the feature vectors, and screen out the behaviors exceeding the threshold as potential APT behaviors by setting a threat determination threshold, and generate a list of potential APT behaviors.

[0106] S222: Analyze the list of potential APT behaviors, determine the similarity, and identify the current security threat to obtain the threat recognition identifier;

[0107] Analyze the list of potential APT behaviors, use the similarity calculation method to determine the matching degree of each behavior with the known APT threat behaviors, judge the threat correlation of the behaviors by calculating the weighted cosine similarity of the eigenvalues, set the similarity threshold, mark the behaviors exceeding the threshold as the current security threats, and at the same time assign a unique identifier to each confirmed security threat, record the identifier and its corresponding behavior characteristics and matching information to obtain the threat recognition identifier.

[0108] Please refer to Figure 6 , and the steps for calculating the risk score are specifically as follows:

[0109] S311: Conduct a risk assessment on the threat recognition identifier, extract the key characteristics of each identifier, including the attack type, frequency, and potential impact, to obtain a set of key characteristics;

[0110] In the process of risk assessment for threat recognition and identification, first extract the key features of each identification, define the attack type, frequency, and potential impact. Identify the features by analyzing past security events and the current security posture. The attack types include, but are not limited to, DDoS, phishing, or malware attacks, etc. The determination of frequency depends on the number of occurrences of such attacks in historical data, while the potential impact evaluates the severity of data loss or system interruption caused. The aggregation of data will form a detailed set of key features, providing a basis for subsequent risk quantification analysis. It not only provides a quantitative basis for risk assessment but also helps predict and prepare for potential security threats.

[0111] S312: Based on the set of key features, use the formula:

[0112] GR score =G a ×G f +G b ×G p

[0113] Calculate the risk score GR for each threat score , where G f represents the occurrence frequency, which refers to the number of times a threat occurs or the expected occurrence frequency. The value is obtained based on the frequency of past observed events or security information prediction. G p represents the potential impact, including the impact assessment of factors such as data loss, system damage, business interruption, etc. G a is the weight coefficient associated with the occurrence frequency G f , used to reflect the importance of frequency in risk assessment. G b is the weight coefficient associated with the potential impact G p , used to illustrate the role of potential impact in the overall risk assessment;

[0114] The occurrence frequency G of a certain threat f is 0.8 (higher frequency, such as a threat with multiple attacks occurring daily), and the potential impact G p is 0.6 (medium impact, such as causing data leakage but not completely destroying the system). The weight coefficients G a =0.5 and G b =0.5. The process of calculating the risk score is as follows:

[0115] GR score =0.5×0.8 + 0.5×0.6

[0116] GR score =0.4 + 0.3

[0117] GR score =0.7

[0118] The result shows that the comprehensive risk score of this threat is 0.7, which is a relatively high risk level, meaning that moderate attention is required and corresponding preventive measures should be formulated.

[0119] Please refer to Figure 7 , and the specific steps for obtaining the security adjustment record are as follows:

[0120] S321: Set up a response mechanism according to the risk level, apply the security adjustment mechanism to the identified high-risk behaviors, and obtain a preliminary mechanism adjustment record;

[0121] Set up a response mechanism according to the risk level. By analyzing the triggering conditions and attack paths of high-risk behaviors, extract relevant behavior characteristics, such as the communication protocols involved, target ports, and resource usage. Design and select corresponding adjustment measures, specifically including blocking access to high-risk ports in the firewall, setting bandwidth limits for abnormal traffic, adding additional authentication mechanisms for critical resource access, etc. Record the operation steps and configuration changes when executing each adjustment measure, and at the same time conduct verification to confirm the effectiveness of the adjustment and its impact on the normal functions of the system, forming a preliminary mechanism adjustment record; for medium risks, send immediate alerts, including detailed information about abnormal behaviors (such as time, IP address, operation content, etc.), temporarily restrict the specific operation permissions of relevant users or processes. For example, for medium-risk data transmission behaviors, limit the network upload rate of the process to avoid potential slow data leakage, and require users to perform additional verification to continue the operation. For low risks, record the low-risk behaviors in the log as basic data for subsequent analysis. For example, record the abnormal time deviation when a user accesses certain folders, set additional monitoring rules for low-risk behaviors, and observe whether they evolve into higher risks in the future. For example, increase the traffic monitoring frequency for the detected slight abnormal traffic and set a trigger threshold to capture the situation of behavior escalation.

[0122] S322: Based on the preliminary mechanism adjustment record, update the security configuration, and optimize the threat response process according to the adjustment effect and feedback information from continuous monitoring to obtain the security adjustment record;

[0123] Based on the preliminary mechanism adjustment record, update the security configuration. Review each item of the adjustment content to ensure that all security policies are correctly applied and synchronized. For the specific rules of adjustment measures such as port blocking, bandwidth limitation, and verification mechanisms, update the firewall rule set, access permission management configuration, and intrusion detection system signature library in sequence. On this basis, obtain adjustment effect data through continuous monitoring of network activities, compare and analyze the adjusted system performance with the expected state, discover potential deficiencies and optimize them in a timely manner, and continuously improve the threat response process by updating response policies and supplementing new rules to generate a detailed security adjustment record.

[0124] Please refer toFigure 8 , the steps for obtaining the dynamic threat adaptation configuration are specifically as follows:

[0125] S411: Based on the security adjustment records, extract the key behavior patterns and abnormal features from the real-time recorded threat detection results and response statuses, analyze the potential threat change trends, and obtain the preliminary parameter set for adaptive adjustment by quantifying the frequency, scope, and impact of threat behaviors;

[0126] Based on the security adjustment records, extract the key behavior patterns and abnormal features from the real-time recorded threat detection results and response statuses. First, classify the threat detection results, screen the high-frequency behaviors according to the occurrence frequency of events, mark the outliers by calculating the frequency distribution of behaviors within a specific time period, then evaluate the scope of influence of the events, quantify it using the coverage parameters of the systems, services, or users involved, and finally combine the historical data to analyze the degree of influence and quantify it as a standardized loss value. Through this process, a parameter set that can describe the threat change trends is generated, and after sorting, the preliminary parameter set for adaptive adjustment is obtained.

[0127] S412: Evaluate the preliminary parameter set for adaptive adjustment, and use the formula

[0128]

[0129] to calculate the optimal threshold θ for dynamic adjustment opt , and obtain the optimized dynamic adjustment parameters. Among them, θ i represents the current value of the monitoring parameter, μ θ is the parameter mean, obtained by calculating the past threat behavior data, is the parameter variance, reflecting the fluctuation range of the parameter, ∈ D is the adjustment coefficient to avoid calculation errors caused by too small denominator values, wd i is the weight coefficient, set according to the importance and frequency of the monitoring parameter, n R is the total number of monitoring parameters;

[0130] There are five values of the monitoring parameter which are θ i = [2, 4, 5, 3, 6], and their corresponding weights wd i = [1, 0.5, 0.3, 1.5, 0.7], the mean μ θ = 4, the variance The adjustment coefficient ∈ D = 0.1, and the calculation process is as follows:

[0131] 1·(2 - 4) 2 +0.5·(4 - 4) 2 +0.3·(5 - 4) 2

[0132]

[0133] The result shows that the optimal threshold for dynamic adjustment is 5.375, and this value will be used to update the monitoring parameters to improve the accuracy and efficiency of threat detection and response.

[0134] S413: Based on the optimized dynamic adjustment parameters, by analyzing the behavior characteristics of new threats and existing adjustment records, matching new threat patterns, updating the monitoring and response rules, a dynamic threat adaptation configuration is obtained;

[0135] Based on the optimized dynamic adjustment parameters, by analyzing the behavior characteristics of new threats and existing adjustment records, first, the behavior characteristics of new threats are analyzed in multiple dimensions, including information such as the attack path, target type, and duration, and compared with the historical adjustment records. By calculating the similarity coefficient to match the threat pattern, it is determined whether the behavior belongs to a new threat type or an extension of a known type. Subsequently, according to the recognition result, the monitoring rules are updated, including adjusting the detection sensitivity of specific events and the priority of response actions. Finally, the updated rules and adjustment records are comprehensively sorted out to generate a dynamic threat adaptation configuration to cope with the changing threat environment.

[0136] A threat analysis system based on big data, the system includes:

[0137] The behavior data synchronization module, based on the collected user behavior data, synchronizes the user activity data and system logs, extracts time tags and behavior metrics, uses the HMM model to calculate the transition frequency and time interval, evaluates the state transition probability, and obtains the state anomaly index;

[0138] The anomaly recognition and analysis module, based on the state anomaly index, filters out abnormal transition sequences, performs time series analysis on the filtered sequences, identifies data patterns that differ from the normal pattern, compares with known threat patterns, marks potential APT behaviors, and generates threat recognition marks;

[0139] The risk response configuration module uses the threat recognition mark to evaluate the associated risks and calculate the risk score, formulates and implements a security adjustment mechanism according to the risk level, updates the security configuration, optimizes the threat response process, and generates a security adjustment record;

[0140] The dynamic monitoring adjustment module, based on the security adjustment record, records the threat detection results and the response status to threats in real time, continuously analyzes the behavior pattern, adjusts the monitoring parameters and thresholds, matches new threat patterns, and obtains a dynamic threat adaptation configuration;

[0141] The performance optimization module uses the dynamic threat adaptation configuration to analyze the current threat recognition performance, evaluates the performance bottleneck and security weaknesses, and adjusts the response configuration and monitoring parameters again to obtain the threat recognition performance optimization result.

[0142] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A threat analysis method based on big data, characterized in that, It includes the following steps: Based on the collected user behavior data, perform time tagging, synchronize user activity data with system logs, extract time tags and behavior metrics, and use the HMM model to evaluate the state transition probability according to the conversion frequency and time interval to obtain the state anomaly index; Based on the state anomaly index, screen abnormal conversion sequences, perform time series analysis on the screened sequences, identify data patterns different from the normal pattern, compare with known threat patterns, identify potential APT behaviors, and generate threat identification labels; Perform risk assessment on the threat identification labels, calculate risk scores, set response mechanisms according to risk levels, implement security adjustment mechanisms for high-risk behaviors, update security configurations, optimize threat response processes, and generate security adjustment records; Based on the security adjustment records, record the threat detection results and response status to threats in real time, perform continuous behavior analysis, adjust monitoring parameters and thresholds, match new threat patterns, and generate dynamic threat adaptation configurations.

2. The threat analysis method based on big data according to claim 1, characterized in that The specific steps for extracting the time tags and behavior metrics are as follows: Based on the collected user behavior data, extract the timestamps of event records and log entries, verify the consistency of the timestamps, and obtain a synchronized data set; Based on the synchronized data set, perform time tagging using the formula: Calculate the time tag T label , and extract the behavior metrics from the log entries simultaneously to obtain a set of behavior metrics with time tags, where T event represents the original timestamp of the event, and δT sync represents the time deviation obtained from log analysis, and γ is the time normalization coefficient used to adjust the granularity of the time tag.

3. The threat analysis method based on big data according to claim 1, wherein The specific steps for obtaining the state anomaly index are as follows: Use the HMM model to perform state sequence analysis, estimate each state by learning and identifying the statistical characteristics of behavior patterns, and obtain the state transition probability; According to the state transition probability, evaluate the frequency and time interval of transitions between multiple states using the formula: Calculate the status anomaly index IS abnormal , where PS ij represents the probability of transitioning from state i to state j, and PS normal is the expected normal state transition probability, and WS ij is the weight of the transition from state i to j.

4. The threat analysis method based on big data according to claim 1, characterized in that The specific steps for identifying the different data patterns are as follows: Based on the state anomaly index, analyze abnormal conversion sequences, identify the behavior characteristics of each data point, and obtain a preliminary set of behavior characteristics; Based on the preliminary set of behavior characteristics, use the formula: Determine which behaviors are different from the normal mode to obtain a set of differential data patterns, where DF cluster is the data variance metric, xf i is the behavior eigenvalue of the data point, μf is the average value of the behavior eigenvalues of the data points, wf i is the weight based on the behavior occurrence frequency, n Q is the total number of data points; Based on the set of different data patterns, perform matching analysis with known normal patterns to determine the data patterns with differences.

5. The threat analysis method based on big data according to claim 1, wherein The specific steps for obtaining the threat identification labels are as follows: Screen potential APT behaviors, compare with current known threat behavior patterns, determine the potential threat of the behaviors, and generate a list of potential APT behaviors; Analyze the list of potential APT behaviors to determine similarity and identify the current security threat to obtain threat identification labels.

6. The threat analysis method based on big data according to claim 1, wherein The specific steps for calculating the risk scores are as follows: Perform risk assessment on the threat identification labels, extract the key characteristics of each label, including attack type, frequency, and potential impact, to obtain a set of key characteristics; Based on the set of key characteristics, use the formula: GR score = G a × G f + G b × G p Calculate the risk score GR for each threat score , where G f represents the occurrence frequency, G p represents the potential impact, G a is the weight coefficient associated with the occurrence frequency G f , and G b is the weight coefficient associated with the potential impact G p .

7. The threat analysis method based on big data according to claim 1, characterized in that, The specific steps for obtaining the security adjustment records are as follows: Set a response mechanism according to the risk level, apply a security adjustment mechanism to behaviors identified as high-risk, and obtain a preliminary mechanism adjustment record; Based on the preliminary mechanism adjustment record, update the security configuration, optimize the threat response process according to the adjustment effect and feedback information from continuous monitoring, and obtain security adjustment records.

8. The threat analysis method based on big data according to claim 1, wherein The specific steps for obtaining the dynamic threat adaptation configuration are as follows: Based on the security adjustment records, extract key behavior patterns and abnormal features from the real-time recorded threat detection results and response statuses, analyze the potential threat change trends, and obtain a preliminary parameter set for adaptive adjustment by quantifying the frequency, scope, and impact of threat behaviors. Evaluate the preliminary parameter set for adaptive adjustment using the formula Calculate the optimal threshold θ for dynamic adjustment opt , and obtain the optimized dynamic adjustment parameter, where θ i represents the current value of the monitoring parameter, μ θ is the parameter mean, is the parameter variance, ∈ D is the adjustment coefficient, wd i is the weight coefficient, n R is the total number of monitoring parameters; Based on the optimized dynamic adjustment parameters, match new threat patterns by analyzing the behavior characteristics of new threats and existing adjustment records, and update the monitoring and response rules to obtain a dynamic threat adaptation configuration.

9. A threat analysis system based on big data, characterized in that, Execute according to the big data-based threat analysis method described in any one of claims 1-8. The system includes: The behavior data synchronization module synchronizes user activity data and system logs based on the collected user behavior data, extracts time tags and behavior metrics, uses the HMM model to calculate transition frequencies and time intervals, and evaluates the state transition probability to obtain a state anomaly index. The anomaly recognition and analysis module filters abnormal transition sequences based on the state anomaly index, performs time series analysis on the filtered sequences, identifies data patterns that differ from the normal patterns, compares with known threat patterns, identifies potential APT behaviors, and generates threat recognition identifiers. The risk response configuration module uses the threat recognition identifiers to evaluate associated risks and calculate risk scores, formulates and implements a security adjustment mechanism according to the risk levels, updates the security configuration, optimizes the threat response process, and generates security adjustment records. The dynamic monitoring and adjustment module, based on the security adjustment records, records the threat detection results and the response status to threats in real time, continuously analyzes behavior patterns, adjusts monitoring parameters and thresholds, and matches new threat patterns to obtain a dynamic threat adaptation configuration. The performance optimization module uses the dynamic threat adaptation configuration to analyze the current threat recognition performance, evaluate performance bottlenecks and security weaknesses, and readjust the response configuration and monitoring parameters to obtain an optimized result of threat recognition performance.