Multi-level dynamic threat monitoring system based on deep learning
By introducing a multi-level dynamic threat monitoring system based on deep learning in the network security protection system, combining network traffic, logs and user behavior data, the problem of difficulty in identifying internal attacks and new threats in the existing technology is solved, and efficient and accurate network security threat monitoring and early warning are achieved.
Patent Information
- Application Number
- CN202510122759.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing network security protection technologies are difficult to effectively identify internal attacks and malicious traffic through legitimate ports, and the intrusion detection system has limited detection capabilities for new attacks, and the false alarm rate is high, making it difficult to accurately distinguish between normal behavior changes and real attack behavior.
A multi-level dynamic threat monitoring system based on deep learning is adopted, through the collaborative work of front-end electronic devices and back-end servers, network traffic, logs and user behavior data are collected, feature extraction and vectorization are performed, cross-semantic relationship diagrams are constructed, and threat prediction is used to use multimodal fusion neural network model.
It realizes comprehensive monitoring and accurate prediction of network security threats, improves the accuracy of threat identification and the system's response speed, and can promptly detect potential threats and make early warnings to ensure the safe and stable operation of the network.
Smart Images

Figure CN120074883A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a multi-level dynamic threat monitoring system based on deep learning. Background Art
[0002] With the rapid development of information technology, the penetration of the network into all fields of modern society is becoming increasingly deep. From the core business operations of enterprises to the daily information interaction of individuals, the network has become an indispensable infrastructure. However, the openness and sharing nature of the network make it face many security threats, such as malicious attacks, data leakage, illegal intrusion, etc. These threats not only pose a serious threat to personal privacy and property security, but also bring huge risks to the economy, security and stability of enterprises and countries.
[0003] Traditional network security protection technologies mainly include firewalls, intrusion detection systems (IDS), antivirus software, etc. Firewalls restrict network access by setting rules and can block external illegal access to a certain extent, but it is often difficult to effectively identify internal attacks and malicious traffic through legitimate ports. IDS discovers potential intrusion behaviors based on feature libraries or anomaly detection. However, its detection ability for new attacks is limited because the features of new attacks are often not included in the existing feature libraries, and the false alarm rate of anomaly detection is relatively high, making it difficult to accurately distinguish normal network behavior changes from real attack behaviors. Antivirus software mainly focuses on preventing known viruses and malware, and its protection effect is greatly reduced for complex and changeable network attack methods, especially attacks based on social engineering, zero-day vulnerability exploitation, etc. Summary of the Invention
[0004] In view of this, the embodiments of the present invention provide a multi-level dynamic threat monitoring system based on deep learning to at least partially solve the above problems.
[0005] According to the first aspect of the embodiments of the present invention, a multi-level dynamic threat monitoring system based on deep learning is provided, which includes: a front-end electronic device and a back-end server. A data acquisition module, a data preprocessing module, a feature extraction module, a data vectorization module, and a semantic sequence construction module are deployed on the front-end electronic device, and a model training module and a model deployment module are deployed on the back-end server; The data acquisition module is used to obtain the first network traffic data sample, the first log data sample, the first user behavior data sample, and the network security threat labels of the first network traffic data sample, the first log data sample, and the first user behavior data sample respectively; The data preprocessing module is used to preprocess the first network traffic data sample, the first log data sample, and the first user behavior data sample to obtain a second network traffic data sample, a second log data sample, and a second user behavior data sample respectively; The feature extraction module is used to extract features from the second network traffic data sample, the second log data sample, and the second user behavior data sample respectively to generate a traffic feature data sample, a log feature data sample, and a behavior feature data sample; The data vectorization module is used to obtain a traffic feature vector sample, a log feature vector sample, and a behavior feature vector sample from the traffic feature data sample, the log feature data sample, and the behavior feature data sample respectively; The semantic sequence construction module is used to perform cross-semantic relationship analysis on the traffic feature vector sample, the log feature vector sample, and the behavior feature vector sample to obtain a cross-semantic relationship graph, and mark the network security threat labels on the cross-semantic relationship graph to form a semantic training data set; The model training module is used to call the semantic training data set to train the deep learning model to be trained, so that the deep learning model learns the relationship between the first network traffic data sample, the first log data sample, the first user behavior data sample and the corresponding network security threat labels, and maps the relationship to the network parameters of the deep learning model to obtain a multi-level dynamic threat monitoring model; The model deployment module is used to deploy the multi-level dynamic threat monitoring model on the cloud server, so that the multi-level dynamic threat monitoring model extracts and processes the operation data features of the target network based on its network parameters to predict the network security threats occurring in the target network.
[0006] In the solution of the embodiment of the present invention, the following technical advantages are achieved: First, in terms of data collection, the system can comprehensively obtain the first network traffic data sample, the first log data sample, and the first user behavior data sample, and simultaneously obtain the corresponding network security threat labels. This multi-source data collection method can capture network activity information from a broader perspective compared with the traditional single data source monitoring technology. For example, network traffic data reflects the basic situation of network connections, log data records the operation events inside the system and applications, and user behavior data reflects the operation tracks of network users. The combination of the three can avoid the missed detection of threats caused by the lack of information in a single data source, greatly improving the comprehensiveness of threat monitoring. In the data preprocessing link, targeted preprocessing operations are carried out for different types of data to obtain the second network traffic data sample, the second log data sample, and the second user behavior data sample. For example, protocol traffic identification and timestamp correction are performed on network traffic data to remove noise data and ensure the accuracy of the data time sequence, laying a foundation for subsequent accurate analysis; semantic understanding, parsing, knowledge graph construction and repair are performed on log data, which helps to mine the deep semantic information and abnormal situations in the log data; multi-source fusion, integrity judgment and filling are performed on user behavior data, which can more completely and accurately depict the user behavior pattern, reduce misjudgment caused by inaccurate or incomplete data, and thus improve the data quality and analysis reliability of the entire system. The feature extraction module extracts features from the preprocessed data of different types respectively to generate traffic feature data samples, log feature data samples, and behavior feature data samples, which can accurately mine the key features related to network security threats in various types of data. For example, feature extraction is performed on the traffic graph structure representation through a graph convolutional neural network, which can effectively identify key nodes, traffic convergence areas, and traffic association patterns between different network segments in the network, thereby discovering potential malicious traffic distribution features; semantic roles are extracted from log data based on a pre-trained language model to generate log feature data samples, which helps to understand the association between key elements such as the subject and operation object of events in the log and security threats; a user behavior trajectory graph is constructed and graph embedding is performed to generate behavior feature data samples, which can clearly present the dynamic change process of user behavior and the correlation between different behaviors, facilitating the detection of abnormal behavior patterns. These feature extraction methods can more accurately locate threat-related features compared with the traditional general feature extraction methods, improving the accuracy of threat identification. The data vectorization module further converts the extracted feature data into vector samples, and innovative technologies such as assigning attention weights according to feature importance weights, calculating a semantic similarity matrix based on a semantic embedding model for weighted expansion, and spatio-temporal fusion vectorization are adopted in the conversion process.This enables the data to be represented in a more appropriate form before entering model training, better preserving the key information and semantic relationships of the data, which helps improve the model's understanding and learning ability of the data. Compared with traditional simple vectorization methods, it allows the model to more efficiently capture subtle threat signals in the data during the training process. The semantic sequence construction module conducts cross-semantic relationship analysis on traffic feature vector samples, log feature vector samples, and behavior feature vector samples, constructs a cross-semantic relationship graph, and simultaneously labels network security threat tags to form a semantic training dataset. The construction of this cross-semantic relationship can deeply explore the hidden semantic connections between different types of data feature vectors. For example, it can discover the internal associations between specific patterns in network traffic, specific events in logs, and specific user behaviors. These associations are often important indicators of complex network security threats. By forming a semantic training dataset, it provides richer and more in-depth training data for the deep learning model, enabling the model to learn more comprehensive and accurate network security threat feature patterns. Compared with traditional training data construction methods lacking semantic association analysis, it can significantly improve the model's recognition ability and generalization ability for complex and changing network security threats. The model training module uses the semantic training dataset to train the deep learning model. Through a multi-branch fusion neural network structure, including a traffic feature processing branch, a log feature processing branch, a behavior feature processing branch, and a cross-modal fusion layer, combined with various neural network structures such as a convolutional neural network and an attention layer, a recurrent neural network, a semantic embedding layer, and a graph neural network, and a series of technical means such as using a bilinear pooling layer to strengthen multi-modal fusion features, a fully connected layer to extract local information, a non-linear transformation to predict the probability distribution of threat types, and a cross-entropy loss function to optimize model parameters, the model can fully learn the complex relationships between the first network traffic data samples, the first log data samples, the first user behavior data samples, and the corresponding network security threat tags, and accurately map them to network parameters. This training method that combines multi-modal data, multi-neural network structures, and an innovative training mechanism can train a multi-level dynamic threat monitoring model with better performance and more accurate judgment of network security threats compared with traditional training methods based on a single modality or a simple model structure. Finally, the model deployment module deploys the trained multi-level dynamic threat monitoring model on a cloud server. Utilizing the powerful computing resources and flexible scalability of the cloud server, it can efficiently extract and process the feature of the running data of the target network in real time and predict network security threats. In the face of large-scale network data and high-concurrency network requests, the system can respond quickly, timely discover potential threats and give early warnings, ensuring the safe and stable operation of the target network, which is an advantage that traditional locally deployed monitoring systems with limited computing power are difficult to achieve. Brief Description of the Drawings
[0007] Figure 1Schematic diagram of a multi-level dynamic threat monitoring system based on deep learning provided by an embodiment of the present invention. Detailed implementation manners
[0008] As Figure 1 shown, it includes: a front-end electronic device and a back-end server. A data acquisition module, a data preprocessing module, a feature extraction module, a data vectorization module, and a semantic sequence construction module are deployed on the front-end electronic device, and a model training module and a model deployment module are deployed on the back-end server; The data acquisition module is used to obtain the first network traffic data sample, the first log data sample, the first user behavior data sample, and the network security threat labels of the first network traffic data sample, the first log data sample, and the first user behavior data sample respectively; The data preprocessing module is used to preprocess the first network traffic data sample, the first log data sample, and the first user behavior data sample to obtain a second network traffic data sample, a second log data sample, and a second user behavior data sample respectively; The feature extraction module is used to extract features from the second network traffic data sample, the second log data sample, and the second user behavior data sample respectively to generate a traffic feature data sample, a log feature data sample, and a behavior feature data sample; The data vectorization module is used to obtain a traffic feature vector sample, a log feature vector sample, and a behavior feature vector sample from the traffic feature data sample, the log feature data sample, and the behavior feature data sample respectively; The semantic sequence construction module is used to perform cross-semantic relationship analysis on the traffic feature vector sample, the log feature vector sample, and the behavior feature vector sample to obtain a cross-semantic relationship graph, and mark the network security threat labels on the cross-semantic relationship graph to form a semantic training data set; The model training module is used to call the semantic training data set to train the deep learning model to be trained, so that the deep learning model learns the relationship between the first network traffic data sample, the first log data sample, the first user behavior data sample and the corresponding network security threat labels, and maps the relationship to the network parameters of the deep learning model to obtain a multi-level dynamic threat monitoring model; The model deployment module is used to deploy the multi-level dynamic threat monitoring model on the cloud server, so that the multi-level dynamic threat monitoring model extracts and processes the operation data features of the target network based on its network parameters to predict the network security threats occurring in the target network.
[0009] Optionally, when the data acquisition module obtains the first network traffic data sample, the first log data sample, the first user behavior data sample, and the network security threat labels of the first network traffic data sample, the first log data sample, and the first user behavior data sample respectively, the steps included are as follows: The data acquisition module sends a collection requirement to the SDN controller, so that the SDN controller, according to the network global view and the flow table rules, guides the traffic meeting the predetermined requirements to the collection port specified by the data acquisition module for the collection of the first network traffic data sample; compresses the collected first network traffic data sample based on the set content-aware lossless compression mechanism for caching in the local cache; based on the annotation component of the constructed multi-source threat intelligence, matches with the first network traffic data sample at the local cache to assign a network security threat label to the first network traffic data sample. Optionally, when the data acquisition module obtains the first network traffic data sample, the first log data sample, the first user behavior data sample, and the network security threat labels of the first network traffic data sample, the first log data sample, and the first user behavior data sample respectively, the steps included are as follows: The data acquisition module calls the containerized log collection node to collect the first log data sample from different log sources to generate a blockchain block containing a timestamp, a device identifier, and a log summary, and performs a consensus verification on the blockchain block to judge the authenticity of the first log data sample; in response to the authenticity of the first log data sample being greater than the set authenticity threshold, parses the first log data sample and stores it in the graph database in the form of a knowledge graph constructed according to entities and relationships; based on the annotation mechanism of the user and device behavior portraits, matches with the graph database to label the corresponding network security threat label in the knowledge graph. Optionally, when the data acquisition module obtains the first network traffic data sample, the first log data sample, the first user behavior data sample, and the network security threat labels of the first network traffic data sample, the first log data sample, and the first user behavior data sample respectively, the steps included are as follows: The data acquisition module calls the monitoring interface based on the cross-platform development framework and the system bottom layer to detect the system bottom layer to capture in real time user gesture operations, system-level application calls, and interactions between different application programs; based on differential privacy, associates the interaction behavior with the corresponding user entity and adds noise to generate the first user behavior data sample; adds the first user behavior data sample to the distributed ledger by pen; based on the set context-aware dynamic annotation mechanism, annotates the network security threat label of the first user behavior data sample in the distributed ledger.
[0010] Specifically, the above steps of the data acquisition module are illustrated by examples:
[0011] I. Detecting and Capturing User Behavior Data
[0012] The data acquisition module detects the system underlying layer to capture in real time user gesture operations, system-level application calls, and interaction behaviors between different applications by invoking monitoring interfaces based on cross-platform development frameworks and the system underlying layer. From an execution perspective, this is equivalent to starting multiple monitoring threads or using the callback mechanism of the system underlying layer to continuously listen for relevant system events.
[0013] Define the following variables to represent relevant data volumes: Let N gesture represent the number of user gesture operations captured per unit time. For example, one finger swipe, click, etc. is counted as one gesture operation. The value of N gesture changes in real time with user operations, and its magnitude reflects the frequency of user interaction with the system in terms of gesture operations. Let N appcall represent the number of system-level application calls per unit time. Whenever a system-level application is started, switched, or its internal function module is invoked, etc., the count of N appcall increases, which reflects the usage activity of each application in the system. Let N interact represent the number of interaction behaviors between different applications per unit time. For example, operations such as application A sending data to application B, requesting services, etc. This value reflects the situation of collaborative work or information transfer between applications.
[0014] When performing this step, these count variables need to be continuously updated, for example, by implementing corresponding system event listening functions. The following is a code example:
[0015] # Initialize count variables
[0016] N_gesture = 0
[0017] N_app_call = 0
[0018] N_interact = 0
[0019] # Define an example of a system underlying layer event listening function (corresponding listening interface)
[0020] def on_gesture_event():
[0021] global N_gesture
[0022] N_gesture += 1
[0023] def on_app_call_event():
[0024] global N_app_call
[0025] N_app_call += 1
[0026] def on_interact_event():
[0027] global N_interact
[0028] N_interact += 1
[0029] # Register event listener (operate according to the registration method of the corresponding framework in practice)
[0030] register_event_listener('gesture_event', on_gesture_event)
[0031] register_event_listener('app_call_event', on_app_call_event)
[0032] register_event_listener('interact_event', on_interact_event)
[0033] # Continuously listen
[0034] while True:
[0035] continue # Keep the listening state and wait for the event to trigger the corresponding function to update the count variable
[0036] II. Generate the first user behavior data sample by adding noise based on differential privacy
[0037] Based on differential privacy, associate the interaction behavior with the corresponding user entity and add noise to generate the first user behavior data sample. Differential privacy protects the privacy information of users by adding specific noise to the data, making it difficult for attackers to infer the sensitive information of specific users by analyzing the data. Let the original user behavior data be represented as a vector X = (x 1 , x 2 , …, x n ), where each element x i can correspond to the quantized value of different types of behavior data captured above (for example, x 1 can be a normalized value of the number of gesture operations within a certain period of time, x 2 is the normalized value of the number of app calls, etc.). The process of adding noise can be represented by the following formula:
[0038] Xnoisy = X + ∈
[0039] Where, X noisy is the first user behavior data sample vector after adding noise, and ∈ is the noise vector, whose elements are usually randomly sampled from values that satisfy a certain probability distribution (such as Laplace distribution, etc.).
[0040] Taking the Laplace distribution as an example, let the probability density function of the Laplace distribution be:
[0041]
[0042] Where, μ is the location parameter (usually set to 0, indicating that the noise is symmetrically distributed around 0), and b is the scale parameter, which determines the size range of the noise. In this application scenario, the selection of the scale parameter b needs to be determined according to the requirements of privacy protection and the trade-off of data availability. If the value of b is larger, more noise is added, and the degree of privacy protection is higher, but it may have a certain impact on the accuracy of subsequent data analysis; on the contrary, if the value of b is smaller, the privacy protection is relatively weaker, but the data is closer to the original real data, which is more conducive to accurate threat monitoring and analysis based on behavior data in the future.
[0043] Code example for generating a noise vector and adding it to the original data (taking Python combined with common scientific computing libraries as an example)
[0044] import numpy as np
[0045] # The original data vector X has been obtained
[0046] X = np.array([0.2, 0.5, 0.3])
[0047] # Set the scale parameter b of the Laplace distribution
[0048] b = 0.1
[0049] # Sample from the Laplace distribution to generate a noise vector epsilon with the same dimension as X
[0050] epsilon = np.random.laplace(0, b, size = X.shape)
[0051] # Generate the first user behavior data sample vector X_noisy after adding noise
[0052] X_noisy = X + epsilon
[0053] III. Adding the first user behavior data sample to the distributed ledger
[0054] Add the first user behavior data sample to the distributed ledger one by one. The distributed ledger (such as the ledger implemented by blockchain technology) can ensure the immutability and traceability of data. From an execution perspective, this involves interacting with the distributed ledger system, such as sending a data addition request to the ledger nodes through the corresponding network communication protocol.
[0055] The distributed ledger system has an interface function for adding data, such as `add_to_ledger(data)`. The operation of adding data can be illustrated by the following code:
[0056] #X_noisy is the vector of the first user behavior data sample after adding noise generated previously
[0057] #Convert it to the format suitable for ledger storage data_to_ledger = {
[0058] 'user_behavior':X_noisy.tolist()
[0059] }
[0060] #Call the ledger addition interface
[0061] add_to_ledger(data_to_ledger)
[0062] IV. Label network security threat tags based on the context-aware dynamic annotation mechanism
[0063] Based on the set context-aware dynamic annotation mechanism, label the network security threat tags of the first user behavior data sample in the distributed ledger. Context awareness means considering various relevant context factors when the user behavior occurs to judge whether there is a threat, such as the security state of the current network, the business operation scenario where the user is located, the resource usage of the system, etc. Define the following variables to measure the relevant context factors: Let S netstatus represent the security state index of the network, and its value range can be [0,1]. For example, 0 means the network is in a secure state, and approaching 1 means that more abnormal traffic, suspected attacks and other insecure situations are detected. Its value can be obtained through real-time network security monitoring tools (such as intrusion detection systems, etc.) and fed back to the annotation mechanism. S businessop represent the business operation scenario index where the user is located. Different business operations may correspond to different normal behavior patterns and potential threat situations. For example, in the financial transfer business scenario, frequently modifying the transfer amount and operating too fast may be abnormal behavior. The relevant module of the business system can be used to judge what business operation scenario the current user is in and assign the corresponding value (also in [0,1], 0 means a normal low-risk business scenario, and 1 means a high-risk business scenario where security threats are likely to occur). Ssysres Indicates the resource usage metrics of the system, such as the CPU usage rate, memory usage rate, etc., which are mapped to values in the range of [0, 1] after comprehensive consideration. An overly high resource usage rate may imply abnormal situations such as malicious programs occupying resources, and can be obtained through system performance monitoring tools.
[0064] Construct a threat assessment function based on context awareness to determine whether it is labeled as a threat and the corresponding threat label. The threat label is represented by L, with a value of 0 indicating no threat, 1 indicating a low-level threat, 2 indicating a medium-level threat, and 3 indicating a high-level threat:
[0065]
[0066] Among them, T 1 、T 2 、T 3 are thresholds set according to experience and actual security requirements. For example, T 1 =0.3, T 2 =0.6, T 3 =0.9. Different threshold intervals correspond to different threat level judgments.
[0067] The code example for implementing this labeling mechanism is as follows:
[0068] # Obtain the values of relevant context metrics
[0069] S_net_status = get_network_status()
[0070] S_business_op = get_business_operation_scenario()
[0071] S_sys_res = get_system_resource_usage()
[0072] # Set the thresholds
[0073] T1 = 0.3
[0074] T2 = 0.6
[0075] T3 = 0.9
[0076] # Determine the threat label according to the evaluation function
[0077] if S_net_status + S_business_op + S_sys_res < T1:
[0078] label = 0
[0079] elif T1 <= S_net_status + S_business_op + S_sys_res < T2:
[0080] label = 1
[0081] elif T2 <= S_net_status + S_business_op + S_sys_res < T3:
[0082] label = 2
[0083] else:
[0084] label = 3
[0085] # Mark the threat label on the corresponding data record in the distributed ledger (the ledger has an interface for updating the label)
[0086] update_label_in_ledger(label)
[0087] Optionally, when the data preprocessing module preprocesses the first network traffic data sample, the first log data sample, and the first user behavior data sample to obtain the second network traffic data sample, the second log data sample, and the second user behavior data sample respectively, the steps included are as follows: Based on the trained protocol classification model, identify the protocol traffic of the first network traffic data sample to remove the noise data in the first network traffic data sample; perform time series analysis on the first network traffic data sample from which the noise data has been removed to correct the time stamp of the first network traffic data sample, and generate the second network traffic data sample accordingly.
[0088] Specifically, the detailed description of the relevant steps of the above data preprocessing module is as follows:
[0089] I. Removal of noise data based on the protocol classification model
[0090] Let the first network traffic data sample be represented as a set T = {t 1 , t 2 , …, t m}, where each element t i represents a network traffic data packet, which contains multiple attributes, such as the source IP address IP src (t i ), the destination IP address IP dst (t i ), the port number Port(t i ), the protocol type Protocol(t i ), the data packet size Size(t i) and timestamp Timestamp(t i ) etc. The trained protocol classification model is a function M protocol (t i ), and its output is the protocol category to which the data packet belongs or a flag indicating whether it is noise data. For example, if M protocol (t i ) = 0, it means that the data packet is determined to be noise data; if M protocol (t i ) ⊙ 0, it means that the data packet belongs to a certain normal network protocol (such as M protocol (t i ) = 1 indicates that it is an HTTP protocol data packet, M protocol (t i ) = 2 indicates that it is an FTP protocol data packet, etc.).
[0091] The process of performing protocol traffic identification and removing noise data can be expressed by the following formula: T filtered = {t i | t i ∈ T ∧ M protocol (t i ) ≠ 0}
[0092] where T filtered is the network traffic data set after removing noise data, that is, the precursor of the second network traffic data sample. From an execution perspective, it needs to traverse the entire first network traffic data sample set T, call the protocol classification model M i for each data packet t protocol (t i ), and decide whether to retain the data packet in T filtered according to the output result. The code example is as follows:
[0093] T_filtered = []
[0094] for t_i in T:
[0095] if M_protocol(t_i)!= 0:
[0096] T_filtered.append(t_i)
[0097] II. Time Series Analysis and Timestamp Correction
[0098] For the network traffic data set T filtered after removing noise data, perform time series analysis to correct the timestamp. Let t prev represent the timestamp of the previous data packet, t curr represent the timestamp of the current data packet, Δtexpected Represents the time interval expected based on the normal network delay model (this model can be trained based on network type, bandwidth, historical traffic data, etc.). Define a function Δt actual (t prev , t curr ) = t curr - t prev to calculate the actual time interval. Then, the timestamp is corrected by comparing the actual time interval with the expected time interval.
[0099] The formula for correcting the timestamp can be expressed as:
[0100]
[0101] where ∈ is an allowable time error threshold, set according to network stability and accuracy requirements. If the difference between the actual time interval and the expected time interval is within the allowable range, it means the timestamp is normal and no correction is needed; otherwise, the timestamp of the current data packet is corrected to the timestamp of the previous data packet plus the expected time interval.
[0102] When traversing T filtered , it is necessary to record the timestamp of the previous data packet and correct the timestamp of each data packet according to the above formula.
[0103] The code example is as follows:
[0104] t_prev = None
[0105] for t_i in T_filtered:
[0106] if t_prev is None:
[0107] t_prev = t_i.Timestamp
[0108] else:
[0109] delta_t_actual = t_i.Timestamp - t_prev
[0110] if abs(delta_t_actual - Delta_t_expected) <= epsilon:
[0111] # The timestamp is normal and no correction is needed
[0112] t_prev = t_i.Timestamp
[0113] else:
[0114] # Correct the timestamp
[0115] t_i.Timestamp = t_prev + Delta_t_expected
[0116] t_prev = t_i.Timestamp
[0117] The network traffic data set T after the above timestamp correction filtered becomes the final second network traffic data sample. Through such a preprocessing process, using the protocol classification model and time series analysis method, noise data is effectively removed and the timestamp is corrected, improving the quality and usability of network traffic data, and laying a good foundation for subsequent feature extraction and network security threat monitoring and analysis.
[0118] Optionally, when the data preprocessing module preprocesses the first network traffic data sample, the first log data sample, and the first user behavior data sample to obtain the second network traffic data sample, the second log data sample, and the second user behavior data sample respectively, the steps included are as follows: Based on the sequence-to-sequence model of the recurrent neural network, perform semantic understanding and parsing on the first log data sample to parse a system log containing various event information into a structured record containing time, event type, event subject, and detailed description; Map the structured record to the knowledge system framework of the log data established based on ontology to generate a log data knowledge graph; Based on log outlier detection combining density clustering and outlier detection, perform semantic repair on the log data knowledge graph to generate the second log data sample.
[0119] Specifically, the detailed description of the preprocessing steps of the above data preprocessing module for the log data sample is as follows:[[]]
[0120] I. Semantic understanding and parsing based on the sequence-to-sequence model of the recurrent neural network
[0121] Let the first log data sample be represented as a text sequence set L = {l 1 , l 2 , …, l n}, where each l i represents a system log text, which is a string containing various event information. The sequence-to-sequence (Seq2Seq) model of the recurrent neural network (RNN) used can be regarded as consisting of an encoder and a decoder. For the encoder part, each word in the input log text l i is represented as a vector x t(where t represents the word order position in the text), the update formula for the hidden state of the encoder can be expressed as (taking the common Long Short-Term Memory Network (LSTM) as an example, here the core state update calculation is simplified):
[0122] i t =σ(W xi x t +W hi h t-1 +b i )
[0123] f t =σ(W xf x t +W hf h t-1 +b f )
[0124] o t =σ(W xo x t +W ho h t-1 +b o )
[0125] g t =tanh(W xg x t +W hg h t-1 +b g )
[0126] c t =f t ⊙c t-1 +i t ⊙g t
[0127] h t =o t ⊙tanh(c t )
[0128] Where: W xi ,W xf ,W xo ,W xg are the weight matrices input to the corresponding gates (input gate i t , forget gate f t , output gate o t , candidate memory cell gate g t ), and they determine the input word vector x tThe degree of influence on each state, whose dimension depends on the word vector dimension and the hidden state dimension, etc., is determined during the model training phase. In this log parsing scenario, different weight matrices capture the associations between different word vectors and each hidden state, and are used to learn the semantic information and sequential relationships in the log text. W hi ,W hf ,W ho ,W hg is the weight matrix of the corresponding gate from the hidden state h t-1 at the previous moment to each state at the current moment, which reflects the transmission and influence of the hidden states at the previous and current moments in the sequence, and helps the model understand the sequential semantics of the log text, such as understanding the manifestation of different events occurring in sequence in the log. b i ,b f ,b o ,b g is the bias vector of the corresponding gate, which is used to adjust the activation of each state. c t is the memory cell state, which summarizes historical information and is continuously updated along with the input log text sequence, accumulating the semantic key information in the log text, h t is the final hidden state, which synthesizes the current input and historical information, representing a semantic encoding representation of the encoder for the log text at the current position. After passing through the encoder, the final hidden state h last (corresponding to the final encoding state after processing the entire log text) is used as the initial state input to the decoder. The decoder part is also based on the recurrent neural network structure. When generating each output structured record element (such as time, event type, event subject, detailed description, etc.), the calculation formula for its output probability distribution can be expressed as (the size of the output vocabulary is V, and here it is simplified to illustrate the case of generating a single element, and actually may be multiple rounds of generating different elements):
[0129]
[0130] where: W hs is the weight matrix from the decoder hidden state h t to the intermediate state s t , which is used to transform the hidden state and extract the semantic representation suitable for generating the output element, b s is the corresponding bias vector.
[0131] W ps is the weight matrix from the intermediate state s t to the output vocabulary probability distribution, which determines the likelihood of different words becoming the current output element, b p is the corresponding bias vector, p t(v) represents the probability of generating the vocabulary v at time t. The intermediate state is converted into a probability distribution through the softmax function, ensuring that the sum of the probabilities of all possible vocabularies is 1. Finally, the appropriate vocabulary is selected according to the probability distribution to form a structured record element. For example, the vocabulary with the highest probability is selected as the output of the event type. For each log text l i , the word vectors need to be input into the encoder one by one in order, updating the hidden state at each time step, and then the final hidden state is passed into the decoder to gradually generate each element of the structured record. The code example is as follows:
[0132] import torch
[0133] import torch.nn as nn
[0134] # Seq2Seq model class, including structures such as Encoder and Decoder
[0135] model = Seq2SeqModel()
[0136] # Traverse each log text in the first log data sample
[0137] for l_i in L:
[0138] # Convert the log text l_i into a sequence of word vectors
[0139] x = word_embedding(l_i)
[0140] # Initialize the encoder hidden state, etc.
[0141] h_prev = init_hidden_state()
[0142] c_prev = init_cell_state()
[0143] for t in range(len(x)):
[0144] # Obtain the word vector at the current time step
[0145] x_t = x[t]
[0146] # Update the hidden state and memory cell state through the encoder (call the function defined in the model, corresponding to the above formula calculation)
[0147] h_t, c_t = model.encoder(x_t, h_prev, c_prev)
[0148] h_prev = h_t
[0149] c_prev = c_t
[0150] # Obtain the final hidden state of the encoder as the initial state of the decoder
[0151] h_last = h_t
[0152] # The decoder generates structured record elements
[0153] s_t = model.decoder.tanh(model.decoder.W_hs * h_last + model.decoder.b_s)
[0154] p_t = nn.functional.softmax(model.decoder.W_ps * s_t + model.decoder.b_p, dim = 0)
[0155] # Select the output word according to the probability distribution output_word = choose_word(p_t)
[0156] After the above process, each first log data sample l i is parsed into a structured record containing elements such as time, event type, event subject, detailed description, etc. Let the set of parsed structured records be S = {s 1 , s 2 , …, s n}, where each s i corresponds to the structured representation of a log.
[0157] II. Map the structured record to the log data knowledge system framework to generate a knowledge graph
[0158] Let the knowledge system framework of log data established based on ontology define some concept sets C = {c 1 , c 2 , …, c k} (for example, concepts such as "login event" and "file access event" belong to event type-related concepts, and "user", "server", etc. belong to event subject-related concepts, etc.) and relation sets R = {r 1 , r 2 , …, r m} (for example, relations such as "occurred before" and "executed by whom"). For each structured record s i , its elements can be mapped according to the concepts and relations of the knowledge system framework. Define a mapping function f map (s i ), which is based on s iThe elements in it are mapped to nodes and edges in the knowledge graph. The elements in the knowledge graph are represented in the form of triples, i.e., G = {(n 1 , r, n 2 )}, where n 1 , n 2 are nodes (corresponding to concept instances), and r is a relationship. For example, if the event type in the structured record s i is "user login" and the event subject is "User A", then the following triple can be generated through the mapping function and added to the knowledge graph: ("User A", "performed", "user login"). It is necessary to traverse the set S of structured records and call the mapping function f i for each s map (s i ) to convert its elements into triple elements of the knowledge graph and add them to the knowledge graph. The code example is as follows:
[0159] knowledge_graph = []
[0160] for s_i in S:
[0161] triples = f_map(s_i)
[0162] knowledge_graph.extend(triples)
[0163] In this way, the knowledge graph G of log data is generated, which integrates the structured information of all logs in the form of a graph and more clearly shows the semantic relationships and associations between log data.
[0164] III. Semantic repair based on the combination of density clustering and outlier detection for log outliers
[0165] Let the set of nodes in the knowledge graph G of log data be N = {n 1 , n 2 , …, n p}, and the set of edges be E = {e 1 , e 2 , …, n q}. For the density clustering part, the density-based spatial clustering algorithm (DBSCAN) can be taken as an example (the principles of other density clustering algorithms are similar), and concepts such as density reachability and density connectivity are defined to divide clusters. Let ∈ be the neighborhood radius parameter, and MinPts be the threshold of the minimum number of points in the neighborhood of core points (these two parameters are empirically set according to the characteristics of log data and clustering effects). The function for calculating the number of nodes in the ∈-neighborhood of node n i can be expressed as:
[0166] N ∈ (ni ) = {n j | d(n i , n j ) ≤ ∈, n j ∈ N}
[0167] where d(n i , n j ) represents the distance metric between node n i and n j . (An appropriate distance function can be defined based on the semantic similarity of nodes in the knowledge graph, such as calculating the cosine distance based on the vector representation of the concepts corresponding to the nodes, etc.). If |N ∈ (n i )| ≥ MinPts, then n i is called a core point.
[0168] By continuously finding sets of nodes that are density-reachable and density-connected, the nodes are divided into different clusters Clusters = {C 1 , C 2 , …, C r}.
[0169] For outlier detection, an outlier score function S outlier (n i ) can be defined, which combines factors such as the density of the cluster where the node is located and the distance to other clusters to measure whether the node is an outlier. For example:
[0170]
[0171] where represents the cluster where node n i is located, is the number of nodes in this cluster, λ is a weight parameter that balances the influence of the intra-cluster distance and the inter-cluster distance (set according to the actual situation). If S outlier (n i ) is greater than a set outlier threshold θ, then n i is determined to be an outlier. For the detected outliers, semantic repair is required. For example, they can be corrected according to the semantic information of other nodes in their clusters. (The repair strategy here can formulate different rules according to the specific semantics and application scenarios of the log knowledge graph). Let the repair function be f repair (n i ), which modifies the outlier n i based on its relevant information to make it more conform to the normal semantic pattern. First, calculate the neighborhood information of each node for density clustering division, then calculate the outlier score to detect outliers, and finally perform repair operations on the outliers. The code example is as follows:
[0172] # Initialize the cluster set and the outlier set
[0173] Clusters = []
[0174] Outliers = []
[0175] # Perform density-based clustering
[0176] for n_i in N:
[0177] neighbors = get_neighbors(n_i, epsilon) # Get the nodes within the epsilon neighborhood
[0178] if len(neighbors) >= MinPts:
[0179] # It is a core point, start expanding the cluster
[0180] cluster = expand_cluster(n_i, neighbors, epsilon, MinPts)
[0181] Clusters.append(cluster)
[0182] else:
[0183] Outliers.append(n_i)
[0184] # Detect outliers and calculate scores
[0185] for n_i in N:
[0186] outlier_score = calculate_outlier_score(n_i, Clusters, lambda)
[0187] if outlier_score > theta:
[0188] # Determine as an outlier
[0189] Outliers.append(n_i)
[0190] # Semantically repair the outliers
[0191] for n_i in Outliers:
[0192] repaired_n_i = f_repair(n_i)
[0193] # Update the repaired nodes to the knowledge graph
[0194] After the above log outlier detection and semantic repair processes, a repaired log data knowledge graph is obtained, which is the final second log data sample. By removing outliers and repairing semantic problems, the quality of the log data and the effectiveness of network security threat analysis are improved, making it more accurately reflect the actual operating state of the network system.
[0195] Optionally, when the data preprocessing module preprocesses the first network traffic data sample, the first log data sample, and the first user behavior data sample to obtain the second network traffic data sample, the second log data sample, and the second user behavior data sample respectively, the steps included are as follows: Based on establishing an association model for multi-source data, cross-verify and fuse the first user behavior data samples of the same user from different channels, and when there are inconsistencies in time or operation sequence between the first user behavior data sample and the user operation logs recorded by the application system, use the user login time and permission information as assistance to comprehensively judge and correct the integrity of the first user behavior data sample. Among them, for the missing data part, according to the user's historical behavior pattern and group behavior statistical law, a behavior sequence prediction filling mechanism based on Markov chain is adopted for filling to generate the second user behavior data sample.
[0196] Specifically, the detailed description of the above preprocessing steps for the user behavior data sample is as follows:
[0197] I. Cross-validation and fusion based on the multi-source data association model
[0198] Let the sets of the first user behavior data samples of the same user obtained from different channels be D 1 ={d 11 ,d 12 ,…,d 1m}, D 2 ={d 21 ,d 22 ,…,d 2n},…, D k ={d k1 ,d k2 ,…,d kp}, where each set represents the data from a specific channel. For example, D 1 may be the operation records of the user on the mobile side, D 2 is the operation records from the web side, etc. And each element d ij represents a specific behavior data record, such as a click operation, a page view, etc., and each record contains multiple attributes, such as the behavior occurrence time t ij , the behavior type typeij , the operation object obj ij , etc. The established association model of multi-source data can be regarded as a function M associate (D 1 , D 2 , …, D k ), and its goal is to perform cross-verification and fusion by analyzing the association relationships between data from different channels.
[0199] First, define a similarity metric function to measure the similarity between two behavior data records from different channels. For example, it can be defined based on attributes such as behavior type and operation object. Let the similarity function be S(d i1j1 , d i2j2 ), and its value range is between [0, 1]. 0 means completely dissimilar, and 1 means completely identical. For the cross-verification part, it is necessary to traverse all possible data pairs from different channels to determine whether they corroborate each other. The verification logic can be expressed by the following formula: Verification passing condition: where θ validate is the set similarity threshold, which is determined according to the characteristics of the data and the requirements for verification accuracy. If the similarity of two records is greater than or equal to this threshold, it is considered that they verify each other's authenticity and rationality to a certain extent. For the fusion part, a weighted average method can be used to fuse similar data records (this is just an example of a common fusion method, and it can be adjusted according to specific requirements in practice). The two similar records to be fused are d a and d b , and their corresponding weights are w a and w b (the weights can be set according to factors such as the reliability of the channel and the timeliness of the data, and satisfy w a +w b =1). The formula for calculating a certain attribute (taking the behavior occurrence time as an example) of the fused record d fusion is as follows: t fusion =w a ×t a +w b ×t b . The same method can be applied to the fusion of other attributes. After processing all the fusible data records, the fused user behavior data set D fused is obtained. From an execution perspective, the code example is as follows:
[0200] D_fused = []
[0201] for i in range(len(D_1)):
[0202] for j in range(len(D_2)):
[0203] similarity = S(D_1[i], D_2[j])
[0204] if similarity >= theta_validate:
[0205] # Determine the weight
[0206] w_a = 0.5
[0207] w_b = 0.5
[0208] # Integrate the behavior occurrence time attribute (taking the time attribute as an example)
[0209] t_fusion = w_a * D_1[i].time + w_b * D_2[j].time
[0210] # Integrate other attributes (specific code is omitted, similar to the time attribute integration method)
[0211] d_fusion = create_fused_data(D_1[i], D_2[j], t_fusion)
[0212] D_fused.append(d_fusion)
[0213] II. Correct data integrity using the user login time and permission information
[0214] Let the set of user behavior data after the above integration be D fused = {d f1 , d f2 , …, d fq}, and the set of user operation logs recorded by the application system be L = {l 1 , l 2 , …, l r}, where each log also contains relevant attributes such as time and operation.
[0215] Let the user login time be T login , and the permission information can be represented by the set P = {p 1 , p 2 , …, p s}, and each element p i represents a permission (such as read - write permission, access to specific function permission, etc.). Define a time difference function Δt(d fi , l j ) = |t fi - t lj|, used to measure the time difference between user behavior data records and operation logs, and then define a permission matching function Match(p i ,d fi ), which is used to determine whether the permission information matches the behavior data (the return value is a boolean value, True for match and False for non-match). When it is found that there are inconsistencies in time or operation order between the user behavior data sample and the user operation logs recorded by the application system, the following rules are used to comprehensively judge whether correction is needed and how to correct it:
[0216] Correction condition 1 (excessive time difference): Δt(d fi ,l j )>θ time , where θ time is the set time difference threshold. If this threshold is exceeded, it is considered that the time difference is too large and the time attribute of the behavior data may need to be corrected. Correction condition 2 (permission mismatch): Match(p i ,d fi ), that is, when the permissions do not match, it is also necessary to consider correcting the relevant attributes of the behavior data (such as judging whether the behavior should be executed with the current permissions, etc.). When the correction conditions are met, the behavior data is corrected according to the user login time and permission information. For example, if the time attribute needs to be corrected and the time of the behavior data is earlier than the login time (unreasonable situation), its time can be corrected to a reasonable time after the login time (corrected to the login time plus a small time interval ΔT), which is expressed by the formula as follows (taking the correction of the time attribute as an example):
[0217]
[0218] where θ time_max is another set maximum reasonable time range threshold, which is used to avoid overcorrection.
[0219] From an execution perspective, the code example is as follows:
[0220] for i in range(len(D_fused)):
[0221] for j in range(len(L)):
[0222] time_diff = Delta_t(D_fused[i],L[j])
[0223] permission_match = Match(P,D_fused[i])
[0224] if time_diff > theta_time or not permission_match:
[0225] if D_fused[i].time < T_login:
[0226] D_fused[i].time = T_login + Delta_T
[0227] elif D_fused[i].time > T_login + theta_time_max:
[0228] D_fused[i].time = D_fused[i].time - Delta_T
[0229] III. Predicting and Complementing Missing Data Based on Markov Chain Behavior Sequences
[0230] After the previous steps of processing, there may still be some missing data. For these missing data parts, a mechanism for predicting and complementing based on Markov chain behavior sequences is adopted. Let the historical behavior sequence of the user be represented as H = {h 1 , h 2 , …, h u}, where each h i is a behavior state (a state description that can be comprehensively represented by behavior types, operation objects, etc.). The statistical law of group behavior can be represented by the transition probability matrix P, where the element P ij represents the probability of transitioning from behavior state i to behavior state j. This matrix can be obtained by analyzing a large amount of historical behavior data of users. Let the behavior state before the current missing data position be h prev . To predict the next behavior state to complement the missing data, according to the principle of Markov chain, the probability calculation formula for predicting the next behavior state is: P(h next = j|h prev ) = P prevj
[0231] That is, the probability that the next behavior state is j is equal to the probability P prev of transitioning from the current previous behavior state h prevj to state j. Then, based on this probability distribution, sampling can be performed (such as using sampling methods like the roulette wheel algorithm) to determine the actual complemented behavior state h fill .
[0232] From an implementation perspective, the code example is as follows:
[0233] for d in D_fused:
[0234] if d.is_missing(): # Function to determine if the data is missing
[0235] prev_state = get_prev_state(d) # Obtain the behavior state before the position of the missing data
[0236] probabilities = P[prev_state] # Obtain the transition probability vector
[0237] fill_state = sample_state(probabilities) # Determine the filled behavior state by sampling according to the probability
[0238] d.fill(fill_state) # Fill the filled state into the position of the missing data
[0239] After the above steps of cross - validation and fusion, integrity correction, and missing data filling, the finally obtained processed user behavior data set is the second user behavior data sample. Through such a pre - processing process, multi - source data, relevant auxiliary information, and historical and group behavior rules can be fully utilized to improve the quality and integrity of user behavior data, providing a more reliable data basis for subsequent operations such as feature extraction and network security threat monitoring.
[0240] Optionally, when the feature extraction module extracts features from the second network traffic data sample, the second log data sample, and the second user behavior data sample respectively to generate traffic feature data samples, log feature data samples, and behavior feature data samples, the steps included are as follows: Take the nodes in the network as the vertices of the graph, the network connections and traffic flow relationships between the nodes as the edges, and the size and transmission frequency of the traffic as the weights of the edges to construct a traffic graph structure representation; Based on the graph convolutional neural network, extract features from the traffic graph structure representation to identify the key nodes, traffic aggregation regions, and traffic association patterns between different network segments in the network, and generate the traffic feature data samples accordingly.
[0241] Therefore, the detailed description of the above steps for extracting features from the network traffic data sample is as follows:
[0242] I. Construct the traffic graph structure representation
[0243] Suppose there are N nodes in the network, and the set V = {v 1 , v 2 , …, v N} is used to represent these nodes, and each node v i corresponds to an entity in the network (such as a router, a server, a terminal device, etc.). The network connections and traffic flow relationships between the nodes form the set of edges E = {e ij |i, j ∈ {1, 2, …, N}}, where e ijDenote the edge from node v i to node v j If there is no direct connection or traffic flow from v i to v j then the corresponding e ij does not exist. For each edge e ij define its weight as the binary tuple where represents the traffic volume from node v i to node v j (which can be measured by, for example, the number of bytes of data packets passing through per unit time, with the unit being bytes / second, etc.), represents the traffic transmission frequency (such as the number of data packets transmitted per second, with the unit being packets / second). Then, the entire traffic graph structure can be represented as G=(V, E, W), where is the set of edge weights. It is necessary to traverse all the nodes in the network and their connection situations, and collect traffic volume and transmission frequency information to construct this graph structure. The code example is as follows:
[0244] V = [] # Initialize the node set
[0245] E = [] # Initialize the edge set
[0246] W = [] # Initialize the edge weight set
[0247] # Obtain node information and connection and traffic situations through the network monitoring interfacefor i in range(len(all_nodes)):
[0248] for j in range(len(all_nodes)):
[0249] if has_connection(i, j): # Function to determine whether nodes i and j are connected
[0250] v_i = all_nodes[i]
[0251] v_j = all_nodes[j]
[0252] V.append(v_i)
[0253] V.append(v_j)
[0254] e_ij = (v_i, v_j)
[0255] E.append(e_ij)
[0256] traffic_size = get_traffic_size(i, j) # Function to obtain traffic size
[0257] traffic_freq = get_traffic_freq(i, j) # Function to obtain traffic transmission frequency
[0258] w_ij = (traffic_size, traffic_freq)
[0259] W.append(w_ij)
[0260] G = (V, E, W) # Construct the traffic graph structure representation
[0261] II. Feature extraction based on the Graph Convolutional Neural Network (GCN)
[0262] The core operation of the Graph Convolutional Neural Network is to perform convolutional operations on the graph structure to aggregate the information of neighboring nodes, and then extract the features of the graph.
[0263] Let the node feature matrix of the l-th layer of the Graph Convolutional Neural Network be where N is the number of nodes (consistent with the number of nodes in the previously defined graph), and d (l) represents the feature dimension of the l-th layer. Initially (i.e., when l = 0), each row of the node feature matrix H (0) can be initialized with some basic attributes of the node itself (such as simple features like the degree of the node, the encoding of the IP address, etc., with the dimension set to d (0) ).
[0264] The formula for the graph convolutional operation is usually expressed as follows:
[0265] where: is the adjacency matrix of the graph G. If e ij ∈E (i.e., there is an edge connection between nodes v i and v j ), then A ij = 1; otherwise, A ij = 0. To consider the node's own information (i.e., self-loop), usually construct where I is the identity matrix, so as to ensure that each node can also aggregate its own information during convolution. is the degree matrix, which is a diagonal matrix, and its diagonal element D ii is equal to the degree of node v i (i.e., the number of edges connected to this node). Similarly, for the convenience of subsequent calculations, construct whose diagonal element is the node degree considering the self-loop. is the learnable weight matrix of the l-th layer, which determines how to map the features of the previous layer to the current layer. Different weight values will learn the importance of different feature combinations for extracting target features (such as key nodes, traffic aggregation regions, etc.) during the training process. By adjusting the weights, the network can focus on extracting features valuable for network security threat judgment. σ(·) is the activation function (such as the commonly used ReLU function, etc.), which is used to introduce non-linearity, enabling the network to learn more complex feature relationships. It non-linearly activates the node features after linear transformation and normalization, helping to uncover non-linear feature patterns hidden in the graph structure, such as complex traffic association patterns between different network segments, etc. After several layers (L layers) of graph convolution operations, the finally obtained feature matrix H (L) contains the extracted high-level features of the graph, which can be used to determine key nodes, traffic aggregation regions, and traffic association patterns between different network segments in the network, etc.
[0266] For example, for judging key nodes, a key node scoring function S node (v i ) can be defined, and the score is calculated based on the feature vector of the corresponding node in the final feature matrix:
[0267] where is the k-th dimensional eigenvalue of node v i in the final feature matrix H (L) , and w node,k is the k-th element of the pre-set weight vector w node used to measure the importance of each dimensional feature for key node judgment. By such weighted summation, the key node score of the node is obtained, and nodes with higher scores can be determined as key nodes. For the judgment of traffic aggregation regions, it can be comprehensively judged based on factors such as feature similarity between nodes and traffic weights. For the traffic association patterns between different network segments, the relationships between node groups corresponding to different network segments in the feature space and the traffic flow direction and magnitude relationships reflected by the edge weights can be analyzed to mine the corresponding patterns (similarly, complex analysis and calculations can be performed based on the feature matrix and edge weights, etc., and the detailed formulas are not specifically listed here).
[0268] According to the above graph convolution operation formula, matrix multiplication, normalization and other operations are sequentially performed on each layer to update the node feature matrix. The code example is as follows:
[0269]
[0270] Finally, based on the above analysis and judgment results of key nodes, traffic aggregation areas, traffic association patterns, etc., corresponding traffic feature data samples are sorted out. These samples will serve as the basic data for subsequent further analysis and network security threat monitoring operations, providing strong support for accurately grasping the security-related features in network traffic.
[0271] Optionally, when the feature extraction module extracts features from the second network traffic data sample, the second log data sample, and the second user behavior data sample respectively to generate traffic feature data samples, log feature data samples, and behavior feature data samples, the steps included are as follows: Based on a pre-trained language model, annotate the corresponding semantic roles for each word in the second log data sample; Based on the annotated semantic roles, extract the subject, object of operation, operation method, and time and location of the event to generate the log feature data sample.
[0272] Specifically, the detailed description of the above steps for feature extraction from the log data sample is as follows:
[0273] I. Annotating semantic roles based on a pre-trained language model
[0274] Let the second log data sample be represented as a text set L = {l 1 , l 2 , …, l n}, where each l i represents a pre-processed log text, which itself is a sequence composed of a series of words and can be expressed as l i = {w i1 , w i2 , …, w im}, where w ij represents the j-th word in the log text l i . The pre-trained language model used is a function M language (w ij ), which can output the corresponding semantic role annotation for the input word. Common semantic role annotations may include different categories such as "Agent", "Patient", "Instrument", "Time", "Location", etc., and are represented by the set R = {r 1 , r 2 , …, r k} to represent all possible semantic role categories.
[0275] For each word w ij , after being processed by the pre-trained language model, a probability distribution vector regarding semantic roles will be obtained Its calculation formula can be expressed as (taking the common way of outputting probability distribution of a neural network-based language model as an example, the actual situation involves complex model structures and parameter operations): p ij = M language (w ij ) = softmax(W·e ij + b)
[0276] Among them: e ij is the word vector representation of the word w ij , which is obtained by mapping the word to a vector space of a fixed dimension through the word embedding layer in the language model. The dimension of the word vector is set to d, that is The word vector contains semantic and syntactic information of the word. For example, words with similar semantics are relatively close in the vector space, which is convenient for the model to perform semantic analysis based on this later. is the weight matrix, which determines the mapping relationship from the word vector space to the semantic role probability distribution space. Different weight values will make different word vector features correspond to different semantic role probabilities, and it is learned during the training process of the pre-trained language model. Through learning a large amount of text data, the model can adjust the weights so that for words in different contexts, it can accurately output a reasonable semantic role probability distribution. is the bias vector, which is used to fine-tune the probability distribution. It is also determined during the model training stage and helps to improve the generalization ability and annotation accuracy of the model on different datasets. softmax(·) is the normalized exponential function, which converts the result after the linear transformation (W·e ij + b) into a probability distribution vector, so that the sum of the elements in the vector is 1. The value range of each element p i r j s (indicating the probability that the word w ij is labeled as the semantic role r s ) is between $[0, 1]$. In this way, it is possible to intuitively determine the most likely semantic role corresponding to the word according to the probability size. Traverse each word in each log text, input it into the pre-trained language model, obtain the corresponding semantic role probability distribution vector, and then select the semantic role with the highest probability as the annotation result of the word (of course, other probability-based decision-making methods can also be used, such as setting thresholds, etc.). The code example is as follows:
[0277]
[0278] II. Extracting Log Feature Data Samples Based on Annotated Semantic Roles
[0279] After completing the semantic role annotation for each word in the log text, based on this annotation information, extract the key features such as the subject, object of operation, operation method, and time and location of the event to generate log feature data samples. Let the set of event subjects extracted from the log text l i be A i ={a i1 , a i2 , …, a ip}, the set of objects of operation be O i ={o i1 , o i2 , …, o iq}, the set of operation methods be M i ={m i1 , m i2 , …, m ir}, the set of times be T i ={t i1 , t i2 , …, t is}, and the set of locations be G i ={g i1 , g i2 , …, g it}. The rules for extracting these features can be defined in the following way: Extract the event subject: If the semantic role annotation of the word w ij is "Agent", then add this word to the set of event subjects, that is: a il =w ij , if and only if the semantic role annotation r ij ="Agent". Extract the object of operation: If the semantic role annotation of the word w ij is "Patient", then add this word to the set of objects of operation, that is: o im =w ij , if and only if the semantic role annotation r ij ="Patient". Extract the operation method: If the semantic role annotation of the word w ij is "Instrument" or other semantic roles related to the operation method (according to specific definitions), then add this word to the set of operation methods, that is: m in =w ij , if and only if the semantic role annotation r ij ∈{the set of semantic roles related to the operation method}. Extract the time: If the semantic role annotation of the word w ij is "Time", then add this word to the set of times, that is: t io =w ij , if and only if the semantic role annotation rij = "Time". Extraction location: If the semantic role annotation of word w ij is "Location", then add this word to the location set, i.e.: g ip = w ij , if and only if the semantic role annotation r ij = "Location". Traverse each log text and the words with annotated semantic roles in it again, and collect the corresponding feature information according to the above extraction rules. The code example is as follows:
[0280]
[0281] Through the above process of annotating semantic roles based on the pre-trained language model and further extracting key features, a set of log feature data samples is finally generated. These samples can focus on the key elements in the log text that are closely related to network security threat analysis, providing a valuable data basis for subsequent accurate network security threat monitoring, abnormal behavior judgment, etc. based on log data.
[0282] Optionally, when the feature extraction module extracts features from the second network traffic data sample, the second log data sample, and the second user behavior data sample respectively to generate traffic feature data samples, log feature data samples, and behavior feature data samples, the steps included are as follows: Taking users as nodes, the operation behaviors of users at different times, on different application systems or devices as edges, and the time intervals between behaviors and the frequency of operations as attributes of the edges, to construct a user behavior trajectory graph; Mapping the user behavior trajectory graph to a low-dimensional vector space through graph embedding to determine the dynamic change process of user behaviors and the correlation between different behaviors, and generating the behavior feature data samples accordingly. Specifically, the detailed description of the above steps for extracting features from the user behavior data sample is as follows:
[0283] I. Constructing a user behavior trajectory graph
[0284] Suppose there are N users in the system, and use the set U = {u 1 , u 2 , …, u N} to represent these users. Each user u i is the node in the behavior trajectory graph to be constructed. The operation behaviors of users at different times, on different application systems or devices form the edge set E = {e ij |i, j ∈ {1, 2, …, N}}, where e ij represents from user u i to user u jThe edges (which can also be the connections between different behaviors of the user himself, i.e., when i = j), represent the associations between users or between different operation behaviors of the user himself. For each edge e ij , define two of its attributes: the time interval attribute of the behavior occurrence, denoted as Δt ij , which represents the time interval from a certain behavior of user u i to the occurrence of the related behavior of user u j . It can be calculated by subtracting the timestamps of the corresponding operation behaviors. The unit can be seconds, minutes, etc. For example, if user u i performs an operation at time T i , and user u j performs a related operation at time T j , then Δt ij = T j - T i . The operation frequency attribute, denoted as f ij , can be measured by counting the number of occurrences of such behaviors from user u i to the related behaviors of user u j within a certain time range. For example, within the past hour, the login operation of user u i to the file download operation of user u j occurred 5 times. Then the corresponding f ij value can be set to 5 (the specific statistical range and measurement method can be determined according to the actual application scenario). In this way, the entire user behavior trajectory graph can be represented by G=(U, E, A), where A ={(Δt ij , f ij ) | e ij ∈ E} is the set of edge attributes.
[0285] From an execution perspective, it is necessary to traverse the user behavior data recorded by the system, identify different users and their operation behaviors at different times, on different application systems or devices, and calculate the corresponding time intervals and operation frequencies to construct this graph structure. The code example is as follows:
[0286]
[0287] II. Mapping the user behavior trajectory graph to a low-dimensional vector space through graph embedding
[0288] The purpose of graph embedding is to represent the nodes (here are users) in the graph structure and the relationships between them (reflected by edges and edge attributes) with low-dimensional vectors, so as to analyze the dynamic changes of user behavior and the correlation between different behaviors in the follow-up. Common graph embedding methods include DeepWalk, Node2Vec, Graph Convolutional Networks (GCN) for graph embedding, etc. Here, Node2Vec is taken as an example to illustrate (its principle and related formulas). Node2Vec is a method based on random walk and combined with word vector learning methods (such as Word2Vec) to generate vector representations of nodes in the graph.
[0289] First, define the relevant parameters of the random walk: Let the walk length be l, which represents the number of consecutive nodes visited starting from the starting node in each random walk. For example, set l = 8, that is, each random walk will pass through 8 nodes in sequence (there may be repeated visits). Let the return parameter be p, which controls the probability of the random walk returning to the previous node and is used to adjust the "breadth" of the walk. A larger p value makes the walk more inclined to explore near the current node, and a smaller p value will make the walk easier to "jump" to farther nodes. Let the in-out parameter be q, which controls the trade-off between exploring new nodes outward and returning to the area of visited nodes inward. Different q values affect the "bias" of the walk path, such as the effect of preferring depth-first search or breadth-first search. For each node u i , a series of node sequences are generated through multiple random walks. Let the set of node sequences obtained by performing r random walks starting from node u i be S i ={s i1 , s i2 ,…, s ir}, where each s ik is a node sequence of length l. For example, s ik =[v k1 , v k2 ,…, v kl , where v km represents the node visited at the m-th step in the k-th random walk. Then, use a Skip-Gram model similar to that in Word2Vec to learn the vector representation of nodes. Let the vector representation of the node be (d is the dimension of the set low-dimensional vector space, for example, d = 128). For each central node and its context nodes in the node sequence (similar to the relationship between a word and its context words in word vector learning), the following objective function is used for optimization learning (here, the core part is simplified and some regularization terms are ignored):
[0290] Among them, c represents the position index of the central node in the sequence (here, the middle position is taken. For example, for a sequence of length l, c = l / 2), m is the set context window size (indicating how many nodes before and after the central node are considered as context. For example, m = 2), and Pr(v k,j+c |v kc ) represents the probability that the context node v kc appears given the central node v k,j+c . This probability can be obtained by converting through the softmax function of the inner product of node vectors. The calculation formula is as follows:
[0291] Among them, z kc is the vector representation of the central node v kc , z k,j+c is the vector representation of the context node v k,j+c . By continuously adjusting the value of the node vector z i , the above objective function is maximized, that is, nodes that often co-occur in the random walk (nodes corresponding to user behaviors with strong correlations in the behavior trajectory graph) are closer in the vector space, so as to learn a low-dimensional vector representation that can reflect the correlation and dynamic changes of user behaviors. Operate according to the above random walk and vector learning processes. The code example is as follows:
[0292] import node2vec
[0293] # The user behavior trajectory graph G has been constructed
[0294] graph=convert_to_graph_object(G) # Convert the custom graph structure to a graph object supported by the library
[0295] # Initialize the Node2Vec model and set relevant parameters (walk length, return parameter, in-out parameter, etc.)
[0296] model=node2vec.Node2Vec(graph,dimensions=128,walk_length=8,num_walks=10,p=1,q=1)
[0297] # Train the model to perform random walk and vector learning
[0298] model.fit()
[0299] # Obtain the vector representation of the nodes, that is, the embedded vectors
[0300] node_embeddings=model.wv.vectors
[0301] # Associate the node vectors with the corresponding users. user_embeddings = {}
[0302] for user, index in user_index_mapping.items():
[0303] user_embeddings[user] = node_embeddings[index]
[0304] Based on the low-dimensional vector representations of each node (user) in the obtained user behavior trajectory graph, the distances between these vectors (e.g., by calculating cosine similarity, etc.) and the changing trends of the vectors (the changes in the vectors corresponding to different behaviors over time) can be further analyzed to determine the dynamic change process of user behaviors and the correlations between different behaviors, and then behavioral feature data samples can be organized and generated. For example, the difference between user vectors at different time points can be calculated to measure the degree of behavior change, or user behaviors can be classified based on vector similarity through clustering analysis, etc. Key behavioral feature information can be extracted according to these analysis results, and finally behavioral feature data samples are formed to provide effective data support for subsequent operations such as network security threat monitoring. For example, by calculating the cosine similarity between two user vectors z i and z j to measure the correlation of their behaviors, the cosine similarity formula is as follows:
[0305] where, z i ·z j represents the inner product of two vectors, ‖z i ‖ and ‖z j ‖ respectively represent the norms (lengths) of the vectors z i and z j . Through such similarity calculations, the degree of tightness of the correlations between different user behaviors can be judged, and relevant correlation information, etc., is organized and collected as part of the behavioral feature data samples for subsequent use.
[0306] From an implementation perspective, the code example is as follows:
[0307]
[0308] Through the above complete process of constructing the user behavior trajectory graph and graph embedding, behavioral feature data samples are generated. These samples can characterize user behavior features from aspects such as behavior correlations and dynamic changes, and help to detect abnormal user behavior patterns and potential security threats in a timely manner in the network security scenario.
[0309] Optionally, when the data vectorization module obtains the traffic feature vector sample, the log feature vector sample, and the behavior feature vector sample from the traffic feature data sample, the log feature data sample, and the behavior feature data sample respectively, the steps included are as follows: Analyze the feature importance weights of each traffic feature data sample in different network security threat scenarios through the constructed knowledge base of network security threat cases and corresponding traffic feature data. Based on the feature importance weights, assign different attention weights to different traffic feature data samples to integrate the traffic feature data samples with attention weights into a traffic feature vector sample. Specifically, the detailed description of the above steps for vectorizing the traffic feature data sample is as follows:
[0310] I. Analyze the feature importance weights of the traffic feature data sample
[0311] Suppose the constructed knowledge base of network security threat cases and corresponding traffic feature data contains M different network security threat scenarios, which are represented by the set S = {s 1 , s 2 , …, s M}, for example, s 1 can represent the "DDoS attack scenario", s 2 can represent the "port scanning attack scenario", etc. Each scenario corresponds to a specific traffic feature manifestation form. At the same time, suppose the traffic feature data sample set is F = {f 1 , f 2 , …, f N}, where each f i represents a traffic feature data sample, which itself contains multiple specific traffic feature attributes, such as traffic size, packet frequency, source port distribution, etc., and can be expressed as f i = {a i1 , a i2 , …, a iK}, where a ij represents the j-th specific traffic feature attribute in the i-th traffic feature data sample. For each traffic feature attribute a ij , it is necessary to analyze its importance weight in different network security threat scenarios s m . Define a weight function w ijm to represent the importance weight of the attribute a ij in the scenario s m , and its value range is between [0, 1]. 0 means that the attribute has little effect on threat judgment in this scenario, and 1 means that the attribute is crucial for threat judgment in this scenario. To calculate this weight, statistical analysis or machine learning methods are used (here, information gain calculation is used as an example to illustrate the basic principle). In the scenario s mUnder the condition, there is a positive sample set (traffic data samples indicating the existence of this security threat) P m and a negative sample set (traffic data samples indicating the non - existence of this security threat) N m , respectively, count the probability distribution of the occurrence of attribute a ij in the positive and negative sample sets. Let the probability of the occurrence of attribute a ij in the positive sample set P m be and the probability of the occurrence of attribute a m in the negative sample set N be m . Then, the importance weight of this attribute for threat judgment in scenario s
[0312] w ijm can be calculated through the information gain formula: ij ,s m ) = Entropy(s m ) - Entropy(s m |a ij )
[0313] where: Entropy(s m ) represents the information entropy of scenario s m , which is used to measure the degree of uncertainty in judging whether there is a threat in this scenario. The calculation formula is as follows:
[0314] Entropy(s m ) = -p + log 2 p + - p - log 2 p -
[0315] Here, p + is the proportion of positive samples in all samples (the sum of positive and negative samples), and p - is the proportion of negative samples. For example, if there are 50 positive samples, 50 negative samples, and a total of 100 samples, then p + = 0.5, p - = 0.5. Through this formula, the initial degree of uncertainty of the scenario can be calculated. Entropy(s m |a ij ) represents the conditional information entropy of scenario s ij under the condition that attribute a m is known. It reflects the degree of reduction in the uncertainty of scenario threat judgment after knowing the value of this attribute. The calculation formula is as follows (taking the case of discrete attributes as an example):
[0316]
[0317] where V ij is the set of all possible values of attribute a ij and s mv is the sample subset corresponding to scenario s ij when the value of attribute a is v m |s mv | is the number of samples in this subset, and |s m | is the total number of samples of scenario s m . By calculating the conditional information entropy and combining it with the previous information entropy, the information gain can be obtained, that is, the importance weight w ij of attribute a m under scenario s ijm .
[0318] From an execution perspective, it is necessary to traverse each network security threat scenario, for the positive and negative sample sets under each scenario, and then traverse each attribute in each traffic feature data sample to calculate its importance weight according to the above formula. The code example is as follows:
[0319]
[0320] II. Allocate attention weights based on feature importance weights and integrate them into traffic feature vector samples
[0321] After obtaining the importance weight w ijm of each traffic feature attribute under different scenarios, different attention weights need to be allocated to different traffic feature data samples based on these weights, and then integrated into traffic feature vector samples. Let the attention weight allocated to traffic feature data sample f i under scenario s m be α im , which can be obtained by normalizing the importance weights of all feature attributes in sample f i under this scenario. The calculation formula is as follows:
[0322]
[0323] Here, the result after summing the weights is converted to a positive number through the exponential function exp(·), and then a normalization operation is performed (the denominator is the sum of similar weights of all traffic feature data samples under this scenario), so that the sum of the attention weights of all samples under this scenario is 1, which can highlight the relatively more important traffic feature data samples under a specific scenario. Then, the traffic feature data samples with attention weights are integrated into traffic feature vector samples. Let the integrated traffic feature vector sample be v, whose dimension is the same as the number of feature attributes K of the traffic feature data sample. The calculation formula for each element of the vector is as follows:
[0324]
[0325] That is, for the k-th element of the vector, it is obtained by traversing all network security threat scenarios and all traffic feature data samples, multiplying the attention weight of each sample in each scenario by the corresponding feature attribute value and then accumulating. In this way, the information in different traffic feature data samples is integrated into a vector according to its importance and relevance to different threat scenarios, and a traffic feature vector sample is obtained.
[0326] Calculate the attention weight of each sample in each scenario according to the above attention weight calculation formula, and then calculate the element values of the traffic feature vector sample according to the integration formula. The code example is as follows:
[0327]
[0328] Through the above operation process combined with specific formulas from the execution perspective, the importance weight of traffic feature data samples can be analyzed according to the network security threat case knowledge base, and the attention weight can be reasonably allocated. Finally, a traffic feature vector sample is integrated. This sample can better reflect the relationship between traffic features and network security threats, and provide an effective data representation form for subsequent model training and network security threat monitoring operations.
[0329] Optionally, when the data vectorization module obtains traffic feature vector samples, log feature vector samples, and behavior feature vector samples from the traffic feature data samples, log feature data samples, and behavior feature data samples respectively, the steps included are as follows: Map keywords and key phrases in the log feature data samples to a low-dimensional vector space based on a pre-trained semantic embedding model, so that words with similar semantics are presented in a close position relationship in the vector space to capture the semantic associations between the log feature data samples and calculate the semantic similarity matrix between the log feature samples accordingly; Weightedly expand the log feature data samples according to the semantic similarity matrix to fuse log feature data samples with related semantics and generate the log feature vector samples accordingly. Specifically, the detailed description of the above solution is as follows:
[0330] I. Mapping Vocabulary to a Low-Dimensional Vector Space and Calculating the Semantic Similarity Matrix Based on a Pre-Trained Semantic Embedding Model
[0331] 1. Mapping Vocabulary to a Low-Dimensional Vector Space
[0332] Let the set of log feature data samples be L = {l 1 ,l 2 ,…,l n}, where each l iRepresents a log feature data sample, and each sample contains a number of keywords and key phrases, which can be expressed as l i ={w i1 ,w i2 ,…,w im}, where w ij represents the j-th keyword or key phrase in the log feature data sample l i .
[0333] Use a pre-trained semantic embedding model to map these words and convert them into vector representations in a low-dimensional vector space. Let the pre-trained semantic embedding model be E(·). For the word w ij , its mapped vector representation is (d is the dimension of the set low-dimensional vector space, for example, d = 300), then there is: v ij =E(w ij )
[0334] The pre-trained semantic embedding model E(·) has rich semantic information trained based on a large amount of text data inside, and realizes the mapping from words to vectors through a complex neural network structure and parameters. The purpose is to make words with similar semantics show similar positional relationships in this low-dimensional vector space, which is convenient for capturing semantic associations subsequently.
[0335] From an execution perspective, it is necessary to traverse each keyword and key phrase in the log feature data sample one by one, and call the pre-trained semantic embedding model to obtain the corresponding vector representation. The code example is as follows:
[0336]
[0337] After obtaining the vector representations of each word in all log feature data samples, it is necessary to calculate the semantic similarity matrix between log feature samples. Let the semantic similarity between the log feature data samples l i and l j be s ij , and its value range is between $[0,1]$. 0 means that the semantics of the two are completely irrelevant, and 1 means that the semantics are highly similar or even equivalent.
[0338] Use cosine similarity to calculate the semantic similarity. The calculation formula is as follows:
[0339] Among them: v i represents the comprehensive vector representation of the log feature data sample l i , which is obtained by aggregating the vectors of all keywords and key phrases in the sample l i . Common aggregation methods such as average aggregation, that is:
[0340] Here, v ik is the vector representation of the k-th word in the sample l i . By averaging and aggregating, a vector representing the semantics of the entire log feature data sample is obtained. v j Similarly, it is the comprehensive vector representation of the log feature data sample l j , obtained in a similar aggregation manner. v i ·v j represents the inner product of two vectors, reflecting the degree of consistency in direction between the vectors. The closer the directions are, the more similar the semantics are, and the larger the inner product value. ‖v i ‖ and ‖v j ‖ respectively represent the norms (lengths) of the vectors v i and v j . Normalization is performed by dividing by the product of the vector norms to ensure that the similarity values are in the range [0, 1]. For n log feature data samples, an n×n semantic similarity matrix S is constructed, where the element S ij is the s ij calculated as above.
[0341] From an execution perspective, it is necessary to traverse all pairs of log feature data samples, calculate their semantic similarities according to the above formula, and fill them into the semantic similarity matrix. The code example is as follows:
[0342]
[0343] II. Generating Log Feature Vector Samples through Weighted Expansion and Fusion Based on the Semantic Similarity Matrix
[0344] 1. Weighted Expansion
[0345] Let the vector after weighted expansion of the log feature data sample l i according to the semantic similarity matrix be . Its calculation method is as follows:
[0346]
[0347] where: s ik is the semantic similarity between the log feature data sample l i and l k (obtained from the semantic similarity matrix S). It serves as a weight, measuring the importance of the sample l k to the sample l i in the fusion process. The higher the semantic similarity, the stronger the correlation between the two. When performing weighted expansion, the vector v k of l k (which is also the comprehensive vector representation of l k , obtained in the aggregation manner as described above) for li The greater the contribution. By weighting and summing the vectors corresponding to all log feature data samples according to their semantic similarity with l i a weighted extended vector is obtained In this way, the vector representation of each sample incorporates the information of samples semantically related to it, further strengthening the manifestation of semantic association at the vector level. From an execution perspective, for each log feature data sample, the weighted extension operation should be performed according to the above formula. The code example is as follows:
[0348]
[0349]
[0350] Let the finally generated log feature vector sample be It is obtained by fusing all the weighted extended vectors. A common fusion method is average fusion, and the calculation formula is as follows:
[0351] That is, all the weighted extended vectors are averaged with equal weights to obtain a vector representation that integrates the semantic related information of all log feature data samples, as the final log feature vector sample, which can more comprehensively reflect the semantic content contained in the log and potential security-related features, etc.
[0352] From an execution perspective, calculate the final log feature vector sample according to the above fusion formula. The code example is as follows:
[0353] V_final = [0 for _ in range(len(L[0].vector_list))]
[0354] for l in range(len(V_final)):
[0355] V_final[l] = sum([V_star[i][l] for i in range(len(L))]) / len(L)
[0356] Therefore, by pre-training a semantic embedding model, keywords and key phrases in the logs are mapped to a low-dimensional vector space, enabling the originally scattered and difficult-to-directly-measure text information with semantic associations to be intuitively represented in the vector space, that is, words with similar semantics are close in position. The semantic similarity matrix calculated on this basis can accurately quantify the degree of semantic association between different log feature data samples, helping to uncover the internal connections between log contents that have different surface expressions but are actually semantically related. For example, logs recorded by different modules about the same system operation but with slightly different wordings can be better associated through the vector space and similarity calculation, providing a more comprehensive perspective for subsequent analysis of the overall operating state and security situation of the system. Additionally, in the process of weighted expansion and fusion of log feature data samples based on the semantic similarity matrix to generate the final vector samples, the semantic association information is fully utilized. Instead of simply piling up log features, the information of each sample is reasonably fused according to the degree of semantic relevance, so that the final log feature vector samples gather the key information of multiple semantically related samples, can more deeply reflect the system behavior represented by the logs and the security-related clues hidden therein, and avoid information omission and misjudgment that may occur when analyzing individual log samples in isolation. Moreover, converting log feature data samples into vector form provides a unified and standardized data representation method, which is more convenient for processing in subsequent various data analysis, machine learning, and deep learning algorithms compared to the original text form. Vectors can easily participate in various mathematical operations, feature extraction, and serve as the input and output of models, improving the efficiency and flexibility of data processing. For example, log feature vector samples can be directly input into a classification model to determine whether there is a network security threat, or used in a clustering model to classify different types of log behavior patterns. Finally, since the log feature vector samples integrate rich semantic-related information and contain more representative and informative features themselves, when used as input for network security-related model training and prediction, they can help the model better capture key information such as abnormal behavior patterns and potential security threats reflected in the logs, thereby improving the accuracy and reliability of the model, enhancing the performance of the entire network security threat monitoring system based on log data, and more effectively ensuring the secure operation of the network system.
[0357] Optionally, when the data vectorization module obtains the traffic feature vector sample, the log feature vector sample, and the behavior feature vector sample from the traffic feature data sample, the log feature data sample, and the behavior feature data sample respectively, the steps included are as follows: For the behavior feature data sample, considering that user behavior has two important dimensions of time and space, a spatio-temporal fusion vectorization technology innovation is adopted. Feature encoding is respectively performed on the time dimension and the space dimension of the behavior feature data sample to obtain a time vector and a space vector. The time dimension includes the moment, duration, and frequency of the behavior occurrence, and the space dimension includes the geographical location where the behavior occurs, the application system or device where it is located; based on the tensor fusion mechanism, the time vector and the space vector are fused to generate the behavior feature vector sample.
[0358] Specifically, the detailed description of the above vectorization steps for the behavior feature data sample is as follows:
[0359] I. Feature encoding for the time dimension and the space dimension of the behavior feature data sample
[0360] 1. Feature encoding for the time dimension
[0361] Let the set of behavior feature data samples be B = {b 1 , b 2 , …, b n}, where each b i represents a behavior feature data sample, which contains multiple time dimension attributes related to user behavior, such as the moment (which can be represented by a timestamp, for example, a certain time point accurate to the second), duration (the unit can be seconds, minutes, etc., representing the time length of a behavior), and frequency (the number of times this behavior occurs within a certain time range, such as the number of operations per hour). To encode these features of the time dimension and convert them into a time vector representation. Using a linear encoding method (which can actually be adjusted according to the specific encoding algorithm), let the encoded time vector be (d t is the set time vector dimension, for example, d t = 3), and the elements of its each dimension can be calculated in the following way:
[0362]
[0363] Where: are respectively the minimum value and the maximum value of the behavior occurrence moment in all behavior feature data samples. By taking the behavior occurrence moment of the specific sample Subtract the minimum value and divide by the range difference (maximum value minus minimum value) for normalization, mapping it to the interval $[0,1]$. In this way, the characteristic of the behavior occurrence time of different samples can be compared and subsequent operations can be carried out on a unified scale. That is the corresponding element in the time vector of the behavior occurrence time after normalization. Similarly, are the minimum and maximum values of the behavior duration among all samples, is the element in the time vector corresponding to the normalized behavior duration of the corresponding sample; are the minimum and maximum values of the behavior frequency, is the vector element after normalizing the behavior frequency.
[0364] From the perspective of execution, it is necessary to traverse each behavior feature data sample, obtain the attribute values of its time dimension, and then perform normalization encoding according to the above formula to generate a time vector. The code example is as follows:
[0365]
[0366] 2. Spatial Dimension Feature Encoding
[0367] For each behavior feature data sample b i , its spatial dimension includes the geographical location where the behavior occurs (which can be represented by longitude and latitude coordinates, denoted as and representing longitude and latitude respectively), and the application system or device where it is located (which can be represented by discrete category numbers, denoted as s i , for example, 1 represents the mobile application system, 2 represents the web application system, etc.). Similarly, encode the spatial dimension features. Let the encoded spatial vector be (d s is the set dimension of the spatial vector, for example, d s = 3). The calculation methods of the elements of each dimension are as follows:
[0368]
[0369] Among them: are respectively the minimum and maximum values of the longitude of the geographical location where the behavior occurs among all behavior feature data samples, is the corresponding element in the spatial vector after normalizing the longitude coordinate of sample b i , so that the position feature can be reflected on a unified scale. is the minimum and maximum values of the latitude, is the vector element after normalizing the corresponding sample latitude. s min , s max are the minimum and maximum values of the category numbers representing the application system or device among all samples, It is an element in the spatial vector after normalizing the device or system category number corresponding to the sample, which facilitates subsequent comprehensive consideration of various characteristics in the spatial dimension.
[0370] From an execution perspective, it is necessary to traverse each behavioral feature data sample, obtain the spatial dimension attribute values, and encode them into a spatial vector according to the above formula. The code example is as follows:
[0371]
[0372] II. Generating Behavioral Feature Vector Samples by Fusing Time Vectors and Spatial Vectors Based on the Tensor Fusion Mechanism
[0373] Let the behavioral feature vector sample generated by fusing through the tensor fusion mechanism be (d is the dimension of the final behavioral feature vector sample. For example, d = 6, which can be the sum of the time vector dimension and the spatial vector dimension, etc., and is specifically determined according to the fusion method).
[0374] The tensor fusion method is to directly splice the time vector and the spatial vector. The calculation formula is as follows:
[0375] That is, the time vector and the spatial vector are spliced together in sequence to form a new vector as the behavioral feature vector sample. For example, if the time vector dimension d t = 3 and the spatial vector dimension d s = 3, then the fused behavioral feature vector sample Of course, other tensor fusion mechanisms can also be used, such as based on operations such as the tensor product. Using tensor product fusion, the calculation formula is as follows:
[0376] Among them represents the tensor product operation. Through this operation, the interaction information between each dimension of the time vector and the spatial vector can be mined to generate a new tensor with a higher dimension, and then it is transformed into the final behavioral feature vector sample through appropriate dimensionality reduction and other processing methods, so that the fused vector not only contains the respective characteristics of time and space, but also reflects the correlation relationship between them.
[0377] From an execution perspective, if it is splicing fusion, only the corresponding time vector and spatial vector need to be combined in sequence; if it is a complex method such as tensor product fusion, it needs to be operated according to the corresponding mathematical operation rules. The code example is as follows:
[0378] for i in range(len(B)):
[0379] v_i_b = B[i].time_vector + B[i].space_vector
[0380] B[i].behavior_vector = v_i_b
[0381] Therefore, by considering that user behavior has two important dimensions of time and space and encoding the features of these two dimensions respectively, the characteristics of user behavior can be comprehensively and meticulously captured. For example, information such as the occurrence time, duration, and frequency of behavior in the time dimension can reflect the user's operation habits, activity level, and the periodic patterns of behavior, etc.; the geographical location and the information of the application system or device in the space dimension can reflect the environment where the user behavior occurs, the platform used, etc. By fusing the feature vectors of these two dimensions, the finally generated sample of the behavior feature vector can completely present the full picture of user behavior, avoiding the problem of losing key information by only focusing on a single dimension, and providing a rich data basis for more accurately analyzing user behavior patterns and judging whether there are abnormalities in the follow-up. In addition, especially when adopting complex tensor fusion mechanisms such as the tensor product, not only are the time and space features simply combined together, but also the interaction correlation information between the two dimensions can be mined. For example, some users always perform specific operations at specific times (time dimension) when in a specific geographical location (space dimension). This cross-dimensional correlation may imply specific behavior patterns or potential security risks. By integrating this correlation information into the sample of the behavior feature vector during the fusion process, it helps to further deeply analyze the hidden rules and possible security hazards behind user behavior. Moreover, converting the behavior feature data sample into a unified vector representation form, that is, the sample of the behavior feature vector, this standardized data representation is convenient for subsequent processing in various machine learning and deep learning models. Whether it is for behavior classification (such as distinguishing normal behavior and abnormal behavior), clustering (grouping similar behavior patterns), or other data analysis tasks, vector-form data is more likely to participate in mathematical operations and serve as the input and output of the model. Compared with the original scattered form containing time and space attributes, it improves the efficiency and versatility of data processing and is convenient for adapting and integrating with various algorithms and models. Finally, due to the fact that the sample of the behavior feature vector synthesizes the multi-dimensions of time and space and integrates the correlated information features, the information it contains is richer and more representative. When applying it as an input to a model related to network security (such as monitoring whether there are malicious behaviors, intrusion behaviors, etc. based on behavior features), it can help the model better capture the complex changes in user behavior patterns and abnormal situations, thereby improving the accuracy and reliability of the model, more effectively realizing the network security guarantee function based on user behavior analysis, and enhancing the monitoring and response capabilities for security threats related to user behavior in the network system.
[0382] Optionally, when the semantic sequence construction module performs cross-semantic relationship analysis on the traffic feature vector samples, log feature vector samples, and behavior feature vector samples to obtain a cross-semantic relationship graph and annotates the network security threat labels on the cross-semantic relationship graph to form a semantic training dataset, the steps included are as follows: Perform effective cross-semantic relationship analysis on the traffic feature vector samples, log feature vector samples, and behavior feature vector samples to construct a semantic alignment network; According to the semantic alignment network, based on the fusion of graph attention mechanisms, construct a cross-semantic relationship fusion graph containing all traffic, log, and behavior feature vector samples, where the nodes are different feature vector samples and the edges represent potential semantic associations between different modality feature vector samples; Based on a predefined network security threat rule knowledge base, a case knowledge graph based on historical network security events, and real-time updated external professional threat intelligence, perform cross-semantic relationship pattern reasoning on the cross-semantic relationship fusion graph to annotate the network security threat labels and form a semantic training dataset accordingly.
[0383] Specifically, the detailed description of the above solution is as follows:
[0384] I. Construct a semantic alignment network for cross-semantic relationship analysis
[0385] 1. Construction of the semantic alignment network
[0386] Let the set of traffic feature vector samples be The set of log feature vector samples be The set of behavior feature vector samples be To construct a semantic alignment network to analyze the cross-semantic relationships between them, this network can be regarded as a function G align (F, L, B), whose goal is to find the semantic correspondence relationships between different modality (traffic, log, behavior) feature vector samples. A common approach is to measure the semantic association degree by calculating the similarity between vectors. For example, for the traffic feature vector sample f i and the log feature vector sample l j , the cosine similarity can be used to calculate their semantic similarity. Let the similarity be The calculation formula is as follows:
[0387]
[0388] Where: f i ·l j Represents the inner product of vectors f i and l j , which reflects the degree of consistency in the direction of the two vectors. The larger the inner product, the closer the vectors are in the vector space. From a semantic perspective, it means that the two may have a stronger semantic association. ‖fi ‖ and ‖l j ‖ are the norms (lengths) of vectors f i and l j respectively. The normalization operation is performed by dividing the product of the vector norms, so that the similarity ranges from $[0,1]$, which is convenient for comparing and measuring the semantic similarity between different sample pairs. Similarly, for the similarity between the traffic feature vector sample and the behavior feature vector sample (denoted as r) and the similarity between the log feature vector sample and the behavior feature vector sample (denoted as s), they are also calculated according to the above cosine similarity formula. Then, according to the set similarity threshold θ (for example, θ = 0.5, empirically set according to the actual data and application scenarios), if then it is considered that there is a potential semantic association between the traffic feature vector sample f i and the log feature vector sample l j . By analogy, the association relationships between different modality samples are constructed, and these association relationships form a semantic alignment network. From an execution perspective, it is necessary to traverse the collection of feature vector samples of different modalities, calculate the similarity between them pairwise, and determine the association relationship according to the threshold. The code example is as follows:
[0389]
[0390] 2. The role of the semantic alignment network
[0391] The semantic alignment network can help initially discover the potential semantic connections between different modality feature vector samples. For example, in the network security scenario, certain specific traffic feature vector samples may be semantically related to the log feature vector samples reflecting specific abnormal operations and the corresponding abnormal behavior feature vector samples. These associations can be discovered through this network, providing a basis for subsequent deeper fusion and analysis, avoiding looking at different modality data in isolation, and helping to integrate multi-source information to more comprehensively grasp the network security situation.
[0392] II. Constructing a cross-semantic relationship fusion graph based on the graph attention mechanism
[0393] 1. The principle and application of the graph attention mechanism
[0394] After constructing the semantic alignment network, a cross-semantic relationship fusion graph is constructed based on the graph attention mechanism (Graph Attention Network, GAT). Let the cross-semantic relationship fusion graph be G fusion=(V, E), where the node set V = F ∪ L ∪ B, that is, it contains all traffic, log, and behavioral feature vector samples. The edge set E represents the potential semantic associations between different modal feature vector samples (associations mined by the previous semantic alignment network and further strengthened and refined by the graph attention mechanism). The core of the graph attention mechanism is to calculate the attention weights of each node in the graph with its neighbor nodes, so as to aggregate the information of neighbor nodes and update its own node representation. For node v i ∈V (here v i can be a certain traffic, log, or behavioral feature vector sample), its update formula in the graph attention layer is as follows (taking one layer of the graph attention mechanism as an example, actually multiple layers can be stacked):
[0395] e ij = LeakyReLU(a T [Wv i ||Wv j )
[0396]
[0397]
[0398] where: e ij represents the intermediate calculation result of the attention coefficient of node v i to neighbor node v j . It is obtained by first passing the vector representations of nodes v i and v j through a shared linear transformation (through the weight matrix d in is the input vector dimension, d out is the dimension after transformation), then concatenating them (using || to represent the concatenation operation), and then passing through a LeakyReLU activation function. This activation function can introduce non-linear factors to help learn more complex attention relationships, is a learnable weight vector used to further adjust the calculation of the attention coefficient. α ij is the finally calculated attention weight of node v i to neighbor node v j . It is obtained by exponentiating and normalizing the intermediate result e ij . Here represents the set of neighbor nodes of node v i . Through such a normalization operation, the sum of the attention weights of node v i to all neighbor nodes is 1, so as to reasonably aggregate the information of neighbor nodes. is node v iThe vector representation updated by the graph attention mechanism, which is obtained by weighted summing the vector representations v of neighbor nodes according to the corresponding attention weights α and performing a non-linear transformation through an activation function σ (such as the ReLU function, etc.). In this way, the node integrates the relevant information of neighbor nodes and focuses on the information of neighbor nodes with closer semantic association with itself according to the attention weights, strengthening the semantic fusion effect between nodes in the cross-semantic relationship fusion graph. j According to the corresponding attention weight α ij It is obtained by weighted summation and undergoes a non-linear transformation through an activation function σ (such as the ReLU function, etc.). In this way, the node integrates the relevant information of neighbor nodes and focuses on the information of neighbor nodes with closer semantic association with itself according to the attention weights, strengthening the semantic fusion effect between nodes in the cross-semantic relationship fusion graph.
[0399] From an execution perspective, it is necessary to perform information update operations on each node in the cross-semantic relationship fusion graph according to the above calculation steps of the graph attention mechanism. The code example is as follows:
[0400]
[0401] 2. Significance of the cross-semantic relationship fusion graph
[0402] The cross-semantic relationship fusion graph constructed by the graph attention mechanism further integrates the semantic information between different modal feature vector samples, enabling each node (feature vector sample) to not only contain its own original information but also incorporate the information of other modal samples semantically related to it. And this fusion is dynamically adjusted based on attention weights, being more able to focus on the associated information that is important for the semantic understanding of the current node. In the network security analysis scenario, for example, it can enable traffic feature vector samples to better combine relevant semantic information in logs and behaviors, jointly reflecting the network security status and providing a richer, more comprehensive, and fused semantic basis for accurately labeling network security threat labels subsequently.
[0403] III. Conduct cross-semantic relationship pattern reasoning based on relevant knowledge and label network security threat labels to form a semantic training dataset
[0404] 1. Cross-semantic relationship pattern reasoning and label annotation
[0405] Based on a pre-defined network security threat rule knowledge base (denoted as K rules ), a case knowledge graph based on historical network security events (denoted as K cases ), and real-time updated external professional threat intelligence (denoted as I threat ), perform cross-semantic relationship pattern reasoning on the cross-semantic relationship fusion graph to label network security threat labels. Let the network security threat label set be For example, t 1 can represent the "DDoS attack" label, t 2It can represent labels such as "malware intrusion". For each node (feature vector sample) in the cross-semantic relationship fusion graph, it is determined whether there is a corresponding cybersecurity threat by looking up matching patterns in the knowledge base, knowledge graph, and threat intelligence, and the corresponding labels are marked. For example, for node v i (which is a node after fusing a traffic feature vector sample with log and behavior feature vector samples), a matching function M(v i ,K rules ,K cases ,I threat ) can be defined, and its return value is a set of labels representing the possible cybersecurity threat labels matching this node. The specific matching process can be a comprehensive application of various methods such as rule matching, graph structure matching, and semantic similarity matching (the matching logic here is customized according to the specific structures of the knowledge base, knowledge graph, and threat intelligence and the application scenario). For example, if there is a rule in the cybersecurity threat rule knowledge base: "When there is a traffic burst in the traffic feature vector, a specific error record appears in the log feature vector, and the behavior feature vector shows abnormal operations, it is determined as a DDoS attack", when the fusion feature vector corresponding to node v i meets the conditions of this rule, M(v i ,K rules ,K cases ,I threat ) will return a set containing the label "DDoS attack".
[0406] From an execution perspective, it is necessary to traverse all nodes in the cross-semantic relationship fusion graph, call the matching function for pattern matching, and mark the corresponding cybersecurity threat labels. The code example is as follows:
[0407]
[0408] 2. Form a semantic training dataset
[0409] After the above process of marking cybersecurity threat labels, the entire marked cross-semantic relationship fusion graph constitutes a semantic training dataset. This dataset can be represented as D = {(v i ,t i )|v i ∈ V, t i ∈ v i .threat lIt is represented by { abels}, where each element is a pair consisting of a feature vector sample and a corresponding network security threat label. It can be used for the training of subsequent machine learning and deep learning models (such as classification models), enabling the model to learn the relationship between the feature vectors after different semantic fusions and network security threats, and then being able to accurately judge the network security threats of new unlabeled data.
[0410] Therefore, by constructing a semantic alignment network and a cross-semantic relationship fusion graph based on the graph attention mechanism, the semantic information between the three different modal feature vector samples of traffic, logs, and behaviors can be fully fused. Instead of looking at the data of each modality in isolation, the potential semantic associations between them are mined and effectively fused, so that the final feature vector sample contains more comprehensive and rich semantic content, and can jointly reflect the operating state of the network system and potential network security threat situations from multiple perspectives, avoiding the problem of information one-sidedness caused by relying only on single-modal data and improving the accuracy of network security situation awareness. In addition, the graph attention mechanism dynamically assigns attention weights according to the semantic associations between nodes during the fusion process, enabling each feature vector sample to focus on the part closely related to its own semantics when fusing other modal information. This can more precisely strengthen important cross-semantic associations, filter out some relatively unimportant noise information, and make the fused semantic information more targeted and effective, helping to more clearly capture the key semantic patterns related to network security threats in a complex network environment. Moreover, by using a pre-defined network security threat rule knowledge base, a case knowledge graph of historical network security events, and real-time updated external professional threat intelligence to perform cross-semantic relationship pattern reasoning and annotate network security threat labels, the domain expert knowledge, historical experience, and the latest security intelligence information are fully combined, making the annotated labels more in line with the actual network security situation, reducing the possibility of misjudgment and missed judgment, improving the accuracy of network security threat recognition, and providing a more reliable basis for subsequent network security management and response measures. Finally, the formed semantic training dataset integrates the feature vectors after multi-modal fusion and accurately annotated network security threat labels, providing high-quality training data for machine learning and deep learning models. The model is trained based on such rich and representative data, can better learn the complex relationship between different semantic features and network security threats, improve the generalization ability of the model and the prediction ability for unknown network security threats, and thus more effectively monitor and prevent network security threats in practical applications, ensuring the stable operation of the network system.
[0411] Optionally, the deep learning model includes a multi-branch fusion neural network structure, including a traffic feature processing branch, a log feature processing branch, a behavior feature processing branch, and a cross-modal fusion layer; the traffic feature processing branch includes a convolutional neural network and an attention layer, the log feature processing branch includes a recurrent neural network and a semantic embedding layer, and the behavior feature processing branch includes a graph neural network; when the model training module calls the semantic training data set to train the deep learning model to be trained, so that the deep learning model learns the relationship between the first network traffic data sample, the first log data sample, the first user behavior data sample and the corresponding network security threat label, and maps the relationship to the network parameters of the deep learning model to obtain a multi-level dynamic threat monitoring model, the steps included are as follows: Extract the local features of the traffic feature vector sample based on the convolutional neural network, and dynamically allocate weights based on the importance of different local features in network security threat judgment based on the attention layer to generate key traffic feature samples; Extract the time series characteristics in the log feature vector sample based on the recurrent neural network to capture the association between log events at different time points and obtain time series features; Map the time series features to a low-dimensional vector space based on the semantic embedding layer, so that log features with similar semantics are close in the vector space to obtain log semantic sequence feature samples; Based on the graph neural network, use user behavior as the nodes of the graph and the association between behaviors as the edges to aggregate neighbor node information to generate behavior propagation features; Based on the cross-modal fusion layer, fuse the key traffic feature samples, log semantic sequence feature samples, and behavior propagation features along different fusion paths to obtain multiple multi-modal fusion feature samples; Based on the bilinear pooling layer, calculate the second-order interaction information between the key traffic feature samples, log semantic sequence feature samples, and behavior propagation features to strengthen the multi-modal fusion feature samples; Input the multi-modal fusion feature samples after feature strengthening into a fully connected layer to extract the local information samples about network security threats contained in each dimension of the different modal fusion feature samples, and perform a non-linear transformation on the local information samples to predict the probability distribution of network security threat types; Based on the cross-entropy loss function, calculate the difference degree between the predicted probability distribution of network security threat categories and the network security threat label, and adjust the network parameters of the deep learning model through backpropagation to learn the relationship between the first network traffic data sample, the first log data sample, the first user behavior data sample and the corresponding network security threat label until a multi-level dynamic threat monitoring model is obtained.
[0412] Specifically, the detailed description of the above deep learning model training scheme is as follows:
[0413] I. Operations of the traffic feature processing branch
[0414] 1. Local feature extraction by convolutional neural network
[0415] Let the set of traffic feature vector samples be F = {f 1 , f 2 , …, f n}, where each traffic feature vector sample (d f represents the dimension of the traffic feature vector sample).
[0416] For a convolutional neural network (CNN), there are m convolutional kernels, each with a size of k×k (taking the two-dimensional case as an example here, if it is a one-dimensional feature vector, then k is the length), a stride of s, and a padding of p. The convolutional kernel weight matrix is represented as (j = 1, 2, …, m,[[]] represents the number of channels of the feature map output by the j-th convolutional kernel, that is, the feature dimension obtained after convolution by this convolutional kernel), and the bias term is For the sample f i , the local feature value obtained at the position (x, y) after convolution by the j-th convolutional kernel is calculated as follows:
[0417]
[0418] where: (x, y) is the coordinate position on the feature map. By sliding on the traffic feature vector sample according to the size, stride, etc. of the convolutional kernel, multiplying the corresponding position elements by the convolutional kernel weights and accumulating, and then adding the bias term, the local feature value at this position is obtained. Performing the above operations on the entire area corresponding to the sample yields the feature map (h and w are the height and width of the feature map calculated according to the input size, convolutional kernel size, stride, and padding).
[0419] From an execution perspective, for each traffic feature vector sample, each convolutional kernel needs to be traversed, and the local feature value is calculated at the corresponding position according to the above formula. The code example is as follows:
[0420]
[0421] 2. Generation of key traffic feature samples by the attention layer
[0422] Let the set of local feature maps obtained after convolution be (corresponding to the m feature maps of the traffic feature vector sample f i ). Define a learnable query vector (d q is related to the convolutional output feature dimension and is usually made to fit the dimension through a linear transformation), and calculate the attention weight To measure the importance of each local feature map in network security threat judgment, the calculation formula is as follows:
[0423]
[0424] Where: It is obtained by performing a dot product operation between the query vector and the feature map vector, reflecting the degree of association between this feature map and the focus of attention represented by the query vector (related to network security threat judgment). The larger the value, the stronger the association. Then, for all the Perform exponential operation and normalization to obtain the attention weight Make the sum of the weights corresponding to all feature maps equal to 1, so as to highlight the important local features. The key traffic feature sample Is obtained by weighted summation of the local feature maps according to the attention weights. The formula is:
[0425]
[0426] From an execution perspective, first calculate the attention weights, and then perform weighted summation to generate the key traffic feature sample. The code example is as follows:
[0427]
[0428] II. Log Feature Processing Branch Operations
[0429] 1. Recurrent neural network extracts time series characteristics
[0430] Let the log feature vector sample set be L = {l 1 , l 2 , …, l n}, where each (d l Is the dimension of the log feature vector sample), and they are arranged in chronological order.
[0431] Taking the long short-term memory network (LSTM) as an example, for the processing of the log feature vector sample l i At time step t, the calculation formula of the LSTM cell is as follows (simplified core part):
[0432]
[0433] Where: l it Is the input vector of the log feature vector sample l i At time step t (corresponding to the elements in the log feature vector), h t-1 Is the hidden state vector of the previous time step (the initial hidden state h 0 Can be randomly initialized), c t-1 Is the cell state vector of the previous time step (the initial cell state c0 can be reasonably initialized). and are weight matrices corresponding to different gating structures (input gate, forget gate, output gate, candidate memory cell). Through these weights, the input vector and the hidden state vector of the previous time step are linearly transformed, and then combined with activation functions, etc. to update each state. is the corresponding bias vector, used to fine-tune the state update. σ(·) is the sigmoid activation function (compressing the value to the interval [0,1]), ⊙ represents element-wise multiplication, and tanh(·) is the hyperbolic tangent activation function (mapping the value to the interval [-1,1]). Through these operations, the LSTM can capture the associations between log events at different time points. After processing the entire log feature vector sample, the final hidden state vector h T (T is the length of the sample time series) is the extracted time series feature, denoted as
[0434] From an execution perspective, according to the above LSTM unit update rules, each time step needs to be processed in sequence. The code example is as follows:
[0435]
[0436] 2. The semantic embedding layer generates log semantic sequence feature samples
[0437] Let the time series feature obtained through the recurrent neural network (corresponding to the time series feature vector of the log feature vector sample l i . The semantic embedding layer has a weight matrix (d embed is the dimension of the low-dimensional vector space) and a bias vector to map the time series feature to the low-dimensional vector space to obtain the log semantic sequence feature sample The calculation formula is:
[0438]
[0439] Through this linear transformation, based on the semantic mapping relationship learned by the semantic embedding layer, logs with similar semantics are close in the vector space, facilitating subsequent fusion processing.
[0440] From an execution perspective, perform the mapping operation according to the above formula. The code example is as follows:
[0441] for i in range(len(L)):
[0442] L[i].semantic_sequence_feature = torch.mm(torch.tensor(W_embed), L[i].time_sequence_feature) + torch.tensor(b_embed)
[0443] III. Behavioral Feature Processing Branch Operations
[0444] Let the set of behavioral feature vector samples be B = {b 1 , b 2 , …, b n}. Using user behaviors as the nodes of the graph and the associations between behaviors as the edges, construct a graph G = (V, E) (where V = B and E represents the set of edges reflecting behavioral associations). Taking the graph convolutional neural network (GCN) as an example, let the node feature matrix of the l-th layer of the graph convolutional neural network be (When initially l = 0, H (0) is the behavioral feature vector sample matrix, with dimension d (0) , and after convolution through each layer, the dimension becomes d (l) ).
[0445] The graph convolution operation formula is as follows:
[0446]
[0447] Where: is the adjacency matrix of graph G. If there is an edge connecting node i and node j (i.e., the behaviors represented by the behavioral feature vector samples b i and b j are related), then A ij = 1; otherwise, A ij = 0. (I is the identity matrix) considers node self-loops. is the degree matrix, and the diagonal element D ii is the degree of node i. The diagonal elements of facilitate subsequent calculations. is the learnable weight matrix of the l-th layer, which determines the feature mapping and can focus on mining the behavioral propagation features related to network security threats by adjustment. σ(·) is the activation function (such as the ReLU function) that introduces non-linearity to help learn complex feature relationships. After L layers of graph convolution operations, each row vector of the final feature matrix H (L) is used as the behavioral propagation feature of the corresponding behavioral feature vector sample. For example, the behavioral propagation feature of the behavioral feature vector sample b i is Update the node feature matrix by performing matrix multiplication, normalization, etc. operations layer by layer according to the above graph convolution operation formula. The code example is as follows:
[0448]
[0449] IV. Cross-modal Fusion Layer Operations
[0450] Let the set of key traffic feature samples be the set of log semantic sequence feature samples be the set of behavior propagation feature samples be
[0451] The cross-modal fusion layer fuses these features along different fusion paths. For example, a splicing fusion method (just an example, it can be more complex in reality). For the multi-modal fusion feature sample m corresponding to the i-th sample i , the calculation formula is as follows:
[0452]
[0453] Here, [;] means concatenating different feature vectors in order to form a new fused feature vector. By different fusion paths (such as different concatenation orders, adding weights, etc.), multiple multi-modal fusion feature samples can be obtained, which fuse information of different modalities, mine the correlations between them, and provide more comprehensive data for subsequent analysis.
[0454] V. Bilinear Pooling Layer Operations
[0455] Let the multi-modal fusion feature sample be (d m is the dimension of the fused feature), and the enhanced feature sample is obtained through bilinear pooling. Bilinear pooling calculates the second-order interaction information between features of different modalities to capture more complex correlation relationships between features and thus enhance the feature representation. First, the key traffic feature sample the log semantic sequence feature sample the behavior propagation feature are subjected to an outer product operation to obtain a third-order tensor T i , and the calculation formula for its elements is as follows:
[0456]
[0457] Among them, that is, they respectively correspond to the dimension indices of the three feature vectors. Through such element-wise multiplication, a third-order tensor containing second-order interaction information is obtained. However, directly using this third-order tensor has too high a dimension and complex subsequent calculations, so dimensionality reduction is usually performed. The dimensionality reduction method is to flatten this third-order tensor into a vector, and then through a linear transformation (such as multiplying by a learnable weight matrix dpool to the set target dimension after dimensionality reduction and add a bias term Obtain the enhanced feature samples, and the calculation formula is as follows:
[0458]
[0459] where vec(T i ) represents the operation of flattening the third-order tensor T i into a vector, that is, arranging all elements in the tensor in a certain order into a one-dimensional vector, so that the second-order interaction information can be fused through a linear transformation and mapped to a suitable dimensional space to achieve the enhancement of the multi-modal fusion feature samples.
[0460] From an execution perspective, it is necessary to first calculate the third-order tensor according to the outer product rule, then perform the flattening operation, and then multiply by the weight matrix and add the bias term to obtain the enhanced feature samples. The code example is as follows:
[0461]
[0462]
[0463] VI. Fully connected layer operation
[0464] Input the multi-modal fusion feature samples after feature enhancement into the fully connected layer to extract the local information samples about network security threats contained in each dimension of the different-modal fusion feature samples, and perform a non-linear transformation on the local information samples to predict the probability distribution of network security threat types. Let the set of enhanced multi-modal fusion feature samples be The fully connected layer has a learnable weight matrix (d in is the dimension of the input enhanced feature, that is, d pool , d out is the output dimension of the fully connected layer, which is set according to the number of network security threat types to be predicted, etc.) and a bias vector
[0465] For the input enhanced feature sample The output vector o i after passing through the fully connected layer (representing the score vector for predicting network security threat types, and each element corresponds to the score situation of a threat type) is calculated as follows:
[0466]
[0467] Then, in order to convert the score vector into a probability distribution, the softmax function is usually used for non-linear transformation to obtain the probability distribution vector of network security threat types The calculation formula is as follows:
[0468]
[0469] Among them, o ij is the j-th element of the output vector o i The elements in the score vector are converted into probability values in the interval $[0,1]$ through the softmax function, and the sum of all probability values is 1. In this way, the possibility of each type of network security threat occurrence can be judged according to the probability size.
[0470] From the execution perspective, linear operations and softmax transformation operations need to be performed according to the above formula. The code example is as follows:
[0471]
[0472] VII. Adjusting Model Parameters Based on the Cross-Entropy Loss Function
[0473] Based on the cross-entropy loss function, calculate the difference degree between the predicted probability distribution of network security threat categories and the network security threat labels, so as to adjust the network parameters of the deep learning model through backpropagation, and learn the relationship between the first network traffic data sample, the first log data sample, the first user behavior data sample and the corresponding network security threat labels until a multi-level dynamic threat monitoring model is obtained.
[0474] Let the predicted probability distribution vector of network security threat categories be (the prediction result corresponding to sample i), and the true network security threat label is represented by one-hot encoding as (for example, if there are k types of threat types, and the true label corresponding to sample i is the m-th threat type, then the m-th element in y i is 1, and the rest of the elements are 0).
[0475] The calculation formula of the cross-entropy loss function is as follows:
[0476] Among them, n is the number of samples. This loss function measures the difference degree between the predicted probability distribution and the true label. The smaller the loss value, the closer the prediction result is to the real situation. During the training process, based on this loss value, the backpropagation algorithm is used to adjust the network parameters of the deep learning model (such as the convolution kernel weights in the convolutional neural network, the weight matrix of the fully connected layer, etc.). The backpropagation algorithm updates the parameters according to the gradients of the loss function with respect to each parameter. For example, for the parameter θ (which can be a general representation of various weight matrices or bias vectors mentioned above), the update formula is as follows:
[0477]
[0478] Among them, η is the learning rate, which is a pre-set positive number used to control the step size of parameter update. By continuously calculating the loss on the training data set and performing backpropagation to update the parameters, the model is continuously optimized, gradually learning the relationship between the input network traffic, logs, user behavior data samples and the corresponding network security threat labels. Until the model achieves good performance on evaluation metrics such as the validation set, a multi-level dynamic threat monitoring model is obtained. From an execution perspective, it is necessary to first calculate the cross-entropy loss of each sample, then calculate the gradients of each parameter according to the backpropagation algorithm, and adjust the parameters according to the update formula. This process will be iterated multiple times on the training data set. The code example is as follows:
[0479]
[0480] To this end, by separately processing traffic, logs, and behavioral characteristics through different branches, the key information contained within each modality of data can be fully explored. For example, the traffic feature processing branch utilizes a convolutional neural network and an attention layer to not only extract the local features of traffic feature vector samples but also, through the attention mechanism, focus on the local features important for network security threat judgment, enabling the model to pay more attention to valuable traffic-related clues; the log feature processing branch uses a recurrent neural network to capture the characteristics of logs in the time series and then strengthens semantic associations through a semantic embedding layer, enabling a better exploration of the evolution of system state changes and potential security issues reflected in the logs over time; the behavioral feature processing branch uses a graph neural network to aggregate the correlation information between behaviors to generate behavioral propagation features, which helps to understand the propagation patterns and abnormal manifestations of user behaviors in the network. Additionally, the design of the cross-modal fusion layer and the bilinear pooling layer enables the organic fusion and enhancement of features from different modalities. The cross-modal fusion layer integrates the key features of each modality through different fusion paths, achieving the convergence of multi-source information and avoiding viewing each modality of data in isolation; the bilinear pooling layer further explores the second-order interaction information between the features of each modality, enabling the model to capture more complex and subtle cross-modal associations, such as the possible security threat situations implied when traffic features appear simultaneously with specific log semantic features and behavioral propagation features, thereby making the fused features contain richer and more comprehensive information about the network security situation and enhancing the model's understanding and analysis capabilities for complex network security scenarios. Moreover, using the cross-entropy loss function in combination with the backpropagation algorithm for model training enables the model to accurately learn the relationships between network traffic, log, user behavior data samples, and the corresponding network security threat labels. By continuously adjusting the network parameters, optimizing the matching degree between the predicted probability distribution of network security threat types and the true labels, reducing the prediction error, the model's recognition accuracy and reliability for different network security threat types are improved, enabling it to more effectively monitor security risks in the network and issue early warnings in practical applications. Finally, the entire deep learning model structure and training process ultimately result in a multi-level dynamic threat monitoring model that can dynamically analyze the network security situation based on the characteristics and associations reflected by different modality data at different levels. Unlike traditional single-dimensional or simple models that can only make static and one-sided judgments, it can comprehensively consider multiple factors and update the assessment of security threats in real time as the network state changes, better adapting to the complex and ever-changing network environment and providing stronger support for network security protection.
[0481] Optionally, when the model deployment module deploys the multi-level dynamic threat monitoring model on the cloud server, so that the multi-level dynamic threat monitoring model extracts and processes the operation data features of the target network based on its network parameters to predict the network security threats occurring in the target network, the steps included are as follows: Input the preprocessed operation data of the target network into the convolutional neural network in the traffic feature processing branch to extract the local features of the network traffic data, and the attention layer dynamically assigns weights according to the current network security threat judgment situation to generate key traffic features; The recurrent neural network extracts the time series characteristics of the log data in the operation data in the log feature processing branch and captures event associations, and the semantic embedding layer maps it into log semantic sequence features; The graph neural network aggregates the neighbor node information with user behavior as nodes and behavior associations as edges for the behavior data in the operation data in the behavior feature processing branch to obtain behavior propagation features; The key traffic features, log semantic sequence features, and behavior propagation features extracted are fused along different fusion paths through the cross-modal fusion layer to obtain multi-modal fusion features, and the bilinear pooling layer is used to strengthen the second-order interaction information between the features to strengthen the multi-modal fusion features; The multi-modal fusion features after feature strengthening are input into the fully connected layer to extract the local information about network security threats contained in the respective dimensions of the different modal fusion features, and perform a non-linear transformation on the local information to predict the probability distribution of network security threat types to determine whether there are security threats in the target network and the specific threat types.
[0482] Here, the prediction process is similar to the above training process, and the details are not repeated here.
Claims
1. A multi-level dynamic threat monitoring system based on deep learning, characterized in that: include: A front-end electronic device and a back-end server, wherein a data acquisition module, a data preprocessing module, a feature extraction module, a data vectorization module, and a semantic sequence construction module are deployed on the front-end electronic device, and a model training module and a model deployment module are deployed on the back-end server; The data collection module is used to obtain a first network traffic data sample, a first log data sample, a first user behavior data sample, and network security threat labels of the first network traffic data sample, the first log data sample, and the first user behavior data sample respectively; The data preprocessing module is used to preprocess the first network traffic data sample, the first log data sample, and the first user behavior data sample to obtain a second network traffic data sample, a second log data sample, and a second user behavior data sample, respectively; The feature extraction module is used to extract features from the second network traffic data sample, the second log data sample, and the second user behavior data sample to generate a traffic feature data sample, a log feature data sample, and a behavior feature data sample; The data vectorization module is used to obtain a flow feature vector sample, a log feature vector sample, and a behavior feature vector sample from the flow feature data sample, the log feature data sample, and the behavior feature data sample respectively; The semantic sequence construction module is used to perform cross-semantic relationship analysis on the traffic feature vector samples, the log feature vector samples, and the behavior feature vector samples to obtain a cross-semantic relationship graph, and to mark the network security threat label on the cross-semantic relationship graph to form a semantic training data set; The model training module is used to call the semantic training data set to train the deep learning model to be trained, so that the deep learning model learns the relationship between the first network traffic data sample, the first log data sample, the first user behavior data sample and the corresponding network security threat label, and maps the relationship to the network parameters of the deep learning model to obtain a multi-level dynamic threat monitoring model; The model deployment module is used to deploy the multi-level dynamic threat monitoring model on the cloud server, so that the multi-level dynamic threat monitoring model extracts and processes the operating data features of the target network based on its network parameters to predict the network security threats occurring to the target network.
2. According to claim 1, a multi-level dynamic threat monitoring system based on deep learning is characterized in that: When the data acquisition module acquires the first network traffic data sample, the first log data sample, the first user behavior data sample, and the network security threat labels of the first network traffic data sample, the first log data sample, and the first user behavior data sample, the steps include the following: The data collection module sends a collection requirement to the SDN controller, so that the SDN controller guides the traffic that meets the predetermined requirements to the collection port specified by the data collection module according to the network global view and flow table rules to collect the first network traffic data sample; Compressing the first network traffic data sample collected based on a set content-aware lossless compression mechanism to cache it in a local cache; Based on the constructed multi-source threat intelligence annotation component, matching is performed with the first network traffic data sample at the local cache to assign a network security threat label to the first network traffic data sample.
3. According to claim 1, a multi-level dynamic threat monitoring system based on deep learning is characterized in that: When the data acquisition module acquires the first network traffic data sample, the first log data sample, the first user behavior data sample, and the network security threat labels of the first network traffic data sample, the first log data sample, and the first user behavior data sample, the steps include the following: The data collection module calls the containerized log collection node to collect the first log data sample from different log sources to generate a blockchain block including a timestamp, a device identifier, and a log summary, and performs consensus verification on the blockchain block to determine the authenticity of the first log data sample; In response to the authenticity of the first log data sample being greater than a set authenticity threshold, parsing the first log data sample, constructing a knowledge graph according to entities and relationships, and storing it in a graph database; A labeling mechanism based on user and device behavior portraits is matched with the graph database to mark corresponding network security threat labels in the knowledge graph.
4. According to claim 1, a multi-level dynamic threat monitoring system based on deep learning is characterized in that: When the data acquisition module acquires the first network traffic data sample, the first log data sample, the first user behavior data sample, and the network security threat labels of the first network traffic data sample, the first log data sample, and the first user behavior data sample, the steps include the following: The data acquisition module calls the monitoring interface based on the cross-platform development framework and the bottom layer of the system to detect the bottom layer of the system to capture in real time such as user gesture operations, system-level application calls, and interactive behaviors between different applications; Based on differential privacy, the interaction behavior is associated with the corresponding user subject and noise is added to generate the first user behavior data sample; Adding the first user behavior data sample to the distributed ledger record by record; Based on a set context-aware dynamic labeling mechanism, a network security threat label of the first user behavior data sample is marked in the distributed ledger.
5. The multi-level dynamic threat monitoring system based on deep learning according to claim 1 is characterized in that: When the data preprocessing module preprocesses the first network traffic data sample, the first log data sample, and the first user behavior data sample to obtain the second network traffic data sample, the second log data sample, and the second user behavior data sample, respectively, the steps included are as follows: Based on the trained protocol classification model, perform protocol traffic identification on the first network traffic data sample to remove noise data in the first network traffic data sample; A time series analysis is performed on the first network traffic data sample from which the noise data is removed to correct the timestamp of the first network traffic data sample, and a second network traffic data sample is generated accordingly.
6. A multi-level dynamic threat monitoring system based on deep learning according to claim 1, characterized in that: When the data preprocessing module preprocesses the first network traffic data sample, the first log data sample, and the first user behavior data sample to obtain the second network traffic data sample, the second log data sample, and the second user behavior data sample, respectively, the steps included are as follows: Based on a sequence-to-sequence model of a recurrent neural network, semantic understanding and parsing of the first log data sample is performed to parse a system log containing multiple event information into a structured record containing time, event type, event subject, and detailed description; Mapping the structured record to a knowledge system framework of log data established based on ontology to generate a knowledge graph of log data; Based on log outlier detection combining density clustering and outlier detection, the log data knowledge graph is semantically repaired to generate the second log data sample.
7. The multi-level dynamic threat monitoring system based on deep learning according to claim 1 is characterized in that: When the data preprocessing module preprocesses the first network traffic data sample, the first log data sample, and the first user behavior data sample to obtain the second network traffic data sample, the second log data sample, and the second user behavior data sample, respectively, the steps included are as follows: Based on establishing an association model for multi-source data, the first user behavior data samples of the same user from different channels are cross-validated and integrated, and when there is an inconsistency in time or operation sequence between the first user behavior data sample and the user operation log recorded by the application system, the user login time and authority information are used as an aid to comprehensively judge and correct the integrity of the first user behavior data sample. Among them, for the missing data part, according to the user's historical behavior pattern and group behavior statistical law, a behavior sequence prediction and completion mechanism based on Markov chain is used to complete the missing data to generate the second user behavior data sample.
8. The multi-level dynamic threat monitoring system based on deep learning according to claim 1 is characterized in that: When the feature extraction module extracts features from the second network traffic data sample, the second log data sample, and the second user behavior data sample to generate a traffic feature data sample, a log feature data sample, and a behavior feature data sample, the steps included are as follows: The nodes in the network are taken as the vertices of the graph, the network connections between the nodes and the traffic flow relationship are taken as the edges, and the traffic size and transmission frequency are taken as the edge weights to construct a traffic graph structure representation; Based on the graph convolutional neural network, feature extraction is performed on the traffic graph structure representation to remove the key nodes in the top network, the traffic convergence area and the traffic correlation pattern between different network segments, and the traffic feature data sample is generated accordingly.
9. The multi-level dynamic threat monitoring system based on deep learning according to claim 1, characterized in that: When the feature extraction module extracts features from the second network traffic data sample, the second log data sample, and the second user behavior data sample to generate a traffic feature data sample, a log feature data sample, and a behavior feature data sample, the steps included are as follows: Based on the pre-trained language model, label each word in the second log data sample with a corresponding semantic role; Based on the annotated semantic roles, the subject, operation object, operation method, time and place of the event are extracted to generate the log feature data sample.
10. The multi-level dynamic threat monitoring system based on deep learning according to claim 1, characterized in that: When the feature extraction module extracts features from the second network traffic data sample, the second log data sample, and the second user behavior data sample to generate a traffic feature data sample, a log feature data sample, and a behavior feature data sample, the steps included are as follows: The user is used as a node, the user's operation behavior at different times, on different application systems or devices is used as an edge, and the time interval between behaviors and the frequency of operations are used as edge attributes to construct a user behavior trajectory graph; The user behavior trajectory graph is mapped to a low-dimensional vector space through graph embedding to determine the dynamic change process of user behavior and the correlation between different behaviors, and the behavior feature data sample is generated accordingly.
Citation Information
Patent Citations
Context-aware feature embedding and anomaly detection of sequential log data using deep recurrent neural networks
CN113190842A
Safe honeypot system and implementation method thereof
CN116760558A
Full-scene network security threat association analysis method and system
CN117478403A
Multi-dimensional internal threat detection method based on deep learning
CN118972113A
Monitoring method of cloud security system
CN119276604A
Cited By
Artificial intelligence-based database security analysis method and device, equipment and medium
CN120257272A
Network security threat intelligent detection method based on big data analysis
CN120498872A
Artificial intelligence early warning and management method for smart ocean
CN120710806A
Network security analysis system based on AI algorithm
CN122204503A