Information security management system based on big data

By designing a big data-based information security management system, using distributed data acquisition, automatic encoder abnormality detection, recurrent neural network threat prediction and random forest risk assessment, the shortcomings of existing systems in data acquisition, abnormality detection, threat prediction and risk assessment are solved, and efficient and accurate information security management and automated response decision-making are achieved.

CN120086768APending Publication Date: 2025-06-03广东晖曜科技有限公司
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510166330.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing information security management system has shortcomings in data collection, abnormality detection, threat prediction and risk assessment, and it is difficult to meet the current needs of information security management.

Method used

A information security management system based on big data is designed, using distributed data acquisition technology, automatic encoder AE algorithm to build anomaly detection model, recurrent neural network RNN ​​algorithm to build a threat prediction model, and a random forest RF algorithm to build a risk assessment model, and a response decision module is provided.

Benefits of technology

Real-time and comprehensive collection of network traffic data, system log data, user behavior data and application log data is realized, the accuracy and timeliness of abnormal detection are improved, threat prediction capabilities are enhanced, the system's security risks are comprehensively and comprehensively evaluated, and security response measures can be automatically implemented.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086768A_ABST
    Figure CN120086768A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network information security, discloses an information security management system based on big data, and aims to solve the defects of an existing system in the aspects of data acquisition, anomaly detection, threat prediction, risk assessment and the like. The system comprises a data acquisition module for collecting various types of data in real time; the abnormal detection model is used for identifying normal and abnormal behaviors by adopting an automatic encoder AE algorithm; the threat prediction module is used for predicting a future security threat type and probability by using a recurrent neural network (RNN) algorithm; the risk assessment module comprehensively assesses the system security risk through a random forest RF algorithm; and the response decision module is used for executing corresponding safety response measures according to the evaluation result. According to the system, the big data technology is utilized, comprehensive monitoring and accurate prediction of the network security state are achieved, the efficiency and accuracy of information security management are improved, and powerful support is provided for coping with novel and complex network attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network information security, and specifically provides an information security management system based on big data. Background Art

[0002] With the rapid development of information technology, the cyber space has become an indispensable part of modern society. However, with the wide popularization of network applications and the massive growth of data, information security issues have become increasingly prominent. Security threats such as cyber attacks, data breaches, and malware emerge in an endless stream, posing severe challenges to the information security of individuals, enterprises, and countries. Traditional information security management methods often rely on technologies such as rule matching and signature recognition. These methods are unable to cope when faced with new and complex attack means and are difficult to meet the current requirements of information security management.

[0003] In the field of information security, the application of big data technology provides new ideas and methods for information security management. Big data has the characteristics of large volume, diverse types, and fast processing speed, and can capture and analyze various data in the network in real time, providing comprehensive data support for information security management. How to effectively utilize big data technology to build an efficient and accurate information security management system is still an urgent problem to be solved currently.

[0004] Existing information security management systems have many deficiencies in aspects such as data collection, anomaly detection, threat prediction, and risk assessment. In terms of data collection, traditional systems can often only collect limited data types and cannot comprehensively reflect the security status of the network. In terms of anomaly detection, traditional methods often rely on known attack patterns or features and are difficult to detect unknown new attacks. In terms of threat prediction, there is a lack of effective means to accurately predict future security threats. In terms of risk assessment, traditional methods often only consider a single security factor and cannot comprehensively and comprehensively evaluate the security risks of the system. Summary of the Invention

[0005] The purpose of the present invention is to provide an information security management system based on big data to solve the problems raised in the above background art.

[0006] To achieve the above purpose, the present invention provides the following technical solution: An information security management system based on big data, the system includes:

[0007] A data collection module, constructed using distributed data collection technology, for collecting network traffic data, system log data, user behavior data, and application program log data in real time;

[0008] Anomaly detection model, constructed using the Autoencoder (AE) algorithm. The input data of the anomaly detection model is the data collected by the data acquisition module, and the output data is the anomaly detection result, including normal behavior and abnormal behavior. The model is trained using a dataset containing normal behavior and known abnormal behavior, with corresponding labels being the data behavior categories.

[0009] Threat prediction module, which constructs a threat prediction model using the Recurrent Neural Network (RNN) algorithm. The input data of the threat prediction model is the historical anomaly detection result and threat intelligence data, and the output data is the predicted types of security threats that may occur in the future and their probabilities within a certain period. The model is trained using a time series dataset containing historical anomaly detection results and threat intelligence, with corresponding labels being the future security threat types.

[0010] Risk assessment module, which constructs a risk assessment model using the Random Forest (RF) algorithm. The input data of the risk assessment model is the detection result of the anomaly detection model and the prediction result of the threat prediction model, and it comprehensively assesses the system security risk by combining the information of both. The model is trained using a combined dataset containing anomaly detection results and threat prediction results, with corresponding labels being the security risk levels.

[0011] Response decision module, connected to the risk assessment module, executes corresponding security response measures according to the assessment result. If the assessment result is high risk, it triggers a security alarm and initiates an emergency response process.

[0012] Preferably, the data acquisition module uses the Kafka message queue to achieve real-time transmission and storage of data.

[0013] Preferably, the anomaly detection model adopts a multi-layer autoencoder structure, reconstructs the input data through the encoder and decoder, and calculates the reconstruction error to identify anomalies. Its network structure includes:

[0014] Input layer: Receives the preprocessed data, and the data format is unified into a vector form.

[0015] Encoder layer: Compresses the input data into a low-dimensional representation through a multi-layer neural network.

[0016] Decoder layer: Decodes the low-dimensional representation back to the original data space to obtain the reconstructed data.

[0017] Loss calculation layer: Calculates the reconstruction error between the input data and the reconstructed data, and uses the Mean Squared Error (MSE) as the loss function.

[0018] Output layer: Judges whether the data is an anomaly according to the reconstruction error, sets a threshold, and if the error exceeds the threshold, it is determined as an anomaly.

[0019] Preferably, the threat prediction model captures the dependencies in time series data through RNN units, and its network structure includes:

[0020] Input layer: Receives historical anomaly detection results and threat intelligence data;

[0021] RNN layer: Contains multiple RNN units, each unit processes one time step in the sequence data and passes the hidden state to the next unit;

[0022] Fully connected layer: Maps the hidden state output by the RNN layer to the probability of the predicted security threat type through the fully connected layer;

[0023] Output layer: Outputs the possible security threat types and their probabilities in a future period of time.

[0024] Preferably, the risk assessment model conducts risk assessment by constructing multiple decision trees and integrating their output results. Its construction steps include:

[0025] S1: Uses the anomaly detection results and threat prediction results as input features and the security risk level as the target variable;

[0026] S2: Randomly selects a feature subset and a data subset to construct a decision tree;

[0027] S3: Repeats step S2 to construct multiple decision trees to form a random forest;

[0028] S4: For new input data, each tree makes a prediction respectively, and uses the majority voting mechanism to determine the final risk level.

[0029] Preferably, the optimization steps of the risk assessment model include:

[0030] Step 1: Adjust the parameters of the random forest, including the number of trees, the maximum features, and the maximum depth;

[0031] Step 2: Uses K-fold cross-validation to evaluate the model performance and selects the optimal parameter combination;

[0032] Step 3: Uses the validation set to verify the optimized model and evaluates the accuracy and AUC value of the model;

[0033] Step 4: When the model performance reaches the preset standard, saves the model parameters to obtain the trained risk assessment model.

[0034] Preferably, the implementation method of the response decision module includes:

[0035] Build a response policy library: Pre-define multiple security response policies, each corresponding to a different security risk level and security threat type. The policy content covers the level of security alerts, the activation conditions of the emergency response process, and specific security protection measures.

[0036] Build a decision-making engine: Receive the evaluation results from the risk assessment module. According to the risk level and security threat type in the evaluation results, match the corresponding security response policy in the response policy library.

[0037] Build an execution unit: Execute specific security response measures according to the security response policy matched by the decision-making engine, including adjusting firewall rules, isolating infected devices, updating security patches, and sending security alert notifications.

[0038] Preferably, the specific method for building the response policy library is as follows:

[0039] Set a decision tree as the data structure for storing security response policies. Among them, each node contains a private attribute set, which is used to record the security risk level and security threat type represented by the path from the root node to the current node; the leaf node contains a target text, which represents a specific security response policy; the decision tree uses a hash table structure to efficiently store and retrieve node information; the decision tree node structure is defined as: D = {(a, b) | a ∈ {risk level, threat type}, b ∈ {null, response policy}}; where, a represents the security risk level and security threat type, and b represents the value of the node, indicating the response policy or null value.

[0040] Encode the security risk level and security threat type. Use a predefined enumeration type to represent the security risk level; use a predefined enumeration type to represent the security threat type.

[0041] Preferably, the response decision module further includes an attribute matcher, which has an enumeration type calculator built in; using the enumeration type calculator, the attribute matcher quickly locates a specific index in the attribute set, and then finds the complete path in the decision tree that matches the encoded security risk level and security threat type.

[0042] Preferably, the response decision module further includes a tree constructor, which provides add and delete functions for users to add or delete nodes and paths in the decision tree according to actual business needs and changes in the security environment.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] This system adopts distributed data acquisition technology and can collect various types of data in real time and comprehensively, such as network traffic data, system log data, user behavior data, and application program log data. This all-round data acquisition method greatly enriches the data foundation required for information security management, enabling the system to more accurately reflect the security status of the network and providing solid data support for subsequent anomaly detection, threat prediction, and risk assessment.

[0045] The system introduces the Autoencoder (AE) algorithm to build an anomaly detection model. This model can automatically learn the normal behavior patterns of data and effectively identify abnormal behaviors that deviate significantly from the normal patterns. Compared with traditional methods that rely on known attack patterns or features, the anomaly detection model of this system has stronger detection capabilities for unknown new attacks, improving the accuracy and timeliness of anomaly detection. Through the Recurrent Neural Network (RNN) algorithm to build a threat prediction model, the system can use historical anomaly detection results and threat intelligence data to accurately predict the types and probabilities of possible security threats in a future period. This prediction ability enables information security managers to take preventive measures in advance and effectively reduce the risk of the system being attacked.

[0046] The system uses the Random Forest (RF) algorithm to build a risk assessment model. This model can comprehensively consider the detection results of the anomaly detection model and the prediction results of the threat prediction model to conduct a comprehensive and integrated assessment of the security risks of the system. This assessment method avoids the limitations of traditional methods that only consider a single security factor and improves the accuracy and reliability of risk assessment. The system is equipped with a response decision-making module that can automatically execute corresponding security response measures according to the results of the risk assessment module. When the assessment result is a high risk, the system will trigger a security alarm and start an emergency response process to ensure that security threats can be responded to in a timely and effective manner and losses can be reduced. Brief Description of the Drawings

[0047] Figure 1 It is the working principle diagram of the big data-based information security management system described in the present invention;

[0048] Figure 2 It is the implementation flow chart of the anomaly detection model and the threat prediction model;

[0049] Figure 3 It is the step diagram of the construction and optimization of the risk assessment model. Detailed Embodiment

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0051] Please refer to Figures 1 - 3 , the present invention provides a technical solution: an information security management system based on big data, and the system includes:

[0052] Data collection module: The data collection module is constructed by using distributed data collection technology to realize the real-time collection of network traffic data, system log data, user behavior data, and application program log data. Specifically, when implemented, data collection agents can be deployed at key nodes in the network. These agents are responsible for collecting various types of data at their respective nodes and transmitting the data to the central data processing center in real time. The central data processing center preprocesses the received data, such as deduplication, cleaning, formatting, etc., for subsequent analysis and processing.

[0053] Anomaly detection model: The anomaly detection model is constructed by using the Autoencoder AE algorithm. The input data of the model is the data collected by the data collection module, and the output data is the anomaly detection result, including normal behavior and abnormal behavior. During the model training process, a data set containing normal behavior and known abnormal behavior is constructed, and each data sample is labeled with its behavior category label. This data set is used to train the autoencoder so that it can learn the feature representations of normal behavior and abnormal behavior. After training, when new data is input into the model, the model will reconstruct the data according to the feature representations it has learned and calculate the reconstruction error. If the reconstruction error is large, it indicates that the input data is quite different from the normal behavior and may be abnormal behavior.

[0054] Threat prediction module: The threat prediction module constructs a threat prediction model by using the Recurrent Neural Network RNN algorithm. The input data of the model is the historical anomaly detection result and threat intelligence data, and the output data is the predicted types of security threats that may occur in the future for a period of time and their probabilities. During the model training stage, a time series data set containing historical anomaly detection results and threat intelligence is constructed, and each time point of data is labeled with the future security threat type label. This data set is used to train the recurrent neural network so that it can learn the temporal relationship between historical anomaly detection results, threat intelligence, and future security threats. After training, when new historical anomaly detection results and threat intelligence are input into the model, the model will predict the possible future security threats according to the temporal relationship it has learned.

[0055] Risk Assessment Module: The risk assessment module constructs a risk assessment model using the Random Forest (RF) algorithm. The input data for the model is the detection results of the anomaly detection model and the prediction results of the threat prediction model. By combining the information from both, a comprehensive assessment of the system security risk is conducted. During the model training process, a combined dataset containing anomaly detection results and threat prediction results is constructed, and each data sample is labeled with its security risk level tag. This dataset is used to train the random forest so that it can learn the mapping relationship between the anomaly detection results, threat prediction results, and the security risk level. After training, when new anomaly detection results and threat prediction results are input into the model, the model will conduct a comprehensive assessment of the system security risk based on the learned mapping relationship.

[0056] Response Decision Module: The response decision module is connected to the risk assessment module and executes corresponding security response measures based on the assessment results. In specific implementation, a set of security response rules can be preset, with each rule corresponding to a security risk level and the corresponding response measures. When the risk assessment module outputs the assessment results, the response decision module will match the corresponding security response rules according to the assessment results and execute the response measures specified in the rules. If the assessment result is a high risk, a security alarm will be triggered and the emergency response process will be initiated to promptly address possible security threats.

[0057] The present invention will be further described below in conjunction with Embodiments 1 to 4:

[0058] Embodiment 1:

[0059] The anomaly detection model adopts a multi-layer autoencoder structure, and the threat prediction model captures the dependencies in time series data through RNN units. The anomaly detection model adopts a multi-layer autoencoder structure, reconstructs the input data through the encoder and decoder, and calculates the reconstruction error to identify anomalies. Its network structure specifically includes:

[0060] Input Layer: Receives the preprocessed data, which has been cleaned, formatted, etc. and unified into vector form. For example, for network traffic data, the features of each data packet (such as source IP, destination IP, port number, protocol type, etc.) can be extracted to form a feature vector as the input.

[0061] Encoder Layer: Compresses the input data into a low-dimensional representation through a multi-layer neural network. Each layer of the neural network contains several neurons, which perform a linear transformation on the input data through weights and biases, and introduce non-linear characteristics through activation functions (such as ReLU, Sigmoid, etc.). After being passed layer by layer, the input data is compressed into a low-dimensional hidden representation.

[0062] Decoder layer: Decodes the low-dimensional representation back to the original data space to obtain the reconstructed data. The structure of the decoder layer is symmetric to that of the encoder layer and also consists of multiple neural networks. Each neural network layer performs a linear transformation on the low-dimensional representation through weights and biases, and restores the features of the original data through an activation function.

[0063] Loss calculation layer: Calculates the reconstruction error between the input data and the reconstructed data, and uses the mean squared error (MSE) as the loss function. That is, it calculates the sum of the squares of the differences between the elements of the input vector and the reconstructed vector, and then takes the average as the loss value.

[0064] Output layer: Determines whether the data is abnormal based on the reconstruction error. A threshold is set. When the reconstruction error exceeds this threshold, the input data is determined to be an abnormal behavior; otherwise, it is determined to be a normal behavior. The selection of the threshold can be determined through experiments or experience to achieve the best anomaly detection effect.

[0065] In practical applications, network traffic data over a period of time can be collected as the training set to train the multi-layer autoencoder model. After training, new network traffic data is input into the model, the reconstruction error is calculated, and it is determined whether it is abnormal. If abnormal behavior is found, the characteristics of the abnormal data can be further analyzed and corresponding security measures can be taken.

[0066] The threat prediction model captures the dependencies in time series data through RNN units, and its network structure specifically includes:

[0067] Input layer: Receives historical anomaly detection results and threat intelligence data. The historical anomaly detection results can be anomaly behavior labels or anomaly probabilities output by the anomaly detection model; the threat intelligence data can be information about security threats such as new attacks and vulnerability exploitations obtained from external sources.

[0068] RNN layer: Contains multiple RNN units, and each unit processes one time step in the sequence data. The RNN units pass information through the hidden state, so that the output of the current time step depends not only on the current input but also on the inputs of previous time steps. In this way, the RNN layer can capture the long-term dependencies in time series data.

[0069] Fully connected layer: Maps the hidden state output by the RNN layer to the probability of the predicted security threat type through the fully connected layer. The fully connected layer contains several neurons, and each neuron corresponds to a possible security threat type. It performs a linear transformation on the hidden state through weights and biases, and converts the output into a probability distribution through the Softmax function.

[0070] Output layer: Output the types of security threats that may occur in the future and their probabilities. Based on the output of the fully connected layer, the predicted probability of each type of security threat can be obtained. The type of security threat with the highest probability can be selected as the prediction result, or a threshold can be set. When the predicted probability of a certain type of security threat exceeds the threshold, it is considered that this threat type may occur.

[0071] In practical applications, historical anomaly detection results and threat intelligence data over a period of time can be collected as the training set to train the RNN threat prediction model. After training, the new historical anomaly detection results and threat intelligence data are input into the model to predict the types of security threats that may occur in the future and their probabilities. According to the prediction results, corresponding defense measures can be taken in advance to reduce the security risks faced by the system.

[0072] Embodiment 2:

[0073] The risk assessment model conducts risk assessment by constructing multiple decision trees and integrating their output results. The steps for constructing the risk assessment model include:

[0074] S1: Feature and target variable determination: Use the anomaly detection results (such as the probability or label of abnormal behavior) and threat prediction results (such as the types of security threats that may occur in the future and their probabilities) as input features. Use the security risk level (such as low, medium, high) as the target variable, which is the output that the model needs to predict.

[0075] S2: Decision tree construction: Randomly select a subset of features and a subset of data to construct a single decision tree. The subset of features is a part of the features randomly selected from all input features, and the subset of data is a part of the data randomly sampled from the training dataset. The construction process of the decision tree includes steps such as selecting the best splitting feature, determining the splitting point, and recursively constructing subtrees until the stopping condition is met (such as reaching the maximum depth, the number of samples in the leaf node is less than the threshold, etc.).

[0076] S3: Random forest formation: Repeat step S2 multiple times, each time using different subsets of features and subsets of data to construct decision trees, thus forming a random forest composed of multiple decision trees. Each tree in the random forest is independent, and they will give different prediction results for the same input data.

[0077] S4: Risk level prediction: For new input data, each tree in the random forest makes a prediction separately to obtain its respective risk level. A majority voting mechanism is used to determine the final risk level. That is, count which level appears the most times among the risk levels given by all the trees, and take that level as the final prediction result.

[0078] The optimization steps of this model include:

[0079] Step 1: Parameter Adjustment: Adjust the parameters of the random forest, including the number of trees (i.e., the total number of decision trees in the random forest), the maximum features (i.e., the maximum number of features that can be considered when each tree splits), and the maximum depth (i.e., the maximum depth of each tree). The settings of these parameters will affect the performance and generalization ability of the random forest, and the optimal parameter combination needs to be determined through experiments.

[0080] Step 2: Model Performance Evaluation: Use K-fold cross-validation to evaluate the performance of the model. Divide the training dataset into K subsets. Each time, use K - 1 subsets to train the model and the remaining 1 subset to test the model. Repeat this K times, and take the average of the K test results as the performance metric of the model. Select the parameter combination that makes the model performance metric (such as accuracy, AUC value, etc.) optimal.

[0081] Step 3: Validation Set Validation: Use an independent validation set to validate the optimized model. The validation set is a dataset that does not overlap with the training dataset and is used to evaluate the generalization ability of the model. Calculate the accuracy and AUC value of the model on the validation set to evaluate whether the performance of the model meets the preset standard.

[0082] Step 4: Model Saving: When the model performance meets the preset standard, save the parameters of the model, including the structure of each tree, the splitting points, the risk levels of the leaf nodes, etc. Obtain the trained risk assessment model, which can be used to predict the risk level of new input data.

[0083] Suppose there is an information security management system that needs to evaluate the security risk level of the system. Collect the anomaly detection results and threat prediction results over a period of time as training data. Build and optimize the risk assessment model according to the above steps. During the building process, randomly select feature subsets (such as the probability of abnormal behavior, the probability of threat types, etc.) and data subsets to build multiple decision trees to form a random forest. For new input data (such as the anomaly detection results and threat prediction results on a certain day), each tree makes a prediction separately, and the majority voting mechanism is used to determine the final risk level. During the optimization process, adjust the parameters of the random forest (such as the number of trees, the maximum features, and the maximum depth), use K-fold cross-validation to evaluate the model performance, and select the optimal parameter combination. Then, use the validation set to validate the optimized model and evaluate the accuracy and AUC value of the model. When the model performance meets the preset standard, save the model parameters to obtain the trained risk assessment model. In this way, the trained risk assessment model can be used to predict the risk level of new input data and provide decision support for the information security management system.

[0084] Example 3:

[0085] The response decision-making module realizes automated response decisions for different security risk levels and security threat types by constructing a response policy library, a decision-making engine, and an execution unit. The specific implementation methods of this module include:

[0086] ① Construct a response policy library: The response policy library is a collection of a series of pre-defined security response policies. Each policy corresponds to a specific security risk level and security threat type to ensure a rapid and accurate response when facing different security events. Specifically, it includes:

[0087] Policy content: Define the alarm levels to be issued when specific security risks or threats occur, such as low, medium, high, or emergency. Clearly define under what circumstances the emergency response process should be initiated, for example, when the risk level reaches high or emergency. List the specific protection measures to be taken for specific threat types, such as adjusting firewall rules, isolating infected devices, updating security patches, etc.

[0088] Policy classification: Classify the policies according to the risk level (such as low, medium, high, emergency) and security threat type (such as DDoS attack, malware infection, data leakage, etc.) to facilitate subsequent matching and selection.

[0089] ② Construct a decision-making engine: The decision-making engine is the core component of the response decision-making module, responsible for receiving the evaluation results of the risk assessment module and matching the corresponding security response policies in the response policy library according to the risk level and security threat type in the results.

[0090] Input: The evaluation results of the risk assessment module, including the risk level and security threat type.

[0091] Processing logic: According to the input risk level and security threat type, search for matching policies in the response policy library. If a matching policy is found, output the policy as the decision result; if no completely matching policy is found, select the closest policy or the default policy according to the preset rules.

[0092] Output: The matching security response policy, including the security alarm level, the activation conditions of the emergency response process, and specific security protection measures.

[0093] ③ Construct an execution unit: The execution unit is responsible for executing specific security response measures according to the security response policy matched by the decision-making engine.

[0094] Modify the firewall configuration according to the requirements in the policy to block or restrict specific types of network traffic. Isolate infected devices from the network to prevent them from further spreading malware or launching attacks. Update security patches for systems or applications in a timely manner according to the requirements in the policy to fix known security vulnerabilities. Send security alert notifications to relevant personnel or systems according to the security alert level in the policy so that they can take timely measures.

[0095] Suppose the risk assessment module evaluates that the current system has a high risk level and the security threat type is a DDoS attack. After receiving this assessment result, the decision engine looks up the policy in the response policy library that matches "high risk" and "DDoS attack". After finding the matching policy, the decision engine outputs the policy to the execution unit. The execution unit adjusts the firewall rules according to the requirements in the policy to restrict traffic from the attack source; isolates infected devices to prevent them from continuing to participate in the attack; updates the security patches of the system to fix possible security vulnerabilities; sends high-level security alert notifications to network administrators and the security team to inform them of the current DDoS attack situation and remind them to take corresponding measures. In this way, the response decision module can achieve automated response decisions for different security risk levels and security threat types, improving the efficiency and accuracy of the information security management system.

[0096] Example 4:

[0097] This example is used to describe the construction method of the response policy library and the implementation methods of the attribute matcher and tree constructor in the response decision module.

[0098] ① Specific method for constructing the response policy library: The response policy library uses a decision tree as the data structure for storing security response policies. Each node of the decision tree contains a private set of attributes, which is used to record the security risk level and security threat type represented by the path from the root node to the current node. The leaf node contains a target text, which represents a specific security response policy.

[0099] Definition of the decision tree node structure:

[0100] D = {(a, b) | a ∈ {risk level, threat type}, b ∈ {null, response policy}}

[0101] Among them, a represents the security risk level and security threat type, which are the attributes of the node;

[0102] b represents the value of the node. For non-leaf nodes, its value is null; for leaf nodes, its value is a specific security response policy.

[0103] Encoding of security risk levels and security threat types: The security risk levels are represented using a predefined enumeration type, such as low risk, medium risk, high risk, etc. The security threat types are represented using a predefined enumeration type, such as DDoS attack, malware, data leakage, etc.

[0104] Storage and retrieval of decision trees: The decision tree uses a hash table structure to efficiently store and retrieve node information. Each node has a unique identifier in the hash table for quick access.

[0105] Suppose there is a decision tree whose root node represents the starting point of all possible security risk levels and security threat types. Starting from the root node, according to the encoding of the security risk level (such as high risk) and the security threat type (such as DDoS attack), one can traverse step by step down the path of the decision tree until reaching a leaf node. The target text contained in this leaf node is a security response strategy for high-risk DDoS attacks.

[0106] ② Implementation of the attribute matcher: The attribute matcher is a component in the response decision module used to quickly locate the complete path in the decision tree that matches the input encoding of the security risk level and security threat type.

[0107] Enumeration type calculator: The attribute matcher has an in-built enumeration type calculator for handling the enumeration types of security risk levels and security threat types. The enumeration type calculator can quickly calculate the specific index in the attribute set based on the input encoding.

[0108] Path search: Using the index calculated by the enumeration type calculator, the attribute matcher can search in the decision tree for the complete path that matches the encoding of the security risk level and security threat type. This path starts from the root node, passes through the nodes that match the input encoding in sequence, and finally reaches the leaf node containing the target security response strategy.

[0109] ③ Implementation of the tree constructor: The tree constructor is another component in the response decision module used to add or delete nodes and paths in the decision tree according to actual business requirements and changes in the security environment.

[0110] Addition and deletion functions: The tree constructor provides addition and deletion functions, allowing users to add new nodes and paths as needed to expand the coverage of the decision tree. Similarly, the tree constructor also allows users to delete nodes and paths that are no longer needed to maintain the simplicity and effectiveness of the decision tree.

[0111] Addition of nodes and paths: When new security response strategies need to be added, users can add the corresponding nodes and paths in the decision tree through the tree constructor. The newly added nodes and paths should reflect the new security risk levels and security threat types, as well as the corresponding security response strategies.

[0112] Deletion of nodes and paths: When a certain security response policy is no longer applicable or has been replaced, the user can delete the corresponding nodes and paths in the decision tree through the tree constructor. The deletion operation should ensure the integrity and consistency of the decision tree and avoid affecting the normal access and retrieval of other nodes.

[0113] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0114] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An information security management system based on big data, characterized in that: The system comprises: The data collection module is built using distributed data collection technology and is used to collect network traffic data, system log data, user behavior data, and application log data in real time; The anomaly detection model is constructed using the autoencoder AE algorithm. The input data of the anomaly detection model is the data collected by the data acquisition module, and the output data is the anomaly detection result, including normal behavior and abnormal behavior. The model training uses a data set containing normal behavior and known abnormal behavior, and the corresponding label is the data behavior category. The threat prediction module uses the recurrent neural network (RNN) algorithm to build a threat prediction model. The input data of the threat prediction model are historical anomaly detection results and threat intelligence data, and the output data are the predicted security threat types and their probabilities that may appear in the future. The model training uses a time series data set containing historical anomaly detection results and threat intelligence, and the corresponding labels are future security threat types. The risk assessment module uses the random forest RF algorithm to build a risk assessment model. The input data of the risk assessment model are the detection results of the anomaly detection model and the prediction results of the threat prediction model. The system security risk is comprehensively assessed by combining the information of the two. The model training uses a combined data set containing anomaly detection results and threat prediction results, and the corresponding label is the security risk level. The response decision module is connected to the risk assessment module and executes corresponding security response measures according to the assessment results. If the assessment result is high risk, a security alarm is triggered and the emergency response process is started.

2. According to the information security management system based on big data in claim 1, it is characterized in that: The data acquisition module uses Kafka message queue to achieve real-time transmission and storage of data.

3. The information security management system based on big data according to claim 1 is characterized in that: The anomaly detection model adopts a multi-layer autoencoder structure, reconstructs input data through encoders and decoders, and calculates reconstruction errors to identify anomalies. Its network structure includes: Input layer: receives preprocessed data, and the data format is unified into vector form; Encoder layer: compresses input data into a low-dimensional representation through a multi-layer neural network; Decoder layer: decodes the low-dimensional representation back to the original data space to obtain reconstructed data; Loss calculation layer: calculates the reconstruction error between the input data and the reconstructed data, using the mean square error (MSE) as the loss function; Output layer: Determine whether the data is abnormal based on the reconstruction error, set a threshold, and determine it as abnormal if the error exceeds the threshold.

4. The information security management system based on big data according to claim 1 is characterized in that: The threat prediction model captures the dependencies in time series data through RNN units, and its network structure includes: Input layer: receives historical anomaly detection results and threat intelligence data; RNN layer: contains multiple RNN units, each unit processes a time step in the sequence data and passes the hidden state to the next unit; Fully connected layer: maps the hidden state output by the RNN layer to the predicted security threat type probability through the fully connected layer; Output layer: Outputs the types of security threats that may occur in the future and their probabilities.

5. The information security management system based on big data according to claim 1 is characterized in that: The risk assessment model performs risk assessment by constructing multiple decision trees and integrating their output results, and the construction steps include: S1: Anomaly detection results and threat prediction results are used as input features, and security risk level is used as the target variable; S2: Randomly select feature subsets and data subsets to build decision trees; S3: Repeat step S2 to construct multiple decision trees to form a random forest; S4: For new input data, each tree makes predictions separately, and a majority voting mechanism is used to determine the final risk level.

6. The information security management system based on big data according to claim 5 is characterized in that: The optimization steps of the risk assessment model include: Step 1: Adjust the parameters of the random forest, including the number of trees, maximum features, and maximum depth; Step 2: Use K-fold cross validation to evaluate model performance and select the optimal parameter combination; Step 3: Use the validation set to validate the optimized model and evaluate the accuracy and AUC value of the model; Step 4: When the model performance reaches the preset standard, save the model parameters and obtain the trained risk assessment model.

7. The information security management system based on big data according to claim 1 is characterized in that: The implementation of the response decision module includes: Build a response strategy library: pre-define multiple security response strategies, each strategy corresponds to different security risk levels and security threat types. The strategy content covers the level of security alerts, the conditions for starting the emergency response process, and specific security protection measures; Build a decision engine: Receive the assessment results of the risk assessment module, and match the corresponding security response strategy in the response strategy library according to the risk level and security threat type in the assessment results; Build an execution unit: According to the security response strategy matched by the decision engine, execute specific security response measures, including adjusting firewall rules, isolating infected devices, updating security patches, and sending security alert notifications.

8. The information security management system based on big data according to claim 7 is characterized in that: The specific method of building a response strategy library is: A decision tree is set as a data structure for storing security response strategies, where each node contains a private attribute set, which is used to record the security risk level and security threat type represented by the path from the root node to the current node; the leaf node contains a target text, which represents a specific security response strategy; the decision tree uses a hash table structure to efficiently store and retrieve node information; the decision tree node structure definition: D = {(a, b) | a ∈ {risk level, threat type}, b ∈ {null, response strategy}}; where a represents the security risk level and security threat type, and b represents the value of the node, indicating the response strategy or a null value; The security risk level and the security threat type are encoded, and a predefined enumeration type is used to represent the security risk level; a predefined enumeration type is used to represent the security threat type.

9. The information security management system based on big data according to claim 8 is characterized in that: The response decision module also includes an attribute matcher, which has an enumeration type operator built in. Using the enumeration type operator, the attribute matcher quickly locates a specific index in the attribute set, and then finds a complete path that matches the security risk level and security threat type code in the decision tree.

10. The information security management system based on big data according to claim 9, characterized in that: The response decision module also includes a tree constructor, which provides an add-delete function for users to add or delete nodes and paths in the decision tree according to actual business needs and changes in the security environment.

Citation Information

Cited By

  • Method for enhancing robustness of artificial intelligence system

    CN120493988A

  • Network security multi-mode intelligent detection system and method

    CN120498904A

  • Low-altitude unmanned aerial vehicle early warning and countering method and system suitable for temporary management and control area

    CN120880598A

  • Internet of Things equipment data interaction system based on edge computing

    CN121309662A

  • IoT device data interaction system based on edge computing

    CN121309662B