Heterogeneous network security detection method based on model interpretation
By constructing a weighted feature matrix and convolutional neural network, the problems of abnormal traffic and malicious program detection in heterogeneous networks are solved, efficient and accurate security detection is achieved, and real-time security analysis is suitable for heterogeneous networks.
Patent Information
- Application Number
- CN202510631465.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to efficiently detect abnormal network traffic and malicious executable programs in heterogeneous networks simultaneously, and the deep learning model is highly complex and difficult to deploy in complex network structures.
The abnormal network traffic analysis module and malicious program analysis module are constructed based on model interpretation, and the weighted feature matrix is generated, and the convolutional neural network is used for real-time detection to optimize the accuracy of the deep learning model.
It realizes efficient and secure detection of heterogeneous networks, can identify new malicious attacks and programs, reduces the complexity of deep learning models, and improves detection efficiency.
Smart Images

Figure CN120378187A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and particularly relates to a heterogeneous network security detection method based on model interpretation. Background Art
[0002] With the rapid development of computer network technology, the Internet has covered all aspects of social life. In the early days of the Internet, due to the small network scale, hosts or terminals usually had only a single network interface and the transmission protocol was mainly a single-path transmission protocol, which was sufficient to meet the user's demand for network transmission. However, with the development of the network, network resources have become increasingly rich, and multi-interface multi-homed terminals have gradually become popular in all aspects. Currently, devices adopt multi-path transmission to improve the overall transmission efficiency of the network. Therefore, the traditional single-path transmission can no longer meet the needs of multiple parties, including users and service providers. A heterogeneous network is a network formed by connecting computers, network devices, and systems manufactured by different manufacturers, and these devices often run different protocol stacks. Through a heterogeneous network, a distributed network system with multi-point interconnection, simplicity, and high efficiency can be built.
[0003] Due to the characteristics of multi-node and multi-link in a heterogeneous network, a large amount of data transmission is involved in the distributed system of the heterogeneous network, and the network structure is complex. Therefore, the security detection of the heterogeneous network is crucial. The security issues that need to be closely concerned during the security detection process are mainly divided into two categories. One is to pay attention to whether the network has been maliciously attacked, resulting in abnormal network traffic patterns. The other is to pay attention to whether malicious viruses or Trojan files have been implanted in the file data transmitted in the network.
[0004] For malicious attack behaviors in the network, the traditional method is to develop an intrusion detection system for network security detection. By collecting and analyzing data samples such as host system and network logs, network traffic, or user activity behaviors, malicious intrusion behaviors are discovered, and then corresponding abnormal conditions are quickly reported to users according to the settings and processed. For malicious program detection, the traditional method is to parse the structure information such as the file header structure of the program to be detected, assuming that there is a certain association between the unknown program and the old program, and comparing it with the existing malicious rule library or malicious program library, and then analyzing whether the program to be detected is a malicious program.
[0005] The disadvantages of the prior art are as follows:
[0006] (1) The heterogeneous network structure is complex, and the existing security detection methods are difficult to simultaneously detect multiple different types of networks. Moreover, the performance of the detection devices required by the existing security detection methods is relatively high, and there may be problems in deployment.
[0007] (2) The existing security detection methods rely too much on existing knowledge. Not only do they need to update the data in the knowledge base in a timely manner, but more importantly, they are unable to detect new types of network attacks and new malicious programs.
[0008] (3) In the existing detection methods based on deep learning, the structure of the deep learning model is often very complex and highly dependent on the computing power of the computing device. It may be difficult to meet such high-performance data analysis on complex heterogeneous network transmission nodes. Summary of the Invention
[0009] To solve the above problems, the present invention proposes a heterogeneous network security detection method based on model interpretation. It captures heterogeneous network traffic and executable programs transmitted between nodes in each transmission node of the heterogeneous network, performs abnormal network traffic detection and malicious executable program detection, so as to ensure the security of the heterogeneous network. The present invention analyzes various numerical features in the network, uses deep learning algorithms for security analysis, and at the same time explains the contribution degree of various features to the model, performs model interpretation on the features, assigns weights beneficial to the model to them, and optimizes the accuracy of deep learning.
[0010] The technical solution of the present invention is as follows:
[0011] A heterogeneous network security detection method based on model interpretation, comprising the following steps:
[0012] Step 1, construct an abnormal network traffic analysis module to generate a network traffic weighted feature matrix;
[0013] Step 2, construct a malicious program analysis module to generate an API feature matrix;
[0014] Step 3, design a deep learning model construction module to perform real-time detection of heterogeneous network security.
[0015] Further, in the above step 1, the specific working process of the abnormal network traffic analysis module is as follows:
[0016] Step 1.1, use a packet capture tool to capture raw network traffic data from the gateway;
[0017] Step 1.2, split the raw network traffic data into multiple network flows, where each network flow contains multiple data packets for communication between two ports of different hosts in the network; each data packet is uniquely identified by a five-tuple consisting of five pieces of information: source IP, destination IP, source port, destination port, and transport layer protocol;
[0018] Step 1.3, perform feature normalization processing on the traffic features;
[0019] Step 1.4, construct the normalized traffic features into a normalized network traffic matrix, where the same columns represent the same features;
[0020] Step 1.5: Based on the model interpretation method, calculate the contribution degree of each feature column, weight the feature columns according to the contribution degree, and assign high weights to the columns with high contribution degrees.
[0021] Step 1.6: Multiply the corresponding columns of the normalized network traffic matrix by the contribution degree values, and finally rearrange the feature matrix in descending order according to the feature contribution ranking. The rearranged matrix is the network traffic weighted feature matrix.
[0022] Further, the specific process of the model interpretation method in Step 1.5 is as follows: Each feature sample corresponds to a one-dimensional vector. The feature samples in the training set are labeled. The supervised learning method is used to train all feature samples, and the accuracy rate of the network security detection and analysis results using this single feature sample is calculated. Take the one-dimensional vector of each feature sample as the input parameter of the random forest. The obtained result represents the accuracy rate of training this single feature. This result is an interpretation of the model and is the contribution degree of this single feature to the accuracy rate of the neural network. The higher this result, the better the influence of this feature on the network security detection model, and a high weight value is assigned to it. On the contrary, the lower this result, the worse the influence of this feature on the detection model, and a low weight value is assigned.
[0023] After obtaining the accuracy rate of each feature, scale these data to the interval [0.1, 1] proportionally, that is, the smallest data is recorded as 0.1, the largest data is recorded as 1, and other data are scaled in this interval. Finally, the contribution degree value corresponding to each contribution degree is obtained.
[0024] Further, the specific process of Step 2 is as follows:
[0025] Step 2.1: Obtain the API data samples of the executable program, and automatically execute the analysis program for sample analysis.
[0026] Step 2.2: Use the virtual machine to simulate the running environment, automatically analyze the program to be detected, and generate a Json analysis report.
[0027] Step 2.3: Parse the Json analysis report to obtain the API sequence.
[0028] Step 2.4: Use the Word2Vec algorithm to convert each API into a one-dimensional vector, that is, each row of API in the sequence corresponds to a row of one-dimensional vectors, so the entire API sequence corresponds to a two-dimensional matrix. After being trained by the Word2Vec algorithm, each identical API corresponds to a unique vector, and there are several APIs in an API sequence. After replacing each API with its corresponding word vector, the API feature matrix is generated.
[0029] Step 2.5: Use the model interpretation method to judge the contribution degree of each API, and only retain the APIs with large contribution degrees to obtain the final API feature matrix.
[0030] Furthermore, in the step 2.5, the model interpretation method specifically uses the TF-IDF algorithm to screen the APIs in the thesaurus; calculate the TF-IDF values of each API in each API sequence, sort the APIs in descending order of the TF-IDF values, screen out the top 2000 APIs, and then for each API sequence, only retain the APIs that exist in the top 2000, and delete all other APIs to obtain a new API sequence; count the lengths of all new API sequences, set the maximum value of the lengths as the length of the API sequence matrix, and pad the API sequences that are shorter than this length with 0s at the end, thereby constructing the final API feature matrix.
[0031] Furthermore, in the step 3, a convolutional neural network is used when constructing the deep learning model; an abnormal network traffic detection model and a malicious program detection model are respectively constructed based on the convolutional neural network; the abnormal network traffic detection model and the malicious program detection model adopt the same convolutional neural network architecture, which includes three consecutive convolutional-pooling modules, a first fully connected layer, a Dropout layer, a second fully connected layer, and a classification layer; each convolutional layer contains several convolutions; the structures of the three convolutional-pooling modules are the same, each convolutional layer is followed by a pooling layer, and the last pooling layer is connected to the first fully connected layer; the structures of the first fully connected layer and the second fully connected layer are the same;
[0032] The specific process of heterogeneous network security real-time detection is as follows:
[0033] Step 3.1: Perform feature analysis and model interpretation on the input network traffic feature data based on the process of the abnormal network traffic analysis module to obtain a network traffic weighted feature matrix; input the network traffic weighted feature matrix into the convolutional neural network architecture, and after passing through three convolutional layers, three pooling layers, a first fully connected layer, a Dropout layer, a second fully connected layer, and a classification layer, obtain the abnormal network traffic detection analysis result;
[0034] Step 3.2: Perform text embedding and model interpretation on the input executable program API data based on the process of the malicious program analysis module to obtain an API feature matrix; input the API feature matrix into the convolutional neural network architecture, and after passing through three convolutional layers, three pooling layers, a first fully connected layer, a Dropout layer, a second fully connected layer, and a classification layer, obtain the malicious program detection analysis result;
[0035] Step 3.3: If both analysis results are secure, it is determined that the current heterogeneous network transmission data is secure; otherwise, it is determined that the current heterogeneous network transmission data is insecure.
[0036] Beneficial technical effects brought by the present invention:
[0037] Compared with the prior art, the present invention completes the security detection of the heterogeneous network by detecting abnormal network traffic and malicious executable programs at the heterogeneous network data transmission nodes. The algorithm of the present invention collects data from the transmission nodes of the heterogeneous network, solving the problem in the prior art of heterogeneous network security detection methods that it is difficult to cope with complex network structures. By analyzing network data based on the deep learning algorithm, the problem in the traditional method of being difficult to detect new malicious attacks and new malicious programs due to untimely database updates is solved. Optimizing network feature data based on the model interpretation method reduces the complexity of the deep learning model and optimizes the detection efficiency. Brief Description of the Drawings
[0038] Figure 1 It is the overall architecture diagram of the heterogeneous network security detection method based on model interpretation of the present invention.
[0039] Figure 2 It is the schematic diagram of the abnormal network traffic analysis process of the present invention.
[0040] Figure 3 It is the schematic diagram of the malicious program analysis process of the present invention.
[0041] Figure 4 It is the schematic diagram of the deep neural network model of the present invention. Detailed Embodiment
[0042] The present invention will be further described in detail below in conjunction with the drawings and the detailed embodiment:
[0043] In order to solve the security detection problem in the heterogeneous network and realize the detection of abnormal network traffic and malicious programs in an automated, trustworthy and efficient manner, it is necessary to conduct research on high-precision network traffic capture and weighting algorithms, automated malicious program feature parsing algorithms, and abnormal network traffic and malicious program detection algorithms based on deep learning. Finally, relying on the above analysis results, the heterogeneous network can be evaluated for security, providing an important reference basis for solving the heterogeneous network security problem.
[0044] The heterogeneous network security detection technology based on model interpretation will build a security detection framework according to the route of data information collection → transmission file capture → feature information extraction → model interpretation analysis → deep learning analysis → test result verification, and realize the high-precision and trustworthy security detection ability.
[0045] Such as Figure 1As shown in the figure, the method of the present invention mainly includes an abnormal network traffic analysis module, a malicious program analysis module, and a deep learning model construction module. The original test data uses the network traffic data and executable file data captured from the test platform, and after being analyzed by model interpretation, artificial intelligence technology is used for analysis to obtain the test results.
[0046] 1. Abnormal network traffic analysis module; this module mainly includes network traffic capture, network traffic analysis, feature weighted analysis, and feature matrix generation;
[0047] In this module, achieving high-precision heterogeneous network traffic collection is the key to realizing heterogeneous network abnormal network traffic detection. Since network traffic data often is transmitted in the form of stream transmission, a shunt operation is adopted to preprocess the traffic data, and the stream format data needs to be converted into numerical format data features. Figure 2 The figure shows the process and main content of this module. This module mainly includes a network packet capture module, a feature extraction module, and a feature analysis module; the network packet capture module mainly uses common packet capture tools to capture the original network traffic data from the gateway and constructs a file in Pcap format; the feature extraction module mainly uses common data stream processing tools to process the original network traffic data into a network traffic feature matrix; the feature analysis module mainly uses model interpretation methods to process the network traffic feature matrix into a weighted feature matrix.
[0048] Network behavior is network traffic, and one network behavior is one network traffic. The original network traffic data is a Pcap file containing all network traffic, and they need to be first split into multiple network flows or sessions, where each network flow or session contains multiple data packets for communication between two ports of different hosts in the network. The original traffic data is converted into a form suitable for input to the deep learning model through steps such as original data segmentation, information desensitization, length fixing, and normalization.
[0049] Network data stream is defined as one or more data packets transmitted between two network addresses using a certain specific protocol. Since the original network traffic data is a file in Pcap format, it first needs to be split. The splitting methods of network traffic can be divided into different methods such as service, flow, TCP connection, session, and host. Different splitting methods will make the final form of the original data set different. In this experiment, the original traffic is split according to the method of splitting the flow. According to the definitions of network flow and network session, a five-tuple consisting of five information items: source IP, destination IP, source port, destination port, and transport layer protocol is used to uniquely identify a data packet, and the traffic packets with the same five-tuple information are divided into one flow. A network flow is composed of multiple data packets with the same five-tuple, and in a network session, the source IP, source port and destination IP, destination port can be swapped.
[0050] The original network traffic packets contain feature information of various networks. However, due to their format characteristics, they cannot be directly analyzed by deep learning. Therefore, data preprocessing of the Pcap file is required. The packets contain various types of network feature information. Numerical format features are parsed from them, and after subsequent data processing, they can be applied to deep learning for detection.
[0051] Since there are numerous protocol parameters in network packets and the numerical situations of the traffic features collected by the data acquisition module are diverse, the units and meanings of each numerical value are different, and the various network flow information is irregularly distributed. In order for the deep network model to quickly iteratively learn the feature information of the traffic packets, the feature normalization method is used to generalize the traffic features to a certain range, so that the model can stably find the global minimum during the gradient descent process. The normalization formula is:
[0052]
[0053] where, X norm is the normalized traffic feature; X is the original traffic feature; X min is the minimum value of the traffic feature; X max is the maximum value of the traffic feature;
[0054] Each column of the extracted features represents a different meaning. In deep learning, these features will be constructed into a matrix. The same column represents the same feature. The deep learning algorithm can learn the differences between abnormal traffic and normal traffic and use this as the basis for judging abnormal traffic. Based on the model interpretation method, this invention performs weighted processing on these feature columns, believing that the contribution degree of each column of features to the deep learning model is different. Some features with obvious recognition effects should be given higher weights, while features with poor recognition effects should have their weights reduced.
[0055] Random forest is an ensemble learning method that combines weak classifiers to generate a strong classifier. A decision tree is a tree structure. Each non-leaf node is a decision, the branches are the results of it, and then each leaf node represents the final predicted result. The decision tree is the weak classifier in the random forest. Each tree has a classification result, and finally these classification results can be integrated. The one with the most votes is the final prediction. In each tree in the random forest, the training samples are selected by the method of random sampling with replacement, which can ensure that the training sets of each tree are different, but the training sets of each tree also contain intersections. This can avoid the situation where the results of each tree are the same and also avoid the situation where the differences between them are too large, thus avoiding overfitting of the decision tree.
[0056] For a specified feature, each sample corresponds to a one-dimensional vector. The samples in the training set are labeled. By using the method of supervised learning to train these samples, the accuracy rate of the network security detection and analysis results using this single sample can be calculated. Taking the one-dimensional vector of each sample as the input parameter of the random forest, the obtained result represents the accuracy rate of training this single feature. This result is an explanation of the model and the contribution degree of this single feature to the accuracy rate of the neural network. The higher this result is, the better the influence of this feature on the network security detection model, and a higher weight should be assigned to it. On the contrary, the lower this result is, the worse the influence of this feature on the detection model, and a lower weight should be assigned to it.
[0057] After assigning weights to each type of feature, the features that have a greater impact on the accuracy rate of the results can play a greater advantage in the deep learning training, further improving the training results of the deep learning.
[0058] After obtaining the accuracy rate of each feature, these data are scaled to the interval [0.1, 1] proportionally, that is, the smallest data is recorded as 0.1, the largest data is recorded as 1, and other data are scaled within this interval. For some features that contribute less to the classification effect, their data are still retained, but these features are not allowed to still have a greater impact on the final training results.
[0059] Multiply the corresponding columns of the normalized feature matrix by the contribution degree values calculated here, and finally re-arrange the feature matrix from largest to smallest according to the feature contribution ranking here. This feature matrix is the input matrix for the deep learning.
[0060] The abnormal network traffic analysis module adopts a high-precision heterogeneous network traffic model interpretation method. This method captures heterogeneous network traffic information, captures network traffic from the original network flow data packets and analyzes the feature information, conducts feature analysis based on the model interpretation method, calculates the contribution degree of each type of feature to the deep learning model, believes that the contribution degree of each column of features to the deep learning model is different, weights each type of feature, and obtains high-precision traffic feature data.
[0061] 2. Malicious program analysis module; this module mainly includes virtual environment construction, API feature parsing, word vector generation, and feature matrix generation;
[0062] Since the dynamic detection of malicious programs requires running the program to be detected, in order to avoid harm to the normal working environment, the detection should be carried out in a closed environment. Advanced malicious programs may detect the environment in which they are running when they are executed. If they identify that the environment is a test or detection environment, they may not show malicious status or affect the accuracy of the data extracted by the detection software, resulting in unreliable collected data. Therefore, this experiment uses a virtual machine to simulate the execution environment of malicious programs, which can not only isolate the program to be detected from the working environment, but also provide a normal running environment for the program to be detected, so that it can fully demonstrate its functions and obtain reliable data. Figure 3 The flow and main contents of this module are demonstrated, which mainly include the feature analysis part and the word vector generation part. The specific process is: first, obtain the executable program API data sample and automatically execute the analysis program. In the feature analysis part, the virtual machine will automatically analyze the program to be detected and generate a Json analysis report. In the word vector generation part, the Json analysis report is parsed to obtain the API call information, and then the Word2Vec algorithm is used to generate word vectors, and then the feature matrix is generated.
[0063] When a program is running, whether it is benign or malicious, it will perform some operations in the system. These operations, whether they are the operations of the program itself or the operations performed by the user, may leave traces in the system. Capturing malicious program parameters is also called the process of "forensics", that is, collecting the traces left by these programs when they are running.
[0064] The API sequence when the program is running is the key path for the program to interact with the system. By analyzing the API sequence, we can understand how the program communicates with the operating system, what operations it performs, and the possible risks. Therefore, the operating system provides many services in the form of APIs, and applications also need to use APIs to use the functions provided by the operating system. These functions are defined as API calls. API is an interface specification for mutual calls between software systems. Malicious programs usually call APIs to perform attacks.
[0065] The API call sequence not only contains the name information of the API, but also its position order is a key feature. When learning deep learning, we not only pay attention to which APIs are called by the program at runtime, but also the order of each API is an important learning indicator. The order of APIs may indicate some associations among the APIs called by the program at runtime. Malicious programs may call multiple APIs in succession to attack the system at runtime, and the API order can also reflect a situation of program calls.
[0066] The API sequence is a series of data in text format, which needs to be converted into a numerical matrix in numerical form before it can be trained by deep learning. Word2Vec is an algorithm that converts words into vectors. It is a type of language model that learns semantic models in an unsupervised manner from a large amount of text. It converts words into vectors, quantifies the relationships between words, and measures the relationships between words. The method of the present invention uses the Word2Vec algorithm to convert each API into a one-dimensional vector, that is, each row of APIs in the sequence corresponds to a row of one-dimensional vectors, so the overall API sequence corresponds to a two-dimensional matrix. After being trained by Word2Vec, each identical API will correspond to a unique vector, and there are several APIs in an API sequence. After each API is replaced with its corresponding word vector, an API feature matrix can be generated.
[0067] After generating word vectors, each sample needs to be converted into a matrix form for representation. Since deep learning requires that the input size of each sample be equal, it is necessary to unify the length of the API sequence, that is, to fix the size of the feature matrix, and a fixed sequence length needs to be set uniformly. However, to ensure the training efficiency of deep learning and avoid memory overflow caused by overly large data volume, for some extremely long sequences, some methods need to be used to select and discard some of the APIs therein to ensure that the sequence length is controlled within a certain range. Therefore, it is necessary to use model interpretation methods to judge the contribution degree of each API and only retain the APIs with a larger contribution degree, which can reduce the sequence length without affecting the sequence accuracy.
[0068] The present invention uses the TF-IDF algorithm to screen the APIs in the thesaurus. The TF-IDF algorithm is a data statistical algorithm in text processing, and its purpose is to calculate the importance of a word in an article. If a certain word appears in only a small part of all documents, but appears many times in these documents, it is considered that it can well represent these documents. Conversely, if a certain word appears in the vast majority of documents, it is considered that this word has a poor representation of the documents. Its calculation method is as follows:
[0069]
[0070] Among them, tfidf i,j represents the TF-IDF value of word j in document i; n i,j represents the number of times word j appears in document i; N i represents the total number of words in document i; D represents the total number of documents; d j represents the number of documents in which word j appears.
[0071] In the present invention, the API is the word of the TF-IDF algorithm, and the API call sequence is the document of the TF-IDF algorithm.
[0072] After calculation by this formula, the TF-IDF value of each word (i.e., API) can be obtained. The APIs are sorted in descending order of TF-IDF values, and the top 2000 APIs are selected. Then, for the API sequence of each sample, only these APIs are retained, and all other APIs are deleted to obtain the final API call sequence. The lengths of all new sequences are counted, and the maximum value of the lengths is set as the length of the API sequence matrix. For the API sequences shorter than this length, 0s are appended at the end.
[0073] The malicious program analysis module adopts an automated dynamic feature analysis method for executable programs. This method provides a dynamic analysis running environment for the executable program, automatically captures various characteristic parameters during program execution from this environment, calculates the relationship dimensions between each type of feature to generate a feature matrix, and filters out high-precision features and optimizes the matrix through a model interpretation method.
[0074] 3. Deep learning model construction module; mainly including neural network construction, model parameter learning, security result generation, and model result verification;
[0075] The deep learning model uses a convolutional neural network to train the above-mentioned feature matrix after feature processing as the basis for evaluating the security of heterogeneous networks. The convolutional neural network was initially used for image processing. Compared with ordinary algorithms, it can reduce preprocessing and directly process raw data, so it is very convenient for distinguishing features. The structural diagram of the neural network model is as Figure 4 shown, and an abnormal network process detection model and a malicious program detection model are respectively constructed;
[0076] The abnormal network process detection model and the malicious program detection model adopt the same convolutional neural network architecture, mainly including three sequentially connected convolutional-pooling modules, a first fully connected layer, a Dropout layer, a second fully connected layer, and a classification layer; each convolutional layer contains several convolutions; the structures of the three convolutional-pooling modules are the same, a pooling layer is connected after each convolutional layer, and the last pooling layer is connected to the first fully connected layer; the structures of the first fully connected layer and the second fully connected layer are the same;
[0077] The specific working process of the deep learning model construction module is as follows:
[0078] The abnormal network process detection model and the malicious program detection model respectively conduct detection and analysis;
[0079] The specific working process of the abnormal network flow detection model is as follows: Based on the process of the abnormal network traffic analysis module, the input network traffic feature data is subjected to feature analysis and model interpretation to obtain a network traffic weighted feature matrix. The network traffic weighted feature matrix is divided into a training set and a validation set in an 8:2 ratio, and their labels are noted; the label refers to safe or unsafe. The network traffic weighted feature matrix is passed into a convolutional neural network architecture. After passing through three convolutional layers, three pooling layers, a first fully connected layer, a Dropout layer, a second fully connected layer, and a classification layer, the abnormal network flow detection analysis result is obtained;
[0080] The specific working process of the malicious program detection model is as follows: Based on the process of the malicious program analysis module, the input executable program API data is subjected to text embedding and model interpretation to obtain an API feature matrix. The API feature matrix is divided into a training set and a validation set in an 8:2 ratio, and their labels are noted. The API feature matrix is passed into a convolutional neural network architecture. After passing through three convolutional layers, three pooling layers, a first fully connected layer, a Dropout layer, a second fully connected layer, and a classification layer, the malicious program detection analysis result is obtained;
[0081] After the input feature data enters the deep learning model, after feature extraction is completed in the convolutional layer and the pooling layer, it undergoes feature integration through the first fully connected layer. The first fully connected layer combines the local features extracted by the convolutional layer into global features to support the final classification task. To prevent overfitting, a Dropout layer is connected after the first fully connected layer, and its parameter is set to 0.5. Then comes the second fully connected layer, which is used to further compress the feature dimension, reduce redundant information, and make the model more efficient. The classification layer uses the Softmax activation function for binary classification tasks, that is, it is divided into "safe" or "unsafe". For the training data of the deep learning model, after the parameters are updated through backpropagation, the model accuracy is automatically evaluated through the validation set data, and the model parameters are optimized to improve the model performance.
[0082] Through the automatic learning of the deep learning model for the feature parameters, the abnormal network traffic analysis result and the malicious program analysis result can be obtained respectively. If both analysis results are "safe", it is determined that the data transmitted by the current heterogeneous network is safe; otherwise, it is determined that the data transmitted by the current heterogeneous network is unsafe;
[0083] If there is abnormal network traffic in the heterogeneous network, it indicates that there may be external malicious attacks during the data transmission between nodes in the heterogeneous network, and situations such as data corruption or data leakage may occur. It is necessary to promptly investigate the possible problems. If the transmitted data is a malicious program, there may be a possibility of node anomalies, and the situation of malicious program propagation to pollute the network may occur. Therefore, abnormal network traffic detection and malicious program detection play different roles in security detection, have different functions, and achieve different functions. The two detection methods jointly create a secure environment, so they should be detected separately, and the detection results are jointly used as the basis for judging the security of the heterogeneous network.
[0084] The deep learning model construction module adopts a heterogeneous network security detection method based on deep learning. This method uses heterogeneous network traffic data and executable program feature data as inputs, builds a deep learning model based on a convolutional neural network, and combines it with the model interpretation layer to achieve high-precision heterogeneous network security prediction as the basis for judging the security of the heterogeneous network.
[0085] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the substantial scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A heterogeneous network security detection method based on model interpretation, characterized in that It includes the following steps: Step 1: Construct an abnormal network traffic analysis module to generate a network traffic weighted feature matrix; Step 2: Construct a malicious program analysis module to generate an API feature matrix; Step 3: Design a deep learning model construction module for real-time detection of heterogeneous network security.
2. The heterogeneous network security detection method based on model interpretation according to claim 1, wherein, In Step 1, the specific working process of the abnormal network traffic analysis module is as follows: Step 1.1: Use a packet capture tool to capture raw network traffic data from the gateway; Step 1.2: Split the raw network traffic data into multiple network flows, where each network flow contains multiple data packets for communication between two ports of different hosts in the network; each data packet is uniquely identified by a five-tuple consisting of five pieces of information: source IP, destination IP, source port, destination port, and transport layer protocol; Step 1.3: Perform feature normalization processing on the traffic features; Step 1.4: Construct a normalized network traffic matrix from the normalized traffic features, where the same columns represent the same features; Step 1.5: Based on the model interpretation method, calculate the contribution degree of each feature column, perform weighted processing on the feature columns according to the contribution degree, and assign a high weight to the column with a high contribution degree; Step 1.6: Multiply the corresponding columns of the normalized network traffic matrix by the contribution degree values, and finally rearrange the feature matrix from largest to smallest according to the feature contribution ranking. The rearranged matrix is the network traffic weighted feature matrix.
3. The heterogeneous network security detection method based on model interpretation according to claim 2, wherein The specific process of the model interpretation method in Step 1.5 is as follows: Each feature sample corresponds to a one-dimensional vector, and the feature samples in the training set are labeled. Use the method of supervised learning to train all feature samples, and calculate the accuracy rate of the network security detection and analysis results using this single feature sample; Take the one-dimensional vector of each feature sample as the input parameter of the random forest, and the obtained result represents the accuracy rate of training this single feature. This result is an interpretation of the model and is the contribution degree of this single feature to the accuracy rate of the neural network. The higher this result, the better the influence of this feature on the network security detection model, and a high weight value is assigned to it. On the contrary, the lower this result, the worse the influence of this feature on the detection model, and a low weight value is assigned; After obtaining the accuracy rate of each feature, scale these data to the interval [0.1, 1], that is, the smallest data is recorded as 0.1, the largest data is recorded as 1, and other data is scaled in this interval; Finally, obtain the contribution degree value corresponding to each contribution degree.
4. The heterogeneous network security detection method based on model interpretation according to claim 3, wherein The specific process of Step 2 is as follows: Step 2.1: Obtain the API data samples of the executable program, and automatically execute the analysis program for sample analysis; Step 2.2: Use a virtual machine to simulate the running environment, automatically analyze the program to be detected, and generate a Json analysis report; Step 2.3: Parse the Json analysis report to obtain the API sequence; Step 2.4: Use the Word2Vec algorithm to convert each API into a one-dimensional vector, that is, each row of APIs in the sequence corresponds to a row of one-dimensional vectors, so the entire API sequence corresponds to a two-dimensional matrix; after being trained by the Word2Vec algorithm, each identical API corresponds to a unique vector, and there are several APIs in an API sequence. After replacing each API with its corresponding word vector, an API feature matrix is generated. Step 2.5: Use the model interpretation method to judge the contribution degree of each API, and only retain the APIs with a large contribution degree to obtain the final API feature matrix.
5. The heterogeneous network security detection method based on model interpretation according to claim 4, wherein In the said Step 2.5, the model interpretation method specifically uses the TF-IDF algorithm to screen the APIs in the thesaurus; calculate the TF-IDF value of each API in each API sequence, sort the APIs in descending order according to the TF-IDF value, screen out the top 2000 APIs, and then for each API sequence, only retain the APIs that exist in the top 2000, and delete all other APIs to obtain a new API sequence; count the lengths of all new API sequences, set the maximum value of the lengths as the length of the API sequence matrix, and fill in 0 at the end for the API sequences that are less than this length, so as to construct and obtain the final API feature matrix.
6. The heterogeneous network security detection method based on model interpretation according to claim 5, wherein In the said Step 3, a convolutional neural network is used when constructing the deep learning model; an abnormal network traffic detection model and a malicious program detection model are respectively constructed based on the convolutional neural network; the abnormal network traffic detection model and the malicious program detection model adopt the same convolutional neural network architecture, which includes three sequentially connected convolutional-pooling modules, a first fully connected layer, a Dropout layer, a second fully connected layer, and a classification layer; each convolutional layer includes several convolutions; the structures of the three convolutional-pooling modules are the same, each convolutional layer is connected to a pooling layer, and the last pooling layer is connected to the first fully connected layer; the structures of the first fully connected layer and the second fully connected layer are the same. The specific process of heterogeneous network security real-time detection is as follows: Step 3.1: Conduct feature analysis and model interpretation on the input network traffic feature data based on the process of the abnormal network traffic analysis module to obtain a network traffic weighted feature matrix; input the network traffic weighted feature matrix into the convolutional neural network architecture, and after passing through three convolutional layers, three pooling layers, a first fully connected layer, a Dropout layer, a second fully connected layer, and a classification layer, obtain the abnormal network traffic detection analysis result. Step 3.2: Conduct text embedding and model interpretation on the input executable program API data based on the process of the malicious program analysis module to obtain an API feature matrix; input the API feature matrix into the convolutional neural network architecture, and after passing through three convolutional layers, three pooling layers, a first fully connected layer, a Dropout layer, a second fully connected layer, and a classification layer, obtain the malicious program detection analysis result. Step 3.3: If both analysis results are secure, it is determined that the data transmitted by the current heterogeneous network is secure; otherwise, it is determined that the data transmitted by the current heterogeneous network is insecure.