Internal Threat Detection Method, Device, Electronic Device and Storage Medium
The data enhancement module enhances and confrontation training of real malicious sequence samples to generate synthetic malicious sequences, solving the problem of sample imbalance in the internal threat detection model and improving detection accuracy.
Patent Information
- Application Number
- CN202211071868.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-08-31
AI Technical Summary
The existing internal threat detection model has low detection accuracy because the number of malicious behavior samples is much smaller than the number of normal behavior samples.
The real malicious sequence samples are enhanced through the data augmentation module, synthesized malicious sequences are generated, the number of normal sequence samples and malicious sequence samples is balanced, and the generator and sequencer are used for adversarial training until the data augmentation module converges, and the detection model is trained in combination with the semantic extraction network and the classifier.
The detection model is realized to avoid overfitting normal sequence samples and underfitting malicious sequence samples during training, which improves the detection accuracy of the detection model.
Smart Images

Figure CN115495732B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to an internal threat detection method, device, electronic device, and storage medium. Background Art
[0002] In recent years, the rapid development of information technology has brought many conveniences to society, but it has also led to an increasing number of attacks in the cyber space. The networks of organizations are not only subject to external attacks but also accompanied by threats posed by internal personnel. Since internal personnel have access rights and know the vulnerabilities of their organizations, compared with external attacks that are difficult to hide their footprints, the subtle and dynamic characteristics of internal threats make detection very difficult.
[0003] Currently, a popular and effective method for detecting internal threats is to analyze the behavior patterns of users accessing organizational resources and obtain results by analyzing the behavior patterns of users through a detection model. The existing detection model is trained by the number of normal behavior samples and malicious behavior samples. However, since the number of malicious behavior samples of internal threats is much smaller than the number of normal behavior samples, the detection accuracy of the trained detection model for user behavior patterns is relatively low. Summary of the Invention
[0004] The present invention provides an internal threat detection method, device, electronic device, and storage medium to solve the defect in the prior art that the detection accuracy of the detection model is low due to the fact that the number of malicious samples is much smaller than the number of normal samples.
[0005] The present invention provides an internal threat detection method, including:
[0006] Determine a sequence to be detected;
[0007] Input the sequence to be detected into a detection model to obtain a detection result output by the detection model;
[0008] The detection model is trained based on normal sequence samples and malicious sequence samples; the malicious sequence samples are obtained by enhancing real malicious sequence samples based on a data enhancement module; the data enhancement module is obtained by performing adversarial training using the real malicious sequence samples.
[0009] According to the internal threat detection method provided by the present invention, the training steps of the data enhancement module are as follows:
[0010] Perform data enhancement on a random noise sequence based on a generator in the data enhancement module to obtain an intermediate enhanced malicious sequence;
[0011] Based on a sorter in the data enhancement module, use a preset reference sequence to sort the intermediate enhanced malicious sequence and the real malicious sequence samples;
[0012] Adjust the parameters of the generator and the sorter based on the sorting result until the data augmentation module converges.
[0013] According to an internal threat detection method provided by the present invention, the adjusting the parameters of the generator and the sorter based on the sorting result until the data augmentation module converges includes:
[0014] When adjusting the parameters of the sorter, fix the parameters of the generator and adjust the parameters of the sorter based on the sorting result.
[0015] When adjusting the parameters of the generator, fix the parameters of the sorter and adjust the parameters of the generator based on the sorting result and by applying the Monte Carlo tree search method to calculate the future reward values of some generated sequences in the intermediate augmented malicious sequence.
[0016] The generator and the sorter perform alternating parameter adjustments until the data augmentation module converges.
[0017] According to an internal threat detection method provided by the present invention, the sorting the intermediate augmented malicious sequence and the real malicious sequence sample by applying a preset reference sequence based on the sorter in the data augmentation module includes:
[0018] In the scenario where the malicious sequence input to the sorter belongs to the real malicious sequence sample, determine the sorting score of the malicious sequence and the sorting score of the intermediate augmented malicious sequence based on the preset reference sequence, and perform sorting based on the sorting score of the malicious sequence and the sorting score of the intermediate augmented malicious sequence.
[0019] In the scenario where the malicious sequence belongs to the intermediate augmented malicious sequence, determine the sorting score of the malicious sequence and the sorting score of the real malicious sequence sample based on the preset reference sequence, and perform sorting based on the sorting score of the malicious sequence and the sorting score of the real malicious sequence sample.
[0020] According to an internal threat detection method provided by the present invention, the training steps of the detection model are as follows:
[0021] Augment the real malicious sequence sample based on the data augmentation module to obtain the malicious sequence sample.
[0022] Train the embedding layer of the detection model based on the normal sequence sample and the malicious sequence sample until the embedding layer converges; the embedding layer is a semantic extraction network.
[0023] Input the normal sequence sample into the embedding layer to obtain the semantic features of the normal sample output by the embedding layer, and input the malicious sequence sample into the embedding layer to obtain the semantic features of the malicious sample output by the embedding layer;
[0024] Based on the semantic features of the normal sample and the semantic features of the malicious sample, train the classifier of the detection model until the classifier converges to obtain the detection model.
[0025] According to an internal threat detection method provided by the present invention, the sequence to be detected is a user behavior sequence, and both the normal sequence sample and the real malicious sequence sample include multiple user behavior sequences; the user behavior sequence is obtained based on the following steps:
[0026] Extract the user behavior pattern data set from the log file based on the extraction date and the extraction user; any data in the user behavior pattern data set includes: user behavior pattern data type, occurrence time, and computer type; the computer type includes the assigned computer and other computers;
[0027] Based on the time type to which the occurrence time in any data belongs and the computer type in any data, determine the offset of any data; the time type includes working hours and non - working hours; the computer type includes the assigned computer and other computers;
[0028] Based on the user behavior pattern data type in any data, apply the offset of any data, and the mapping relationship between the user behavior pattern data type and the user behavior index to encode the user behavior pattern data of any data;
[0029] Based on the occurrence time and the user behavior pattern data encoding of each data in the user behavior pattern data set, determine the user behavior sequence of the extraction user on the extraction date.
[0030] The present invention also provides an internal threat detection device, including:
[0031] A determination module for determining the sequence to be detected;
[0032] A detection module for inputting the sequence to be detected into the detection model to obtain the detection result output by the detection model;
[0033] The detection model is trained based on normal sequence samples and malicious sequence samples; the malicious sequence samples are obtained by enhancing real malicious sequence samples by a data enhancement module; the data enhancement module is obtained by performing adversarial training on the real malicious sequence samples.
[0034] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the internal threat detection method described in any one of the above is implemented.
[0035] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the internal threat detection method described in any one of the above is implemented.
[0036] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the internal threat detection method described in any one of the above is implemented.
[0037] The internal threat detection method, device, electronic device, and storage medium provided by the present invention enhance real malicious sequence samples through a data enhancement module to obtain synthetic malicious sequences, achieving data balance between the number of normal sequence samples and malicious sequence samples for training the detection model. Thus, the problems of overfitting of normal sequence samples and underfitting of malicious sequence samples during the training process of the detection model are avoided, and the detection accuracy of the detection model is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 is a flowchart of the internal threat detection method provided by the present invention;
[0040] Figure 2 is a flowchart of the training method of the data enhancement module provided by the present invention;
[0041] Figure 3 is a flowchart of the data enhancement module provided by the present invention;
[0042] Figure 4 is a flowchart of the method for obtaining a user behavior model provided by the present invention
[0043] Figure 5 is a data index diagram of the user behavior pattern provided by the present invention;
[0044] Figure 6 is an overall architecture diagram of the training of the detection model provided by the present invention;
[0045] Figure 7It is a schematic structural diagram of the internal threat detection device provided by the present invention;
[0046] Figure 8 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] Currently, since the number of malicious behavior samples of internal threats is much smaller than the number of normal behavior samples, it will cause overfitting on normal behavior samples and underfitting on malicious behavior samples during training, resulting in a low detection accuracy of the detection model obtained after training.
[0049] Therefore, how to improve the detection accuracy of the detection model is a technical problem that needs to be solved urgently by those skilled in the art.
[0050] In view of the above technical problems, an embodiment of the present invention provides an internal threat detection method. Figure 1 It is a schematic flow diagram of the internal threat detection method provided by the present invention. As Figure 1 shown, the method includes:
[0051] Step 110, determining a sequence to be detected;
[0052] It should be noted that the sequence to be detected is a sequence of behavior pattern data for a certain user. The behavior pattern data can be the data corresponding to the user extracted from a log file within a specified time period, or the data corresponding to the user extracted from a database within a specified time period. The embodiments of the present invention do not limit this. Among them, the specified time period can be 1 day or multiple days, and the embodiments of the present invention do not limit this.
[0053] Step 120, inputting the sequence to be detected into the detection model to obtain a detection result output by the detection model;
[0054] The detection model is trained based on normal sequence samples and malicious sequence samples; the malicious sequence samples are obtained by enhancing real malicious sequence samples based on a data enhancement module; the data enhancement module is obtained by performing adversarial training using real malicious sequence samples.
[0055] Considering that internal threat detection is a behavioral pattern detection for internal personnel, and the number of normal behavioral pattern data of internal personnel is much larger than that of malicious behavioral pattern data. Therefore, in the embodiments of the present invention, a data augmentation module is used to generate malicious sequence samples to supplement the number of malicious sequence samples, so as to balance the number of normal sequence samples and malicious sequence samples.
[0056] Specifically, the data augmentation module is adversarially trained with real malicious sequence samples until the data augmentation module converges. Then, the real malicious sequence samples are input into the data augmentation module, and the data augmentation module enhances the real malicious sequence samples to generate a large number of synthetic malicious sequence samples. Then, a specified number of synthetic malicious sequence samples and real malicious sequence samples are selected to form malicious sequence samples, so that the number of malicious sequence samples is the same as or within a preset error range of the number of normal sequence samples. Then, the detection model is trained with the malicious sequence samples and normal sequence samples until the detection model converges. After the training of the detection model is completed, the sequence to be detected can be input into the detection model, and the detection model will output the detection result of the sequence to be detected.
[0057] It should be noted that the data augmentation module can be the generator in the generative adversarial model. Among them, the generator and discriminator in the generative adversarial model can be adversarially trained with real malicious sequence samples to adjust the parameters until both the generator and discriminator converge, and then the generator is used as the data augmentation module; the data augmentation module can also be a generative adversarial model with the discriminator changed to a sorter, that is, the data augmentation module can include a generator and a sorter. The generator generates synthetic malicious sequences, and the sorter sorts the synthetic malicious sequences and real malicious sequence samples. The generator aims to generate synthetic malicious sequences with a sorting score higher than that of the real malicious sequence samples, while the sorter aims to sort the synthetic malicious sequences with a sorting score lower than that of the real malicious sequence samples. The generator and the sorter are adversarially trained to adjust the parameters of the generator and the sorter until the data augmentation module converges. At this time, the data augmentation module outputs the synthetic malicious sequence samples sorted by the sorter. The embodiments of the present invention do not limit this.
[0058] The detection model is a classification model that can be directly trained with normal sequence samples and malicious sequence samples until the detection model converges. It can also be trained with normal sequence samples and malicious sequence samples relative to the embedding layer first. Then, after the embedding layer converges, the classifier is trained with the normal sequence sample features and malicious sample sequence features output by the embedding layer until the classifier converges. The embedding layer and the classifier then form the detection model. The embodiments of the present invention do not limit this. Among them, the normal sequence sample features and malicious sample sequence features can be word vectors or word vectors with context semantics. The embodiments of the present invention do not limit this. In addition, the determination methods of the behavior pattern data in the real malicious sequence sample and the behavior pattern data in the sequence to be detected are the same.
[0059] The internal threat detection method provided by the embodiments of the present invention enhances the real malicious sequence samples through a data enhancement module to obtain synthetic malicious sequences, making the data balance between the number of normal sequence samples and malicious sequence samples for training the detection model. This avoids the problems of overfitting of normal sequence samples and underfitting of malicious sequence samples during the training process of the detection model, thereby improving the detection accuracy of the detection model.
[0060] Based on the above embodiments, Figure 2 is a schematic flowchart of the training method of the data enhancement module provided by the present invention. As Figure 2 shown, the training steps of the data enhancement module are as follows:
[0061] Step 210: Perform data enhancement on the random noise sequence based on the generator in the data enhancement module to obtain an intermediate enhanced malicious sequence;
[0062] Step 220: Based on the sorter in the data enhancement module, apply a preset reference sequence to sort the intermediate enhanced malicious sequence and the real malicious sequence samples;
[0063] Step 230: Adjust the parameters of the generator and the sorter based on the sorting result until the data enhancement module converges.
[0064] Considering that the real malicious sequence samples are discrete sequence data and the standard generative adversarial model has limitations for discrete sequence data, the discriminator can only determine whether it is a real sequence and cannot score discrete sequence data well, that is, the discriminator cannot pass the gradient to the generator, making it difficult to train the generator. Therefore, in the embodiments of the present invention, the discriminator in the generative adversarial model is changed to a sorter so that discrete sequence data can be scored.
[0065] Specifically, a random noise sequence is input into the generator in the data augmentation module for data augmentation to obtain an intermediate augmented malicious sequence; the intermediate augmented malicious sequence and the real malicious sequence samples are used as the inputs of the sorter in the data augmentation module, and the sorter sorts the intermediate augmented malicious sequence and the real malicious sequence samples according to a preset reference sequence to obtain a sorting result. In the initial stage of training, the augmentation ability of the generator is poor, and the augmented malicious sequence does not conform to the distribution of the real malicious sequence. After the random noise sequence passes through the generator, the output intermediate augmented malicious sequence and the real malicious sequence samples are input into the sorter, and the sorter can easily rank the intermediate augmented malicious sequence lower than the real malicious sequence samples. The parameters of the generator and the sorter are adjusted according to the sorting result, and the random noise sequence is iteratively augmented based on the adjusted generator, and the intermediate augmented malicious sequence and the real malicious sequence samples are sorted based on the adjusted sorter. As the training progresses, the augmentation ability of the generator becomes stronger and stronger, and the generated intermediate augmented malicious sequence becomes more and more in line with the distribution of the real malicious sequence, and it is difficult for the sorter to sort the intermediate augmented malicious sequence and the real malicious sequence samples. When the data augmentation module converges, the probability that the sorting result output by the sorter ranks the intermediate augmented malicious sequence lower than the real malicious sequence samples is close to the probability that the sorting of the intermediate augmented malicious sequence is higher than the real malicious sequence samples. The final intermediate augmented malicious sequence output by the generator in the data augmentation module converges to be close to the distribution of the real malicious sequence.
[0066] It should be noted that the generator in the data augmentation module aims to generate an intermediate augmented malicious sequence with a sorting score higher than that of the real malicious sequence samples, and the sorter in the data augmentation module aims to sort with the sorting score of the intermediate augmented malicious sequence output by the generator being lower than that of the real malicious sequence samples. The policy gradient is used during the training process of the data augmentation module to solve the problem that the numerical gradient cannot be updated.
[0067] Based on the above embodiments, step 230 includes:
[0068] Step 231, when adjusting the parameters of the sorter, fix the parameters of the generator and adjust the parameters of the sorter based on the sorting result;
[0069] When adjusting the parameters of the generator, fix the parameters of the sorter, and based on the sorting result and applying the Monte Carlo tree search method to calculate the future reward values of some generated sequences in the intermediate augmented malicious sequence, adjust the parameters of the generator;
[0070] Step 232, the generator and the sorter perform alternating parameter adjustments until the data augmentation module converges.
[0071] Considering that it is difficult for the generator in the data augmentation module to smoothly receive and update gradients, therefore, the embodiments of the present invention introduce the Monte Carlo tree search method to calculate the reward value of the intermediate augmented malicious sequence generated by the generator, and adjust the parameters of the generator in combination with the sorting result, so as to solve the problem that the generator is difficult to smoothly receive and update gradients.
[0072] Specifically, the generator and the sorter perform alternating parameter adjustments, that is, when adjusting the parameters of the sorter, the parameters of the generator are fixed, the generator generates an intermediate augmented malicious sequence, the sorter sorts the intermediate augmented malicious sequence and the real malicious sequence samples, and adjusts the parameters of the sorter according to the sorting result; when adjusting the parameters of the generator, the Monte Carlo tree search method is applied to calculate the future reward value of some generated sequences in the intermediate augmented malicious sequence generated by the generator, and in combination with the sorting result of the intermediate augmented malicious sequence output by the sorter, the parameters of the generator are adjusted until the data augmentation module converges.
[0073] It should be noted that the formula for calculating the future reward value F of some generated sequences in the intermediate augmented malicious sequence is:
[0074]
[0075] In the formula, is the expectation operator, θ is the parameter of the generator, φ is the parameter of the sorter, s r is the complete sequence simulated using the Monte Carlo tree search method based on some generated sequences s in the intermediate augmented malicious sequence 0:t-1 t is the length of the intermediate augmented malicious sequence, s r can be regarded as a path synthesized according to the current policy, U is the preset reference sequence, G θ is the generator, R φ is the sorter, C + is composed of real malicious sequence samples. Simulate n paths with corresponding sorting scores. And use the average sorting score of the n paths as the future reward value F of some generated sequences in the intermediate augmented malicious sequence.
[0076] Based on the above embodiments, step 220 includes:
[0077] In the scenario where the malicious sequence input by the sorter belongs to the real malicious sequence sample, based on the preset reference sequence, determine the sorting score of the malicious sequence and the sorting score of the intermediate augmented malicious sequence, and perform sorting based on the sorting score of the malicious sequence and the sorting score of the intermediate augmented malicious sequence;
[0078] In the scenario where the malicious sequence belongs to the intermediate enhanced malicious sequence, based on the preset reference sequence, determine the sorting score of the malicious sequence and the sorting score of the true malicious sequence sample, and perform sorting based on the sorting score of the malicious sequence and the sorting score of the true malicious sequence sample.
[0079] Considering that when the sorter sorts, the input malicious sequence can be the intermediate enhanced malicious sequence generated by the generator or the true malicious sequence sample, therefore, in the embodiments of the present invention, by distinguishing the type of the input malicious sequence and selecting the corresponding comparison sequence set for comparison and sorting, the deviation of the sorting result caused by comparing the same malicious sequence type is reduced, and the sorting result output by the sorter can be made more accurate.
[0080] Specifically, in the scenario where the malicious sequence input to the sorter belongs to the true malicious sequence sample, that is, when the malicious sequence type is the true sequence, use the intermediate enhanced malicious sequence as the comparison set, apply the preset reference sequence, score the input malicious sequence respectively to obtain the sorting score of the input malicious sequence, and score the intermediate enhanced malicious sequence to obtain the sorting score of the intermediate enhanced malicious sequence, and finally sort according to the sorting score of the malicious sequence and the sorting score of the intermediate enhanced malicious sequence. In the scenario where the malicious sequence input to the sorter belongs to the intermediate enhanced malicious sequence, that is, when the malicious sequence type is the synthetic malicious sequence, use the true malicious sequence sample as the comparison set, apply the preset reference sequence, score the input malicious sequence respectively to obtain the sorting score of the input malicious sequence, and score the true malicious sequence sample to obtain the sorting score of the true malicious sequence sample, and finally sort according to the sorting score of the malicious sequence and the sorting score of the true malicious sequence sample.
[0081] It should be noted that Figure 3 is the schematic flowchart of the data enhancement module provided by the present invention. As Figure 3 shown in, R represents the true malicious sequence sample, the generator G θ generates the intermediate enhanced malicious sequence, and the sorter R φ sorts according to the sorting score of the input malicious sequence and the preset reference sequence U, and the sorting score of the true malicious sequence sample and the preset reference sequence U. G θ and R φ perform a minimax game, and the mathematical formula of the objective function is:
[0082]
[0083] In the formula, is the expectation operator, ρ r is the set of real-world sequences, and U is the preset reference sequence. If the input malicious sequence belongs to the true malicious sequence sample, the comparison set C -It consists of intermediate enhanced malicious sequences. If the input malicious sequence belongs to the intermediate enhanced malicious sequence, compare set C + It consists of real malicious sequence samples.
[0084] Based on any of the above embodiments, the training steps of the detection model are as follows:
[0085] Step S1, based on the data augmentation module, augment the real malicious sequence samples to obtain malicious sequence samples;
[0086] Step S2, based on the normal sequence samples and the malicious sequence samples, train the embedding layer of the detection model until the embedding layer converges; the embedding layer is a semantic extraction network;
[0087] Step S3, input the normal sequence samples into the embedding layer to obtain the normal sample semantic features output by the embedding layer, and input the malicious sequence samples into the embedding layer to obtain the malicious sample semantic features output by the embedding layer;
[0088] Step S4, based on the normal sample semantic features and the malicious sample semantic features, train the classifier of the detection model until the classifier converges to obtain the detection model.
[0089] Considering that in order to obtain the correlation between the behavioral pattern data in the normal sequence samples and the malicious sequence samples, therefore, the embodiment of the present invention uses an embedding layer based on a semantic extraction network. At the same time, in order to reduce the training difficulty of the detection model, the embedding layer is trained first, and then the classifier is trained according to the trained embedding layer.
[0090] Specifically, first, the real malicious sequence samples are augmented through the data augmentation module to obtain malicious sequence samples that are balanced with the number of normal sequence samples. Then, the embedding layer of the detection model is trained with a sequence sample set with an equalized class distribution composed of normal sequence samples and malicious sequence samples until the embedding layer converges; the normal sequence samples are input into the embedding layer, and after passing through an input layer, a hidden layer, and an output layer in the embedding layer, the normal sample semantic features are obtained; the malicious sequence samples are input into the embedding layer, and after passing through an input layer, a hidden layer, and an output layer in the shallow layer, the malicious sample semantic features are obtained, and the classifier of the detection model is trained through the normal sample semantic features and the malicious sample semantic features until the classifier converges to obtain the detection model.
[0091] It should be noted that the embedding layer is the continuous bag-of-words model (CBOW) of Word2Vec. Assuming that the normal sequence samples and the malicious sequence samples are sequences containing n tokens, denoted as TS=(w1, w2, …, w n ), the embedding layer takes maximizing the probability as the training objective, and the probability formula is:
[0092]
[0093] In the formula, C w represents the set of tokens around w, and W w represents the weight matrix connecting the hidden layer and the output layer (softmax layer). is the sum of the vectors of the tokens around the target. Therefore, the objective function of the embedding layer is:
[0094]
[0095] In the formula, C w represents the set of tokens around w, and W w represents the weight matrix connecting the hidden layer and the output layer (softmax layer). is the sum of the vectors of the tokens around the target.
[0096] The embedding layer maps each token in each sequence in the sequence sample set to a unique N - tuple vector. Assume the sequence sample set is represented by where, represents the user behavior sequence of user k on the j - th day, represents the single - user behavior pattern data at time t, and u k,j represents the j - th day of user k, and T is the length of each user behavior sequence. Therefore, for each input user behavior sequence, the embedding layer outputs a matrix M T×N with a fixed size, where T is the number of rows and N is the number of columns.
[0097] The classifier consists of a one - dimensional CNN (Convolutional Neural Network), a Bi - LSTM (Bidirectional Long Short - Term Memory Network), an attention mechanism layer, and a SoftMax layer. The role of the one - dimensional CNN is to reduce the input dimension and extract the local abstract features of the pattern sequence; the Bi - LSTM captures the temporal semantics of user behavior and models the user's resource access profile; the attention mechanism layer filters out trivial information and assigns higher weights to key behaviors. At the beginning of the classifier, a one - dimensional convolutional layer and a one - dimensional pooling layer are used to extract abstract features, reduce the size of the feature matrix, and reduce the computational complexity. The input of the one - dimensional convolutional layer is M T×N and the output is M T′×N (T′ < T), and the input of the one - dimensional pooling layer is M T′×N and the output is M T″×N (T″ < T′ < T).
[0098] The Bi-LSTM trains two LSTM models, one for the input sequence and the other for the reverse copy of the input sequence, that is, forward feedback and backward feedback, which will increase the reverse information of the sequence and bring more sufficient training. The mathematical expression is as follows:
[0099] i t =σ(W i e t +U i h t-1 +b i ),
[0100] f t =σ(W f e t +U f h t-1 +b f )
[0101] o t =σ(W o e t +U o h t-1 +b o )
[0102] g t =tanh(W g e t +U g h t-1 +b g )
[0103] c t =i t ·g t +f t ·c t-1
[0104]
[0105] In the formula, W i , W f , W o , W g , U i , U f , U o , U g are weight matrices, b i , b f , b o , b g are biases, which are continuously updated during the classifier training. σ, i t , f t , o t , g trepresents the sigmoid function, input gate, forget gate, output gate and hidden representation respectively, et is the semantic feature of the user behavior pattern data at time t, h t is the hidden layer state at time t, h t-1 is the hidden layer state at time t-1, → represents the forward direction, ← represents the reverse direction, and Bi-LSTM is based on M T″×N Each row is taken as input and outputs a series of hidden states that reflect the user's behavior pattern data.
[0106] The input of the attention mechanism layer is the hidden state of each user’s behavior pattern data 1≤t≤T″, the output is its hidden representation Then calculate and The similarity of the randomly initialized context vector is used to measure the importance weight of each user's behavior pattern data. The normalized behavior weight is calculated by the SoftMax function. Therefore, the hidden state is calculated. The semantic feature v of the weighted sum of the user behavior pattern data can be used to understand which user behavior pattern data of the user contributes more to the internal threat. Finally, v is input to the SoftMax layer for classification, and the SoftMax layer outputs the classification probability of each class.
[0107] Based on the above embodiments, Figure 4 FIG. 1 is a flow chart of the method for obtaining a user behavior model provided by the present invention. Figure 4 As shown, the sequence to be detected is user behavior pattern data, normal sequence samples and real malicious sequence samples all include multiple user behavior pattern data; the user behavior sequence is obtained based on the following steps:
[0108] Step 410, extracting a user behavior pattern data set from the log file based on the extraction date and the extraction user; any data in the user behavior pattern data set includes: user behavior pattern data type, occurrence time and computer type; computer type includes the assigned computer and other computers;
[0109] Step 420, based on the time type to which the occurrence time in the data belongs and the computer type in the data, determine the offset of the data; the time type includes working time and non-working time; the computer type includes the assigned computer and other computers;
[0110] Step 430, encoding user behavior pattern data of the data based on the data type in the data, applying the offset of the data, and the mapping relationship between the data type and the user behavior index;
[0111] Step 440: Determine the user behavior sequence of the user on the extraction date based on the occurrence time of each data in the user behavior pattern data set and the user behavior pattern data encoding.
[0112] Specifically, the user behavior sequence takes the extraction date and the extracted user as the time range of each day, that is, the time range from 0:00:00 of a certain day to 23:59:59 of this day, and extracts the user behavior pattern data set from the date file. Among them, any data in the user behavior pattern data set includes the user behavior pattern data type, the occurrence time, and the computer type; the computer type includes the assigned computer and other computers. A specified user will have his own assigned computer. If the operation is performed on the assigned computer, the computer type is the assigned computer, otherwise it is other computers.
[0113] Then, according to the time type of the occurrence time in the data and the computer type in the data, determine the offset of the data. Among them, the time type includes working hours and non-working hours. It should be noted that the working hours and non-working hours can be set by the management personnel, and the embodiments of the present invention do not make specific limitations.
[0114] Then, according to the user behavior pattern data type in the data, apply the offset of the data, and the mapping relationship between the user behavior pattern data type and the user behavior index preset in advance, to encode the user behavior pattern data of the data.
[0115] Finally, based on the occurrence time of each data in the user behavior pattern data set and the user behavior pattern data encoding, determine the user behavior sequence of the user on the extraction date.
[0116] It should be noted that considering that the same resource access behavior will show different attributes under different time factors, the embodiments of the present invention set four offsets to represent spatio-temporal offsets, including (i) operating on the assigned computer during working hours (ii) operating on other computers during working hours (iii) operating on the assigned computer during non-working hours (iv) operating on other computers during non-working hours. Referring to the modeling method of natural language, a user behavior pattern data corresponds to a single character, and a user behavior sequence corresponds to a sentence. The log file can be the log file in the organization. Assume that U = {u1, u2,..., u k} is a set of k users. For a user u k (1 ≤ k ≤ K), the user behavior sequence set of this user from the 1st day to the Jth day can be obtained, and is represented by represents the set of sequence user behavior sequences of user k from the 1st day to the Jth day, Denote the user behavior sequence of user k on the j-th day. The labels are assigned according to facts. The malicious sequence is denoted as S′, and the normal sequence is denoted as S″, where S′∪S″=S. Among them, the log file can include login log files, device operation log files, file operation log files, website access log files, and email operation log files.
[0117] Figure 5 is the user behavior pattern data index diagram provided by the present invention. As Figure 5 shown, the working hours are from 8 am to 8 pm, and the non-working hours are from 8 pm to 8 am the next day. Set the offset of the time type as working hours and the computer type as the assigned computer to 0, set the offset of the time type as working hours and the computer type as other computers to 1, set the offset of the time type as non-working hours and the computer type as the assigned computer to 2, set the offset of the non-time type as working hours and the computer type as other computers to 3. At the same time, the user behavior pattern data type can be divided into login type (corresponding to the Logon.csv login log file), device type (corresponding to the Device.csv device operation log file), file type (corresponding to the File.csv file operation log file), website type (corresponding to the Http.csv website access log file), and email type (corresponding to the Email.csv email operation log file). The login type includes login (the corresponding code (operation coefficient) is 1) and logout (the corresponding code (operation coefficient) is 5); the device type includes connection (the corresponding code (operation coefficient) is 9) and disconnection (the corresponding code (operation coefficient) is 13); the file type includes executable file (the corresponding code (operation coefficient) is 17), text file (the corresponding code (operation coefficient) is 29), Word file (the corresponding code (operation coefficient) is 27), picture file (the corresponding code (operation coefficient) is 33), PDF file (the corresponding code (operation coefficient) is 25), and compressed file (the corresponding code (operation coefficient) is 37); the website type includes neutral website (the corresponding code (operation coefficient) is 41), hacker website (the corresponding code (operation coefficient) is 45), cloud storage (the corresponding code (operation coefficient) is 49), and job hunting website (the corresponding code (operation coefficient) is 53); the email type includes internal to internal (the corresponding code (operation coefficient) is 57), external to external (the corresponding code (operation coefficient) is 61), external to internal (the corresponding code (operation coefficient) is 65), and internal to external (the corresponding code (operation coefficient) is 69). Therefore, the encoding of the user behavior pattern data can be obtained based on the following process:
[0118] Step S10, extract a line of user behavior pattern data from the log
[0119] Step S20: If the occurrence time in the user behavior pattern data is within working hours and the computer type in the user behavior pattern data is the assigned computer, then the offset 0 is added to the encoding corresponding to the user behavior pattern data type;
[0120] Step S30: If the occurrence time in the user behavior pattern data is within working hours and the computer type in the user behavior pattern data is other computers, then the offset 1 is added to the encoding corresponding to the user behavior pattern data type;
[0121] Step S40: If the occurrence time in the user behavior pattern data is outside working hours and the computer type in the user behavior pattern data is the assigned computer, then the offset 3 is added to the encoding corresponding to the user behavior pattern data type;
[0122] Step S50: If the occurrence time in the user behavior pattern data is outside working hours and the computer type in the user behavior pattern data is other computers, then the offset 4 is added to the encoding corresponding to the user behavior pattern data type;
[0123] Step S50: Determine whether the log has been completely read. If not, execute Step S10; otherwise, execute S60.
[0124] Step S60: Return the user behavior sequence.
[0125] In addition, the sequence to be detected is the user behavior data of the user to be detected on the day of the date to be detected; the normal sequence sample and the real malicious sequence sample can be obtained based on the following steps: First, obtain the set of user behavior sequences of all users for consecutive days within a certain time period from the log, and then mark the normal sequence samples and the real malicious sequence samples in the set of user behavior sequences.
[0126] Figure 6 It is the overall architecture diagram of the detection model training provided by the present invention. As Figure 6As shown in the figure, the training architecture of the detection model includes a pattern extraction module, a data augmentation module, an embedding layer, and a classifier. Among them, the pattern extraction module is used to extract the user behavior sequence of each user every day from the log file to obtain a set of user behavior sequences within a specified time range, and label the real malicious sequence samples (threat patterns) and normal sequence samples (normal patterns) from the data behavior sequence set. The data augmentation module is used to enhance the real malicious sequence samples through an advanced generative adversarial network to obtain malicious sequence samples, and train the normal sequence samples and malicious sequence samples on the embedding layer until the embedding layer converges. Then, the normal sequence samples and malicious sequence samples are input into the trained embedding layer. After passing through an input layer, a hidden layer, and an output layer in the embedding layer, the normal sample semantic features and malicious sample semantic features are obtained. Then, the classifier is trained according to the normal sample semantic features and malicious sample semantic features until the classifier converges. The classifier includes a 1D CNN layer for reducing the dimension of the input features, a Bi-LSTM model for capturing the temporal semantics of user behavior, an attention mechanism layer for filtering out trivial information and assigning higher weights to key behaviors, and a SoftMax layer for outputting the classification probability of each class.
[0127] The internal threat detection device provided by the present invention will be described below. The internal threat detection device described below can be correspondingly referred to the internal threat detection method described above.
[0128] Figure 7 is a schematic structural diagram of the internal threat detection device provided by the present invention. As Figure 7 shown, the device includes: a determination module 710 and a detection module 720.
[0129] Among them,
[0130] The determination module 710 is used to determine the sequence to be detected;
[0131] The detection module 720 is used to input the sequence to be detected into the detection model to obtain the detection result output by the detection model;
[0132] The detection model is trained based on normal sequence samples and malicious sequence samples; the malicious sequence samples are obtained by enhancing the real malicious sequence samples based on the data augmentation module; the data augmentation module is obtained by applying adversarial training to the real malicious sequence samples.
[0133] The internal threat detection device provided by the embodiment of the present invention includes a determination module for determining a sequence to be detected, a detection module for inputting the sequence to be detected into a detection model to obtain a detection result output by the detection model. The detection model is trained based on normal sequence samples and malicious sequence samples. The malicious sequence samples are obtained by enhancing real malicious sequence samples based on a data enhancement module. The data enhancement module is obtained by performing adversarial training using real malicious sequence samples, achieving data balance between the number of normal sequence samples and malicious sequence samples for training the detection model, thereby avoiding the problems of overfitting of normal sequence samples and underfitting of malicious sequence samples during the training of the detection model, and further improving the detection accuracy of the detection model.
[0134] Based on any of the above embodiments, the detection module 720 includes a data enhancement module training sub-module, which includes:
[0135] An enhancement sub-module for enhancing a random noise sequence based on a generator in the data enhancement module to obtain an intermediate enhanced malicious sequence;
[0136] A sorting sub-module for sorting the intermediate enhanced malicious sequence and real malicious sequence samples based on a sorter in the data enhancement module and applying a preset reference sequence;
[0137] An adjustment sub-module for adjusting the parameters of the generator and sorter based on the sorting result until the data enhancement module converges.
[0138] Based on any of the above embodiments, the adjustment sub-module is specifically configured to:
[0139] When adjusting the parameters of the sorter, fix the parameters of the generator and adjust the parameters of the sorter based on the sorting result;
[0140] When adjusting the parameters of the generator, fix the parameters of the sorter and adjust the parameters of the generator based on the sorting result and by applying the Monte Carlo tree search method to calculate the future reward values of some generated sequences in the intermediate enhanced malicious sequence;
[0141] Alternately adjust the parameters of the generator and sorter until the data enhancement module converges.
[0142] Based on any of the above embodiments, the sorting sub-module is specifically configured to:
[0143] In a scenario where the malicious sequence input to the sorter belongs to real malicious sequence samples, determine the sorting scores of the malicious sequence and the intermediate enhanced malicious sequence based on the preset reference sequence, and perform sorting based on the sorting scores of the malicious sequence and the intermediate enhanced malicious sequence;
[0144] In the scenario where the malicious sequence belongs to the intermediate enhanced malicious sequence, based on a preset reference sequence, determine the sorting scores of the malicious sequence and the true malicious sequence sample, and perform sorting based on the sorting scores of the malicious sequence and the true malicious sequence sample.
[0145] Based on any of the above embodiments, the detection module 720 includes a detection model training sub-module, which includes:
[0146] A sample enhancement sub-module for enhancing the true malicious sequence sample based on the data enhancement module to obtain a malicious sequence sample;
[0147] An embedding layer training sub-module for training the embedding layer of the detection model based on the normal sequence sample and the malicious sequence sample until the embedding layer converges; the embedding layer is a semantic extraction network;
[0148] A semantic extraction sub-module for inputting the normal sequence sample into the embedding layer to obtain the normal sample semantic features output by the embedding layer, and inputting the malicious sequence sample into the embedding layer to obtain the malicious sample semantic features output by the embedding layer;
[0149] A classifier training sub-module for training the classifier of the detection model based on the normal sample semantic features and the malicious sample semantic features until the classifier converges to obtain the detection model.
[0150] Based on any of the above embodiments, the internal threat detection device further includes a pattern extraction module, which includes:
[0151] A log reading sub-module for extracting a set of user behavior pattern data from the log file based on the extraction date and the extraction user; any data in the set of user behavior pattern data includes: user behavior pattern data type, occurrence time, and computer type; the computer type includes the assigned computer and other computers;
[0152] An offset sub-module for determining the offset of the data based on the time type to which the occurrence time in the data belongs and the computer type in the data; the time type includes working hours and non-working hours; the computer type includes the assigned computer and other computers;
[0153] An encoding sub-module for encoding the user behavior pattern data of the data based on the user behavior pattern data type in the data, applying the offset of the data, and the mapping relationship between the user behavior pattern data type and the user behavior index;
[0154] A sequence determination sub-module for determining the user behavior sequence of the extraction user on the extraction date based on the occurrence time and the user behavior pattern data encoding of each data in the set of user behavior pattern data.
[0155] Figure 8 Illustrates a schematic diagram of the physical structure of an electronic device, as Figure 8 shown. The electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute an internal threat detection method, which includes: determining a sequence to be detected; inputting the sequence to be detected into a detection model to obtain a detection result output by the detection model; the detection model is trained based on normal sequence samples and malicious sequence samples; the malicious sequence samples are obtained by enhancing real malicious sequence samples based on a data enhancement module; the data enhancement module is obtained by performing adversarial training using real malicious sequence samples.
[0156] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, and other various media that can store program codes.
[0157] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the internal threat detection method provided by the above-mentioned various methods. The method includes: determining a sequence to be detected; inputting the sequence to be detected into a detection model to obtain a detection result output by the detection model; the detection model is trained based on normal sequence samples and malicious sequence samples; the malicious sequence samples are obtained by enhancing real malicious sequence samples based on a data enhancement module; the data enhancement module is obtained by performing adversarial training using real malicious sequence samples.
[0158] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the internal threat detection method provided by the above-mentioned various methods. The method includes: determining a sequence to be detected; inputting the sequence to be detected into a detection model to obtain a detection result output by the detection model; the detection model is trained based on normal sequence samples and malicious sequence samples; the malicious sequence samples are obtained by enhancing real malicious sequence samples based on a data enhancement module; and the data enhancement module is obtained by performing adversarial training using real malicious sequence samples.
[0159] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An internal threat detection method, characterized in that, Including: Determine the sequence to be detected; Input the sequence to be detected into the detection model to obtain the detection result output by the detection model; The detection model is trained based on normal sequence samples and malicious sequence samples; The malicious sequence samples are obtained by enhancing real malicious sequence samples based on a data enhancement module; the data enhancement module is obtained by performing adversarial training using the real malicious sequence samples; The training steps of the data enhancement module are as follows: Based on the generator in the data enhancement module, perform data enhancement on a random noise sequence to obtain an intermediate enhanced malicious sequence; Based on the sorter in the data enhancement module, apply a preset reference sequence to sort the intermediate enhanced malicious sequence and the real malicious sequence samples; Adjust the parameters of the generator and the sorter based on the sorting result until the data enhancement module converges.
2. The internal threat detection method according to claim 1, wherein The adjusting the parameters of the generator and the sorter based on the sorting result until the data enhancement module converges includes: When adjusting the parameters of the sorter, fix the parameters of the generator and adjust the parameters of the sorter based on the sorting result; When adjusting the parameters of the generator, fix the parameters of the sorter and adjust the parameters of the generator based on the sorting result and by applying the Monte Carlo tree search method to calculate the future reward values of some generated sequences in the intermediate enhanced malicious sequence; The generator and the sorter perform alternating parameter adjustments until the data enhancement module converges.
3. The internal threat detection method according to claim 1, wherein The applying a preset reference sequence to sort the intermediate enhanced malicious sequence and the real malicious sequence samples based on the sorter in the data enhancement module includes: In the scenario where the malicious sequence input to the sorter belongs to the real malicious sequence samples, based on the preset reference sequence, determine the sorting score of the malicious sequence and the sorting score of the intermediate enhanced malicious sequence, and perform sorting based on the sorting score of the malicious sequence and the sorting score of the intermediate enhanced malicious sequence; In the scenario where the malicious sequence belongs to the intermediate enhanced malicious sequence, based on the preset reference sequence, determine the sorting score of the malicious sequence and the sorting score of the real malicious sequence samples, and perform sorting based on the sorting score of the malicious sequence and the sorting score of the real malicious sequence samples.
4. The internal threat detection method according to any one of claims 1 to 3, characterized in that The training steps of the detection model are as follows: Based on the data enhancement module, enhance the real malicious sequence samples to obtain the malicious sequence samples; Based on the normal sequence samples and the malicious sequence samples, train the embedding layer of the detection model until the embedding layer converges; The embedding layer is a semantic extraction network; Input the normal sequence samples into the embedding layer to obtain the normal sample semantic features output by the embedding layer, and input the malicious sequence samples into the embedding layer to obtain the malicious sample semantic features output by the embedding layer; Based on the semantic features of the normal samples and the semantic features of the malicious samples, train the classifier of the detection model until the classifier converges to obtain the detection model.
5. The internal threat detection method according to claim 1, wherein The sequence to be detected is a user behavior sequence, and both the normal sequence samples and the real malicious sequence samples include multiple user behavior sequences; the user behavior sequence is obtained based on the following steps: Extract a set of user behavior pattern data from the log file based on the extraction date and the extracted user; any data in the set of user behavior pattern data includes: the type of user behavior pattern data, the occurrence time, and the computer type; the computer type includes the assigned computer and other computers. Based on the time type to which the occurrence time in any data belongs and the computer type in any data, determine the offset of any data; the time type includes working hours and non-working hours; the computer type includes the assigned computer and other computers. Based on the type of user behavior pattern data in any data, apply the offset of any data, and the mapping relationship between the type of user behavior pattern data and the user behavior index, to encode the user behavior pattern data of any data. Based on the occurrence time and the encoded user behavior pattern data of each data in the set of user behavior pattern data, determine the user behavior sequence of the extracted user on the extraction date.
6. An internal threat detection device, characterized in that, Including: A determination module, configured to determine the sequence to be detected; A detection module, configured to input the sequence to be detected into the detection model to obtain a detection result output by the detection model; The detection model is trained based on normal sequence samples and malicious sequence samples; The malicious sequence samples are obtained by enhancing real malicious sequence samples through a data enhancement module; the data enhancement module is obtained by performing adversarial training using the real malicious sequence samples. The training steps of the data enhancement module are as follows: Based on the generator in the data enhancement module, perform data enhancement on a random noise sequence to obtain an intermediate enhanced malicious sequence; Based on the sorter in the data enhancement module, apply a preset reference sequence to sort the intermediate enhanced malicious sequence and the real malicious sequence samples; Adjust the parameters of the generator and the sorter based on the sorting result until the data enhancement module converges.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the internal threat detection method according to any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the internal threat detection method according to any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the internal threat detection method according to any one of claims 1 to 5.