Sensitive application identification and flow data processing method capable of resisting concept drift
By using the ET-BERT model and a small sample vector database, the problem of difficult to identify the source of network abnormal traffic data in the prior art is solved, and the accurate identification and supervision of sensitive applications is achieved, and the ability to resist concept drift is achieved.
Patent Information
- Application Number
- CN202510257917.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to identify the source of network abnormal traffic data, especially in the case of small samples of sensitive application types, and it is impossible to effectively supervise unknown traffic session samples.
The ET-BERT model is used for sensitive application recognition, and the token sequence is generated by collecting multiple traffic session samples, secondary pre-training is performed, and the small sample vector database is regularly updated to achieve sensitive application recognition that resists concept drift.
It improves the accuracy of the ET-BERT model in identifying sensitive applications, and can identify sensitive application types in the open world under a small sample training set, achieving effective identification and supervision of unknown traffic session samples.
Smart Images

Figure CN120086373A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sensitive application identification and data monitoring, and in particular, to a method for identifying sensitive applications resistant to concept drift and processing traffic data. Background Art
[0002] Artificial intelligence technology is a very good tool for ensuring the security of sensitive information. Due to its ability to quickly process data and perform predictive analysis, artificial intelligence is widely used in fields such as information security. In fact, ensuring data security is also an actual application of current artificial intelligence technology, while there are also hackers using artificial intelligence technology for attack activities. In the past, when people discovered abnormal network traffic data (such as pornographic, fraud-related, cybercrime, etc.), they could only intercept and block the abnormal traffic data, but did not know which application it originated from and could not supervise it from the source. Especially when the number of known sensitive application type samples is small, it is more difficult to identify the source of sensitive applications to which unknown traffic session samples belong. Summary of the Invention
[0003] In view of the above analysis, embodiments of the present invention aim to provide a method for identifying sensitive applications resistant to concept drift and processing traffic data to solve the problem of inconvenient supervision caused by the difficulty in identifying which application the abnormal traffic data originates from when discovering abnormal network traffic data in the prior art.
[0004] Embodiments of the present invention provide a method for identifying sensitive applications resistant to concept drift and processing traffic data, the method comprising:
[0005] Collecting a plurality of traffic session samples, differentiating each data packet in the traffic session samples and generating a token sequence to form a first training set;
[0006] Performing secondary pre-training on the ET-BERT model with the samples in the first training set to obtain a trained ET-BERT model;
[0007] Inputting the token sequences of a plurality of traffic session samples with known sensitive application types into the trained ET-BERT model to obtain an output vector v of the traffic session sample corresponding to each token sequence i , and storing the output vector v i of each traffic session sample and the sensitive application type into a vector database; regularly collecting new traffic session samples with known sensitive application types to update the vector database;
[0008] Inputting the token sequence of the traffic session sample of the sensitive application to be identified into the trained ET-BERT model to obtain a corresponding output vector u;
[0009] Determine the type of the sensitive application to be recognized based on the output vector u corresponding to the sensitive application to be recognized, the output vectors in the vector database, and their corresponding sensitive application types;
[0010] Monitor the traffic data of the sensitive application whose type has been recognized, and process the traffic data of the sensitive application when the traffic data meets the preset conditions.
[0011] Further, when the traffic data meets the preset conditions, processing the traffic data of the sensitive application includes: when the traffic data meets the preset conditions, intercepting and blocking the traffic data.
[0012] Further, intercepting and blocking the traffic data includes: parsing and restoring the traffic session sample of the sensitive application to be recognized; obtaining the source port and source IP address from the parsed and restored data, and intercepting the outgoing traffic of the source port and source IP address and all outgoing traffic returned by the server to the source port and source IP address.
[0013] Further, collecting multiple traffic session samples and generating a token sequence based on each data packet in the traffic session samples includes: processing each data packet payload part in the traffic session samples using a double-byte token segmentation method to obtain a double-byte token sequence, adding a [SEP] token at the end of the token sequence segmented from each data packet, and splicing the token sequences of each data packet with the [SEP] token added in the order of the data packets to obtain the token sequence of the traffic session sample.
[0014] Further, processing each data packet payload part in the traffic session samples using a double-byte token segmentation method includes: using a sliding window with a length of 2 and a step size of 1 to segment each data packet payload part to form a double-byte token sequence.
[0015] Further, splicing the double-byte token sequences of each data packet with a [SEP] token added at the end to form a traffic session sample token sequence includes: if the length of the spliced token sequence is less than 512, using a [PAD] token to complete it; if the length of the spliced token sequence exceeds 512, removing the redundant tokens with a length exceeding 512 in the token sequence, and replacing the 512th token with a [SEP] token. If two [SEP] tokens appear at the end of the token sequence after replacement, the 512th token is replaced with a [PAD] token.
[0016] Further, determining the type of the sensitive application to be recognized based on the output vector u corresponding to the sensitive application to be recognized, the output vectors in the vector database, and their corresponding sensitive application types includes:
[0017] Calculate the similarity between the output vector u corresponding to the sensitive application to be recognized and each output vector in the vector database, sort the similarities from largest to smallest, and based on the sensitive application types to which each traffic session sample corresponding to several of the top similarities belongs, and the output vector v in the vector database corresponding to the sensitive application type i The similarity value between the output vector u corresponding to the sensitive application to be recognized, and determine the type of the traffic session sample of the sensitive application to be recognized.
[0018] Further, sort the similarities from largest to smallest, and based on the sensitive application types to which each traffic session sample corresponding to several of the top similarities belongs, and the output vector v in the vector database corresponding to the sensitive application type i The similarity value between the output vector u corresponding to the sensitive application to be recognized, and determine the type of the traffic session sample of the sensitive application to be recognized, including: selecting m of the top similarities, and for the sensitive application types cls corresponding to the m traffic session samples corresponding to the m similarities i Conduct statistics, select the sensitive application type with the largest proportion as the statistical category, and check the output vector v corresponding to the statistical category i Whether the maximum similarity between the output vector u of the traffic session sample of the sensitive application to be recognized is greater than the first threshold. If so, it is considered that the type of the traffic session sample of the sensitive application to be recognized is the same as the statistical category; otherwise, it is considered that the type of the traffic session sample of the sensitive application to be recognized has not been identified.
[0019] Further, the secondary pre-training of the ET-BERT model with the samples in the first training set includes calculating the similarity between the traffic feature representations output each time after the token sequence of the same traffic session sample in the same batch is input into the ET-BERT model that has undergone primary training as the similarity of the positive sample pair, and calculating the similarity between the traffic feature representations output after the token sequences of different traffic session samples in the same batch are input into the ET-BERT model that has undergone primary training as the similarity of the negative sample pair; calculating the secondary pre-training loss function of the ET-BERT model based on the similarity of the positive sample pair and the similarity of the negative sample pair.
[0020] Further, the formula for the secondary pre-training loss function is:
[0021]
[0022] where τ is the scale coefficient, N is the number of traffic session samples included in each batch; l i is the loss function of the i-th traffic session sample; sim is the similarity function, x i is the token sequence of the i-th traffic session sample included in each batch, xj The token sequence of the j-th traffic session sample included in each batch; is x i The feature representation extracted by ET-BERT for a certain input, is x i The feature representation extracted by ET-BERT for another input; is x j The feature representation extracted by ET-BERT for a certain input, is x j The feature representation extracted by ET-BERT for another input; L is the value of the secondary pre-training loss function for the current batch; is the similarity of the positive sample pair; is the similarity of the negative sample pairs in the same iteration, is the similarity of the negative sample pairs in different iterations.
[0023] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects:
[0024] 1. A method for identifying sensitive applications resistant to concept drift and processing traffic data according to the present invention regularly updates a small sample vector database based on the token sequences of traffic session samples of multiple known sensitive application types, and realizes a method for identifying sensitive applications resistant to concept drift.
[0025] 2. A method for identifying sensitive applications resistant to concept drift and processing traffic data according to the present invention differentiates each data packet during the processing of traffic session samples, and adds the [SEP] token at the end of the token sequences segmented from each data packet in each traffic session sample, which improves the ability of the ET-BERT model to extract the features of traffic session samples, thereby improving the accuracy of sensitive application identification.
[0026] 3. A method for identifying sensitive applications resistant to concept drift and processing traffic data according to the present invention uses contrastive learning technology to perform secondary pre-training on ET-BERT, which improves the traffic characterization ability of ET-BERT, enables the ET-BERT model after secondary pre-training to accurately extract the features of traffic session samples, and thereby improves the accuracy of sensitive application identification.
[0027] 4. A method for identifying sensitive applications resistant to concept drift and processing traffic data according to the present invention obtains the output vector u corresponding to the sensitive application to be identified and the vector v in the vector database based on the ET-BERT model after secondary pre-training i The similarity between them is sorted from large to small, m similarities ranked in the front are selected, and the sensitive application types cls corresponding to the m traffic session samples corresponding to the m similarities iPerform statistics, select the most prevalent sensitive application type as the statistical category, and check whether the one with the highest similarity under this statistical category is greater than the first threshold. If so, determine that the type of the traffic session sample of the sensitive application to be identified is the same as the statistical category. Through the above algorithm, the type recognition of open-world sensitive applications is achieved based on a small sample training set with 20 to 100 traffic session samples known for each sensitive application type.
[0028] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combined solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can be made obvious from the description, or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the content specifically pointed out in the description and the drawings. Brief Description of the Drawings
[0029] The drawings are only for the purpose of showing specific embodiments, and are not considered as limiting the present invention. Throughout the drawings, the same reference signs represent the same components.
[0030] Figure 1 It is a schematic flow chart for converting traffic session samples of a method for identifying sensitive applications and processing traffic data against concept drift into token sequences;
[0031] Figure 2 It is a schematic diagram of inputting the token sequence of each traffic session sample of a method for identifying sensitive applications and processing traffic data against concept drift into an ET-BERT model and outputting vectors;
[0032] Figure 3 It is a schematic flow chart of a secondary pre-training process based on contrastive learning for a method for identifying sensitive applications and processing traffic data against concept drift;
[0033] Figure 4 It is a schematic flow chart for identifying open-world unknown sensitive applications based on a small sample training set in a method for identifying sensitive applications and processing traffic data against concept drift. Detailed Embodiments
[0034] The following will specifically describe the preferred embodiments of the present invention in conjunction with the drawings. The drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, and are not used to limit the scope of the present invention.
[0035] A specific embodiment of the present invention discloses a method for identifying sensitive applications and processing traffic data against concept drift, which specifically includes S1-S4.
[0036] S1. Collect multiple traffic session samples, distinguish each data packet in the traffic session samples, generate a token sequence, and form a first training set.
[0037] Collect multiple traffic session samples and generate a token sequence based on each data packet in the traffic session samples, including: For the payload part of each data packet in the traffic session samples, use a double-byte token segmentation method to process and obtain a double-byte token sequence. Add a [SEP] token at the end of the token sequence segmented from each data packet, and splice the token sequences of each data packet after adding the [SEP] token in the order of the data packets to obtain the token sequence of the traffic session sample.
[0038] Specifically, encrypted traffic data can be divided into multiple traffic sessions through a five-tuple of (source IP address, destination IP address, source port, destination port, protocol type). A complete traffic session consists of several data packets and contains all the data generated during a network access interaction process.
[0039] In the processing of the input token sequence of ET-BERT, a special [CLS] token will be added at the very front of the token sequence, and the numerical vector output at this token position is regarded as the feature representation of the entire input traffic session. If the [SEP] token is not added at the end of the token sequence segmented from each data packet and directly spliced, the differences between different data packets in the same session are ignored, confusing the feature levels of the session and the data packet at the input level, which is not conducive to the model to learn the feature representation of the traffic session. Therefore, adding a [SEP] token at the end of the token sequence segmented from each data packet in each traffic session sample can distinguish each data packet in each traffic session sample.
[0040] The process of converting a traffic session sample into a token sequence is as Figure 1 shown.
[0041] For the payload part of each data packet in the traffic session samples, use a double-byte token segmentation method to process, including: Use a sliding window with a length of 2 and a step size of 1 to segment the payload part of each data packet to form a double-byte token sequence.
[0042] Splice the double-byte token sequences of each data packet with a [SEP] token added at the end to form a traffic session sample token sequence, including: If the length of the spliced token sequence is less than 512, use the [PAD] token to complete it; if the length of the spliced token sequence exceeds 512, remove the extra tokens with a length exceeding 512 in the token sequence, and replace the 512th token with a [SEP] token. If two [SEP] tokens appear at the end of the token sequence after replacement, the 512th token is replaced with a [PAD] token.
[0043] S2. Use the samples in the first training set to perform secondary pre-training on the ET-BERT model to obtain a trained ET-BERT model.
[0044] A pre-trained model is a deep learning model that has been pre-trained on a large amount of unlabeled data. It usually learns basic features such as language or vision on general tasks and can then be fine-tuned to adapt to specific domain tasks, such as answering questions, translation, or text classification. This can reduce the time and computational resources required to train a model on small-scale data and achieve better results.
[0045] Specifically, the schematic diagram of the token sequence in each traffic session sample input into the ET-BERT model and outputting a vector is as Figure 2 shown, and this process can be expressed as:
[0046] Z = ET-BERT(x) = {z 1 , z 2 , …, z 512}}, x ∈ R 512 , z ∈ R 512×768 , z i ∈ R 768 , x represents the 512-token sequence formed by each traffic session sample, and the output result corresponding to the i-th token x i in the token sequence x is z i , the 768-dimensional numerical vector z i indicates the features extracted by the ET-BERT model after global attention mechanism modeling for the token at this position. Among the 512 vectors z i that the ET-BERT model can output for each traffic session sample, in this embodiment, the first vector z 1 (i.e., the output corresponding to the position of the [CLS] token) is input into the subsequent mapping layer.
[0047] The secondary pre-trained ET-BERT model includes a mapping layer; the mapping layer includes a first fully connected layer and a second fully connected layer.
[0048] The mapping layer extracts features through the following formula:
[0049] h = (σ(z 1 ·W 1 + b 1 ))·W 2 + b 2 , (1)
[0050]
[0051] z 1 ∈ R 768 , h ∈ R 768, W 1 ∈R 768×3072 , b 1 ∈R 3072 , W 2 ∈R 3072×768 , b 2 ∈R 768 ,
[0052] Among them, W 1 and W 2 are the weight matrices of the first and second fully connected layers respectively, and b 1 and b 2 are the bias vectors of the first and second fully connected layers respectively. z 1 is the input vector of the first fully connected layer, h is the output vector of the second fully connected layer, and a is each element in the vector z 1 ·W 1 + b 1 .
[0053] Specifically, in the secondary pre-training, all the model parameters of ET-BERT, as well as W 1 , W 2 , b 1 , b 2 will be updated.
[0054] The secondary pre-training refers to the process of performing re-pre-training on the basis of the ET-BERT model after the initial pre-training. The purpose is to enable the model to better represent the features of the traffic session samples on the basis of the existing general capabilities.
[0055] In the embodiments of the present invention, the main body of the ET-BERT model that has been initially pre-trained is retained, and the main body is used to extract the traffic session feature representation. After being processed by the algorithm described in S1, a single traffic session data has been converted into a sequence composed of 512 tokens, and this sequence is directly input into the ET-BERT model. ET-BERT already includes the structure for converting the token sequence and extracting traffic features. It will represent the 512 tokens of the input single traffic session as a 512×768 numerical matrix. However, in order to enhance the representation ability of ET-BERT using contrastive learning during the subsequent secondary pre-training, it is necessary to map the traffic features into the contrastive loss space. Therefore, a mapping layer composed of 2 fully connected layers is used to perform feature mapping on the output vector at the [CLS] token position, and the numerical vector calculated by this mapping layer, that is, h in formula (1), is used as the final traffic feature representation. W 1 , W 2 are weight matrices that project the input feature vector into a new vector space; b 1 , b 2It is a bias vector used to adjust the values of each dimension of the output vector, helping the model better fit the training data and improve the model performance.
[0056] The token sequences in the first training set are divided into multiple batches, and the token sequences of each batch are repeatedly input into the pre-trained ET-BERT model for secondary pre-training based on contrastive learning to obtain the trained ET-BERT model.
[0057] The schematic diagram of the secondary pre-training process based on contrastive learning is as Figure 3 shown.
[0058] Calculate the similarity between the traffic feature representations of each output when the token sequences of the same traffic session samples in the same batch are repeatedly input into the pre-trained ET-BERT model as the similarity of the positive sample pairs, and calculate the similarity between the traffic feature representations of the token sequences of different traffic session samples in the same batch input into the pre-trained ET-BERT model as the similarity of the negative sample pairs; calculate the secondary pre-training loss function of the ET-BERT model based on the similarity of the positive sample pairs and the similarity of the negative sample pairs.
[0059] The formula for the secondary pre-training loss function is:
[0060]
[0061] where τ is the scale coefficient, N is the number of traffic session samples included in each batch; l i is the loss function of the i-th traffic session sample; sim is the similarity function, x i is the token sequence of the i-th traffic session sample included in each batch, x j is the token sequence of the j-th traffic session sample included in each batch; is the feature representation extracted by x i when input into ET-BERT for a certain time, is the feature representation extracted by x i when input into ET-BERT another time; is the feature representation extracted by x j when input into ET-BERT for a certain time, is the feature representation extracted by x j when input into ET-BERT another time; L is the value of the secondary pre-training loss function of the current batch; is the similarity of the positive sample pairs; is the similarity of the negative sample pairs in the same iteration, is the similarity of the negative sample pairs in different iterations.
[0062] Specifically, during the training process, minimize \(l\) calculated by formula (2) as much as possible. i When the similarity of positive samples and the total similarity tends to be maximized, while the sum of similarities of all negative sample pairs
[0063] tends to be minimized, so that the similarity of positive sample pairs is as large as possible, while the similarity of negative sample pairs is as small as possible, achieving the effect of being able to distinguish similar features and different features.
[0063] Calculate the similarity of positive sample pairs and the similarity of negative sample pairs. The similarity is the cosine similarity, and the formula is:
[0064]
[0065] where, "·" represents the dot product, and |||| represents calculating the length of the vector.
[0066] S3. Input the token sequences of multiple traffic session samples with known sensitive application types into the trained ET-BERT model to obtain the output vector \(v\) of each traffic session sample corresponding to the token sequence i , store the output vector \(v\) i of each traffic session sample and the sensitive application type in the vector database; regularly collect new traffic session samples with known sensitive application types to update the vector database; input the token sequence of the traffic session sample of the sensitive application to be recognized into the trained ET-BERT model to obtain the corresponding output vector \(u\).
[0067] Open-world recognition means that a trained model needs to recognize all possible target categories in the actual application environment, rather than just a predefined set of categories, that is, not only to process input data of known categories, but also to be able to recognize new categories that do not appear in the training set.
[0068] Sensitive applications: mainly refer to application software that hides or bypasses censorship through encrypted communication methods, mainly including various VPN software. Such software usually uses custom encryption protocols for encrypted communication or disguises itself as protocols such as TCP and TLS for encrypted communication.
[0069] The model parameters are adjusted using the Adam optimization algorithm. The encrypted traffic characterization model obtained after the training is called ET-Sim. Currently, ET-BERT performs poorly in open-world recognition, few-shot recognition, etc. The reason is that there is no way to learn traffic characterization from sample pairs of similar traffic data and dissimilar traffic data in the original pre-training of ET-BERT. Therefore, in this embodiment, a contrastive learning method is proposed to retrain ET-BERT. Based on the existing general traffic characterization ability, by inputting sample pairs of similar data and dissimilar data, the model is encouraged to learn the ability to compare and distinguish different features, effectively improving the characterization ability of ET-BERT and enhancing the performance of ET-BERT in open-world recognition and few-shot recognition.
[0070] Specifically, in the prior art, it is necessary to input a large number of training samples with known sensitive application types into the ET-BERT model to identify the categories of sensitive applications. In this embodiment, since the ET-BERT model is subjected to the secondary pre-training described in S2, the number of traffic session samples with known types of each sensitive application required is greatly reduced.
[0071] Specifically, the token sequences of traffic session samples with known types of multiple sensitive applications are input into the trained ET-BERT model. The number of traffic session samples with known types of each sensitive application is between 20 and 100.
[0072] Concept drift refers to the phenomenon that the distribution between the training data and the actual application data of a machine learning model changes. In machine learning, the model learns rules and patterns based on the training data. Once these rules and patterns are no longer applicable in the new data, concept drift may occur. The phenomenon of concept drift is very common in traffic analysis. Taking sensitive application analysis as an example, VPN providers regularly change the communication methods between the client and the server, and the traffic generated by the VPN is likely to change the existing rules and patterns. The model trained on the original VPN data will become invalid when facing new VPN data.
[0073] To overcome concept drift, new traffic session samples with known types of sensitive applications are regularly collected to update the vector database; specifically, new traffic session samples with known types of sensitive applications are regularly collected and token sequences are generated according to the method described in S1. The token sequences are input into the ET-BERT model that has been pre-trained twice to generate the output vectors of the corresponding traffic session samples. Data sample pairs are generated from the generated output vectors and the corresponding sensitive application types and stored in the vector database to update the output vectors of the same sensitive application type stored previously. The number of traffic session samples with known types of each sensitive application in the output vector database collected is between 20 and 100.
[0074] AsFigure 4 The ET-Sim shown includes an ET-BERT with a mapping layer for identifying traffic session samples of sensitive applications to be identified based on a small-sample training set.
[0075] S4. Determine the type of the sensitive application to be identified based on the output vector u corresponding to the sensitive application to be identified, the output vectors in the vector database, and their corresponding sensitive application types; monitor the traffic data of the identified sensitive application, and process the traffic data of the sensitive application when the traffic data meets the preset conditions.
[0076] Calculate the similarity between the output vector u corresponding to the sensitive application to be identified and each output vector in the vector database, sort the similarities from largest to smallest, and based on the sensitive application types to which each traffic session sample corresponding to several of the top similarities belongs, and the output vector v in the vector database corresponding to the sensitive application type i Determine the type of the traffic session sample of the sensitive application to be identified based on the similarity value between the output vector u corresponding to the sensitive application to be identified.
[0077] Select m of the top similarities, and count the sensitive application types cls corresponding to the m traffic session samples corresponding to the m similarities i Conduct statistics, select the sensitive application type with the largest proportion as the statistical category, and check whether the maximum similarity between the output vector v corresponding to the statistical category and the output vector u of the traffic session sample of the sensitive application to be identified is greater than the first threshold. If so, it is considered that the type of the traffic session sample of the sensitive application to be identified is the same as the statistical category; otherwise, it is considered that the type of the traffic session sample of the sensitive application to be identified has not been recognized. i i
[0078] In a specific embodiment of the present invention, calculating the similarity between the output vector u corresponding to the sensitive application to be identified and the output vector v of each traffic session sample corresponding to the small-sample training set i using the cosine similarity of formula (3) in step 2 has achieved good results.
[0079] Select m of the top similarities, and count the sensitive application types cls corresponding to the m traffic session samples corresponding to the m similarities i Conduct statistics, select the sensitive application type with the largest proportion as the statistical category, and check whether the maximum similarity between the output vector v corresponding to the statistical category and the output vector u of the traffic session sample of the sensitive application to be identified is greater than the first threshold. If so, it is considered that the type of the traffic session sample of the sensitive application to be identified is the same as the statistical category; otherwise, it is considered that the type of the traffic session sample of the sensitive application to be identified has not been recognized. i i
[0080] The open recognition of small samples is based on the trained ET-Sim. ET-Sim is used as a traffic characterization model to extract the feature representation of traffic samples, and the similarity measurement algorithm based on contrast learning in step 2 is used to determine the sensitive applications to which the traffic samples collected in the open world belong.
[0081] Before using ET-Sim for the traffic recognition of sensitive applications in the open world, it is necessary to determine the list of sensitive applications concerned by the business according to the actual functional requirements, that is, only recognize the sensitive applications in the list and distinguish other traffic that does not belong to these sensitive applications. In a specific embodiment of the present invention, 50 traffic session samples of known sensitive application types can be collected for each sensitive application in the list. ET-Sim is used to extract the feature representation of these traffic samples, and the extracted feature representation and the sensitive application label to which each feature representation belongs are stored in the vector database.
[0082] When using ET-Sim for the traffic recognition of sensitive applications in the open world, for each newly collected traffic session sample of the sensitive application to be recognized, first use ET-Sim to extract the feature vector u of the traffic sample, and then query the set S of the 10 <feature vector, sensitive application label> tuples with the largest similarity to u in the vector database that already stores the list of sensitive applications concerned by the business. S = (<v 1 ,cls 1 >,<v 2 ,cls 2 >,…,<v 10 ,cls 10 >). Classify and count the sensitive application labels carried by the tuples in the set S, and the sensitive application category with the largest number of labels is regarded as the candidate category cls. If there is at least one feature vector v i in the set S whose cosine similarity to the feature u of the traffic sample to be recognized is higher than 0.5, then the traffic sample to be recognized is regarded as belonging to the sensitive application type indicated by cls, otherwise the traffic sample to be recognized is regarded as an unknown type on the grounds that the matching degree between the traffic sample to be recognized and the known sensitive application traffic features is not high.
[0083] When the traffic data meets the preset conditions, intercept and block the traffic data.
[0084] Intercepting and blocking the traffic data includes: parsing and restoring the traffic session sample of the sensitive application to be recognized; obtaining the source port and source IP address from the parsed and restored data, and intercepting the outgoing traffic of the source port and source IP address and all outgoing traffic returned by the server to the source port and source IP address.
[0085] Specifically, the preset conditions include using a machine learning method (extracting features from traffic data of known malicious software such as network fraud, porn, and hackers based on a neural network and constructing a classification model, and inputting the monitored traffic data after parsing and restoration into the neural networks of each classification model to obtain probability scores belonging to each type of data traffic) or a heuristic method (such as classifying traffic data based on application layer protocols, comparing the monitored traffic data after parsing and restoration with the characteristics of known malicious traffic data categories, and manually judging the probability scores belonging to various types of malicious data traffic such as network fraud, porn, and hackers) to obtain the probability scores of the traffic data belonging to each type of traffic, assigning corresponding scores to the damage degrees of each type of traffic, multiplying the damage degree scores of this type of traffic by the probability scores to obtain a product result. If the product result of any one type of traffic is greater than the set threshold, it is considered that the preset conditions are met.
[0086] Compared with the prior art, the method for identifying sensitive applications resistant to concept drift and processing traffic data provided in this embodiment regularly updates the small sample vector database based on the token sequences of traffic session samples of multiple known sensitive application types, realizing a method for identifying sensitive applications resistant to concept drift. The method for identifying sensitive applications resistant to concept drift and processing traffic data provided in this embodiment differentiates each data packet during the processing of traffic session samples, adds a [SEP] token at the end of the token sequence segmented from each data packet in each traffic session sample, improving the ability of the ET-BERT model to extract the features of traffic session samples, thereby improving the accuracy of sensitive application identification. The method for identifying sensitive applications resistant to concept drift and processing traffic data provided in this embodiment uses contrastive learning technology to perform secondary pre-training on ET-BERT, improving the traffic representation ability of ET-BERT, enabling the ET-BERT model after secondary pre-training to accurately extract the features of traffic session samples, thereby improving the accuracy of sensitive application identification. The method for identifying sensitive applications resistant to concept drift and processing traffic data provided in this embodiment calculates the similarity between the output vector u corresponding to the sensitive application to be identified and the vector database vector v i based on the ET-BERT model after secondary pre-training, sorts the similarities from largest to smallest, selects m similarities ranked in the front, and for the m sensitive application types cls i corresponding to the m traffic session samples corresponding to the m similarities, perform statistics, select the sensitive application type with the largest proportion as the statistical category, and check whether the similarity with the largest value under this statistical category is greater than the first threshold. If so, determine that the type of the traffic session sample of the sensitive application to be identified is the same as the statistical category. Through the above algorithm, the type identification of open-world sensitive applications based on a small sample training set with the number of traffic session samples of each known sensitive application type ranging from 20 to 100 is realized.
[0087] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, a random access memory, etc.
[0088] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A sensitive application identification and flow data processing method resistant to concept drift, characterized in that: The method comprises: Collecting multiple traffic session samples, distinguishing each data packet in the traffic session samples and generating a token sequence to form a first training set; The ET-BERT model is pre-trained twice using the samples in the first training set to obtain a trained ET-BERT model; The token sequences of multiple traffic session samples with known sensitive application types are input into the trained ET-BERT model to obtain the output vector v of the traffic session samples corresponding to each token sequence. i , the output vector v of each traffic session sample i , storing sensitive application types in a vector database; regularly collecting new known traffic session samples of sensitive application types to update the vector database; Input the token sequence of the traffic session sample of the sensitive application to be identified into the trained ET-BERT model to obtain the corresponding output vector u; Determine the type of the sensitive application to be identified based on the output vector u corresponding to the sensitive application to be identified and the output vector of the vector database and its corresponding sensitive application type; The flow data of the sensitive application of the identified type is monitored, and the flow data of the sensitive application is processed when the flow data meets the preset conditions.
2. The sensitive application identification and flow data processing method according to claim 1 is characterized in that: When the traffic data meets the preset conditions, the traffic data of the sensitive application is processed, including: when the traffic data meets the preset conditions, the traffic data is intercepted and blocked.
3. The sensitive application identification and flow data processing method according to claim 2 is characterized in that: Intercepting and blocking traffic data includes: parsing and restoring traffic session samples of sensitive applications to be identified; obtaining source ports and source IP addresses from the parsed and restored data, intercepting the egress traffic of the source port and source IP address and all egress traffic returned by the server to the source port and source IP address.
4. The sensitive application identification and flow data processing method according to claim 1 is characterized in that: Collect multiple traffic session samples, and generate a token sequence based on each data packet in the traffic session samples, including: using a double-byte token segmentation method to process the payload portion of each data packet in the traffic session samples to obtain a double-byte token sequence, adding a [SEP] token at the end of the token sequence segmented from each data packet, and splicing the token sequences of each data packet after adding the [SEP] token in the order of the data packets to obtain the token sequence of the traffic session sample.
5. The sensitive application identification and flow data processing method according to claim 4 is characterized in that: The payload part of each data packet in the traffic session sample is processed by a double-byte token segmentation method, including: using a sliding window with a length of 2 and a step size of 1 to segment the payload part of each data packet to form a double-byte token sequence.
6. The sensitive application identification and flow data processing method according to claim 5 is characterized in that: The two-byte token sequence of each data packet with the [SEP] token added at the end is concatenated to form a traffic session sample token sequence, including: if the length of the concatenated token sequence is less than 512, the [PAD] token is used to complete it; if the length of the concatenated token sequence exceeds 512, the redundant tokens with a token sequence length exceeding 512 are removed, and the 512th token is replaced with a [SEP] token; if two [SEP] tokens appear at the end of the token sequence after the replacement, the 512th token is replaced with a [PAD] token.
7. The sensitive application identification and flow data processing method according to claim 1 is characterized in that: Determining the type of the sensitive application to be identified based on the output vector u corresponding to the sensitive application to be identified and the output vector of the vector database and the sensitive application type corresponding thereto, including: Calculate the similarity between the output vector u corresponding to the sensitive application to be identified and each output vector in the vector database, sort the similarities from large to small, and select the sensitive application type to which each traffic session sample belongs and the output vector v in the vector database corresponding to the sensitive application type based on the similarities in the front. i The similarity value between the output vector u corresponding to the sensitive application to be identified determines the type of the traffic session sample of the sensitive application to be identified.
8. The sensitive application identification and flow data processing method according to claim 7 is characterized in that: Sort each similarity from large to small, and output the vector v in the vector database corresponding to each traffic session sample to which the sensitive application type belongs and the vector corresponding to the sensitive application type based on the similarities ranked in front i The similarity value between the output vector u corresponding to the sensitive application to be identified is used to determine the type of the traffic session sample of the sensitive application to be identified, including: selecting m similarities in the front, and calculating the sensitive application type cls corresponding to the m traffic session samples corresponding to the m similarities. i Perform statistics, select the most sensitive application type as the statistical category, and check the output vector v corresponding to the statistical category i Whether the maximum similarity between the output vector u of the traffic session sample of the sensitive application to be identified is greater than the first threshold, if so, it is considered that the type of the traffic session sample of the sensitive application to be identified is the same as the statistical category, otherwise it is considered that the type of the traffic session sample of the sensitive application to be identified is not identified.
9. The sensitive application identification and flow data processing method according to claim 1, characterized in that: The secondary pre-training of the ET-BERT model with samples in the first training set includes calculating the similarity between the traffic feature representations output each time after the token sequence of the same traffic session sample in the same batch is input multiple times into the ET-BERT model that has undergone primary training as the similarity of the positive sample pair, and calculating the similarity between the traffic feature representations output after the token sequences of different traffic session samples in the same batch are input into the ET-BERT model that has undergone primary training as the similarity of the negative sample pair; and obtaining the secondary pre-training loss function of the ET-BERT model based on the similarity of the positive sample pair and the similarity of the negative sample pair.
10. The sensitive application identification and flow data processing method according to claim 9, characterized in that: The secondary pre-training loss function formula is: Where τ is the scale factor, N is the number of traffic session samples contained in each batch; l i is the loss function of the ith traffic session sample; sim is the similarity function, x i is the token sequence of the i-th traffic session sample contained in each batch, x j is the token sequence of the jth traffic session sample contained in each batch; For x i The feature representation extracted by ET-BERT is input once. For x i Another input is the feature representation extracted by ET-BERT; For x j The feature representation extracted by ET-BERT is input once. For x j Another input is the feature representation extracted by ET-BERT; L is the secondary pre-training loss function value of the current batch; is the similarity of the positive sample pair; is the similarity of negative sample pairs in the same iteration, is the similarity of negative sample pairs in different iterations.
Citation Information
Cited By
Online activity prediction method fusing process constraint
CN121051559A