A method for realizing network security threat prediction based on large language models

Through traffic prediction and malicious detection model based on large language model, the problems of limited detection fields and low efficiency in network security threat prediction are solved, efficient and accurate threat prediction are achieved, and the initiative and effectiveness of network security are improved.

CN119996082BActive Publication Date: 2025-07-25BEIJING XINLIAN SHUAN TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510457620.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-25
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The prior art has limited detection fields in the prediction of network security threats, low prediction efficiency, low accuracy, and numerous data collection requirements.

Method used

The method based on the large language model is adopted to build a sample traffic set, and the traffic timing detection model and malicious traffic detection model are trained through time-series adjacent and non-adjacent traffic groups. It combines BERT and BART models for fine-tuning to generate traffic prediction and malicious detection models.

Benefits of technology

It improves the accuracy and efficiency of network security threat prediction, can analyze potential malicious traffic in real time, take defensive measures in advance, and enhances the initiative and effectiveness of network security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996082B_ABST
    Figure CN119996082B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for realizing network security threat prediction based on a large language model. Based on a sample traffic set, a positive sample traffic group composed of two adjacent sample traffics in each time series is constructed. With the traffic influence eigenvalue as the input and the traffic key eigenvalue as the output, it is fine-tuned in cooperation with a traffic time series detection model for detecting whether the two traffics are time series continuous, and a traffic prediction model is trained. Then, combined with the malicious traffic detection model obtained from the training for malicious detection, the traffic prediction model is applied to predict traffic, and the malicious traffic detection model is applied to detect the predicted traffic for maliciousness. Through advanced deep learning technology, the present invention can analyze and predict potential malicious traffic in real time, so as to take preventive measures in advance, greatly improving the initiative and effectiveness of network security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for realizing network security threat prediction based on a large language model, and belongs to the technical field of traffic prediction and detection. Background Art

[0002] Machine learning is an artificial intelligence technology that enables a computer to learn from data and improve its performance without explicit programming. The core idea of machine learning is to build a model that allows the computer to automatically learn the rules and patterns from the data, so as to be able to make predictions or decisions on new data. A specific application under machine learning is deep learning, which adopts a multi-layer neural network structure to learn complex features and patterns from a large amount of data. A deep learning model usually consists of multiple layers, each layer contains a large number of neurons (or nodes), and these layers are connected by weights. The main advantage of deep learning is its ability to automatically extract high-level abstract features from the original data without manual feature design.

[0003] A large language model (LLM) refers to a class of natural language processing (NLP) models based on deep learning. They usually have billions or even hundreds of billions of parameters and can perform well in a wide range of natural language tasks. These models learn to capture the complex structure and semantic information of language through pre-training on large-scale text data, and can generate coherent and grammatically correct text, answer questions, provide explanations, and even perform creative writing. The core advantage of large language models lies in their strong generalization ability and multi-task adaptability, which can handle various downstream tasks without specialized training.

[0004] Threat prediction is a key technology in the field of network security, aiming to identify potential security threats in advance by analyzing historical data, real-time monitoring of network traffic and behavior patterns, and taking corresponding preventive measures. The core goal of threat prediction is to improve the initiative of network security, changing from the traditional "post-response" mode to the "prevention-first" mode, so as to intervene before or in the early stage of an attack and reduce losses and risks. Regarding the existing technologies for threat prediction, in terms of system performance: the existing technologies adopt the method of combining CNN and LSTM to identify and predict potential threats to network security, but the method is outdated and the predicted fields are limited; in terms of data collection, the network activity data required by the existing technologies includes traffic data, log data, security event data, protocol data, timestamps, anomaly metrics, terminal device data, honeypot data, and external threat intelligence, requiring a large amount of data, with low prediction efficiency and low prediction accuracy. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for realizing network security threat prediction based on a large language model, expanding the detection fields, and improving the accuracy and efficiency of traffic detection and malicious traffic detection.

[0006] The present invention adopts the following technical solutions to solve the above technical problems: The present invention designs a method for realizing network security threat prediction based on a large language model. Based on a sample traffic set composed of sample traffic corresponding to malicious labels or non-malicious labels respectively according to a preset quantity, the following steps A to C are executed to obtain a traffic prediction model and a malicious traffic detection model, and then step i is executed to predict and perform malicious detection on the next traffic adjacent to the current traffic in the future time direction.

[0007] Step A. In the manner of forming positive sample traffic groups with two adjacent sample traffic in time series and negative sample traffic groups with two non-adjacent sample traffic in time series, each positive sample traffic group and each negative sample traffic group are extracted from the sample traffic set, and the positive sample traffic group corresponds to a traffic continuous label, and the negative sample traffic group corresponds to a traffic non-continuous label, and then step B is entered.

[0008] Step B. First, based on each positive sample traffic group and each negative sample traffic group, a traffic time series detection model for detecting whether the two traffic correspond to a traffic continuous label or a traffic non-continuous label is trained. Then, based on each positive sample traffic group and fine-tuning of the traffic time series detection model, a traffic prediction model for predicting the next traffic in the adjacent future time direction is trained, and then step C is entered.

[0009] Step C. Based on the sample traffic set, a malicious traffic detection model for detecting whether the traffic corresponds to a malicious label or a non-malicious label is trained.

[0010] Step i. Apply the traffic prediction model to predict the next traffic adjacent to the current traffic in the future time direction, and apply the malicious traffic detection model to perform malicious detection on the next traffic.

[0011] As a preferred technical solution of the present invention: The step B includes the following steps B1 to B3;

[0012] Step B1. For each sample traffic in the sample traffic set, extract the feature values corresponding to each preset malicious traffic detection feature of the sample traffic as each key feature value, and extract the feature values corresponding to each feature that affects the change of the malicious traffic detection feature of the sample traffic as each influencing feature value; then step B2 is entered.

[0013] Step B2. Based on each positive sample traffic group and each negative sample traffic group, using the key feature values of two sample traffics in the sample traffic group as inputs and the corresponding traffic continuous label or traffic discontinuous label of the sample traffic group as the output, train the BERT model to obtain a traffic time series detection model, and then proceed to Step B3;

[0014] Step B3. Based on each positive sample traffic group, using the impact feature values of the first sample traffic in sequence in the positive sample traffic group as inputs and the key feature values of the second sample traffic in sequence in the positive sample traffic group as the output, train the first BART model. At the same time, apply the traffic time series detection model to perform the detection of the corresponding traffic continuous label or traffic discontinuous label for the key feature values of the input traffic and the key feature prediction values of the output predicted traffic during the training of the first BART model. Use the detection result to fine-tune the training of the first BART model, and then obtain the trained first BART model, which constitutes the traffic prediction model.

[0015] As a preferred technical solution of the present invention: in the said Step B3, apply the traffic time series detection model to perform the detection of the corresponding traffic continuous label or traffic discontinuous label for the key feature values of the input traffic and the key feature prediction values of the output predicted traffic during the training of the first BART model. If the detection result is that the input traffic and the output predicted traffic correspond to a traffic discontinuous label, then fine-tune the training of the first BART model; if the detection result is that the input traffic and the output predicted traffic correspond to a traffic continuous label, then do not fine-tune the training of the first BART model.

[0016] As a preferred technical solution of the present invention: in the said Step B3, first update the first BART model by adding a regression head module at the end of its structure, and then based on each positive sample traffic group, using the impact feature values of the first sample traffic in sequence in the positive sample traffic group as inputs and the key feature values of the second sample traffic in sequence in the positive sample traffic group as the output, train the first BART model.

[0017] As a preferred technical solution of the present invention: in the said Step C, based on each sample traffic in the sample traffic set, using the key feature values of the sample traffic as inputs and the corresponding malicious label or non-malicious label of the sample traffic as the output, train the second BART model to obtain a malicious traffic detection model for detecting the corresponding malicious label or non-malicious label of the traffic.

[0018] As a preferred technical solution of the present invention: in step C, first, a classification head module is added to the end of the structure of the second BART model for updating, and then, based on each sample traffic in the sample traffic set, using each key feature value of the sample traffic as input and the corresponding malicious label or non-malicious label of the sample traffic as output, the second BART model is trained to obtain a malicious traffic detection model for detecting the corresponding malicious label or non-malicious label of the traffic.

[0019] As a preferred technical solution of the present invention: step i includes the following steps i1 to i3;

[0020] Step i1. Extract the feature values of the features corresponding to the preset features that affect the change of malicious traffic detection for the current traffic, as the respective influencing feature values corresponding to the current traffic, and then enter step i2;

[0021] Step i2. Apply the traffic prediction model to process the respective influencing feature values corresponding to the current traffic to obtain the respective key feature values corresponding to the next adjacent traffic in the future time direction of the current traffic, and then enter step i3;

[0022] Step i3. Apply the malicious traffic detection model to process the respective key feature values corresponding to the next traffic to obtain the corresponding malicious label or non-malicious label of the next traffic, and perform malicious detection on the next traffic.

[0023] As a preferred technical solution of the present invention: for the sample traffic set composed of the preset number of sample traffics corresponding to malicious labels or non-malicious labels respectively, the following preprocessing is performed for updating;

[0024] First, remove the invalid or duplicate sample traffics in the sample traffic set; then perform noise filtering on each sample traffic in the sample traffic set; then correct the format errors of each sample traffic in the sample traffic set and fill in the missing values according to the preset traffic format; finally, perform standardization or normalization processing on each sample traffic in the sample traffic set and convert it to the preset target traffic format.

[0025] As a preferred technical solution of the present invention: through a network sniffer set at the target network interface position in the target network environment, listen to and capture the traffics corresponding to the preset number of malicious labels or non-malicious labels respectively at the target network interface position as each sample traffic to form a sample traffic set.

[0026] For the method for realizing network security threat prediction based on a large language model of the present invention, compared with the prior art by adopting the above technical solutions, it has the following technical effects:

[0027] The present invention designs a method for realizing network security threat prediction based on a large language model. Based on a sample traffic set, a positive sample traffic group composed of two adjacent samples in each time series is constructed. With the traffic impact eigenvalue as the input and the traffic key eigenvalue as the output, it is fine-tuned in cooperation with a traffic time series detection model for detecting whether two traffic flows are time series continuous, and a traffic prediction model is obtained through training. Then, in combination with the malicious traffic detection model obtained through training for malicious detection, the traffic prediction model is used for traffic prediction, and the malicious traffic detection model is used to detect the maliciousness of the predicted traffic. Through advanced deep learning technology, the present invention can analyze and predict potential malicious traffic in real time, so as to take defensive measures in advance, greatly improving the initiative and effectiveness of network security. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is the flowchart of the fine-tuning stage in the method for realizing network security threat prediction based on a large language model designed by the present invention;

[0029] Figure 2 is the flowchart of the deployment stage in the method for realizing network security threat prediction based on a large language model designed by the present invention;

[0030] Figure 3 is the application schematic diagram of the traffic prediction model in the design of the present invention;

[0031] Figure 4 is the application schematic diagram of the malicious traffic detection model in the design of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following further details the specific embodiments of the present invention with reference to the accompanying drawings of the specification.

[0033] The present invention designs a method for realizing network security threat prediction based on a large language model. In practical applications, as Figure 1 shown, first, for the target network interface position in the target network environment, a network sniffer such as Wireshark or a similar tool is set up. The network sniffer listens to and captures a preset number of traffic flows corresponding to malicious tags or non-malicious tags respectively at the target network interface position as each sample traffic to form a sample traffic set, or by integrating network logs, a preset number of traffic flows corresponding to malicious tags or non-malicious tags respectively within a period of time are obtained. The captured data packets contain rich information such as source IP address, destination IP address, transport layer protocol (TCP / UDP), port number, and data payload, etc. And the data collection process should ensure that it will not affect network performance and comply with laws and regulations, especially when involving personal privacy data. Then, for the sample traffic set, the following preprocessing is carried out for updating to ensure the accuracy and effectiveness of subsequent analysis.

[0034] First, remove the invalid or duplicate sample flows in the sample flow set; then perform noise filtering on each sample flow in the sample flow set; next, correct the format errors of each sample flow in the sample flow set and fill in the missing values according to the preset flow format; finally, perform standardization or normalization processing on each sample flow in the sample flow set and convert it into a preset target flow format such as a CSV file.

[0035] Based on the obtained sample flow set above, as Figure 1 shown, perform the following steps A to C to obtain a flow prediction model and a malicious flow detection model.

[0036] Step A. In the way that two adjacent sample flows in time series form a positive sample flow group and two non - adjacent sample flows in time series form a negative sample flow group, extract each positive sample flow group and each negative sample flow group from the sample flow set. And the positive sample flow group corresponds to a flow continuous label, and the negative sample flow group corresponds to a flow non - continuous label. In practical applications, control the quantity ratio of the positive sample flow group to the negative sample flow group to be 1:1, and then enter Step B.

[0037] Step B. First, based on each positive sample flow group and each negative sample flow group, train a flow time - series detection model for detecting the flow continuous label or flow non - continuous label corresponding to two flows. Then, based on each positive sample flow group and fine - tuning of the flow time - series detection model, train a flow prediction model for predicting the next flow in the adjacent future time direction, and then enter Step C.

[0038] In practical applications, the above Step B is specifically designed as the following Steps B1 to B3.

[0039] Step B1. For each sample flow in the sample flow set respectively, extract the feature values corresponding to each preset malicious flow detection feature of the sample flow as each key feature value, and extract the feature values corresponding to each feature that affects the change of the malicious flow detection feature of the sample flow as each influencing feature value; then enter Step B2.

[0040] Step B2. Based on each positive sample flow group and each negative sample flow group, use the key feature values of the two sample flows in the sample flow group as the input and the flow continuous label or flow non - continuous label corresponding to the sample flow group as the output to train the BERT model to obtain a flow time - series detection model, and then enter Step B3.

[0041] BART is a powerful pre-trained language model. By appropriately fine-tuning BART, we can apply it to tasks in non-linguistic fields, such as network traffic prediction. BART's bidirectional encoder architecture enables it to consider both past and future traffic information simultaneously, which is particularly important for traffic prediction. For example, when predicting the traffic for a certain period, BART will not only refer to the previous traffic data but also use the traffic trends in subsequent time periods for correction. This bidirectional modeling ability makes BART perform well in dealing with complex traffic patterns, especially suitable for traffic scenarios with periodic fluctuations or long-term dependencies.

[0042] In addition, BART's self-attention mechanism can dynamically weight and combine traffic information at different time points, capturing long-range temporal dependencies, which is very helpful for identifying long-term trends hidden behind short-term fluctuations. In this way, BART can make accurate traffic predictions in a short time, helping the network security system to make early resource allocation and defense preparations.

[0043] Therefore, further design step B3 as follows.

[0044] Step B3. First, update the first BART model by adding a regression head module at the end of its structure. Then, based on each positive sample traffic group, use the influence feature values of the first sample traffic in the positive sample traffic group in order as the input, and use the key feature values of the second sample traffic in the positive sample traffic group in order as the output to train the first BART model. Specifically, send the input to the tokenizer of BART for tokenization, convert the traffic information into embedded tokens, and then perform fine-tuning training on the loaded BART model.

[0045] The BART model aims to optimize the probability of the next token in the sequence by using its autoregressive decoder and considering the context established by the previous tokens, in order to generate an accurate next data packet based on the current data packet and ensure context consistency. Its calculation formula is as follows:

[0046] ;

[0047] where, by adjusting the model parameters in the entire training dataset containing tokens , represents the traffic information before the predicted traffic information.

[0048] While training the first BART model, apply the traffic time series detection model to perform the detection of the corresponding traffic continuous label or traffic discontinuous label for each key feature value of the input traffic and each key feature prediction value of the predicted traffic output during the training of the first BART model, and fine-tune the training of the first BART model with the detection result. Among them, if the detection result is that the input traffic and the predicted traffic output correspond to the traffic discontinuous label, then fine-tune the training of the first BART model. If the detection result is that the input traffic and the predicted traffic output correspond to the traffic continuous label, then do not fine-tune the training of the first BART model.

[0049] In this way, in cooperation with the fine-tuning of the traffic time series detection model, the trained first BART model is obtained to form a traffic prediction model. In the process of separately training the above-mentioned designed traffic time series detection model and traffic prediction model, in the specific implementation of the training data involved respectively, the labels corresponding to each sample traffic are not involved, that is, both training processes here are under the training without considering label data.

[0050] In the fine-tuning of BART for traffic prediction, the input traffic information is transformed into embedded data and embedded into a high-dimensional vector to capture semantic information. These embedded data pass through the encoder and decoder layers of the BART model and use the multi-head self-attention mechanism and the position feed-forward layer to effectively process the traffic data. Its schematic diagram is as Figure 3 shown.

[0051] Step C. First, update the second BART model by adding a classification head module at the end of its structure, and then train the second BART model based on each sample traffic in the sample traffic set, with each key feature value of the sample traffic as the input and the corresponding malicious label or non-malicious label of the sample traffic as the output, to obtain a malicious traffic detection model for detecting the corresponding malicious label or non-malicious label of the traffic.

[0052] After generating the predicted next traffic data through the traffic prediction model, fine-tune the second BART model for traffic classification. The dataset it uses is a dataset considering labels, that is, a dataset containing malicious labels or non-malicious labels. The fine-tuned BART model for supervising traffic data classification uses the same network data packet input for the encoder and decoder and is represented by the final output. The goal is to maximize the likelihood of the correct classification label y given the traffic data features. Its calculation formula is as follows:

[0053] ;

[0054] During the embedding process, BART utilizes various noise masking techniques and introduces controlled noise into the input sequence to help the model learn robust representations by inferring missing or altered tokens in the sequence context. This approach ensures a comprehensive understanding of the forward and backward temporal context in network activities, enabling the model to distinguish abnormal patterns. During the entire fine-tuning process, BART adjusts its weights through backpropagation to minimize the loss function, aligning the predicted class probabilities with the ground truth labels. This iterative refinement process helps accurately classify network packets, as shown in Figure 4 shown.

[0055] In the specific implementation of the training data involved in the above process of obtaining the training of the malicious traffic detection model, the labels corresponding to each sample traffic are specifically considered, that is, the training process here is a training under the consideration of label data.

[0056] In practical applications, that is, by performing the above steps A to C, a traffic prediction model and a malicious traffic detection model are obtained, and then as shown in Figure 2 shown, the network sniffer monitors and captures the current traffic at the target network interface location, then preprocesses and updates the current traffic, and then performs the following step i to predict and detect maliciousness for the next traffic adjacent to the current traffic in the future time direction.

[0057] Step i. Apply the traffic prediction model to predict the next traffic adjacent to the current traffic in the future time direction, and apply the malicious traffic detection model to detect maliciousness for the next traffic.

[0058] In practical applications, the above step i is specifically designed and executed as the following steps i1 to i3.

[0059] Step i1. Extract the feature values of the features corresponding to each preset feature that affects the change of malicious traffic detection for the current traffic as the respective influencing feature values corresponding to the current traffic, and then proceed to step i2.

[0060] Step i2. Apply the traffic prediction model to process the respective influencing feature values corresponding to the current traffic to obtain the respective key feature values corresponding to the next traffic adjacent to the current traffic in the future time direction, and then proceed to step i3.

[0061] Step i3. Apply the malicious traffic detection model to process the respective key feature values corresponding to the next traffic to obtain the malicious label or non-malicious label corresponding to the next traffic, and achieve malicious detection for the next traffic.

[0062] In practical applications, once suspicious malicious traffic is detected, an alarm is immediately triggered and corresponding protection measures are taken, such as blocking connections and isolating infected devices. In this way, the system can maintain stable operation during traffic peaks and effectively resist various types of attacks.

[0063] The present invention designs a method for realizing network security threat prediction based on a large language model. Based on a sample traffic set, a positive sample traffic group composed of two adjacent samples in each time series is constructed. With the traffic influence eigenvalue as the input and the traffic key eigenvalue as the output, it is fine-tuned in cooperation with a traffic time series detection model for detecting whether two traffic flows are time series continuous, and a traffic prediction model is obtained through training. Then, combined with the malicious traffic detection model obtained through training for malicious detection, the traffic prediction model is applied for traffic prediction, and the malicious traffic detection model is applied to detect the maliciousness of the predicted traffic. Through advanced deep learning technology, the present invention can analyze and predict potential malicious traffic in real time, so as to take preventive measures in advance, greatly improving the initiative and effectiveness of network security.

[0064] In the application, the design of the present invention can effectively capture the dynamic patterns and time-dependent features in traffic data, which not only includes the time series characteristics of the original data, but also may reveal hidden patterns and trends, providing strong support for subsequent threat prediction. And it is proposed to use the BERT model to evaluate the prediction results of the first BART model, making use of the bidirectional context modeling ability of BERT to improve the accuracy of the prediction results of the first BART model, that is, the traffic prediction model. And based on the application of the traffic prediction model, the second BART model, that is, the malicious traffic detection model, is used to perform malicious detection on the data of the predicted traffic, improving the accuracy of the threat prediction results. At the same time, the prediction requirements for multiple fields are integrated, achieving a comprehensive improvement in the prediction accuracy and the traceability accuracy. The specific advantages are as follows.

[0065] 1. Improve detection accuracy: The threat prediction method of the large language model can more accurately predict future traffic fields;

[0066] 2. Real-time prediction: The design scheme has a fast calculation speed, low requirements for devices, and can give real-time predictions for real-time network traffic data;

[0067] 3. Strong adaptability: The design scheme can adapt to different network environments and attack patterns, analyze its attack pattern while receiving new network traffic, and fine-tune system parameters to continuously improve the prediction accuracy;

[0068] 4. Easy to integrate and deploy: The modular design makes each component easy to integrate into the existing network security system, and the deployment is flexible.

[0069] From the perspective of market demand, a real-time and powerful threat prediction system has a very broad market prospect. With the acceleration of digital transformation, the demand for network security from enterprises and individuals is increasing day by day, especially in the context of the increasingly complex and frequent network attack means. Enterprises and institutions not only need security solutions that can detect known threats, but also need advanced protection measures that can predict and prevent unknown threats. The threat prediction system designed by the present invention just fills this market gap. Through advanced deep learning technology, it can analyze and predict potential malicious traffic in real time, so as to take defensive measures in advance, greatly improving the initiative and effectiveness of network security. In applications, such as being integrated with a security solution platform, it can provide more comprehensive threat detection and prevention capabilities, enhancing customers' trust and satisfaction; in addition, with the rapid development of the Internet of Things (IoT) and the industrial Internet, more and more devices and systems are connected to the network, which not only increases potential entry points for network attacks, but also poses higher requirements for network security. In terms of application scenarios, the threat prediction design can not only be applied to traditional IT environments, but also play an important role in fields such as the IoT and the industrial Internet, protecting critical infrastructure and devices from attacks.

[0070] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.

Claims

1. A method for realizing network security threat prediction based on a large language model, characterized in that: Based on a sample traffic set composed of sample traffic corresponding to malicious labels or non-malicious labels respectively according to a preset quantity, perform the following steps A to C to obtain a traffic prediction model and a malicious traffic detection model, and then perform step i to predict and detect maliciousness for the next traffic adjacent to the current traffic in the future time direction; Step A. In the way that two adjacent sample traffics in time sequence form a positive sample traffic group and two non-adjacent sample traffics in time sequence form a negative sample traffic group, extract each positive sample traffic group and each negative sample traffic group from the sample traffic set, and the positive sample traffic group corresponds to a traffic continuous label, and the negative sample traffic group corresponds to a traffic non-continuous label, and then enter step B; Step B. First, based on each positive sample traffic group and each negative sample traffic group, train a traffic time sequence detection model for detecting whether the two traffics correspond to a traffic continuous label or a traffic non-continuous label, and then, based on each positive sample traffic group and fine-tuning of the traffic time sequence detection model, train a traffic prediction model for predicting the next traffic in the adjacent future time direction, and then enter step C; The above step B includes the following steps B1 to B3; Step B1. For each sample traffic in the sample traffic set, extract the feature values corresponding to each preset malicious traffic detection feature of the sample traffic as each key feature value, and extract the feature values corresponding to each feature that affects the change of the malicious traffic detection feature of the sample traffic as each influencing feature value; Then enter step B2; Step B2. Based on each positive sample traffic group and each negative sample traffic group, use the key feature values of the two sample traffics in the sample traffic group as the input and the traffic continuous label or traffic non-continuous label corresponding to the sample traffic group as the output to train the BERT model to obtain a traffic time sequence detection model, and then enter step B3; Step B3. Based on each positive sample traffic group, use the influencing feature values of the first sample traffic in sequence in the positive sample traffic group as the input and the key feature values of the second sample traffic in sequence in the positive sample traffic group as the output to train the first BART model, and at the same time apply the traffic time sequence detection model to perform detection on the traffic continuous label or traffic non-continuous label corresponding to the key feature values of the input traffic and the key feature prediction values of the predicted traffic output during the training of the first BART model, and use the detection result to fine-tune the training of the first BART model, and then obtain the trained first BART model, which constitutes the traffic prediction model; Step C. Based on the sample traffic set, train a malicious traffic detection model for detecting whether the traffic corresponds to a malicious label or a non-malicious label; Step i. Apply the traffic prediction model to predict the next traffic adjacent to the current traffic in the future time direction, and apply the malicious traffic detection model to perform malicious detection on the next traffic.

2. The method for realizing network security threat prediction based on a large language model according to claim 1, wherein: In step B3, the flow time series detection model is applied to perform the detection of the corresponding flow continuous label or flow discontinuous label for each key feature value of the input flow and each key feature prediction value of the output predicted flow during the training of the first BART model. If the detection result is that the input flow and the output predicted flow correspond to the flow discontinuous label, the training of the first BART model is fine-tuned. If the detection result is that the input flow and the output predicted flow correspond to the flow continuous label, the training of the first BART model is not fine-tuned.

3. The method for realizing network security threat prediction based on a large language model according to claim 1, characterized in that: In step B3, first, the regression head module is added and updated at the end of the structure of the first BART model. Then, based on each positive sample flow group, with the influence feature values of the first sample flow in the positive sample flow group in order as the input and the key feature values of the second sample flow in the positive sample flow group in order as the output, the first BART model is trained.

4. The method for implementing network security threat prediction based on a large language model according to claim 1, wherein: In step C, based on each sample flow in the sample flow set, with the key feature values of the sample flow as the input and the corresponding malicious label or non-malicious label of the sample flow as the output, the second BART model is trained to obtain a malicious traffic detection model for detecting the corresponding malicious label or non-malicious label of the traffic.

5. The method for implementing network security threat prediction based on a large language model according to claim 4, wherein: In step C, first, the classification head module is added and updated at the end of the structure of the second BART model. Then, based on each sample flow in the sample flow set, with the key feature values of the sample flow as the input and the corresponding malicious label or non-malicious label of the sample flow as the output, the second BART model is trained to obtain a malicious traffic detection model for detecting the corresponding malicious label or non-malicious label of the traffic.

6. The method for implementing network security threat prediction based on a large language model according to claim 1, wherein: Step i includes the following steps i1 to i3; Step i1. Extract the feature values of the features corresponding to each preset influence malicious traffic detection feature change of the current traffic as the influence feature values corresponding to the current traffic, and then enter step i2; Step i2. Apply the traffic prediction model to process the influence feature values corresponding to the current traffic to obtain the key feature values corresponding to the next adjacent traffic in the future time direction of the current traffic, and then enter step i3; Step i3. Apply the malicious traffic detection model to process the key feature values corresponding to the next traffic to obtain the corresponding malicious label or non-malicious label of the next traffic, and perform malicious detection on the next traffic.

7. The method for realizing network security threat prediction based on a large language model according to claim 1, characterized in that: For the sample flow set composed of the preset number of sample flows corresponding to malicious labels or non-malicious labels respectively, update it according to the following preprocessing; First, remove the invalid or duplicate sample flows in the sample flow set; then perform noise filtering on each sample flow in the sample flow set; then correct the format errors of each sample flow in the sample flow set and fill in the missing values according to the preset flow format; Finally, perform standardization or normalization processing on each sample flow in the sample flow set and convert it to the preset target flow format.

8. The method for implementing network security threat prediction based on a large language model according to claim 4, wherein: A network sniffer is set at the target network interface location in the target network environment to monitor and capture traffic with a preset number corresponding to malicious tags or non-malicious tags at the target network interface location as each sample traffic, thus forming a sample traffic set.

Citation Information

Patent Citations

  • Self-supervised model training method and device for encrypted traffic threat detection

    CN117375897A

  • Traffic anomaly detection method and system based on improved BERT fusion contrast learning

    CN118013201A