Website fingerprint attack method and system based on LSTM model prediction, medium and program product
By adopting a prediction method based on LSTM model in the field of website fingerprint attack, we can capture the dynamic changes in data distribution and predict future model parameters, and solve the problem of identifying and predicting website fingerprint attacks in dynamically changing network environments, achieving higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510078013.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-30
AI Technical Summary
In the face of dynamically changing network environments, it is difficult to effectively identify and predict website fingerprint attacks, especially in the case of severe concept drift.
The prediction method based on the LSTM model is adopted to capture the dynamic changes in the data distribution through sequence learning, and to predict future model parameters, the model's performance for future data processing is improved.
It significantly improves the recognition accuracy and robustness of the website fingerprint attack model in dynamically changing network environments, and enhances the system's attack prediction capabilities and adaptability to the network environment.
Smart Images

Figure CN120074871A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and particularly relates to a website fingerprint attack method, system, medium and program product based on LSTM model prediction. Background Art
[0002] In the early research in the field of website fingerprint attacks, researchers generally used traditional machine learning methods to construct attack models. These methods usually relied on manually extracting features and combining appropriate machine learning algorithms for model training to identify the websites visited by users. Although this feature extraction method relying on expert knowledge and intuition is to a certain extent dependent on the accuracy of features, and is often more critical than the machine learning algorithms used, its inherent limitation is that it is difficult to accurately describe complex traffic patterns through a fixed set of features, especially the attack effect is limited in cross-protocol scenarios.
[0003] With the rise of deep learning technology and the emergence of its powerful ability to automatically extract abstract features from data, more and more researchers have begun to explore the application of deep learning to website fingerprint attacks, which provides a new perspective for improving attack performance. Deep learning can reveal more complex pattern relationships by learning deep features in data, thus overcoming the limitations of traditional machine learning methods to a certain extent.
[0004] However, the dynamics of the real network environment lead to changes in data distribution over time, and sometimes this change is very drastic. This phenomenon is called concept drift. Website fingerprint attacks mainly analyze encrypted traffic to train and classify models. If there is a large time interval between the data used for model training and the actual test data, the difference in data distribution will lead to a decline in model performance. Especially in the case of fewer training samples, the concept drift problem is particularly serious.
[0005] In the current research, for traffic data containing concept drift, Cheng Guang and Jiang Zhendong adopted a statistics-based method to analyze the basic characteristics of traffic data. Cheng Guang developed a method called ECDD. This method compares the traffic data within the sliding windows of two different time periods, calculates the JS divergence difference between these two windows, and then compares it with a set threshold to identify concept drift. After detecting concept drift, they utilized the labeled drift data and combined incremental learning and ensemble learning techniques to update those underperforming models. Jiang Zhendong adopted a similar strategy, identifying concept drift through the KS test within the sliding window. Once drift is detected, they would adopt a multi-view learning strategy, selecting classifiers with significant differences in different views to replace the existing classifier. On the other hand, the foreign researcher Attarian also adopted a similar technique, detecting concept drift by applying the Hoeffding inequality and monitoring concept drift by regularly evaluating the confidence interval of the sample mean of the network data stream to ensure that the model can still maintain good performance in the face of concept drift.
[0006] Although statistical methods are effective in detecting concept drift, these methods usually require a large number of labeled samples to train and update the model, which not only increases resource consumption but may also affect the generalization ability of the model. To overcome these limitations, recent research has begun to adopt strategies based on transfer learning. These strategies do not rely on a large number of labeled samples but utilize the knowledge of existing models and a small number of samples for rapid adaptation of the model.
[0007] Sirinam et al. proposed a technique called Triplets Fingerprinting (TF). This technique first constructs a triplet pre-training model through a large amount of source domain data and then uses a K-NN classifier for prediction and classification. This method can perform well even when the number of labeled samples is extremely small. Wang et al. developed a technique called AF. This technique first adjusts the feature extractor using the DANN method in the adversarial domain and then achieves model transfer through a small number of samples in the target domain, and can also achieve good results even when the number of labeled samples is extremely small. Maohua Guo proposed a method called DNNF, which extracts the deep layout fingerprint features of websites through a convolutional neural network (CNN) and uses a K-NN classifier for classification prediction. This technique can still maintain high accuracy when providing a small number of samples for each website. When dealing with transfer learning and concept drift problems, this method performs excellently, allowing the already trained model to adapt to new website fingerprint classification tasks without repeated training.
[0008] Currently, many concept drift detection methods based on transfer learning adopt a fine-tune strategy to update the model. Specifically, these methods usually first train a preliminary model using the labeled data in the source domain, and then apply this model to the target domain (i.e., the dataset where concept drift occurs) for prediction. If there is a deviation between the prediction result and the actual situation, a small amount of labeled data from the target domain will be introduced into the training set to retrain the model to adapt to the new data distribution. This fine-tune method effectively utilizes the similarity between the source domain and the target domain, avoids completely training the model from scratch, reduces the annotation cost, and improves the generalization ability of the model. In addition, the fine-tune strategy may also include measures such as dynamically adjusting the learning rate, expanding the model structure, or adding regularization to further enhance the model performance. However, the premise of adopting fine-tune is that there must be a certain similarity between the source domain and the target domain; if the difference between the two is too large, fine-tune may lead to the deterioration of the model performance. And if there is no labeled data in the target domain, the model will not work. Summary of the Invention
[0009] The object of the present invention is to provide a website fingerprint attack method, system, medium and program product based on LSTM model prediction, which adopts a sequence learning method to adaptively capture the dynamic changes of data distribution and improves the performance of the model in processing future data by predicting future model parameters.
[0010] The object of the present invention is achieved by the following technical solutions:
[0011] A website fingerprint attack prediction method based on the LSTM model, comprising the following steps:
[0012] Step 1: Website fingerprint collection;
[0013] When a user accesses a website, a website fingerprint of accessing the website will be generated. Collect the encrypted traffic generated by the user accessing the website and save it as a PCAP file for subsequent processing;
[0014] Step 2: Website fingerprint cleaning;
[0015] Clean all the traffic in the PCAP traffic file generated in Step 1, including removing noise and irrelevant data, and at the same time extracting and purifying useful information;
[0016] Step 3: Website fingerprint feature extraction;
[0017] Utilize deep learning technology to extract general features from the preprocessed and reorganized auxiliary website fingerprint data, so that the model can accurately identify and distinguish different websites;
[0018] Step 4: Website fingerprint dynamic prediction;
[0019] Use the long short-term memory (LSTM) model to predict the parameters of the future website fingerprint model; by analyzing historical data and current features, anticipate potential website fingerprint changes, enhance the adaptability to the dynamically changing network environment, and improve the anticipation ability of the website fingerprint attack prediction system.
[0020] Further, step 1 specifically includes:
[0021] Step 1.1: The traffic collection device collects traffic from the local area network environment using tshark and tcpdump network probes.
[0022] Step 1.2: The traffic collection device saves the large amount of collected traffic into multiple PCAP files according to certain rules.
[0023] Further, step 2 specifically includes:
[0024] Step 2.1: Clean all the data in the PCAP traffic files generated by the traffic collection device, removing the noise and irrelevant data, including filtering out the data of non-target websites and removing the error data packets caused by network interference.
[0025] Step 2.2: Extract and purify useful information from the already cleaned traffic data; including identifying and saving the key packet features, including the timestamp, packet length, and transmission direction; and keeping the packet length as 5000, deleting the excess and padding with 0 if less.
[0026] Further, step 3 specifically includes:
[0027] Step 3.1: Obtain the preprocessed and reorganized website fingerprint data from the auxiliary dataset and convert it into a format suitable for model input; assume that a traffic sample A is preprocessed into a packet direction sequence This sequence will be used as the input to the feature extractor.
[0028] Step 3.2: Use deep learning techniques to train the feature extractor module Φ to extract initial feature embeddings from high-dimensional data; the input sequence a is passed into the feature extractor Φ, and the output feature embedding is represented as The feature embedding Φ(a) is passed to the BDC module Its output is used for loss calculation; define the multi-similarity loss function Multi-Similarity Loss to optimize the feature extractor so that it can more effectively extract and distinguish features; the function formula of the feature extractor is as follows:
[0029]
[0030] where q represents the batch size, θ, σ, and λ are hyperparameters, and p i and N i represent the positive and negative sample sets in the samples of the i-th batch of anchor samples, and S ik represents the mutual similarity between the i-th batch and the k-th batch of samples.
[0031] Furthermore, step 4 includes the following steps:
[0032] Step 4.1: Obtain the processed and reorganized feature embedding from the pre-trained feature extractor to provide basic data for subsequent dynamic prediction;
[0033] Step 4.2: Use a dynamic graph model to simulate the evolution of the neural network over time and maintain the maximum expressive ability of the model; Let each node v represent a neuron in the neural network, and each edge e represent the connection between neurons; In the recurrent model, each unit f θ is driven by the parameter θ and predicts the current state ω i by analyzing the historical weights {w s : i < s}; Use the decoding function F ξ (·) to generate the dynamic graph ω s from the latent probability distribution h s ; Introduce a residual connection to connect the training process with previous information and alleviate the forgetting phenomenon. The specific formula is as follows:
[0034]
[0035] where ω s represents the state or feature corresponding to the current time s, τ represents the size of the sliding window, and ω s-τ:s-1 represents the historical state or feature from time s - τ to s - 1, and λ is a regularization coefficient, represents the process of weighted summation of the states or features over a past period (from s - τ to s - 1);
[0036] Step 4.3: Obtain the processed and reorganized feature embedding from the pre-trained feature extractor to provide basic data for subsequent dynamic prediction; Use the dynamic graph model to simulate the evolution of the neural network over time, combine the prediction results, and represent the change trend of the website fingerprint; Finally, predict the change of the website fingerprint based on the dynamic prediction results, identify potential attack trends in advance, and enhance the attack prediction ability of the system and the adaptability to the network environment.
[0037] A computer device / system, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the website fingerprint attack prediction system attack method based on the LSTM model.
[0038] A computer-readable storage medium stores computer programs / instructions thereon, and when the computer programs / instructions are executed by a processor, the steps of an attack method of a website fingerprint attack prediction system based on an LSTM model are implemented.
[0039] A computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of an attack method of a website fingerprint attack prediction system based on an LSTM model are implemented.
[0040] The beneficial effects of the present invention are as follows:
[0041] 1. By using two key features, namely the timestamp and direction in the data packet, virtual samples are generated through data augmentation methods in the spatial dimension and the time dimension. The spatial dimension includes operations such as rotation, masking, and erasing to simulate the situation of data packet loss in a real network environment; the time dimension includes operations such as adding random noise and scaling the time interval to simulate the random fluctuations in the data packet arrival time.
[0042] 2. A set of diversity evaluation metrics is designed, including in-sample diversity and inter-sample diversity, to comprehensively evaluate the diversity of the augmented samples. By iteratively optimizing the data augmentation parameters, the diversity of the data set is maximized, providing richer and more favorable training data for model training.
[0043] 3. By combining the features in the spatial dimension and the time dimension, through strategies such as multiplicative fusion, etc., fused features integrating spatio-temporal information are obtained, enhancing the adaptability and robustness of the model to changes in complex network environments.
[0044] . The present invention adopts the LSTM model prediction technology, adapts to concept drift through dynamic prediction, and significantly improves the recognition accuracy and robustness of the website fingerprint attack model in the face of a dynamically changing network environment. At the same time, the introduced multi-similarity loss function and residual connection strategy further optimize the performance and generalization ability of the model.. Description of the Drawings
[0045] Figure 1 Flowchart of the attack method of the website fingerprint attack prediction system based on the LSTM model of the present invention;
[0046] Figure 2 Overall framework diagram of the model;
[0047] Figure 3 Schematic diagram of the pre-training feature extraction stage;
[0048] Figure 4 Schematic diagram of the feature extractor module;
[0049] Figure 5 Schematic diagram of the multi-similarity loss example;
[0050] Figure 6 Schematic diagram of the dynamic prediction adaptation concept drift stage;
[0051] Figure 7 Attack success rate at different time nodes;
[0052] Figure 8 Analysis diagram of the influence of the depth of the classifier model. Detailed implementation manners
[0053] The present invention will be further described below with reference to the accompanying drawings.
[0054] Figure 1 A website fingerprint attack prediction method based on the LSTM model according to the present invention comprises the following steps:
[0055] Step 1: Website fingerprint collection;
[0056] When a user accesses a website, website fingerprints for accessing the website are generated. The encrypted traffic generated by the user accessing the website is collected and saved as a PCAP file for subsequent processing;
[0057] Step 2: Website fingerprint cleaning;
[0058] All traffic in the PCAP traffic file generated in Step 1 is cleaned, including removing noise and irrelevant data, and at the same time extracting and purifying useful information;
[0059] Step 3: Website fingerprint feature extraction;
[0060] Using deep learning technology, general features are extracted from the preprocessed and reorganized auxiliary website fingerprint data, enabling the model to accurately identify and distinguish different websites;
[0061] Step 4: Website fingerprint dynamic prediction;
[0062] The long short-term memory (LSTM) model is used to predict the parameters of the future website fingerprint model; by analyzing historical data and current features, potential changes in website fingerprints are predicted, enhancing the adaptability to the dynamically changing network environment and improving the prediction ability of the website fingerprint attack prediction system.
[0063] A website fingerprint attack prediction system based on the LSTM model according to the present invention mainly comprises four devices: a website fingerprint collection device, a website fingerprint cleaning device, a website fingerprint feature extraction device, and a website fingerprint dynamic prediction device.
[0064] Figure 2It shows the overall framework diagram of the model, which is mainly divided into two stages. In the first stage, the pre-trained feature extractor is trained using the auxiliary dataset and multi-similarity loss. In the second stage, based on the pre-trained feature extractor, the model parameters at multiple time nodes are learned to predict the model parameters at future time nodes, thus solving the concept drift problem existing in the field of website fingerprint attacks.
[0065] Figure 3 The following shows the framework diagram of the pre-training stage of the present invention. The feature extractor part is trained through the auxiliary dataset and multi-similarity loss. The BDC module only performs data conversion and does not participate in the training.
[0066] Figure 4 The following shows the schematic diagram of the feature extractor of the present invention, which mainly includes eight groups of convolution, pooling, and activation.
[0067] Figure 6 It shows the training stage and the testing stage in Stage 2. In the training stage, the model parameters at several time points in the training stage are learned through LSTM to predict the model parameters in the future time domain.
[0068] To verify the effectiveness of the present invention in implementing website fingerprint attacks, a dataset of several time points of the same website collected in a real network environment is used.
[0069] The present invention designs four different experiments to verify the website fingerprint attack ability of the present invention. It mainly includes:
[0070] 1. Experiment 1: WFLP effectiveness experiment: The experiment aims to test the effectiveness of the WFLP model and compare it with the baseline experiments TF, DNNF, and WFBDC. The AWFTime dataset is used. 100 . The training dataset uses the sum of data at each time point of 3 days, 10 days, 2 weeks, and 4 weeks. The number of samples per class at each time point is 1 or 5. The testing dataset uses the data at the 6-week time point. The number of samples per class is 70. The evaluation metric is the model attack success rate.
[0071] The experimental results are shown in Table 1 below. In the WFLP effectiveness experiment, four different attack models were compared: TF, DNNF, WFBDC, and WFLP. The WFLP model performed optimally when the sample size was either 1 or 5. When the sample size was 1, the attack success rate of WFLP reached 89.65%, and when the sample size was 5, it reached 91.84%. This indicates that the WFLP model can maintain a high recognition accuracy under various sample size configurations, effectively improving the applicability and robustness of the model. The attack success rate of the TF model was 82.48% when the sample was 1 and 86.06% when the sample size was 5. Although its performance improved with an increase in the sample size, there was still a certain gap compared to the WFLP model. The performance of the DNNF model was relatively low, with an attack success rate of 76.28% when the sample size was 1 and 84.19% when the sample size was 5. The attack success rate of the WFBDC model was 80.52% when the sample size was 1, and it increased significantly to 90.76% when the sample size was 5. This shows that the WFBDC model can significantly improve its recognition ability when processing more samples, especially performing close to WFLP at high sample sizes.
[0072] Table 1
[0073]
[0074] It can be seen from the experimental results that the WFLP model outperforms other models under different sample size configurations, demonstrating its effectiveness in the field of website fingerprinting attacks. Compared with other models, the WFLP model has obvious advantages in processing single samples, which is particularly important for scenarios with limited data volume in practical applications.
[0075] 2. Experiment 2: WFLP Robustness Experiment: This experiment tested the performance of the WFLP model at different time points and compared it with the baseline experiments TF, DNNF, and WFBDC. The test times set in the experiment were 2 weeks, 4 weeks, and 6 weeks respectively. For each test time node, the training samples were the cumulative samples of all previous time stages before the test time node, and the sample size for each time stage was 1 or 5. The number of samples for each category at the test time node was 70, and the evaluation index was the model attack success rate to examine the stability of the model under time changes.
[0076] The experimental results are shown in Table 2 below. The WFLP method always maintained a high accuracy rate, demonstrating its excellent adaptability and robustness in combating concept drift. Figure 7 This view can also be proven. The horizontal axis represents the sample size, and the vertical axis represents the model attack success rate. Among them, (a) shows the results at the 2-week time point; (b) shows the results at the 4-week time point; (c) shows the results at the 6-week time point.
[0077] At the 2-week time point, the accuracy rate of the WFLP method was 88.31% when the sample size was 1, and it increased to 89.66% when the sample size increased to 5. In contrast, although the performance of TF, DNNF, and WFBDC showed competitiveness, the performance of WFLP was more excellent. This indicates that WFLP can effectively cope with slight changes in data distribution in a relatively short time and improve the accuracy of website fingerprint recognition.
[0078] By the 4-week time point, the WFLP method continued to maintain its advantage in accuracy rate, reaching 88.81% when the sample size was 1 and 90.76% when the sample size increased to 5. Although WFBDC had similar performance in some tests, the overall performance of WFLP was more stable under different sample sizes, showing its effective adaptation to concept drift within a medium time range.
[0079] In the long-term time interval test of 6 weeks, WFLP continued to demonstrate its adaptability and robustness, with the accuracy rate reaching 89.65% and 91.84% when the sample sizes were 1 and 5 respectively. Although the performance of WFLP was similar to that of WFBDC when the sample size was 5, overall, WFLP still showed better performance.
[0080] Through comprehensive analysis of the experimental results at each time point, it can be clearly seen that the WFLP method demonstrated better adaptability and robustness than TF, DNNF, and WFBDC in the face of concept drift challenges with different time spans. This finding not only emphasizes the advantages of the WFLP method in dealing with dynamically changing environments but also provides strong experimental support for the further development of future website fingerprint recognition technologies.
[0081] Table 2
[0082]
[0083]
[0084] 3. Experiment 3: Experiment on the impact of the depth L of the classifier model on the model performance: The experimental results are as Figure 8 , the optimal hidden layer depth is 2, and the changing trend of the model performance with the number of hidden layers is an inverted "U" shape. The results show that increasing a small number of hidden layers can significantly improve the expression ability of the model. This is mainly because it enhances the model's learning ability for data, allowing it to capture more complex data features and internal patterns, thereby improving the classification accuracy. This performance improvement is attributed to the ability of deeper network structures to identify and simulate complex structures and relationships in the data, which is particularly crucial when solving highly nonlinear problems.
[0085] However, as the number of hidden layers further increases, the performance starts to decline, which may indicate that the model begins to overfit. Overfitting occurs because the model structure is too complex and starts to capture random noise in the training data rather than the true patterns of the underlying data distribution. This not only reduces the generalization ability of the model on new data but also may make the model overly sensitive to specific noise or anomalies in the training data. Additionally, overly deep models may lead to problems such as vanishing or exploding gradients during training, which seriously affect the stability and efficiency of model training.
[0086] In summary, choosing an appropriate model depth is the key to optimizing the model structure to balance its expressive ability and generalization ability. An appropriate number of layers can ensure that the model is complex enough to fully learn and represent data features while avoiding overfitting and training stability problems caused by excessive complexity. In practical applications, through careful experiments and verification, determining the optimal number of layers becomes an important task in model design.
[0087] 4. Experiment 4: Ablation Experiment: This experiment evaluates the role of each module in the WFLP model through ablation analysis. Specifically, the experiment design removes the pre-training module (labeled as WFLP-C), the LSTM sequential learning mechanism (labeled as WFLP-L), and the residual connection function (labeled as WFLP-S). Each model will be compared with the complete model with all functions (labeled as WFLP) in terms of performance to determine the contribution of each module to the overall model performance. The dataset used is AWFTime. 100 The training dataset covers data at various time points from 3 days to 4 weeks, with the number of samples per class at each time point being 1, and the test set is data at the 6-week time point, with the number of samples per class being 70. The evaluation metric is also the attack success rate of the model.
[0088] The experimental results are shown in Table 3 below. The result analysis shows that after removing the pre-training module (WFLP-C), the performance of the model drops significantly, with accuracies of 14.24%, 30.71%, and 40.22% when the number of samples is 1, which highlights the importance of the pre-trained feature extractor. In the standard setting, the feature extractor learns the key features in the training data and freezes its parameters, enabling the classifier to make effective predictions by adjusting only a few parameters. If this step is omitted, the model needs to learn from scratch without the guidance of initial knowledge, which significantly increases the learning difficulty and thus reduces the overall performance.
[0089] The WFLP-L model, which removes the sequential learning mechanism of LSTM, shows obvious limitations in its performance. This result emphasizes the key role of the sequential learning mechanism in tasks that process time-dependent or sequential data. This mechanism helps the model better understand and predict sequential dynamic changes by considering the temporal order of the data.
[0090] Table 3
[0091]
[0092] On the other hand, the WFLP-S model removes the residual connections but retains the sequential learning mechanism, and its performance is better than that of WFLP-L. This indicates that the sequential learning mechanism can effectively improve the model performance even in the absence of direct long-distance information transmission with residual connections. This shows that the combination of pre-training and sequential processing still provides the model with sufficient information to adapt to the dynamic changes of the data.
[0093] The complete model WFLP, integrating all the modules, shows the best performance under all settings. This further verifies the integrated advantages of the pre-training module, the sequential learning mechanism, and the residual connections in enhancing the model's adaptation to dynamic data changes. The collaborative work of these three components significantly enhances the adaptability and robustness of the model, especially in dealing with complex or continuously changing environments.
[0094] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A website fingerprint attack prediction method based on LSTM model, characterized by: The following steps are involved: Step 1: Website fingerprint collection; When a user visits a website, a fingerprint of the website is generated. The encrypted traffic generated by the user visiting the website is collected and saved as a PCAP file for subsequent processing; Step 2: Website fingerprint cleaning; Clean all traffic in the PCAP traffic file generated in step 1, including removing noise and irrelevant data, and extracting and purifying useful information; Step 3: Website fingerprint feature extraction; Using deep learning technology, common features are extracted from the pre-processed and reorganized auxiliary website fingerprint data, so that the model can accurately identify and distinguish different websites; Step 4: Dynamic prediction of website fingerprint; Use the long short-term memory (LSTM) model to predict future website fingerprint model parameters; by analyzing historical data and current features, predict potential website fingerprint changes, enhance the ability to adapt to dynamically changing network environments, and improve the predictive ability of the website fingerprint attack prediction system.
2. According to a website fingerprint attack prediction method based on LSTM model according to claim 1, it is characterized in that: The step 1 specifically includes: Step 1.1: The traffic collection device will use tshark and tcpdump network probes to collect traffic from the LAN environment; Step 1.2: The traffic collection device saves the collected large amount of traffic into multiple PCAP files according to certain rules.
3. According to a website fingerprint attack prediction method based on LSTM model according to claim 1, it is characterized in that: The step 2 specifically includes: Step 2.1: Clean all data in the PCAP traffic file generated by the traffic collection device to remove noise and irrelevant data, including filtering out data from non-target websites and removing erroneous data packets caused by network interference; Step 2.2: Extract and purify useful information from the cleaned traffic data; including identifying and saving key packet features, including timestamp, packet length and transmission direction; and keep the packet length at 5000, deleting any excess and adding 0 if it is less.
4. The website fingerprint attack prediction method based on the LSTM model according to claim 1 is characterized in that: The step 3 specifically includes: Step 3.1: Obtain the preprocessed and reorganized website fingerprint data from the auxiliary dataset and convert it into a format suitable for model input; suppose a traffic sample A is preprocessed into a packet direction sequence This sequence will serve as input to the feature extractor; Step 3.2: Use deep learning techniques to train a feature extractor module Φ to extract initial feature embeddings from high-dimensional data; the input sequence a is passed to the feature extractor Φ, and the output feature embedding is represented as The feature embedding Φ(a) is passed to the BDC module Its output It is used for loss calculation; Multi-Similarity Loss is defined to optimize the feature extractor so that it can extract and distinguish features more effectively; the function formula of the feature extractor is as follows: Among them, q represents the batch size, θ, σ, λ are hyperparameters, and p i and N i represents the set of positive and negative samples in the i-th batch of anchor samples, S ik It represents the mutual similarity between the i-th batch and the k-th batch of samples.
5. According to a website fingerprint attack prediction system attack method based on LSTM model in claim 1, it is characterized in that: The step 4 comprises the following steps: Step 4.1: Obtain processed and reorganized feature embeddings from the pre-trained feature extractor to provide basic data for subsequent dynamic prediction; Step 4.2: Use a dynamic graph model to simulate the evolution of the neural network over time, maintaining the maximum expressive power of the model; let each node v represent a neuron in the neural network, and each edge e represent the connection between neurons; in the recurrent model, each unit f θ is driven by the parameter θ, and predicts the current state ω i by analyzing the historical weights {w s : i < s}; use the decoding function F ξ (·) to generate the dynamic graph ω s from the latent probability distribution h s ; introduce residual connections to connect the training process with previous information and alleviate the forgetting phenomenon. The specific formula is as follows: Among them, ω s represents the state or feature corresponding to the current moment s, τ represents the size of the sliding window, ω s-τ:s-1 represents the historical state or feature from time s-τ to s-1, λ is a regularization coefficient, It represents the process of weighted summation of the state or features over a period of time (from s-τ to s-1); Step 4.3: Obtain processed and reorganized feature embeddings from the pre-trained feature extractor to provide basic data for subsequent dynamic predictions; use the dynamic graph model to simulate the evolution of the neural network over time, and combine the prediction results to represent the changing trend of the website fingerprint; finally, predict the changes in website fingerprints based on the dynamic prediction results, identify potential attack trends in advance, and enhance the system's attack prediction capabilities and adaptability to the network environment.
6. A computer device / equipment / system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 5.
7. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer program product comprising a computer program / instructions, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.