Multi-label data traffic classification method, model training method and device

By adopting a multi-label data traffic classification model of a dual-channel label semantic guide encoder and a dynamic weighted label sequence decoder in data traffic classification, the problem of low accuracy of data traffic classification in the prior art is solved, and higher classification accuracy and adaptability are achieved.

CN120180299APending Publication Date: 2025-06-20SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237535.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing data traffic classification methods require a lot of manual annotation and time, making it difficult to mine deep semantic information, and fail to effectively consider the interactive relationship between label semantic information and data traffic semantic information, resulting in low classification accuracy.

Method used

The multi-label data traffic classification model is adopted, and the multi-level semantic guidance encoder and dynamic weighted label sequence decoder of the data traffic are interacted with multi-level semantic information and tag semantic information, dynamically adjust the label weight, and improve classification accuracy.

Benefits of technology

The performance and adaptability of the multi-label data traffic classification model are improved, especially in the network environment with a lot of changes and noise, which effectively improves the accuracy and robustness of data traffic classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180299A_ABST
    Figure CN120180299A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-label data traffic classification method and device and a model training method and device, and relates to the technical field of data traffic classification, and the method comprises the steps: constructing a multi-label data traffic classification model composed of a dual-channel label semantic guidance encoder and a dynamic weighted label sequence decoder; a channel label semantic guidance encoder is used for interacting multi-level semantic information corresponding to original data traffic in the multi-label data traffic classification data set with label semantic information, and label-guided multi-semantic-level data traffic information is obtained; and performing label prediction and weight dynamic adjustment by using a dynamic weighting label sequence decoder to obtain a trained multi-label data traffic classification model. According to the method, the performance of the multi-label data traffic classification model is improved, the adaptability of the multi-label data traffic classification model to complex data traffic is enhanced, and the accuracy and robustness of data traffic classification can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data traffic classification, and particularly relates to a multi-label data traffic classification method, a model training method and a device. Background Art

[0002] The Internet technology revolution (such as fiber optic and 5G connections), convenient access to multiple devices, competitive pricing strategies, and the diverse services provided by the Internet have made it an indispensable part of the daily operations of individuals and enterprises. This has imposed a heavy burden on Internet service providers, who need to find solutions to network traffic classification and identification that trigger applications, protocols, or services to meet goals such as quality of service, content filtering, legal interception, and malicious behavior identification.

[0003] Related data traffic classification methods require a large amount of manual prior knowledge and time for annotation. The extracted features are difficult to mine deep semantic information, lacking data traffic semantic information, unable to meet applications in a large-scale network traffic environment, and not considering the interaction relationship between label semantic information and local data traffic semantic information, resulting in low data traffic classification accuracy. Summary of the Invention

[0004] The present application provides a multi-label data traffic classification method, a model training method and a device to at least solve the problem of low data traffic classification accuracy in related technologies.

[0005] The present application provides a method for training a multi-label data traffic classification model, including:

[0006] Constructing a multi-label data traffic classification model; the multi-label data traffic classification model is composed of a dual-channel label semantic guided encoder and a dynamic weighted label sequence decoder;

[0007] Obtaining a multi-label data traffic classification data set, and using the channel label semantic guided encoder to interact the multi-level semantic information and label semantic information corresponding to the original data traffic in the multi-label data traffic classification data set to obtain label-guided multi-semantic level data traffic information;

[0008] Based on the label-guided multi-semantic level data traffic information, using the dynamic weighted label sequence decoder for label prediction and weight dynamic adjustment to obtain the trained multi-label data traffic classification model.

[0009] The present application also provides a multi-label data traffic classification method, which is implemented based on any of the above-mentioned multi-label data traffic classification model training methods. The method includes:

[0010] Collecting the data traffic to be classified and preprocessing the data traffic to be classified;

[0011] Input the preprocessed data traffic to be classified into the trained multi-label data traffic classification model to obtain the multi-label data traffic classification result.

[0012] This application also provides a training device for a multi-label data traffic classification model, including:

[0013] A construction module for constructing a multi-label data traffic classification model; the multi-label data traffic classification model is composed of a dual-channel label semantic guidance encoder and a dynamic weighted label sequence decoder;

[0014] An interaction module for obtaining a multi-label data traffic classification data set, and interacting the multi-level semantic information and label semantic information corresponding to the original data traffic in the multi-label data traffic classification data set by using the channel label semantic guidance encoder to obtain label-guided multi-semantic-level data traffic information;

[0015] An adjustment module for performing label prediction and weight dynamic adjustment by using the dynamic weighted label sequence decoder based on the label-guided multi-semantic-level data traffic information to obtain the trained multi-label data traffic classification model.

[0016] This application also provides a multi-label data traffic classification device, including:

[0017] A preprocessing module for collecting the data traffic to be classified and preprocessing the data traffic to be classified;

[0018] A classification module for inputting the preprocessed data traffic to be classified into the trained multi-label data traffic classification model to obtain the multi-label data traffic classification result.

[0019] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above multi-label data traffic classification model training methods or the steps of the above multi-label data traffic classification method when executing the computer program.

[0020] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of any of the above multi-label data traffic classification model training methods or the steps of the above multi-label data traffic classification method.

[0021] Through this application, a multi-label data traffic classification model is constructed; the multi-label data traffic classification model consists of a dual-channel label semantic-guided encoder and a dynamic weighted label sequence decoder; a multi-label data traffic classification dataset is obtained, and the multi-level semantic information and label semantic information corresponding to the original data traffic in the multi-label data traffic classification dataset are interacted using the channel label semantic-guided encoder to obtain label-guided multi-semantic-level data traffic information; based on the label-guided multi-semantic-level data traffic information, the dynamic weighted label sequence decoder is used for label prediction and weight dynamic adjustment to obtain the trained multi-label data traffic classification model; the dual-channel label semantic-guided encoder interacts with label semantic information using semantic information at different levels to obtain multi-level data traffic representations, effectively integrating the rich semantic information in the data traffic, and capturing and expressing features at different semantic levels through the semantic guidance of labels; secondly, the dynamic weighted label sequence decoder adopts a dynamic weighting strategy in the dynamic classification layer, which helps to alleviate the exposure bias problem. By dynamically adjusting the influence of each label in the classification process, the dynamic weighted label sequence decoder can treat the classification decisions of different labels more balancedly, making the model more objective and stable when processing multi-label data traffic; it not only improves the performance of the multi-label data traffic classification model, but also enhances the adaptability of the multi-label data traffic classification model to complex data traffic. Especially in the face of an actual network environment with more changes and noises, it can effectively improve the accuracy and robustness of data traffic classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a schematic flowchart of a training method for a multi-label data traffic classification model provided by an embodiment of the present application;

[0024] Figure 2 It is a schematic flowchart of another training method for a multi-label data traffic classification model provided by an embodiment of the present application;

[0025] Figure 3 It is a model framework diagram of a multi-label data traffic classification model provided by an embodiment of the present application;

[0026] Figure 4 It is a schematic flowchart of yet another training method for a multi-label data traffic classification model provided by an embodiment of the present application;

[0027] Figure 5Schematic flowchart of a multi-label data traffic classification method provided by an embodiment of the present application;

[0028] Figure 6 Schematic flowchart of the multi-label data traffic classification processing provided by an embodiment of the present application;

[0029] Figure 7 Schematic flowchart of a multi-label data traffic classification method with dual-channel label semantics guidance provided by an embodiment of the present application;

[0030] Figure 8 Block diagram of the structure of a training device for a multi-label data traffic classification model provided by an embodiment of the present application;

[0031] Figure 9 Block diagram of the structure of a multi-label data traffic classification device provided by an embodiment of the present application;

[0032] Figure 10 Schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0034] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0035] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0036] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the multi-label data traffic classification method and the training method of the multi-label data traffic classification model depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0037] This application is applied to an electronic device, which can be a server or a terminal in fields such as the network security field, the medical and health field, and the financial field. For example, in the network security field, this application can improve the accuracy and efficiency of data traffic classification; in the medical and health field, this application can be used for the classification and diagnosis of medical images. By integrating multi-level medical information and semantic guidance of disease tags, it helps doctors analyze and diagnose the patient's condition more accurately; this application can also be applied to risk assessment and fraud detection in the financial field. By combining multi-dimensional data of customers and semantic information of transaction tags, it improves the ability to identify abnormal behaviors and fraud activities, thereby effectively protecting the security and stability of the financial system. Therefore, this application not only helps to improve network security but also brings intelligent and precise application solutions in multiple fields.

[0038] As the economy turns to digital form, malicious behavior recognition becomes particularly crucial. It is necessary to prevent cybercriminals from conducting malicious attacks through the Internet for profit. Therefore, the security of the network system is of utmost importance. It is very necessary to accurately and timely classify the traffic in the network environment and identify potential threats.

[0039] The related traffic classification methods are mainly divided into problem transformation methods and algorithm adaptation methods.

[0040] In the technical solution of the problem transformation method for data traffic classification, the main idea of the problem transformation method for data traffic classification is: converting the multi-label data traffic classification task into other task problems, mainly converting the multi-label data traffic classification task into a binary classification problem, a multi-classification problem, and converting the multi-label data traffic classification task into a data traffic sequence to label sequence modeling task; its core idea is to modify the data structure to adapt to the input and output of existing machine learning algorithms, so as to achieve the effectiveness of task conversion and problem solving.

[0041] The disadvantages of the problem transformation method for data traffic classification are: when data needs to be reconstructed to adapt to the data of machine learning algorithms, it usually requires a large amount of manual prior knowledge and time for annotation; in addition, during the process of data feature construction, different features often show an orthogonal state in the spatial vector, which leads to a large number of zero values in the constructed feature space. As a result, the data often faces the problem of feature sparsity, and it is difficult to completely avoid even through feature engineering expansion.

[0042] In the technical solution of the algorithm adaptation method for data traffic classification, the core idea of the algorithm adaptation method for data traffic classification lies in improving the existing machine learning algorithms to make them adapt to the data of the multi-label data traffic classification task. The main methods include: Ranking Support Vector Machine (Rank-SVM) and Multi-Label K-Nearest-Neighborhood (ML-KNN). Rank-SVM introduces the idea of support vector machine into multi-label data traffic classification, processes multi-label data traffic classification data by improving the maximum margin strategy, optimizes a set of linear classifiers using the minimization of empirical ranking loss, and evaluates the model's ranking ability for relevant and irrelevant labels. ML-KNN draws on the idea of K-nearest neighbors, finds K nearest neighbor samples, uses Bayesian conditional probability to calculate the probability of each label, and takes the label with the highest probability as the classification label of the sample. This method does not consider the correlation between labels.

[0043] The disadvantages of the algorithm adaptation method for data traffic classification are as follows: Although Rank-SVM attempts to handle the label ranking problem by optimizing the ranking loss, its complex non-convex optimization has a high computational cost and strong dependence on data quality and distribution, and it fails to directly model the correlation between labels. In contrast, although ML-KNN is simple, intuitive, and easy to understand, its computational complexity is high, and based on the label independence assumption, it does not consider the potential correlation between labels. Therefore, in practical applications, it may face performance degradation problems when the data distribution is unbalanced or there is too much noise.

[0044] In summary, the problem transformation method usually requires a large amount of manual prior knowledge and time for annotation. The data often faces the problem of feature sparsity and cannot meet the application requirements in a large-scale network traffic environment. The algorithm adaptation method fails to directly model the correlation between labels, does not consider the potential correlation between labels, and obtains biased data traffic semantic information related to the classification task, resulting in category understanding deviation.

[0045] Therefore, accurate data traffic classification is crucial for ensuring network information security. In-depth research on data traffic classification methods has important research significance and can effectively help prevent the occurrence of network security incidents.

[0046] According to an embodiment of the present invention, an embodiment of a training method for a multi-label data traffic classification model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0047] In this embodiment, a training method for a multi-label data traffic classification model is provided, which can be used in the above-mentioned electronic device. Figure 1 It is a flowchart of a training method for a multi-label data traffic classification model according to an embodiment of the present invention, as Figure 1 shown. The process includes the following steps:

[0048] Step S101, construct a multi-label data traffic classification model; the multi-label data traffic classification model is composed of a dual-channel label semantic guidance encoder and a dynamic weighted label sequence decoder.

[0049] Step S102, obtain a multi-label data traffic classification data set, and use the channel label semantic guidance encoder to interact the multi-level semantic information and label semantic information corresponding to the original data traffic in the multi-label data traffic classification data set to obtain label-guided multi-semantic-level data traffic information.

[0050] Specifically, the multi-label data traffic classification data set D is expressed as:

[0051] D = {(x (1) , y (1) ),..., (x (N) , y (N) )} (1)

[0052] Among them, N represents the number of sample groups in the multi-label data traffic classification data set, Y = {y1,..., y n} represents the label space with n labels, x (z) = (W1,..., W s ) represents that the z-th data traffic in the multi-label data traffic classification data set has s features, and W s represents the word vector corresponding to the s-th feature in the z-th data traffic, and y (z) is the label subset scalar of Y.

[0053] Furthermore, the dual-channel label semantic guidance encoder uses semantic information at different levels to interact with label semantic information to obtain multi-level data traffic representations and obtain label-guided multi-semantic-level data traffic information.

[0054] Step S103, based on the label-guided multi-semantic-level data traffic information, use the dynamic weighted label sequence decoder to perform label prediction and weight dynamic adjustment to obtain the trained multi-label data traffic classification model.

[0055] Specifically, in order to reduce the dependence on the label order and mitigate the exposure bias problem, a method different from using a fully connected layer and linear addition on the encoder output and decoder hidden layer state to obtain the output vector is redesigned. By dynamically adjusting the weights, some hidden layers are involved in the operation to mitigate the exposure bias problem.

[0056] The training method of a multi-label data traffic classification model provided in this embodiment enables the dual-channel label semantic-guided encoder to perform label interaction at two different semantic granularities of local data traffic semantic information and global data traffic semantic information, obtaining label-guided multi-semantic-level data traffic information. In addition, the dynamic weighted label sequence decoder reduces the dependence on the label order by adjusting the output weights of the decoder hidden layer, mitigating the exposure bias problem. This not only improves the performance of the multi-label data traffic classification model but also enhances its adaptability to complex data traffic. Especially in the face of an actual network environment with more changes and noise, it can effectively improve the accuracy and robustness of data traffic classification.

[0057] In this embodiment, a training method of a multi-label data traffic classification model is provided, which can be used in the above-mentioned electronic device. Figure 2 It is a flowchart of a training method of a multi-label data traffic classification model according to an embodiment of the present invention, as Figure 2 shown. The process includes the following steps:

[0058] Step S201, construct a multi-label data traffic classification model; the multi-label data traffic classification model is composed of a dual-channel label semantic-guided encoder and a dynamic weighted label sequence decoder. For details, please refer to Figure 1 step S101 of the shown embodiment, which will not be elaborated here.

[0059] Step S202, obtain a multi-label data traffic classification data set, and use the channel label semantic-guided encoder to interact the multi-level semantic information and label semantic information corresponding to the original data traffic in the multi-label data traffic classification data set to obtain label-guided multi-semantic-level data traffic information.

[0060] Specifically, as Figure 3 shown, the dual-channel label semantic-guided encoder includes a local filtering channel, a global filtering channel, and an adaptive gate. The local filtering channel includes a convolutional neural network (Convolutional Neural Networks, abbreviated as CNN) and a local interaction attention layer. The global filtering channel includes a bidirectional long short-term memory network (Bidirectional LongShort-Term Memory, abbreviated as Bi-LSTM) and a global interaction attention layer; the above step S202 includes:

[0061] Step S2021, use a convolutional neural network to extract semantic features from the original data traffic to obtain local data traffic semantic information.

[0062] Specifically, the convolutional neural network is used to model the local data traffic semantic information of the data traffic. That is, the convolutional neural network captures the local data traffic semantic information from the window context through the convolutional layer and the non-linear activation function. The local data traffic semantic information A obtained by the convolutional neural network is as follows:

[0063] A = CNN(W) (2)

[0064] Where CNN(.) represents one-dimensional convolution, batch normalization, and non-linear ReLu activation function operations, and W represents the original data traffic.

[0065] Step S2022, use a local interaction attention layer to establish the relationship between the label semantic information and the local data traffic semantic information to obtain label-guided local data traffic semantic information.

[0066] Specifically, the local interaction attention layer can filter out the local data traffic semantic information related to label description and classification tasks, use the label semantic information to guide the local data traffic semantic information captured by the convolutional neural network, obtain the local data traffic semantic information most relevant to the classification task, and then realize filtering out irrelevant and redundant content in the local data traffic semantic information. Among them, the local interaction attention layer uses the multi-head attention mechanism to capture the relationship between the label semantic information and the local data traffic semantic information.

[0067] In some optional embodiments, the above step S2022 includes:

[0068] Step a1, use the parameter matrix of the multi-head attention in the local interaction attention layer to transform the local data traffic semantic information and the label semantic information to obtain the query vectors corresponding to the multi-head attention respectively.

[0069] Step a2, use the parameter matrix of the multi-head attention to transform the label semantic information to obtain the key vectors and value vectors corresponding to the multi-head attention respectively.

[0070] Step a3, based on the query vectors, key vectors, and value vectors corresponding to the multi-head attention respectively, map the label semantic information and the local data traffic semantic information to obtain the mapped feature vectors corresponding to the multi-head attention respectively.

[0071] Step a4, splice the mapped feature vectors corresponding to the multi-head attention respectively to obtain label-guided local data traffic semantic information.

[0072] Specifically, the label-guided local data traffic semantic information L can be expressed as:

[0073]

[0074] L = Concat(head1, head2,..., head r )W O (4)

[0075] where head r represents the r-th attention head in the local interactive attention layer, A represents the local data traffic semantic information, C represents the label semantic information, respectively represent the query vector, value vector, and key vector corresponding to the multi-head attention in the local interactive attention layer, R represents a real number matrix, d represents the dimension of the vector, r represents the number of attention heads, and W O represents the parameter matrix of the multi-head attention in the local interactive attention layer.

[0076] Step S2023, use a bidirectional long short-term memory network to extract semantic features from the original data traffic to obtain global data traffic semantic information.

[0077] Specifically, use a bidirectional long short-term memory network to capture the global data traffic semantic information, that is, input the original data traffic into the bidirectional long short-term memory network. At time q, the forward and backward hidden states obtained are as follows:

[0078]

[0079] where represents the backward semantic feature at time q corresponding to any input feature vector, represents the backward semantic feature at time p - 1 corresponding to any input feature vector, w q represents the input feature vector at time q in the original data traffic, represents the forward semantic feature at time q corresponding to any input feature vector, represents the forward semantic feature at time p + 1 corresponding to any input feature vector, h q represents the hidden state of the q-th word in the original data traffic, where the word is a single data packet or data unit in the original data traffic, and each word contains a specific input feature vector w q , and the input feature vector can represent specific attributes of the data packet, such as source IP address, destination IP address, port number, protocol type, etc.

[0080] Furthermore, concatenate the hidden states of each word to obtain the global data traffic semantic information. The global data traffic semantic information H can be expressed as:

[0081] H = {h1, h2,.., h q} (8)

[0082] Step S2024: Use the global interactive attention layer to establish the relationship between the label semantic information and the global data flow semantic information, so as to obtain the label-guided global data flow semantic information.

[0083] Specifically, the steps to obtain the label-guided global data flow semantic information using the global interactive attention layer are the same as steps a1 - a4 above, where the label-guided global data flow semantic information G can be expressed as:

[0084]

[0085] G = Concat(head′1, head′2,..., head′ r )W O′ (10)

[0086] where head′ r represents the r-th attention head in the global interactive attention layer, H represents the global data flow semantic information, respectively represent the query vector, value vector, and key vector corresponding to the multi-head attention in the global interactive attention layer, W O′ represents the parameter matrix of the multi-head attention in the global interactive attention layer.

[0087] Step S2025: Use the adaptive gate to fuse the label-guided local data flow semantic information and the label-guided global data flow semantic information to obtain the label-guided multi-semantic level data flow information.

[0088] Specifically, use the adaptive gate to fuse the label-guided global data flow semantic information G and the label-guided local data flow semantic information L to obtain the label-guided multi-semantic level data flow information. The label-guided multi-semantic level data flow information U can be expressed as:

[0089] U = μ⊙L + (1 - μ)⊙G (11)

[0090] where μ represents the adaptive gate coefficient, and the adaptive gate coefficient controls the fusion ratio of G and L and is automatically learned during the training phase.

[0091] Step S203: Based on the label-guided multi-semantic level data flow information, use the dynamic weighted label sequence decoder to perform label prediction and weight dynamic adjustment to obtain the trained multi-label data flow classification model. For details, please refer to Figure 1 Step S103 of the embodiment shown, which will not be elaborated here.

[0092] A training method for a multi-label data traffic classification model provided in this embodiment. The dual-channel label semantic guidance encoder uses semantic information at different levels to interact with label semantic information to obtain multi-level data traffic representations, and finally forms label-guided multi-semantic-level data traffic information. The dual-channel label semantic guidance encoder can comprehensively utilize the rich semantic information in the data traffic and combine the semantic guidance of the labels to effectively capture and express the characteristics of the data traffic at different semantic levels, thereby improving the accuracy and generalization ability of the data traffic classification task.

[0093] In this embodiment, a training method for a multi-label data traffic classification model is provided, which can be used in the above-mentioned electronic device. Figure 4 It is a flowchart of a training method for a multi-label data traffic classification model according to an embodiment of the present invention, as Figure 4 shown. The process includes the following steps:

[0094] Step S401, construct a multi-label data traffic classification model; the multi-label data traffic classification model is composed of a dual-channel label semantic guidance encoder and a dynamic weighted label sequence decoder. For details, please refer to Figure 2 step S201 of the shown embodiment, which will not be elaborated here.

[0095] Step S402, obtain a multi-label data traffic classification data set, and use the channel label semantic guidance encoder to interact the multi-level semantic information and label semantic information corresponding to the original data traffic in the multi-label data traffic classification data set to obtain label-guided multi-semantic-level data traffic information. For details, please refer to Figure 2 step S202 of the shown embodiment, which will not be elaborated here.

[0096] Step S403, based on the label-guided multi-semantic-level data traffic information, use the dynamic weighted label sequence decoder to perform label prediction and weight dynamic adjustment to obtain the trained multi-label data traffic classification model.

[0097] Specifically, as Figure 3 shown, the dynamic weighted label sequence decoder includes an attention mechanism layer, a hidden layer, and a dynamic classification layer; the above step S403 includes:

[0098] Step S4031, use the attention mechanism layer to establish an association relationship between the label-guided multi-semantic-level data traffic information and the hidden layer state at the current moment; wherein, the hidden layer state at the current moment is a hidden layer state determined based on the context vector at the previous moment, the hidden layer state at the previous moment, and the label vector at the previous moment.

[0099] Specifically, attention is used to establish the correlation between the output of the dual-channel label semantic guidance encoder and the hidden layer state of the dynamically weighted label sequence decoder. The hidden layer state s at time t t and the i-th feature u in the more label-discriminative data traffic representation U (i.e., label-guided multi-semantic level data traffic information) i The correlation e between ti can be expressed as:

[0100]

[0101] where {v a , W a , W a} represents the trained parameter matrix.

[0102] Furthermore, the hidden layer state s of the dynamically weighted label sequence decoder at time t is obtained through the long short-term memory network t , and the hidden layer state s at time t t can be expressed as:

[0103] s t = LSTM(s t-1 , [z(y t-1 ); b t-1 ) (13)

[0104] where z(y t-1 ) represents the word embedding of the label with the highest probability in the probability distribution y t-1 (i.e., the label vector of the previous moment), y t-1 is the prediction probability distribution of all possible labels by the long short-term memory network at time t - 1, s t-1 represents the hidden layer state of the previous moment, and b t-1 represents the context vector of the previous moment.

[0105] Step S4032: Determine the context vector at the current moment using the correlation between the label-guided multi-semantic level data traffic information and the hidden layer state at the current moment.

[0106] Specifically, the calculation formula for the context vector b t at the current moment is as follows:

[0107]

[0108] where α ti represents the weight assigned to the i-th feature u of the label-guided multi-semantic level data traffic information at the t-th moment, and e i represents the hidden layer state s at time t tj t ​and the correlation between the j-th feature u in the label-guided multi-semantic level data traffic information, where m represents the total number of features in the label-guided multi-semantic level data traffic information U, and i and j are index variables used to traverse all features in the label-guided multi-semantic level data traffic information U. j Between them, m represents the total number of features in the label-guided multi-semantic level data traffic information U, and i and j are index variables used to traverse all features in the label-guided multi-semantic level data traffic information U.

[0109] Step S4033, use the parameter matrix in the dynamic classification layer to map the context vector and the hidden layer state at the current moment to an intermediate representation, and obtain the hidden layer weight at the current moment.

[0110] Specifically, the hidden layer weights θ1 and θ2 at the current moment can be expressed as:

[0111] θ1 = sigmoid(W1b t ) (16)

[0112] θ2 = sigmoid(W2s t ) (17)

[0113] θ1 += 1 (18)

[0114] Among them, W1 and W2 represent the parameter matrices in the dynamic classification layer, which are used to map the context vector and the hidden layer state at the current moment to an intermediate representation.

[0115] Step S4034, determine the output vector at the current moment based on the context vector, the hidden layer state, and the hidden layer weight at the current moment.

[0116] Specifically, the output vector o t at the current moment can be expressed as:

[0117] o t = W o f(θ1b t + θ2s t ) (19)

[0118] Among them, W o represents the trained parameter matrix, and f represents the non-linear activation function.

[0119] Step S4035, perform label prediction on the output vector at the current moment based on the mask vector at the current moment, and obtain the data traffic prediction label corresponding to the original data traffic at the current moment.

[0120] Specifically, the data traffic prediction label y t at the t-th moment can be expressed as:

[0121] y t = softmax(o t + It ) (20)

[0122] Among them, I t represents a mask vector, which is used to prevent the prediction of duplicate labels.

[0123] In step S4036, when there is a prediction deviation in the data traffic prediction label, the dynamic weighted fusion strategy is used to adjust the hidden layer weights at the current moment, and a multi-label data traffic classification model after training is obtained.

[0124] In some optional implementation manners, the above step S4036 includes:

[0125] Step b1, obtain the true label of the data traffic, and calculate the cross-entropy loss function value based on the data traffic prediction label and the data traffic true label.

[0126] Specifically, use cross-entropy loss as the loss function of the multi-label data traffic classification model, compare each data traffic prediction label with its corresponding data traffic true label one by one, and calculate the cross-entropy loss function value. The cross-entropy loss function value L T is calculated as follows:

[0127]

[0128] Among them, n represents the number of the label set, y i ′ j ∈ [0, 1] represents the data traffic prediction label, and y ij ∈ [0, 1] represents the data traffic true label.

[0129] Step b2, compare the cross-entropy loss function value with a preset threshold. If the cross-entropy loss function value is greater than the preset threshold, adjust the hidden layer weights at the current moment.

[0130] Specifically, if the cross-entropy loss function value is greater than the preset threshold, and the difference between the cross-entropy loss function value and the preset threshold is greater than the first difference and less than the second difference, increase the hidden layer weights at the current moment to make the hidden layer at the current moment play the maximum role; if the cross-entropy loss function value is greater than the preset threshold, and the difference between the cross-entropy loss function value and the preset threshold is greater than the second difference, reduce the hidden layer weights at the current moment to reduce the exposure bias; among them, the first difference is less than the second difference.

[0131] Step b3, reclassify the original data traffic by using the adjusted hidden layer weights at the current moment until the cross-entropy loss function value is less than the preset threshold, and obtain a multi-label data traffic classification model after training.

[0132] Specifically, the original data traffic is reclassified using the adjusted hidden layer weights at the current moment to obtain an updated data traffic prediction label. Then, the cross-entropy loss function is recalculated and compared with a preset threshold. If the value of the cross-entropy loss function is still greater than the preset threshold, the hidden layer weights are further adjusted, and the multi-label data traffic classification model is iteratively trained using the adjusted hidden layer weights until the value of the cross-entropy loss function is less than the preset threshold, obtaining a trained multi-label data traffic classification model.

[0133] Furthermore, during the process of label prediction in the dynamic classification layer, the beam search algorithm is used to obtain the optimal prediction sequence, alleviating the exposure bias problem of the sequence-to-sequence model. Beam search is a heuristic search algorithm used to find the optimal possible label combination in multi-label classification instead of exhausting all possible combinations, which can save computing resources.

[0134] Furthermore, the specific steps of obtaining the optimal prediction sequence using the beam search algorithm include: initializing the beam width, which represents the number of candidates retained at each step, and initializing the candidate sequence set, where the candidate sequence set contains different combinations of label vectors; then gradually expanding the candidates for each time step of the candidate sequences. For each current candidate sequence, all possible next tokens are generated and the candidate scores are calculated, and the top k with the highest scores are selected; when the preset maximum generation length is reached, the generation stops, completing the beam search. The candidate sequence with the candidate score is used as the prediction sequence, and the labels in the prediction sequence are used as the data traffic prediction labels.

[0135] In the training method of a multi-label data traffic classification model provided in this embodiment, the dynamically weighted label sequence decoder adopts a dynamic weighting strategy in the dynamic classification layer to alleviate the exposure bias problem. Through the dynamic weighting strategy, the dynamically weighted label sequence decoder can dynamically adjust its influence in the classification process according to the importance of each label, and can treat the classification decisions of different labels more balancedly, making the classification decisions for different labels more balanced and objective, which helps to improve the processing ability of the classification model for complex data traffic. Especially when facing the multi-label data traffic classification task, it can effectively improve the classification accuracy and stability of the model.

[0136] According to an embodiment of the present invention, there is also provided an embodiment of a multi-label data traffic classification method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0137] In this embodiment, a multi-label data traffic classification method is provided, which can be used in the above-mentioned electronic device.Figure 5 is a flowchart of a multi-label data traffic classification method according to an embodiment of the present invention. As Figure 5 shown, the process includes the following steps:

[0138] Step S501, collect the data traffic to be classified and preprocess the data traffic to be classified.

[0139] Specifically, perform data cleaning on the data traffic to be classified to remove invalid data from the data traffic to be classified.

[0140] Step S502, input the preprocessed data traffic to be classified into the trained multi-label data traffic classification model to obtain the multi-label data traffic classification result.

[0141] The following uses a specific embodiment to illustrate the specific steps of a multi-label data traffic classification method.

[0142] Embodiment 1:

[0143] As Figure 6 shown, the process of multi-label data traffic classification includes three main processes: traffic preprocessing, traffic feature representation, and classifier. The goal of multi-label data traffic classification is to assign a label subset y (z) from the label set Y to x (1) . If the multi-label data traffic classification task is regarded as a sequence-to-sequence learning task, then for the i-th data traffic x (z) , the goal of the sequence-to-sequence learning task is to find the optimal label sequence y that maximizes the probability, where the maximized probability can be expressed as:

[0144]

[0145] where y b represents the b-th label in the label subset y associated with the data traffic x (z) , l represents the number of labels in the label subset y; similar to the data traffic features, the labels are also vectorized. After vectorization, the label set is represented as C = {c1,..., c n} ∈ R n*d , c z represents the label vector corresponding to the z-th label y z in the label set; the dimension of the label vector is the same as the dimension of the word vector; for the sorting method of the labels in the label set, they are arranged in descending order of frequency, and the label vector and the word vector of the data traffic are obtained through pre-training.

[0146] However, the above method for multi-label data traffic classification requires a large amount of manual prior knowledge and time for annotation. The extracted features are difficult to mine deep semantic information, lacking data traffic semantic information, and unable to meet the application in a large-scale network traffic environment. The interaction relationship between label semantic information and local data traffic semantic information is not considered. Therefore, in this embodiment, the traffic feature representation and classifier of the multi-label data traffic classification process are improved, and a multi-label data traffic classification model is used for data traffic classification. The multi-label data traffic classification model is based on the encoder-decoder mode, as Figure 7 shown. The specific steps of the multi-label data traffic classification method guided by dual-channel label semantics include:

[0147] Construct a multi-label data traffic classification model, which consists of a dual-channel label semantics-guided encoder and a dynamically weighted label sequence decoder. The dual-channel label semantics-guided encoder mainly includes: a local filtering channel, a global filtering channel, and an adaptive gate; the dynamically weighted label sequence decoder includes an attention mechanism layer, a hidden layer, and a dynamic classification layer. The local filtering channel is used to extract the local data traffic semantic information most relevant to the label, the global filtering channel is used to capture the global data traffic semantic information most relevant to the label, and the adaptive gate is used to weight and aggregate the outputs of the local filtering channel and the global filtering channel to obtain the final data traffic representation. The attention mechanism layer is used to establish the association between the output of the dual-channel label semantics-guided encoder and the hidden layer state of the dynamically weighted label sequence decoder, thereby obtaining a context vector; the hidden layer is used to mine the association relationship between labels; the dynamic classification layer is used to obtain the final classification result.

[0148] Obtain a traffic dataset for training the multi-label data traffic classification model, and perform data cleaning on the data in the traffic dataset to obtain a standardized traffic dataset.

[0149] Input the standardized traffic dataset into the dual-channel label semantics-guided encoder, perform weight marking and information extraction on the data in the standardized traffic dataset to obtain a data traffic representation vector (i.e., label-guided multi-semantic-level data traffic information).

[0150] Input the data traffic representation vector into the dynamically weighted label sequence decoder. Through the dynamic weighting strategy, the decoder can dynamically adjust its influence in the classification process according to the importance of each label to obtain the trained multi-label data traffic classification model.

[0151] Input the data traffic with data traffic interaction into the trained multi-label data traffic classification model for processing to obtain the label classification result corresponding to the data traffic.

[0152] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0153] An embodiment of the present application also provides a training device for a multi-label data traffic classification model, as Figure 8 shown, including:

[0154] A construction module 801, configured to construct a multi-label data traffic classification model; the multi-label data traffic classification model is composed of a dual-channel label semantic-guided encoder and a dynamic weighted label sequence decoder;

[0155] An interaction module 802, configured to obtain a multi-label data traffic classification data set, and use the channel label semantic-guided encoder to interact the multi-level semantic information and label semantic information corresponding to the original data traffic in the multi-label data traffic classification data set to obtain label-guided multi-semantic-level data traffic information;

[0156] An adjustment module 803, configured to perform label prediction and weight dynamic adjustment based on the label-guided multi-semantic-level data traffic information by using the dynamic weighted label sequence decoder to obtain the trained multi-label data traffic classification model.

[0157] In some alternative embodiments, the interaction module 802 includes:

[0158] A first extraction unit, configured to extract semantic features from the original data traffic by using a convolutional neural network to obtain local data traffic semantic information;

[0159] A first establishment unit, configured to establish a relationship between the label semantic information and the local data traffic semantic information by using a local interaction attention layer to obtain label-guided local data traffic semantic information;

[0160] A second extraction unit, configured to extract semantic features from the original data traffic by using a bidirectional long short-term memory network to obtain global data traffic semantic information;

[0161] A second establishment unit, configured to establish a relationship between the label semantic information and the global data traffic semantic information by using a global interaction attention layer to obtain label-guided global data traffic semantic information;

[0162] A fusion unit, configured to fuse the label-guided local data traffic semantic information and the label-guided global data traffic semantic information by using an adaptive gate to obtain label-guided multi-semantic-level data traffic information.

[0163] In some alternative embodiments, the first establishment unit includes:

[0164] The first transformation subunit is configured to transform the local data traffic semantic information and the label semantic information by using the parameter matrices of the multi-head attention in the local interaction attention layer, so as to obtain query vectors corresponding to the multi-head attention respectively;

[0165] The second transformation subunit is configured to transform the label semantic information by using the parameter matrices of the multi-head attention, so as to obtain key vectors and value vectors corresponding to the multi-head attention respectively;

[0166] The mapping subunit is configured to map the label semantic information and the local data traffic semantic information based on the query vectors, key vectors and value vectors corresponding to the multi-head attention respectively, so as to obtain mapping feature vectors corresponding to the multi-head attention respectively;

[0167] The splicing subunit is configured to splice the mapping feature vectors corresponding to the multi-head attention respectively, so as to obtain the label-guided local data traffic semantic information.

[0168] In some alternative embodiments, the adjustment module 803 includes:

[0169] The third establishment unit is configured to establish an association relationship between the label-guided multi-semantic-level data traffic information and the hidden layer state at the current moment by using the attention mechanism layer; wherein, the hidden layer state at the current moment is a hidden layer state determined based on the context vector at the previous moment, the hidden layer state at the previous moment and the label vector at the previous moment;

[0170] The first determination unit is configured to determine the context vector at the current moment by using the association relationship between the label-guided multi-semantic-level data traffic information and the hidden layer state at the current moment;

[0171] The mapping unit is configured to map the context vector at the current moment and the hidden layer state at the current moment to an intermediate representation by using the parameter matrix in the dynamic classification layer, so as to obtain the hidden layer weight at the current moment;

[0172] The second determination unit is configured to determine the output vector at the current moment based on the context vector at the current moment, the hidden layer state at the current moment and the hidden layer weight at the current moment;

[0173] The prediction unit is configured to perform label prediction on the output vector at the current moment based on the mask vector at the current moment, so as to obtain the data traffic prediction label corresponding to the original data traffic at the current moment;

[0174] The adjustment unit is configured to, when there is a prediction deviation in the data traffic prediction label, adjust the hidden layer weight at the current moment by using the dynamic weighted fusion strategy, so as to obtain the trained multi-label data traffic classification model.

[0175] In some alternative embodiments, the adjustment unit includes:

[0176] A calculation subunit, configured to obtain the true label of the data traffic, and calculate the value of the cross-entropy loss function based on the predicted label and the true label of the data traffic;

[0177] A comparison subunit, configured to compare the value of the cross-entropy loss function with a preset threshold. If the value of the cross-entropy loss function is greater than the preset threshold, adjust the weights of the hidden layer at the current moment;

[0178] A classification subunit, configured to re-classify the original data traffic by using the adjusted weights of the hidden layer at the current moment until the value of the cross-entropy loss function is less than the preset threshold, so as to obtain a trained multi-label data traffic classification model.

[0179] For the description of the features in the corresponding embodiment of a training device for a multi-label data traffic classification model, reference can be made to the relevant description in the corresponding embodiment of a training method for a multi-label data traffic classification model, which will not be elaborated here one by one.

[0180] An embodiment of the present application further provides a multi-label data traffic classification device, as Figure 9 shown, including:

[0181] A preprocessing module 901, configured to collect the data traffic to be classified and preprocess the data traffic to be classified;

[0182] A classification module 902, configured to input the preprocessed data traffic to be classified into the trained multi-label data traffic classification model to obtain a multi-label data traffic classification result.

[0183] For the description of the features in the corresponding embodiment of a multi-label data traffic classification device, reference can be made to the relevant description in the corresponding embodiment of a multi-label data traffic classification method, which will not be elaborated here one by one.

[0184] An embodiment of the present application further provides an electronic device, as Figure 10 shown, including a processor 10 and a memory 20. The memory 20 stores a computer program, and the processor 10 is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the training method for a multi-label data traffic classification model or the multi-label data traffic classification method.

[0185] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. The computer program is configured to execute the steps in any of the above-mentioned embodiments of the training method for a multi-label data traffic classification model or the multi-label data traffic classification method when running.

[0186] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM), random access memory (RAM), external hard drives, magnetic disks, or optical discs that can store computer programs.

[0187] The embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described training methods or classification methods of the multi-label data traffic classification model.

[0188] Those skilled in the art can further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0189] The above has introduced in detail a multi-label data traffic classification method, a model training method, and a device provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A training method for a multi-label data flow classification model, characterized in that: include: Build a multi-label data traffic classification model; The multi-label data traffic classification model consists of a dual-channel label semantics-guided encoder and a dynamic weighted label sequence decoder; Acquire a multi-label data flow classification data set, and use the channel label semantic guidance encoder to interact the multi-level semantic information corresponding to the original data flow in the multi-label data flow classification data set with the label semantic information to obtain label-guided multi-semantic level data flow information; Based on the label-guided multi-semantic data flow information, the dynamic weighted label sequence decoder is used to perform label prediction and weight dynamic adjustment to obtain a multi-label data flow classification model after training.

2. The method according to claim 1, characterized in that The dual-channel label semantics-guided encoder includes a local filtering channel, a global filtering channel and an adaptive gate, the local filtering channel includes a convolutional neural network and a local interactive attention layer, and the global filtering channel includes a bidirectional long short-term memory network and a global interactive attention layer; the channel label semantics-guided encoder interacts the multi-level semantic information and label semantic information corresponding to the original data flow in the multi-label data flow classification data set to obtain label-guided multi-semantic level data flow information, including: Using the convolutional neural network to extract semantic features from the original data traffic to obtain local data traffic semantic information; Using the local interactive attention layer to establish a relationship between the label semantic information and the local data flow semantic information to obtain the label-guided local data flow semantic information; Using the bidirectional long short-term memory network to extract semantic features from the original data flow, to obtain global data flow semantic information; Using the global interactive attention layer to establish a relationship between the label semantic information and the global data flow semantic information to obtain the label-guided global data flow semantic information; The adaptive gate is used to fuse the label-guided local data flow semantic information and the label-guided global data flow semantic information to obtain the label-guided multi-semantic level data flow information.

3. The method according to claim 2, characterized in that The using the local interactive attention layer to establish a relationship between the label semantic information and the local data flow semantic information to obtain the label-guided local data flow semantic information includes: The local data flow semantic information and the label semantic information are changed by using the parameter matrix of the multi-head attention in the local interactive attention layer to obtain query vectors corresponding to the multi-head attention respectively; Using the parameter matrix of the multi-head attention to change the label semantic information, obtaining the key vector and value vector corresponding to the multi-head attention respectively; Mapping the label semantic information and the local data flow semantic information based on the query vector, key vector, and value vector respectively corresponding to the multi-head attention, to obtain the mapping feature vectors respectively corresponding to the multi-head attention; The mapping feature vectors corresponding to the multiple attention heads are concatenated to obtain the label-guided local data flow semantic information.

4. The method according to claim 1, characterized in that: The dynamic weighted label sequence decoder includes an attention mechanism layer, a hidden layer and a dynamic classification layer; the multi-semantic level data flow information guided by the label is based on the dynamic weighted label sequence decoder to perform label prediction and weight dynamic adjustment, and obtain a multi-label data flow classification model after training, including: The attention mechanism layer is used to establish an association relationship between the label-guided multi-semantic data flow information and the hidden layer state at the current moment; wherein the hidden layer state at the current moment is a hidden layer state determined based on the context vector at the previous moment, the hidden layer state at the previous moment, and the label vector at the previous moment; Determine the context vector at the current moment by using the association relationship between the multi-semantic level data flow information guided by the label and the hidden layer state at the current moment; Mapping the context vector at the current moment and the hidden layer state at the current moment to an intermediate representation using a parameter matrix in the dynamic classification layer to obtain a hidden layer weight at the current moment; Determine the output vector at the current moment based on the context vector at the current moment, the hidden layer state at the current moment, and the hidden layer weight at the current moment; Performing label prediction on the output vector at the current moment based on the mask vector at the current moment to obtain the data flow prediction label corresponding to the original data flow at the current moment; When there is a prediction deviation in the data traffic prediction label, a dynamic weighted fusion strategy is used to adjust the hidden layer weights at the current moment to obtain the multi-label data traffic classification model after the training is completed.

5. The method according to claim 4, characterized in that When the data traffic prediction label has a prediction deviation, a dynamic weighted fusion strategy is used to adjust the hidden layer weight at the current moment to obtain the multi-label data traffic classification model after the training is completed, including: Obtaining a true label of the data traffic, and calculating a cross entropy loss function value based on the predicted label of the data traffic and the true label of the data traffic; Compare the cross entropy loss function value with a preset threshold, and if the cross entropy loss function value is greater than the preset threshold, adjust the hidden layer weight at the current moment; The original data traffic is reclassified using the adjusted hidden layer weights at the current moment until the cross entropy loss function value is less than the preset threshold, thereby obtaining the multi-label data traffic classification model after the training is completed.

6. A multi-label data traffic classification method, characterized in that: The training method of the multi-label data traffic classification model according to any one of claims 1 to 5 is implemented, and the method comprises: Collecting data traffic to be classified, and preprocessing the data traffic to be classified; The preprocessed data traffic to be classified is input into the trained multi-label data traffic classification model to obtain the multi-label data traffic classification result.

7. A training device for a multi-label data flow classification model, characterized in that: include: Building module for building multi-label data flow classification model; The multi-label data traffic classification model consists of a dual-channel label semantics-guided encoder and a dynamic weighted label sequence decoder; An interaction module is used to obtain a multi-label data flow classification data set, and use the channel label semantic guidance encoder to interact the multi-level semantic information corresponding to the original data flow in the multi-label data flow classification data set with the label semantic information to obtain the label-guided multi-semantic level data flow information; The adjustment module is used to perform label prediction and weight dynamic adjustment based on the multi-semantic level data flow information guided by the label, using the dynamic weighted label sequence decoder to obtain a multi-label data flow classification model after training.

8. A multi-label data flow classification device, characterized in that: include: A preprocessing module, used for collecting data traffic to be classified and preprocessing the data traffic to be classified; The classification module is used to input the preprocessed data traffic to be classified into the trained multi-label data traffic classification model to obtain the multi-label data traffic classification result.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the training method of the multi-label data traffic classification model as described in any one of claims 1 to 5, or the steps of the multi-label data traffic classification method as described in claim 6 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the training method of the multi-label data traffic classification model as described in any one of claims 1 to 5, or the steps of the multi-label data traffic classification method as described in claim 6.