A Terminal Device Classification Method and System Based on Traffic Analysis
Through a traffic analysis method, using device prediction network, fuzzy decision-making and attention mechanisms to classify terminal devices, solving the problem of difficult to accurately identify and classify devices in dynamic network environments in the prior art, and achieving highly accurate and robust device classification.
Patent Information
- Application Number
- CN202411146094.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-08-20
AI Technical Summary
The prior art is difficult to accurately identify and classify terminal devices in dynamically changing network environments, especially without the need for additional hardware support.
Through a traffic analysis method, pre-trained equipment prediction network uses a pre-trained device to extract the network traffic data, and classifies terminal devices in combination with fuzzy decision-making and attention mechanisms.
It improves the accuracy and robustness of terminal device classification, enhances the interpretability and adaptability of the model, optimizes the computing efficiency, and improves the user experience.
Smart Images

Figure CN119128678B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet technologies, and in particular, to a method and system for classifying terminal devices based on traffic analysis. Background Art
[0002] In the current network environment, the diversity of terminal devices is increasing day by day. From traditional personal computers to smart phones, tablets, Internet of Things devices, etc., each device has its unique network behavior characteristics. Traditional device identification methods mainly rely on the physical identification of the device or manual configuration by the user, and these methods appear cumbersome and inflexible in a dynamically changing network environment. In addition, with the continuous evolution of network attack means, accurately identifying and classifying terminal devices has become crucial for network security management.
[0003] Existing automatic identification technologies, such as identification methods based on MAC addresses or IP addresses, although simple and easy to implement, are vulnerable to network configuration changes or address spoofing, and cannot provide continuous and accurate device identification. Therefore, developing a method that can automatically and accurately identify the type of terminal device, especially without the need for additional hardware support, has become an important research direction in the current network management field. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for classifying terminal devices based on traffic analysis to solve the above problems existing in the prior art.
[0005] An embodiment of the present invention provides a method for classifying terminal devices based on traffic analysis, and the method includes:
[0006] Obtain network traffic data of a device to be tested;
[0007] Input the network traffic data into a pre-trained device prediction network, and the device prediction network extracts features of the network traffic data to obtain network traffic features; obtain a prediction probability vector based on the network traffic features, the prediction probability vector includes multiple prediction probability values, and each prediction probability value corresponds to a device type;
[0008] Convert the network traffic features into fuzzy values;
[0009] Perform fuzzy inference on the fuzzy values to obtain an inference probability vector, the inference probability vector includes multiple inference probability values, and each inference probability value corresponds to a device type;
[0010] Obtain an attention weight vector based on the prediction probability vector and the inference probability vector; adjust the prediction probability vector based on the attention weight vector to obtain a device classification probability vector; the device classification probability vector includes multiple classification probability values, and each classification probability value corresponds to a device type;
[0011] Use the device type corresponding to the maximum classification probability value in the device classification probability vector as the type of the device to be measured.
[0012] Optionally, the training method of the device prediction network includes:
[0013] The device prediction network includes an input layer, multiple convolutional layers, multiple pooling layers, a fully connected layer, and an output layer, and the multiple convolutional layers and multiple pooling layers are arranged crosswise;
[0014] Obtain a training set, where the training set includes the network traffic data of multiple training devices; there are training devices with different device types among the multiple training devices, and each training device is pre-labeled with a device type; each device type corresponds to a device label value; the device label values of the multiple training devices form a training type vector;
[0015] Perform data alignment processing on the training set through the input layer to obtain aligned training data;
[0016] Perform a convolution operation on the aligned training data through the convolutional layer to obtain convolution features;
[0017] Perform a pooling operation on the convolution data through the pooling layer to obtain pooling number features;
[0018] Convert the pooling number features into a training prediction vector through the fully connected layer, and output the training prediction vector through the output layer; the training prediction vector includes multiple training probabilities, and each training probability corresponds to a device type; the training prediction vector corresponds to the training device one by one;
[0019] Convert the pooling data features into training fuzzy values;
[0020] Perform fuzzy inference on the training fuzzy values to obtain a training inference vector; the training inference vector includes multiple training probability values, and each training probability value corresponds to a device type; the training inference vector corresponds to the training device one by one;
[0021] Obtain a training attention weight vector based on the training prediction vector and the training inference vector; adjust the training prediction probability vector based on the training attention weight vector to obtain a training device classification probability vector; the training device classification probability vector includes multiple training classification probability values, and each training classification probability value corresponds to a device type; the training device classification probability vector corresponds to the training device one by one;
[0022] Obtain a model loss function based on the training device classification probability vector;
[0023] If the model loss function converges, determine that the training of the device prediction network is completed.
[0024] Optionally, obtaining a model loss function based on the training device classification probability vector includes:
[0025] Multiple training devices respectively obtain multiple training device classification probability vectors; using the multiple training device classification probability vectors as columns of a training prediction probability matrix to form the training prediction probability matrix;
[0026] Obtain the correlation index between the training type vector and the training prediction probability matrix;
[0027] Use the correlation index as the loss function.
[0028] Optionally, obtaining the correlation index between the training type vector and the training prediction probability matrix includes:
[0029] Obtain the information entropy of each column of the training type vector and the training prediction probability matrix, and one information entropy is obtained corresponding to each column of the training prediction probability matrix; there are multiple rows in the training prediction probability matrix and multiple information entropies are obtained correspondingly;
[0030] Use the mean of the multiple information entropies as the correlation index.
[0031] Optionally, obtaining the training attention weight vector based on the training prediction vector and the training inference vector includes:
[0032] Obtain the dot product similarity between the training prediction vector and the training inference vector: S(i, j) = P i ·I j where P i represents the i-th element of the training prediction vector, and I j represents the j-th element of the training inference vector, and the values of i and j are positive integers; S(i, j) represents the dot product similarity between the i-th element of the training prediction vector and the j-th element of the training inference vector; obtain the attention score based on the dot product similarity: A(i, j) = f(S(i, j)), where A(i, j) is the attention score between the i-th element of the training prediction vector and the j-th element of the training inference vector; f(S(i, j)) is the RELU function;
[0033] Normalize the weights: where α j is the attention weight of the j-th element I j of the training inference vector, and the value of k is 1, 2,..., N, and N is the number of elements of the training inference vector; N is a positive integer;
[0034] Use the multiple normalized weights as the elements of the training attention weight vector to form the training attention weight vector.
[0035] Optionally, adjusting the training prediction probability vector based on the training attention weight vector to obtain the training device classification probability vector includes:
[0036] Perform a cross - multiplication operation on the trained attention weight vector and the trained prediction probability vector to obtain a training device classification probability vector.
[0037] An embodiment of the present invention also provides a terminal device classification system based on traffic analysis. The system includes:
[0038] An acquisition module, configured to acquire network traffic data of a device to be measured;
[0039] A prediction module, configured to input the network traffic data into a pre - trained device prediction network. The device prediction network extracts features of the network traffic data to obtain network traffic features; and based on the network traffic features, obtains a prediction probability vector. The prediction probability vector includes multiple prediction probability values, and each prediction probability value corresponds to a device type;
[0040] A classification module, configured to convert the network traffic features into fuzzy values; perform fuzzy inference on the fuzzy values to obtain an inference probability vector. The inference probability vector includes multiple inference probability values, and each inference probability value corresponds to a device type; based on the prediction probability vector and the inference probability vector, obtain an attention weight vector; adjust the prediction probability vector based on the attention weight vector to obtain a device classification probability vector. The device classification probability vector includes multiple classification probability values, and each classification probability value corresponds to a device type; and use the device type corresponding to the largest classification probability value in the device classification probability vector as the type of the device to be measured.
[0041] Optionally, in the terminal device classification system based on traffic analysis, the training method of the device prediction network includes:
[0042] The device prediction network includes an input layer, multiple convolutional layers, multiple pooling layers, a fully - connected layer, and an output layer. The multiple convolutional layers and the multiple pooling layers are arranged alternately;
[0043] Obtain a training set. The training set includes network traffic data of multiple training devices. Among the multiple training devices, there are training devices with different device types, and each training device is pre - marked with a device type; each device type corresponds to a device marking value; the device marking values of the multiple training devices form a training type vector;
[0044] Perform data alignment processing on the training set through the input layer to obtain aligned training data;
[0045] Perform a convolution operation on the aligned training data through the convolutional layer to obtain convolution features;
[0046] Perform a pooling operation on the convolution data through the pooling layer to obtain pooling features;
[0047] Convert the pooled feature into a training prediction vector through a fully connected layer, and output the training prediction vector through an output layer; the training prediction vector includes multiple training probabilities, each training probability corresponding to a device type; the training prediction vector corresponds to a training device one by one;
[0048] Convert the pooled data feature into a training fuzzy value;
[0049] Perform fuzzy inference on the training fuzzy value to obtain a training inference vector; the training inference vector includes multiple training probability values, each training probability value corresponding to a device type; the training inference vector corresponds to a training device one by one;
[0050] Obtain a training attention weight vector based on the training prediction vector and the training inference vector; adjust the training prediction probability vector based on the training attention weight vector to obtain a training device classification probability vector; the training device classification probability vector includes multiple training classification probability values, each training classification probability value corresponding to a device type; the training device classification probability vector corresponds to a training device one by one;
[0051] Obtain a model loss function based on the training device classification probability vector;
[0052] If the model loss function converges, determine that the training of the device prediction network is completed.
[0053] Optionally, in the terminal device classification system based on traffic analysis, obtaining a model loss function based on the training device classification probability vector includes:
[0054] Multiple training devices respectively obtain multiple training device classification probability vectors; use the multiple training device classification probability vectors as the columns of a training prediction probability matrix to form a training prediction probability matrix;
[0055] Obtain the correlation index between the training type vector and the training prediction probability matrix;
[0056] Use the correlation index as the loss function.
[0057] Optionally, in the terminal device classification system based on traffic analysis, obtaining the correlation index between the training type vector and the training prediction probability matrix includes:
[0058] Obtain the information entropy of each column of the training type vector and the training prediction probability matrix, and one information entropy is obtained corresponding to each column of the training prediction probability matrix; there are multiple rows in the training prediction probability matrix, and multiple information entropies are obtained correspondingly;
[0059] Use the mean value of the multiple information entropies as the correlation index.
[0060] Compared with the prior art, the embodiments of the present invention achieve the following beneficial effects:
[0061] An embodiment of the present invention provides a method and system for classifying terminal devices based on traffic analysis. The method includes: by combining the feature extraction ability of CNN, the robustness of fuzzy decision-making, and the focusing ability of the attention mechanism, classifying the device to be tested based on network traffic data, which improves the accuracy of classifying the device to be tested. Specifically: obtaining network traffic data of the device to be tested; inputting the network traffic data into a pre-trained device prediction network, and the device prediction network extracts features of the network traffic data to obtain network traffic features; obtaining a prediction probability vector based on the network traffic features, the prediction probability vector includes multiple prediction probability values, and each prediction probability value corresponds to a device type; converting the network traffic features into fuzzy values; performing fuzzy reasoning on the fuzzy values to obtain an inference probability vector, the inference probability vector includes multiple inference probability values, and each inference probability value corresponds to a device type; obtaining an attention weight vector based on the prediction probability vector and the inference probability vector; adjusting the prediction probability vector based on the attention weight vector to obtain a device classification probability vector; the device classification probability vector includes multiple classification probability values, and each classification probability value corresponds to a device type; using the device type corresponding to the largest classification probability value in the device classification probability vector as the type of the device to be tested. It not only improves the accuracy and robustness of classification, but also enhances the interpretability and adaptability of the model, optimizes the computational efficiency, and improves the user experience. Through this comprehensive method, the terminal device classification technology can better meet the increasingly complex application requirements and lay a solid foundation for the development of future intelligent systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 FIG. is a flowchart of a method for classifying terminal devices based on traffic analysis provided by an embodiment of the present invention.
[0063] Figure 2 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present invention.
[0064] Reference numerals in the figure: 500 - bus; 501 - receiver; 502 - processor; 503 - transmitter; 504 - memory; 505 - bus interface. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] The present invention will be described in detail below with reference to the accompanying drawings.
[0066] Embodiment
[0067] An embodiment of the present invention relates to a method for classifying terminal devices based on traffic analysis. This method realizes the automatic classification of terminal devices by analyzing the feature data in network traffic. This method uses advanced data mining techniques and machine learning algorithms to accurately identify different types of terminal devices, such as smart phones, tablets, laptops, etc., thereby improving the efficiency and security of network management.
[0068] Specifically, an embodiment of the present invention provides a method for classifying terminal devices based on traffic analysis. As Figure 1 shown, the method includes:
[0069] S101: Obtain the network traffic data of the device to be tested. The network traffic data can be captured by network monitoring tools such as network sniffers and network analyzers (such as Wireshark, tcpdump, etc.). The network traffic data includes the traffic packet size (traffic value), traffic direction, timestamp, protocol type, etc. Optionally, the network traffic data includes: basic traffic information, traffic characteristics, timestamp, traffic session information, application layer data, traffic type, regular traffic and security events, and QoS (Quality of Service) indicators, etc.
[0070] Basic traffic information: Source IP address: The IP address of the device that sends the data packet. Destination IP address: The IP address of the device that receives the data packet. Source port: The port number for sending data. Destination port: The port number for receiving data. Protocol type: Such as TCP, UDP, ICMP, etc. Traffic characteristics: Packet size: The number of bytes of each data packet. Traffic rate: The amount of data transmitted per unit time (such as bps, Kbps, Mbps). Traffic direction: The direction of the data stream, such as inbound (download) and outbound (upload). Timestamp: Capture time: The time when the data packet is captured. Session duration: The length of time from the establishment to the end of a session. Traffic session information: Session ID: The ID that uniquely identifies a network session. Session status: Such as established, terminated, active, etc. status. Transmission control information: Such as TCP handshake, connection closure, etc. information. Application layer data: URL: The web page address accessed. HTTP header information: The header information in the request and response. DNS query and response: The domain name resolution request and its result. Traffic type: Application type: Such as web traffic, FTP, streaming media, VoIP, etc. User agent information: Client information such as browser, operating system, etc. Abnormal traffic and security events: Abnormal traffic pattern: Detected abnormal traffic (such as traffic surge). Attack pattern: Identified network attacks (such as DDoS, scanning, etc.). QoS (Quality of Service) indicators: Latency: The transmission latency of data from the source to the destination. Packet loss rate: The proportion of data packets lost during transmission. Jitter, etc.
[0071] S102: Input the network traffic data into a pre-trained device prediction network. The device prediction network extracts the characteristics of the network traffic data to obtain network traffic characteristics; and obtains a prediction probability vector based on the network traffic characteristics. The prediction probability vector includes multiple prediction probability values. Among them, each prediction probability value corresponds to a device type.
[0072] In the embodiments of the present invention, features are extracted from the original network traffic data. Commonly used features include: the number of packets, the number of bytes, the average packet size, etc. The temporal features of the traffic, such as peak traffic hours. The usage of specific protocols (such as HTTP, FTP, etc.).
[0073] S103: Convert the network traffic features into fuzzy values.
[0074] Specifically, S103 includes: Defining fuzzy rules: Define fuzzy rules according to traffic features. For example, "If the packet size is large and the protocol is HTTP, then the device type may be a PC".
[0075] Definition of fuzzy sets: Define fuzzy sets for each feature and set membership functions (such as triangular or trapezoidal membership functions). That is, define fuzzy sets for each input and output variable. For example:
[0076] Input: The traffic size can be defined as "small", "medium", "large".
[0077] Output: The device type can be defined as "mobile phone", "computer", "IoT device", etc.
[0078] Rule example: If the traffic size is small (the traffic size is less than the threshold, and the threshold value is 20G) and the protocol is HTTP, then the device type is a mobile phone. If the traffic size is large (the traffic size is greater than the threshold, and the threshold value is 20G) and the protocol is TCP, then the device type is a computer.
[0079] Definition of the inference engine: Use a fuzzy inference mechanism (such as the Mamdani or Sugeno method) to combine fuzzy rules and obtain a decision.
[0080] S104: Perform fuzzy inference on the fuzzy values to obtain an inference probability vector. The inference probability vector includes multiple inference probability values, and each inference probability value corresponds to a device type.
[0081] Among them, using a fuzzy inference mechanism to combine fuzzy rules with fuzzy inputs usually includes the following steps:
[0082] Rule evaluation: Evaluate each fuzzy rule. Calculate the activation degree of each rule (that is, the degree of satisfaction of the conditions), and usually use the minimum value or multiplication to combine the membership degrees of the inputs. For example: For "If the traffic size is small and the protocol is HTTP", calculate the minimum value of the membership degrees of the traffic size and the protocol as the activation degree of this rule.
[0083] Fuzzy Inference: Determine the fuzzy value of the output based on the activation degree of the rules. For example: Determine the fuzzy value of the output according to the activation degree of the rules. For the activation degree of each rule, combine it with the fuzzy set of the corresponding output. Common combination methods include: Mamdani Inference: For each rule, use the activation degree (such as the minimum value) to scale the output membership function. For example: If the membership degree of "small" for the traffic volume is 0.7, and the rule is "If the traffic is small and the protocol is HTTP, then the device type is mobile phone", then the membership function of the output "mobile phone" will be fuzzified according to 0.7. Sugeno Inference: In the Sugeno model, the output is a linear combination or a constant, and usually the output value is directly given after the rule is activated, without further fuzzification.
[0084] Aggregate Output: Aggregate the output fuzzy values of all rules. Common aggregation methods include:
[0085] Maximum Value Method: For each output variable, select the maximum value of the output membership degrees of all rules as the membership degree of the final output. Weighted Average Method: Perform a weighted average on the output values of all rules to obtain the final fuzzy output. Defuzzification: Convert the aggregated fuzzy output into an exact value. Common defuzzification methods include: Centroid Method: Calculate the centroid of the fuzzy output membership function to obtain the final output value. The centroid method is one of the most commonly used methods because it can effectively handle complex fuzzy outputs. Maximum Membership Degree Method: Select the output value with the maximum membership degree as the final result.
[0086] Illustrate the implementation process of the above technical solution with a detailed example:
[0087] 1. Data Preparation:
[0088] Collect data related to network traffic, including the following features: Traffic Features: Packet size, transmission rate, transmission protocol, etc. Session Information: Source IP, destination IP, source port, destination port, session duration, etc. Application Layer Data: HTTP requests, domain name queries, etc.
[0089] 2. Feature Extraction:
[0090] Extract features from the original network traffic data. Commonly used features include: The number of packets, the number of bytes, the average packet size, etc. The temporal features of the traffic, such as peak traffic hours, etc. The usage of specific protocols (such as HTTP, FTP, etc.).
[0091] 3. Fuzzy Logic System Design:
[0092] Design a fuzzy logic system, including the following steps:
[0093] Define input and output variables: Input variables: traffic characteristics (such as packet size, traffic rate, etc.). Output variable: probability of device type (such as the probability that the device type is a mobile phone, computer, IoT device).
[0094] Fuzzification: Convert the input features into fuzzy values. For example, divide the packet size into three categories: "small", "medium", and "large", and divide the traffic rate into "low", "medium", and "high".
[0095] Establish fuzzy rules: Create fuzzy rules to define the relationship between the input and output. For example: If "the packet size is large" and "the traffic rate is high", then the probability that "the device type is a computer" is high. If "the packet size is small" and "the traffic rate is low", then the probability that "the device type is a mobile phone" is high.
[0096] 4. Inference:
[0097] Use a fuzzy inference method (such as Mamdani inference or Sugeno inference) to calculate the output: Infer the input according to the fuzzy rules to obtain the fuzzy values of each device type.
[0098] 5. Defuzzification:
[0099] Convert the fuzzy output into specific probability values: Use a defuzzification method (such as the centroid method) to convert the fuzzy output into a definite probability value. For example, the process of predicting the characteristics of network traffic data through fuzzy decision-making and obtaining the probability of the device type is described in detail as follows:
[0100] To convert the fuzzy output into specific probability values, the following steps can be taken:
[0101] Select a defuzzification method: Centroid Method: Calculate the centroid of the fuzzy output to obtain a specific value, which can be used as the probability of each device type. Max Membership Method: Select the output value with the largest membership degree as the final result.
[0102] Calculate the probability: For each device type, calculate its corresponding probability in combination with the inference result. For example, for the three categories of output "mobile phone", "computer", and "IoT device", obtain the corresponding probability values respectively.
[0103] 6. Output the result:
[0104] The final output will be a probability distribution representing the probability of the device type corresponding to the network traffic data. For example:
[0105] The probability of the device type "mobile phone" is 0.70. The probability of the device type "computer" is 0.20. The probability of the device type "IoT device" is 0.10. These probability values form an inference probability vector.
[0106] Optionally, the method further includes:
[0107] 7. Model evaluation and optimization:
[0108] Evaluate the model performance: Use methods such as cross-validation to evaluate the performance of the fuzzy decision model, and calculate metrics such as accuracy, recall, and F1-score.
[0109] Optimize the fuzzy rules: According to the evaluation results, adjust the fuzzy rules, input features, and membership functions to improve the accuracy and robustness of the model.
[0110] 8. Application and deployment:
[0111] Apply the trained fuzzy decision model to a real-time network traffic monitoring system to analyze and predict the device type in real time.
[0112] Regularly update the model to adapt to changes in network traffic characteristics and the emergence of new device types.
[0113] As an example:
[0114] Suppose there is the following characteristic data:
[0115] Packet size: 1500 bytes; traffic rate: 200 kbps.
[0116] These features can be fuzzified as: Packet size: large. Traffic rate: medium.
[0117] According to the set fuzzy rules, it may be deduced that: If the packet size is "large" and the traffic rate is "medium", then the probability of the device type being "computer".
[0118] S105: Obtain an attention weight vector based on the prediction probability vector and the inference probability vector; adjust the prediction probability vector based on the attention weight vector to obtain a device classification probability vector. Among them, the device classification probability vector includes multiple classification probability values, and each classification probability value corresponds to a device type.
[0119] Specifically, obtain the dot product similarity between the predicted probability vector and the inferred probability vector; obtain the attention score based on the dot product similarity between the predicted probability vector and the inferred probability vector. Normalize the weights, and then use the multiple normalized weights as the elements of the attention weight vector to form the attention weight vector. Perform a cross product operation on the attention weight vector and the predicted probability vector to obtain the device classification probability vector. For the specific implementation, please refer to the description of the part of obtaining the model loss function based on the training device classification probability vector in the following training method of the device prediction network.
[0120] S106: Use the device type corresponding to the maximum classification probability value in the device classification probability vector as the type of the device to be tested.
[0121] Among them, the training method of the device prediction network includes:
[0122] The device prediction network includes an input layer, multiple convolutional layers, multiple pooling layers, a fully connected layer, and an output layer, and the multiple convolutional layers and multiple pooling layers are arranged crosswise.
[0123] Obtain a training set, where the training set includes the network traffic data of multiple training devices; there are training devices with different device types among the multiple training devices, and each training device is pre-labeled with a device type; each device type corresponds to a device label value; the device label values of the multiple training devices form a training type vector; for example, the device label values of mobile phones, computers, and IoT devices in the device type are 1, 2, and 3 respectively.
[0124] Perform data alignment processing on the training set through the input layer to obtain aligned training data. Specifically include: Resampling: Resample the network traffic data in the training set to the same time interval. For example, use methods such as mean, sum, or interpolation to fill in missing values. Interpolation: Use methods such as linear interpolation and spline interpolation to fill in missing values. Timestamp alignment: Ensure that all network traffic data has the same timestamp, and interpolation or other methods can be used for alignment if necessary.
[0125] Perform a convolution operation on the aligned training data through the convolutional layer to obtain convolution features. Perform a pooling operation on the convolutional data through the pooling layer to obtain pooling number features.
[0126] Among them, the output of each layer of convolution is used as the input of the next layer of pooling layer, and the output of the pooling layer is used as the input of the next layer of convolutional layer. For example, if there are 3 convolutional layers and 3 pooling layers, then the input of the first convolutional layer is the aligned training data, the input of the first pooling layer is the output of the first convolutional layer, the input of the second convolutional layer is the output of the first pooling layer, the input of the second pooling layer is the output of the second convolutional layer, the input of the third convolutional layer is the output of the second pooling layer, and the input of the third pooling layer is the output of the third convolutional layer.
[0127] Convert the pooled number features into training prediction vectors through a fully connected layer, and output the training prediction vectors through an output layer; the training prediction vectors include multiple training probabilities, and each training probability corresponds to a device type; the training prediction vectors correspond one-to-one with the training devices.
[0128] That is, convert the output of the third pooling layer (pooled number features) into training prediction vectors through a fully connected layer. Specifically, the pooled number features can be classified (specifically, a convolutional neural network can be used for classification), and then a dimensionality reduction operation is performed to obtain the training prediction vectors.
[0129] Convert the pooled data features into training fuzzy values. Perform fuzzy inference on the training fuzzy values to obtain training inference vectors; the training inference vectors include multiple training probability values, and each training probability value corresponds to a device type; the training inference vectors correspond one-to-one with the training devices. The specific method can refer to the above method and will not be elaborated here.
[0130] Obtain the training attention weight vectors based on the training prediction vectors and the training inference vectors; adjust the training prediction probability vectors based on the training attention weight vectors to obtain the training device classification probability vectors. The training device classification probability vectors include multiple training classification probability values, and each training classification probability value corresponds to a device type. The training device classification probability vectors correspond one-to-one with the training devices.
[0131] Obtain the model loss function based on the training device classification probability vectors. If the model loss function converges, determine that the training of the device prediction network is completed.
[0132] Among them, obtaining the model loss function based on the training device classification probability vectors includes:
[0133] Multiple training devices correspond to obtaining multiple training device classification probability vectors; use the multiple training device classification probability vectors as the columns of the training prediction probability matrix to form the training prediction probability matrix. Obtain the correlation index between the training type vector and the training prediction probability matrix. Use the correlation index as the loss function.
[0134] Specifically, obtaining the correlation index between the training type vector and the training prediction probability matrix includes:
[0135] Obtain the information entropy (or cross-entropy) of each column of the training type vector and the training prediction probability matrix. Each column of the training prediction probability matrix corresponds to obtaining an information entropy; there are multiple rows in the training prediction probability matrix corresponding to obtaining multiple information entropies;
[0136] Use the mean of the multiple information entropies as the correlation index.
[0137] As a further step, obtaining the training attention weight vectors based on the training prediction vectors and the training inference vectors includes:
[0138] Obtain the dot product similarity between the training prediction vector and the training inference vector: S(i, j) = P i ·I j , where P i represents the i-th element of the training prediction vector, and I j represents the j-th element of the training inference vector, and the values of i and j are positive integers; S(i, j) represents the dot product similarity between the i-th element of the training prediction vector and the j-th element of the training inference vector;
[0139] Obtain the attention score based on the dot product similarity: A(i, j) = f(S(i, j)), where A(i, j) is the attention score between the i-th element of the training prediction vector and the j-th element of the training inference vector; f(S(i, j)) is the RELU function;
[0140] Normalized weights: where α j is the attention weight of the j-th element I j of the training inference vector, and the value of k is 1, 2,......, N, where N is the number of elements of the training inference vector; N is a positive integer;
[0141] Use multiple normalized weights as the elements of the training attention weight vector to form the training attention weight vector.
[0142] Finally, adjust the training prediction probability vector based on the training attention weight vector to obtain the training device classification probability vector. Specifically: perform a cross product operation on the training attention weight vector and the training prediction probability vector to obtain the training device classification probability vector.
[0143] The terminal device classification method based on traffic analysis provided in this application combines a Convolutional Neural Network (CNN), fuzzy decision-making, and an attention mechanism, showing various beneficial effects in terminal device classification. The following is a detailed description of these effects:
[0144] 1. Improve classification accuracy: Feature extraction and selection: The powerful feature extraction ability of CNN combined with the attention mechanism can automatically identify and focus on the most discriminative features. This combination reduces the interference of irrelevant features, thereby improving the classification accuracy. Handling ambiguity: Fuzzy decision-making can effectively handle the overlap and uncertainty between categories. When the input data has fuzzy labels, the model can still make reasonable judgments through fuzzy logic.
[0145] 2. Enhance model robustness: Anti-interference ability against noise and deformation: CNN has strong robustness to noise and deformation in input data. After combining with fuzzy logic, the model can operate more stably in the face of uncertainty and changes. Dynamic adaptation ability: The attention mechanism enables the model to dynamically adjust the feature regions of interest according to different inputs, improving the adaptability of classification in various environments.
[0146] 3. Improve decision interpretability: Enhance model transparency: The fuzzy decision-making system can provide clear explanations about the decision-making process, improve the accuracy of classification, help users understand why the model classifies a certain terminal device into a specific category, and enhance user trust.
[0147] 4. Optimize computational efficiency: Reduce redundant calculations: The attention mechanism can automatically screen key information during processing, reduce the calculation of unnecessary features, and improve the overall efficiency of the model. Speed up training: By focusing on important features, the model converges faster during training, thus speeding up the training time and the utilization efficiency of resources.
[0148] 5. Support multi-modal data processing: Integrate multiple data sources: By combining CNN and the attention mechanism, it can effectively process multi-modal data (such as images, text, and sensor data), providing a more comprehensive perspective and information for terminal device classification. This multi-modal processing ability enables the model to comprehensively utilize data from different sources in the face of complex scenarios, improving the accuracy and reliability of classification.
[0149] 6. Improve adaptability and flexibility: Multi-scenario applications: The new method combining fuzzy decision-making and the attention mechanism can be flexibly applied to the classification of various types of terminal devices, meeting the requirements in different fields (such as smart homes, industrial devices, mobile terminals, etc.). Real-time update ability: Due to the flexibility of fuzzy decision-making, the model can be quickly adjusted according to new data and rules, continuously optimizing the classification performance to adapt to the rapid development and changes of technology.
[0150] 7. Enhance user experience: Intelligent recommendations and personalized services: Through more accurate classification, the system can provide more personalized recommendations and services for users, enhancing the user experience. Reduce the risk of misjudgment: The introduction of fuzzy decision-making reduces misjudgments caused by uncertainty, enhances the reliability of classification results, and enables users to rely more confidently on the system's judgment.
[0151] In summary, classifying terminal devices by combining CNN, fuzzy decision-making, and attention mechanism shows significant advantages in terminal device classification. This method not only improves the accuracy and robustness of classification, but also enhances the interpretability and adaptability of the model, optimizes the computational efficiency, and improves the user experience. Through this comprehensive method, terminal device classification technology can better meet the increasingly complex application requirements and lay a solid foundation for the development of future intelligent systems.
[0152] Based on the above terminal device classification method based on traffic analysis, an embodiment of the present invention provides a terminal device classification system based on traffic analysis, and the system includes:
[0153] An acquisition module for acquiring network traffic data of a device to be tested;
[0154] A prediction module for inputting the network traffic data into a pre-trained device prediction network. The device prediction network extracts the features of the network traffic data to obtain network traffic features; based on the network traffic features, a prediction probability vector is obtained. The prediction probability vector includes multiple prediction probability values, and each prediction probability value corresponds to a device type;
[0155] A classification module for converting the network traffic features into fuzzy values; performing fuzzy inference on the fuzzy values to obtain an inference probability vector. The inference probability vector includes multiple inference probability values, and each inference probability value corresponds to a device type; obtaining an attention weight vector based on the prediction probability vector and the inference probability vector; adjusting the prediction probability vector based on the attention weight vector to obtain a device classification probability vector. The device classification probability vector includes multiple classification probability values, and each classification probability value corresponds to a device type; using the device type corresponding to the largest classification probability value in the device classification probability vector as the type of the device to be tested.
[0156] Optionally, in the terminal device classification system based on traffic analysis, the training method of the device prediction network includes:
[0157] The device prediction network includes an input layer, multiple convolutional layers, multiple pooling layers, a fully connected layer, and an output layer, and the multiple convolutional layers and the multiple pooling layers are arranged crosswise;
[0158] Obtaining a training set, the training set includes network traffic data of multiple training devices; there are training devices with different device types among the multiple training devices, and each training device is pre-labeled with a device type; each device type corresponds to a device label value; the device label values of the multiple training devices form a training type vector;
[0159] Performing data alignment processing on the training set through the input layer to obtain aligned training data;
[0160] Align the training data through a convolutional layer to perform a convolutional operation to obtain convolutional features;
[0161] Perform a pooling operation on the convolutional data through a pooling layer to obtain pooled feature;
[0162] Convert the pooled feature into a training prediction vector through a fully connected layer, and output the training prediction vector through an output layer; the training prediction vector includes multiple training probabilities, and each training probability corresponds to a device type; the training prediction vector corresponds to the training device one by one;
[0163] Convert the pooled data feature into a training fuzzy value;
[0164] Perform fuzzy inference on the training fuzzy value to obtain a training inference vector; the training inference vector includes multiple training probability values, and each training probability value corresponds to a device type; the training inference vector corresponds to the training device one by one;
[0165] Obtain a training attention weight vector based on the training prediction vector and the training inference vector; adjust the training prediction probability vector based on the training attention weight vector to obtain a training device classification probability vector; the training device classification probability vector includes multiple training classification probability values, and each training classification probability value corresponds to a device type; the training device classification probability vector corresponds to the training device one by one;
[0166] Obtain a model loss function based on the training device classification probability vector;
[0167] If the model loss function converges, determine that the training of the device prediction network is completed.
[0168] The specific implementation manners of the functions of each module of the above system are the same as those described in the above method for classifying terminal devices based on traffic analysis, and will not be elaborated here.
[0169] An embodiment of the present invention further provides an electronic device for integrating the above terminal device classification system based on traffic analysis, as Figure 2 shown, the electronic device includes a memory 504, a processor 502, and a computer program stored on the memory 504 and executable on the processor 502. When the processor 502 executes the program, it implements the steps of any one of the above methods for classifying terminal devices based on traffic analysis.
[0170] Among them, in Figure 2Among them, a bus architecture (represented by bus 500), the bus 500 may include any number of interconnected buses and bridges. The bus 500 links together various circuits of one or more processors represented by the processor 502 and the memory represented by the memory 504. The bus 500 may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art, and thus will not be further described herein. The bus interface 505 provides an interface between the bus 500 and the receiver 501 and the transmitter 503. The receiver 501 and the transmitter 503 may be the same element, i.e., a transceiver, which provides a unit for communicating with various other devices on the transmission medium. The processor 502 is responsible for managing the bus 500 and general processing, while the memory 504 may be used to store data used by the processor 502 when performing operations.
[0171] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of any one of the methods for classifying terminal devices based on traffic analysis described above.
[0172] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings based herein. The structure required to construct such a system is obvious from the above description. In addition, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of the specific language above is for disclosing the best mode of the present invention.
[0173] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.
[0174] Similarly, it should be understood that, in order to streamline this disclosure and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present invention.
[0175] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and set in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be adopted to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise explicitly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature providing the same, equivalent or similar purpose.
[0176] In addition, those skilled in the art can understand that although some of the embodiments herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0177] Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the device according to the embodiments of the present invention. The present invention can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
Claims
1. A terminal device classification method based on traffic analysis, characterized in that: The method comprises: Obtain network traffic data of the device under test; The network traffic data is input into a pre-trained device prediction network, and the device prediction network extracts network traffic data features to obtain network traffic features; a prediction probability vector is obtained based on the network traffic features, and the prediction probability vector includes multiple prediction probability values, and each prediction probability value corresponds to a device type; Convert network traffic features into fuzzy values; Perform fuzzy reasoning on the fuzzy value to obtain a reasoning probability vector, the reasoning probability vector includes multiple reasoning probability values, and each reasoning probability value corresponds to a device type; Obtain an attention weight vector based on the prediction probability vector and the inference probability vector; adjust the prediction probability vector based on the attention weight vector to obtain a device classification probability vector; the device classification probability vector includes multiple classification probability values, each classification probability value corresponds to a device type; The device type corresponding to the largest classification probability value in the device classification probability vector is used as the type of the device to be tested; The training methods for the device prediction network include: The device prediction network includes an input layer, multiple convolutional layers, multiple pooling layers, a fully connected layer, and an output layer. Multiple convolutional layers and multiple pooling layers are cross-set. A training set is obtained, the training set includes network traffic data of multiple training devices; the multiple training devices include training devices of different device types, and each training device is pre-labeled with a device type; each device type corresponds to a device label value; the device label values of the multiple training devices constitute a training type vector; Perform data alignment processing on the training set through the input layer to obtain aligned training data; Align the training data through the convolution layer to perform convolution operations and obtain convolution features; The convolution data is pooled through the pooling layer to obtain the pooling number features; The pooled number features are converted into training prediction vectors through the fully connected layer, and the training prediction vectors are output through the output layer; the training prediction vectors include multiple training probabilities, each training probability corresponds to a device type; the training prediction vectors correspond to the training devices one by one; Convert the pooled data features into training fuzzy values; Perform fuzzy reasoning on the training fuzzy value to obtain a training reasoning vector; the training reasoning vector includes multiple training probability values, each training probability value corresponds to a device type; the training reasoning vector corresponds to the training device one by one; Based on the training prediction vector and the training inference vector, a training attention weight vector is obtained; based on the training attention weight vector, the training prediction probability vector is adjusted to obtain a training device classification probability vector; the training device classification probability vector includes multiple training classification probability values, each training classification probability value corresponds to a device type; the training device classification probability vector corresponds to the training device one-to-one; based on the training device classification probability vector, a model loss function is obtained; If the model loss function converges, it is determined that the device prediction network training is completed; The training attention weight vector is obtained based on the training prediction vector and the training inference vector, including: Get the dot product similarity between the training prediction vector and the training inference vector: S(i, j) = P i I j , where P i Represents the i-th element of the training prediction vector, I j represents the jth element of the training inference vector, i and j are positive integers; S(i, j) represents the dot product similarity between the i-th element of the training prediction vector and the j-th element of the training inference vector; Attention scores are obtained based on dot product similarity: A(i, j) = f(S(i·j)), where A(i, j) is the attention score of the i-th element of the training prediction vector and the j-th element of the training inference vector; f(S(i, j)) is the RELU function; Normalized weights: Among them, α j is the jth element of the training inference vector I j The attention weight of k is 1, 2, ..., N, where N is the number of elements in the training reasoning vector; N is a positive integer; A training attention weight vector is formed by using a plurality of normalized weights as elements of the training attention weight vector; The training prediction probability vector is adjusted based on the training attention weight vector to obtain the training device classification probability vector, including: Perform a cross product operation on the training attention weight vector and the training prediction probability vector to obtain the training device classification probability vector.
2. The terminal device classification method based on traffic analysis according to claim 1 is characterized in that: The model loss function is obtained based on the classification probability vector of the training device, including: A plurality of training device classification probability vectors are obtained corresponding to the plurality of training devices; the plurality of training device classification probability vectors are used as columns of a training prediction probability matrix to form a training prediction probability matrix; Obtain the correlation index between the training type vector and the training prediction probability matrix; The correlation index is used as the loss function.
3. The terminal device classification method based on traffic analysis according to claim 2 is characterized in that: Get the correlation index between the training type vector and the training prediction probability matrix, including: The information entropy of each column of the training type vector and the training prediction probability matrix is obtained, and each column of the training prediction probability matrix corresponds to an information entropy; if the training prediction probability matrix has multiple rows, multiple information entropies are obtained correspondingly; The mean of multiple information entropies is used as the correlation index.
4. A terminal device classification system based on traffic analysis, characterized in that: The system comprises: An acquisition module is used to obtain network traffic data of the device under test; A prediction module is used to input network traffic data into a pre-trained device prediction network, and the device prediction network extracts network traffic data features to obtain network traffic features; based on the network traffic features, a prediction probability vector is obtained, and the prediction probability vector includes multiple prediction probability values, each of which corresponds to a device type; A classification module is used to convert network traffic features into fuzzy values; perform fuzzy reasoning on the fuzzy values to obtain a reasoning probability vector, the reasoning probability vector includes multiple reasoning probability values, and each reasoning probability value corresponds to a device type; obtain an attention weight vector based on the prediction probability vector and the reasoning probability vector; adjust the prediction probability vector based on the attention weight vector to obtain a device classification probability vector; the device classification probability vector includes multiple classification probability values, and each classification probability value corresponds to a device type; the device type corresponding to the largest classification probability value in the device classification probability vector is used as the type of the device to be tested; The training methods for the device prediction network include: The device prediction network includes an input layer, multiple convolutional layers, multiple pooling layers, a fully connected layer, and an output layer. Multiple convolutional layers and multiple pooling layers are cross-set. A training set is obtained, the training set includes network traffic data of multiple training devices; the multiple training devices include training devices of different device types, and each training device is pre-labeled with a device type; each device type corresponds to a device label value; the device label values of the multiple training devices constitute a training type vector; Perform data alignment processing on the training set through the input layer to obtain aligned training data; Align the training data through the convolution layer to perform convolution operations and obtain convolution features; The convolution data is pooled through the pooling layer to obtain the pooling number features; The pooled number features are converted into training prediction vectors through the fully connected layer, and the training prediction vectors are output through the output layer; the training prediction vectors include multiple training probabilities, each training probability corresponds to a device type; the training prediction vectors correspond to the training devices one by one; Convert the pooled data features into training fuzzy values; Perform fuzzy reasoning on the training fuzzy value to obtain a training reasoning vector; the training reasoning vector includes multiple training probability values, each training probability value corresponds to a device type; the training reasoning vector corresponds to the training device one by one; Obtain a training attention weight vector based on the training prediction vector and the training inference vector; adjust the training prediction probability vector based on the training attention weight vector to obtain a training device classification probability vector; the training device classification probability vector includes multiple training classification probability values, each training classification probability value corresponds to a device type; the training device classification probability vector corresponds to the training device one by one; Obtaining a model loss function based on the classification probability vector of the training device; If the model loss function converges, it is determined that the device prediction network training is completed; The training attention weight vector is obtained based on the training prediction vector and the training inference vector, including: Get the dot product similarity between the training prediction vector and the training inference vector: S(i, j) = P i I j , where P i Represents the i-th element of the training prediction vector, I j represents the jth element of the training inference vector, i and j are positive integers; S(i, j) represents the dot product similarity between the i-th element of the training prediction vector and the j-th element of the training inference vector; Attention scores are obtained based on dot product similarity: A(i, j) = f(S(i, j)), where A(i, j) is the attention score of the i-th element of the training prediction vector and the j-th element of the training inference vector; f(S(i, j)) is the RELU function; Normalized weights: Among them, α j is the jth element of the training inference vector I j The attention weight of k is 1, 2, ..., N, where N is the number of elements in the training reasoning vector; N is a positive integer; A training attention weight vector is formed by using a plurality of normalized weights as elements of the training attention weight vector; The training prediction probability vector is adjusted based on the training attention weight vector to obtain the training device classification probability vector, including: Perform a cross product operation on the training attention weight vector and the training prediction probability vector to obtain the training device classification probability vector.
5. The terminal device classification system based on traffic analysis according to claim 4 is characterized in that: The model loss function is obtained based on the classification probability vector of the training device, including: A plurality of training device classification probability vectors are obtained corresponding to the plurality of training devices; the plurality of training device classification probability vectors are used as columns of a training prediction probability matrix to form a training prediction probability matrix; Obtain the correlation index between the training type vector and the training prediction probability matrix; The correlation index is used as the loss function.
6. The terminal device classification system based on traffic analysis according to claim 5 is characterized in that: Get the correlation index between the training type vector and the training prediction probability matrix, including: The information entropy of each column of the training type vector and the training prediction probability matrix is obtained, and each column of the training prediction probability matrix corresponds to an information entropy; if the training prediction probability matrix has multiple rows, multiple information entropies are obtained correspondingly; The mean of multiple information entropies is used as the correlation index.
Citation Information
Patent Citations
Event intention reasoning method and device, equipment and storage medium
CN112488316A
Internet-of-Things equipment fingerprint identification method based on deep learning
CN112564974A