A Method for Identifying End Devices in NATed Networks Based on Clustering of TCP Timestamp Sequences
By analyzing the sequence of TCP timestamp field TSval, using quasi-linear regression analysis and unsupervised clustering technology, the equipment identification problem in the NAT environment is solved, and the accurate identification and counting of terminal devices in the home network is achieved, and the accuracy of network management and traffic analysis is improved.
Patent Information
- Application Number
- CN202510452394.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-11
AI Technical Summary
In the NAT environment, home devices share the same IP address, making it difficult to accurately identify and distinguish terminal devices through traditional MAC addresses or IP addresses, affecting the accuracy of network management and traffic analysis.
By analyzing the sequence of TCP timestamp field TSval, using quasi-linear regression analysis and unsupervised clustering technology, an identification model of the terminal device is established, so as to accurately identify and count terminal devices in the NAT environment and distinguish their TCP flow.
It realizes effective identification of individual devices in the home network in the NAT environment, improves the accuracy and reliability of device identification, provides operators with efficient and accurate device counting and traffic analysis tools, and improves network management and user experience.
Smart Images

Figure CN119996266B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of broadband networks in data communication, and particularly relates to a method for identifying terminal devices in a NATed network based on clustering of TCP timestamp sequences. Background Art
[0002] With the rapid progress of information technology, the types and quantities of devices relied on by household users have increased rapidly, covering personal computers, mobile devices (such as smart phones), surveillance cameras, and smart home products (such as smart speakers, automatic floor sweepers, etc.). The wide use of these devices not only improves the convenience of household life, but also poses higher requirements for household network bandwidth. In order to effectively manage and optimize the household network, operators need to be able to accurately identify the devices used by users through network traffic, and perform correlation analysis on the traffic of user devices, so as to better understand the behavior patterns and preferences of users, and thus provide more accurate services and optimize network resource allocation. Therefore, the present invention proposes a method for identifying terminal devices based on clustering of TCP (Transfer Control Protocol) timestamp sequences. This method can effectively utilize network traffic data, and by analyzing the clustering characteristics of TCP timestamp sequences, accurately identify the terminal devices in the household network and correlate their traffic data, providing a powerful tool for device counting and traffic analysis for operators, which helps to improve the user experience and network operation efficiency.
[0003] Due to technical and practical requirements limitations, multiple household devices often share the same IP address. Due to the scarcity of IPv4 addresses, household devices usually adopt the Network Address Translation (NAT) protocol to map multiple internal network addresses to one external network address, enabling the router to aggregate the traffic of multiple devices behind its internal network IP address. NAT maps the device (IP, port) pairs in the local network to a single port, which is used together with the external IP address of the network. As long as this internal (IP, port) pair is used, this internal network address mapping will be maintained. The NAT router hides the original IP address, which is generally regarded as a privacy protection measure, making it more difficult to identify the communication of individual devices behind the NAT and unable to distinguish which device the packets of the original mixed traffic of multiple terminals belong to through the MAC (Media Access Control) address and IP address of the terminal device.
[0004] Based on the existence of a certain relationship between certain specific fields and specific terminals, the present invention establishes a linear mapping relationship between the sequence of the TCP timestamp field TSval (the Timstampvalue value in the Options field of the TCP data packet) in the transport layer of the original data traffic and the packet arrival time sequence, and extracts the slope and intercept of the linear function for unsupervised clustering, and finally outputs the number of terminal devices and the corresponding associated TCP flows.
[0005] The TCP timestamp refers to the digital representation used to mark the specific time point of the TCP packet data record, usually defined as the number of seconds elapsed since a fixed starting point (such as 00:00:00 UTC on January 1, 1970). The origin of the timestamp can be traced back to the development needs of computer science and network communication. It is used to ensure the consistency and accuracy of data. Especially in a distributed system, the timestamp becomes a key tool for synchronizing the operations of different nodes and recording the order of events, providing a method to verify the authenticity and validity of data, helping to verify the integrity of data, and ensuring that the data has not been tampered with during transmission.
[0006] The Timestamp value (TSval) in the Options field of the TCP data packet is a key timestamp information, which is used to record the current timestamp of the sender when sending this TCP segment. As part of the timestamp option, TSval follows the option type and length fields and occupies 4 bytes of space. This value records the timestamp at the sending moment in a monotonically increasing manner, usually in milliseconds or microseconds, depending on the implementation of the operating system. Through TSval, the TCP protocol can more accurately measure the round-trip time (RTT), which is of great significance for dynamically adjusting transmission parameters and optimizing network performance. At the same time, in a high-speed network environment, TSval also helps to distinguish new and old segments, prevent confusion caused by sequence number wraparound, and enhance the reliability and security of TCP. Therefore, TSval is an indispensable part of the TCP protocol and plays an important role in ensuring the efficient, accurate and stable data transmission. Summary of the Invention
[0007] The objectives of this invention are as follows: 1. Improve the accuracy of home network monitoring and user behavior analysis: With the rapid growth in the types and quantities of home devices, it is crucial for operators to conduct security monitoring of home networks and research on the resource usage and preferences of user devices. The objective of this invention is to accurately identify various devices used by users from network traffic through innovative methods, thereby improving the accuracy of monitoring and analysis and providing more valuable user behavior data for operators. 2. Solve the problem of device identification in NAT environments: Due to the scarcity of IPv4 addresses and the widespread application of NAT technology, home devices often share the same IP address, rendering traditional methods of differentiating devices by MAC address or IP address ineffective. One of the core objectives of this invention is to solve this technical problem by analyzing and clustering the sequence of the TCP timestamp field TSval to achieve effective identification of individual devices in NAT environments and provide operators with new means of device identification. 3. Provide a new method for calculating the number of terminal devices: Given the challenge in accurately calculating the number of devices in existing technologies when dealing with multi-terminal mixed traffic, this invention proposes an innovative method that utilizes the characteristics of the TCP timestamp field TSval to calculate the number of terminal devices from the original multi-terminal mixed traffic. This method not only overcomes the identification obstacles brought by NAT but also provides operators with an efficient and accurate means of device counting, which helps optimize network resource allocation and management. 4. Provide a method for differentiating TCP flows of terminal devices: In addition to being able to identify the number of terminal devices, this invention can also distinguish TCP flows belonging to different devices from the original mixed traffic data, extract TCP flows belonging to the same device, which can serve as the basis for subsequent extended traffic identification (such as different service application categories accessed by users when surfing the Internet).
[0008] To solve the above technical problems, the specific technical solution of a method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering of the present invention is as follows:
[0009] A method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering includes the following steps:
[0010] Step 1: Collect mixed traffic under a router to form a packet capture file, and extract all TCP flows from the packet capture file. The TCP flows are required to have the IPv4 protocol at the L4 network layer and the TCP protocol at the L5 transport layer; the extracted TCP flows include the following three fields: TCP flow sequence number, packet arrival timestamp, and TCP timestamp;
[0011] Step 2: Extract the above three fields from the TCP stream, filter and delete the data with an empty TCP timestamp value and the streams with a packet count less than the set number. The remaining data is then used to extract the packet arrival timestamp and TCP timestamp according to the same TCP stream sequence number to form different sample sets;
[0012] Step 3: Assume there are N TCP streams, and each TCP stream has data packets, corresponding to packet arrival timestamp values and TCP timestamp values . Therefore, for the i-th TCP stream, there are samples for quasi-linear regression analysis;
[0013] For the j-th packet arrival timestamp value and TCP timestamp value , perform standard normalization respectively to obtain , where represents the normalized packet arrival timestamp value, and represents the normalized TCP timestamp value;
[0014] For the i-th TCP stream, by determining the slope and intercept , make the sum of squared residuals of the following formula minimum:
[0015] ;
[0016] where e is the natural constant, and ln() represents the logarithmic function with the natural constant e as the base;
[0017] Step 4: Obtain a total of N parameter pairs corresponding to N TCP streams from Step 3. As clustering samples, perform unsupervised clustering on these N samples to obtain the number of clusters and the corresponding category set, where each category represents a terminal device, and the number of clusters is the number of terminals; It is also possible to associate all TCP streams corresponding to each category, so as to distinguish all TCP streams of the terminal device.
[0018] Furthermore, the formula for the standard normalization is as follows:
[0019] ;
[0020] where represents the j-th normalized packet arrival timestamp value, and represents the j-th normalized TCP timestamp value; is the number of data packets, and is the j-th packet arrival timestamp value, is the j-th TCP timestamp value ; is the arrival timestamp value of the k-th packet, is the k-th TCP timestamp value, and the subscript min represents the minimum value of this parameter.
[0021] Furthermore, in the step 1, all TCP flows are extracted from the packet capture file according to the five-tuple, and the five-tuple is the source IP, destination IP, source port, destination port, and the protocol group above the L4 layer.
[0022] Furthermore, the set number is 8.
[0023] Furthermore, a linear regression function is used to determine the slope and the intercept .
[0024] Furthermore, the unsupervised clustering is specifically the spectral clustering or density clustering algorithm.
[0025] The method of the present invention utilizes the sequence characteristics of the timestamp field (TSval) in the TCP protocol. By analyzing and clustering it, effective identification of each personal device in the home network under the NAT (Network Address Translation) environment is achieved. It breaks through the limitations of traditional device identification based on MAC addresses or IP addresses, and is particularly suitable for environments where IPv4 addresses are scarce and NAT is widely used. By analyzing the change rules of the timestamp field in the TCP connection, a device identification model is established, thereby accurately distinguishing and counting the terminal devices in the network. The accuracy and reliability of device identification in complex network environments are improved. At the same time, a new and efficient means of device identification and counting is provided for operators, which helps to optimize network resource allocation and management.
[0026] Combined with precise device identification technology, in-depth analysis of the traffic in the home network is carried out to obtain user behavior data, and further distinguish the TCP flows belonging to different devices. Based on device identification, refined analysis of user behavior and precise differentiation of TCP traffic are realized, providing strong support for subsequent network management, service optimization, and business expansion. By monitoring network traffic and combining the device identification results, in-depth analysis of the user's Internet access behavior, device usage preferences, etc. is carried out, and specific algorithms are used to distinguish the TCP flows of different devices. More in-depth and detailed user behavior data is provided for operators, which helps to formulate more personalized service strategies and improve the user experience. At the same time, the accuracy and efficiency of network traffic analysis are improved, which helps to discover potential network problems and take corresponding optimization measures, providing rich data support and technical reserves for subsequent network research and technological innovation.
[0027] The present invention has the following technical effects:
[0028] 1. Significantly improve the accuracy of home network monitoring and user behavior analysis: By innovatively and precisely identifying various devices used by users from network traffic, the present invention overcomes the identification difficulties caused by the large variety and rapid increase in the number of device types in traditional methods, and realizes refined processing of home network monitoring and user behavior analysis. This can not only provide operators with more in-depth and detailed user behavior data, but also significantly improve the accuracy and application value of data analysis. It enhances the operator's insight into the home network environment and user habits, helps it formulate more personalized service strategies, improve the user experience, and at the same time strengthen network security monitoring and prevent potential risks. The algorithm of this application does not use a machine learning model, but a fixed-mode algorithm, which consumes less computing resources, has faster calculation results, and is more suitable for the scenario of real-time identification of the number of terminal devices; at the same time, since the algorithm itself does not rely on complex machine learning models, it avoids the uncertainty and potential overfitting problems in the model training process, further enhances the stability and reliability of the system, reduces false alarms or missed alarms caused by algorithm errors, and provides users with more accurate and stable terminal device number identification services.
[0029] 2. Effectively solve the technical problem of device identification in the NAT environment:
[0030] In response to the challenges of device identification brought about by the scarcity of IPv4 addresses and the widespread application of NAT, the present invention realizes the effective identification of personal devices in the NAT environment by analyzing the sequence of the TCP timestamp field TSval. This innovative method breaks through the limitations of traditional identification based on MAC addresses or IP addresses, and provides new ideas and ways for device identification. It enables operators to accurately grasp the specific information of each device in the home network in the NAT environment, provides a solid data foundation for subsequent network management, resource allocation, and service optimization, and improves the overall network management efficiency.
[0031] 3. Provide a new method for efficient and accurate calculation of the number of terminal devices and TCP flow differentiation: The present invention utilizes the characteristics of the TCP timestamp field TSval to propose a novel method for calculating the number of devices in multi-terminal mixed traffic and can distinguish TCP flows belonging to different devices from the original data. This method not only solves the problem of device counting in the NAT environment, but also realizes refined differentiation of TCP traffic, providing convenience for subsequent more advanced traffic analysis (such as business application category identification). It provides operators with an efficient and accurate device counting and traffic analysis tool, helps them optimize network resource allocation and management, and improves the overall quality and efficiency of network services. At the same time, it also provides strong data support and technical reserves for subsequent network research and technological innovation. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic flow chart of the method of the present invention;
[0033] Figure 2 It is a schematic diagram of the clustering result after quasi-linear regression fitting in the embodiment of the present invention;
[0034] Figure 3 In the three categories of the clustering result in the embodiment of the present invention, an example of a quasi-linear regression fitting graph is drawn by randomly selecting a point (slope - intercept) for each category. Detailed implementation manners
[0035] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering of the present invention with reference to the accompanying drawings.
[0036] The present invention aims at the scenario of multiple downstream device terminals after NAT conversion by a router in a home optical network environment, and is based on the following principle as a prerequisite for the TCP timestamp TSval value sequence clustering method: The timestamp value sent in TSval is obtained from a virtual clock, which is called a "timestamp clock". Its value must be at least approximately proportional to the real time to measure the actual RTT.
[0037] Based on the above conditions, the present invention uses the Quasi-Linear Regression algorithm to establish an approximate linear mapping relationship model for the packet arrival timestamp and the TCP timestamp TSval value of the same TCP flow. The specific method is as follows: First, a linear function is used to map the real time to obtain a linear output, and then the linear output passes through a univariate quasi-linear (non-linear but approximately linear) function to obtain the final fitting result. The key point of the present invention is to find a suitable quasi-linear function. Inspired by the neural network activation function Softplus (this function is approximately linear in the part where x > 0), the present invention designs a quasi-linear function suitable for non-negative approximate linear regression fitting analysis as follows:
[0038]
[0039] Where ln() represents the natural logarithm function with the natural constant e as the base, and the parameters a and b correspond to the slope and intercept of the linear regression function. The parameterless quasi-linear function It satisfies f(0)=0 and is approximately a pure linear function y = x when x>0. This quasi-linear model has two parameters: slope a and intercept b. At this time, a binary tuple (a, b) corresponding to this TCP flow can be obtained. By this method, binary tuples (a, b) corresponding to all TCP flows of the original mixed traffic can be obtained, thereby obtaining samples for clustering (the number of samples is equal to the number of TCP flows, and the number of features is 2). Use spectral clustering or density clustering methods to cluster the samples to determine the number of terminal devices and the corresponding associated TCP flows. Specifically, as Figure 1 shown, this method includes the following steps:
[0040] Step S1: The original multi-terminal mixed traffic extracts TCP flows;
[0041] The mixed traffic collected simultaneously from multiple terminals under a router forms a PCAP (Packet Capture) file. All TCP flows are extracted from this PCAP file according to the five-tuple (source IP, destination IP, source port, destination port, protocol group above L4 layer). The TCP flow requires that the L4 layer network layer is the IPv4 protocol and the L5 layer transport layer is the TCP protocol.
[0042] The corresponding fields of the extracted TCP flows are prepared for subsequent clustering methods. Each data packet corresponds to a packet arrival timestamp value and at most one TCP timestamp value (the data packet may not have a TCP timestamp value). The corresponding fields of the extracted TCP flows only need to extract 3 types of fields:
[0043] 1) TCP flow sequence number, which provides a basis for sample set division for subsequent quasi-linear regression analysis;
[0044] 2) Packet arrival timestamp, which is used as the input value for subsequent quasi-linear regression analysis;
[0045] 3) TCP timestamp, which is used as the output value for subsequent quasi-linear regression analysis.
[0046] Specifically, the tshark.exe executable file included in Wireshark (a professional software name for capturing and analyzing network traffic) can be used to implement network traffic data collection. The function of this file is to quickly and conveniently read the TCP flows of the original mixed traffic data file through the command line in the console (such as Windows PowerShell), extract the TCP flow sequence number (tcp.stream), packet arrival timestamp (frame.time_epoch), and TCP timestamp (tcp.options.timestamp.tsval) from the TCP flows and output them to a CSV file. Subsequently, the packet arrival timestamp and TCP timestamp of the same TCP flow can be read from this CSV file for quasi-linear regression analysis.
[0047] Step 2: Packet filtering;
[0048] Read the above three types of fields from the CSV file output in Step 1. It is necessary to filter and delete the data with empty TCP timestamp values and the flows with too few packets (less than 8). The remaining data is then extracted according to the same TCP flow sequence number to form different sample sets of packet arrival timestamps and TCP timestamps. The quasi-linear regression model is as follows:
[0049]
[0050] Each TCP flow sequence number corresponds to a sample set of a quasi-linear regression model (1 TCP flow corresponds to multiple data packets (not less than 8), and each data packet corresponds to a packet arrival timestamp (input value of the quasi-linear regression model )) and TCP timestamp (output value of the quasi-linear regression model ). The number of samples for the quasi-linear regression analysis of each TCP flow is the number of data packets in that flow.
[0051] Step 3: Obtain clustering feature values, i.e., slope and intercept, through quasi-linear regression analysis;
[0052] Suppose there are N TCP flows, and each TCP flow has data packets, corresponding to packet arrival timestamp values and TCP timestamp values . Thus, for the i-th TCP flow, there are samples for quasi-linear regression analysis. The quasi-linear regression analysis method finds the best function match for the data by minimizing the sum of the squares of the errors of a linear function. Its goal is to solve the slope a and intercept b of this linear function so that the sum of the squares of the packet arrival timestamp values and TCP timestamp values reaches the minimum.
[0053] Since the timestamp is the number of seconds since January 1, 1970, it is generally a specific value greater than 1 billion. Therefore, it is necessary to normalize the two types of timestamps before quasi-linear regression analysis. Since the approximate linearity of the non-linear function is only obvious in the non-negative number domain, the traditional standard normalization (Z-score) method cannot be used for normalization. Therefore, this patent improves the standard normalization method and proposes a normalization method applicable to the positive number domain, replacing the average value of the standard normalization with the minimum value, while the calculation method of the denominator variance remains unchanged. For the j-th packet arrival timestamp value , the standard normalization calculation formula is as follows:
[0054]
[0055] For obtained after normalization calculation , where represents the normalized packet arrival timestamp value , represents the normalized TCP timestamp value; is the k-th packet arrival timestamp value, is the k-th TCP timestamp value, and the subscript min represents the minimum value of this parameter.
[0056] The goal of the quasi-linear regression analysis is to determine the slope and intercept of the linear function for the i-th TCP flow, such that the sum of squared residuals of the following formula is minimized:
[0057]
[0058] The quasi-linear regression analysis algorithm to determine the slope and intercept can be quickly implemented using the linear regression function of the Skit-learn package in Pyhton. The example is as follows:
[0059] from sklearn.linear_model import LinearRegression
[0060] model = LinearRegression()
[0061] model.fit(s, t)
[0062] print(model.coef_[0], model.intercept_)
[0063] where model.coef_[0] and model.intercept_ correspond to the slope a and intercept b respectively.
[0064] Step 4: Cluster based on the quasi-linear regression analysis parameters;
[0065] From Step 3, the unique pair of parameters for the quasi-linear regression analysis generated by the -th TCP flow can be obtained. In this way, a total of N TCP flows generate N parameter pairs , which is equivalent to generating N samples with 2 eigenvalue. Performing unsupervised clustering on these N samples to obtain the number of clusters and the corresponding
[0066] A set of categories, where each category represents a terminal device, and the number of clusters is the number of terminals. In addition, since each corresponds to a unique TCP flow sequence number, each category (terminal device) can be associated with all corresponding TCP flows, thus distinguishing all TCP flows of the terminal device.
[0067] Since the clustering algorithm needs to output the number of categories (number of devices) through the algorithm and cannot be specified in advance, only clustering algorithms that do not require pre-determining the number of clusters can be used, such as spectral clustering and density clustering algorithms. The present invention adopts the spectral clustering algorithm. Spectral clustering is a method of converting a data set into a graph structure and using the eigenvalues and eigenvectors of the Laplacian matrix of the graph for dimensionality reduction and clustering. It constructs the edge weights of the graph by calculating the similarity between data points, then represents the data points in a low-dimensional space using eigenvectors, and finally applies the K-nearest neighbor algorithm to achieve clustering in the dimensionality-reduced space to discover clusters with complex shapes and structures.
[0068] Spectral clustering can be quickly implemented by calling the spectral clustering algorithm of the Skit-learn package in Pyhton (from sklearn.cluster import SpectralClustering).
[0069] Experimental verification: One optical modem is connected to one router, and this router is connected to 3 mobile phones via WiFi. 10 minutes of data is collected through router mirroring, about 900,000 packets in total, and 537 TCP flows. After filtering through the packet filtering rules (only retaining TCP flows with the number of packets with non-empty TCP timestamps not less than 8), 304 valid TCP flows are retained. The clustering results after fitting by the quasi-linear regression analysis method are as follows Figure 2 shown, divided into 3 categories; points (outliers) that are far from the center point in each category are deleted, and the optimized clustering results are as Figure 3 shown. For each of the 3 categories in spectral clustering, a point (slope - intercept) is randomly selected to draw a quasi-linear regression fitting graph. It can be seen from the graph that the method of the present invention can accurately identify 3 mobile phone devices. In addition, the boundaries between the 3 categories are relatively large, indicating that the method of the present invention can effectively distinguish TCP flows of different devices. An example of the quasi-linear regression analysis results is drawn at a point near the center of these 3 clusters as follows. The quasi-linear regression analysis shows that the TCP timestamp of the mobile phone is positively correlated with the packet arrival time.
[0070] It will be understood that the present invention is described by way of some embodiments, and those skilled in the art will be aware that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the present invention. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering, characterized in that: The following steps are involved: Step 1: Collect mixed traffic from a router to form a data packet capture file, and extract all TCP flows from the data packet capture file. The TCP flow requires that the L4 network layer is the IPv4 protocol and the L5 transport layer is the TCP protocol; the extracted TCP flow includes the following three fields: TCP flow sequence number, packet arrival timestamp, and TCP timestamp; Step 2: Extract the above three fields in the TCP flow, filter and delete the data with empty TCP timestamp value and the flow with less than the set number of packets, and then extract the packet arrival timestamp and TCP timestamp of the remaining data according to the same TCP flow sequence number to form different sample sets; Step 3: Assume there are N TCP flows, each with Data packets, corresponding to Packet arrival timestamp value and TCP timestamp value , so for the i-th TCP flow, we have The samples were used for quasi-linear regression analysis; Arrival timestamp value of the jth packet 、TCP timestamp value After standard normalization, we get ,in, Represents the normalized packet arrival timestamp value, Indicates the normalized TCP timestamp value; For the i-th TCP flow, by determining the slope and the intercept , so that the following residual sum of squares is minimized: ; Wherein, e is a natural constant, and ln() represents a logarithmic function with the natural constant e as the base; Step 4: Obtain N parameter pairs corresponding to a total of N TCP flows from step 3 , as cluster samples, perform unsupervised clustering on these N samples to obtain the number of clusters and the corresponding A set of categories, where each category represents a terminal device and the number of clusters is the number of terminals.
2. According to claim 1, a method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering, characterized in that: The formula for the standard normalization is as follows: ; in, represents the jth normalized packet arrival timestamp value, represents the j-th normalized TCP timestamp value; is the number of packets, is the arrival timestamp value of the jth packet, is the jth TCP timestamp value ; is the arrival timestamp value of the kth packet, is the kth TCP timestamp value. The subscript min indicates the minimum value of this parameter.
3. A method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering according to claim 2, characterized in that: In the step 1, all TCP flows are extracted from the data packet capture file according to five tuples, where the five tuples are source IP, destination IP, source port, destination port, and protocol groups above L4 layer.
4. The method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering according to claim 3, characterized in that: The set number is 8.
5. The method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering according to claim 4, characterized in that: Determine the slope using the linear regression function and the intercept .
6. The method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering according to claim 5, characterized in that: The unsupervised clustering is specifically a spectral clustering algorithm or a density clustering algorithm.
7. A method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering according to claim 6, characterized in that: By associating all corresponding TCP flows with each category, all TCP flows of the terminal device can be distinguished.
Citation Information
Patent Citations
Method and device for counting number of Linux hosts after NAT and electronic equipment
CN117176407A
Counting method for system terminal equipment in NATed network based on optical network
CN119277242A