Method for identifying terminal equipment in NATed network based on TCP timestamp sequence clustering

By analyzing the sequence of TCP timestamp fields TSval, a quasi-linear regression analysis model was established, and the problem of device identification in the NAT environment was solved, efficient and accurate terminal device identification and TCP flow distinction were achieved, and network management and user experience were improved.

CN119996266AActive Publication Date: 2025-05-13UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510452394.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

In the NAT environment, home devices share the same IP address, making it difficult to accurately identify and distinguish terminal devices through traditional MAC addresses or IP addresses, affecting the accuracy of network management and traffic analysis.

Method used

By analyzing the sequence of the TCP timestamp field TSval, a device recognition model is established using quasi-linear regression analysis to realize terminal device recognition and TCP flow distinction in the NAT environment.

Benefits of technology

It effectively solves the problem of device identification in the NAT environment, improves the accuracy and reliability of terminal device identification, provides operators with efficient and accurate device counting and traffic analysis tools, and improves network management and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996266A_ABST
    Figure CN119996266A_ABST
Patent Text Reader

Abstract

The invention discloses a method for identifying terminal equipment in an NATed network based on TCP timestamp sequence clustering, and belongs to the technical field of broadband networks in data communication. The method comprises the following steps: firstly, extracting all TCP streams from a PCAP file formed by mixed traffic collected by a plurality of terminals under a router at the same time according to a quintuple, and carrying out packet filtering to delete invalid data rows; secondly, performing quasi-linear regression analysis on the residual data to obtain slope and intercept parameters required by clustering, and performing special normalization processing on a timestamp in the process; and finally, clustering the generated samples by using a spectral clustering algorithm so as to determine the number of terminal devices and the correspondingly associated TCP streams. The method is suitable for scenes with scarce IPv4 addresses and wide application of NAT (Network Address Translation). The precision of home network monitoring and user behavior analysis can be obviously improved, and more valuable data can be provided for operators; the equipment identification problem in the NAT environment is solved; an efficient and accurate equipment counting and TCP flow distinguishing method is provided, and reasonable distribution and management of network resources are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of broadband networks in data communications, and in particular relates to a method for identifying terminal equipment in a NATed network based on TCP timestamp sequence clustering. Background Art

[0002] With the rapid progress of information technology, the types and number of devices that home users rely on have grown rapidly, covering personal computers, mobile devices (such as smart phones), surveillance cameras, and smart home products (such as smart speakers, automatic sweepers, etc.). The widespread use of these devices has not only improved the convenience of family life, but also put forward higher requirements for home network bandwidth. In order to achieve effective management and optimization of home networks, operators need to be able to accurately identify the devices used by users through network traffic and perform correlation analysis on the traffic of user devices in order to better understand the user's behavior patterns and preferences, thereby providing more accurate services and optimizing network resource allocation. Therefore, the present invention proposes a terminal device identification method based on TCP (Transfer Control Protocol) timestamp sequence clustering. The method can effectively utilize network traffic data, accurately identify terminal devices in the home network by analyzing the clustering characteristics of TCP timestamp sequences, and associate their traffic data, providing operators with powerful device counting and traffic analysis tools, which helps to improve user experience and network operation efficiency.

[0003] Due to technical and practical limitations, multiple home devices often share the same IP address. Due to the scarcity of IPv4 addresses, home devices often use the Network Address Translation (NAT) protocol to map multiple internal network addresses to one external network address, allowing the router to aggregate the traffic of multiple devices behind its internal network IP address. NAT maps a device (IP, port) pair in the local network to a port that is used with the network's external IP address, and this internal network address mapping is maintained as long as this internal (IP, port) pair is used. NAT routers hide the original IP address, which is often seen as a privacy protection measure, making it more difficult to identify the communications of personal devices behind NAT, and it is impossible to distinguish which device the original data packet of multiple terminal mixed traffic belongs to through the terminal device's MAC (Media Access Control) address and IP address.

[0004] The present invention is based on the existence of a certain relationship between certain specific fields and specific terminals. It establishes a linear mapping relationship between the TCP timestamp field TSval (Timstampvalue in the Options field of the TCP data message) sequence in the TCP protocol of the transport layer in the original data traffic and the packet arrival time series, and extracts the slope and intercept of the linear function for unsupervised clustering, and finally outputs the number of terminal devices and the corresponding associated TCP flows.

[0005] TCP timestamp refers to the digital representation used to mark the specific time point of TCP message data record, usually defined as the number of seconds that have passed since a fixed starting point (such as 00:00:00 UTC on January 1, 1970). The origin of timestamp can be traced back to the development needs of computer science and network communication. It is used to ensure the consistency and accuracy of data, especially in distributed systems. Timestamps have become a key tool for synchronizing operations of different nodes and recording the order of events. It provides a method to verify the authenticity and validity of data, which can help verify the integrity of data and ensure that data has not been tampered with during transmission.

[0006] The Timestamp value (TSval) in the Options field of the TCP datagram is a key timestamp information, which is used to record the current timestamp of the sender when sending the TCP segment. As part of the timestamp option, TSval follows the option type and length fields and occupies 4 bytes of space. This value records the timestamp of the sending moment in a monotonically increasing manner, usually in milliseconds or microseconds, depending on the implementation of the operating system. Through TSval, the TCP protocol can measure the round-trip time (RTT) more accurately, which is of great significance for dynamically adjusting transmission parameters and optimizing network performance. At the same time, in a high-speed network environment, TSval also helps to distinguish between new and old segments, prevent confusion caused by sequence number wraparound, and enhance the reliability and security of TCP. Therefore, TSval is an indispensable part of the TCP protocol and plays an important role in ensuring efficient, accurate and stable data transmission. Summary of the invention

[0007] The purpose of the invention is: 1. Improve the accuracy of home network monitoring and user behavior analysis: With the rapid growth of the types and number of home devices, it is crucial for operators to conduct security monitoring of home networks and research on user device usage resources and preferences. The purpose of the invention is to accurately identify various devices used by users from network traffic through innovative methods, thereby improving the accuracy of monitoring and analysis and providing operators with more valuable user behavior data. 2. Solve the problem of device identification under NAT environment: Due to the scarcity of IPv4 addresses and the widespread application of NAT technology, home devices often share the same IP address, resulting in the failure of the method of distinguishing devices by traditional MAC addresses or IP addresses. One of the core purposes of the invention of the present invention is to solve this technical problem. By analyzing and clustering the sequence of the TCP timestamp field TSval, effective identification of personal devices in a NAT environment is achieved, providing operators with new means of device identification. 3. Provide a new method for calculating the number of terminal devices: In view of the challenge that the existing technology is difficult to accurately calculate the number of devices when processing multi-terminal mixed traffic, the present invention proposes an innovative method that uses the characteristics of the TCP timestamp field TSval to calculate the number of terminal devices from the original multi-terminal mixed traffic. This method not only overcomes the identification barriers brought by NAT, but also provides operators with an efficient and accurate device counting method, which helps to optimize network resource allocation and management. 4. Provide a method for distinguishing TCP flows of terminal devices: In addition to identifying the number of terminal devices, the present invention can also distinguish TCP flows belonging to different devices from the original mixed traffic data, and can extract TCP flows belonging to the same device, which can serve as the basis for subsequent expanded traffic identification (such as different business application categories accessed by users online).

[0008] In order to solve the above technical problems, the specific technical solution of a method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering of the present invention is as follows:

[0009] A method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering comprises the following steps:

[0010] Step 1: Collect mixed traffic from a router to form a data packet capture file, and extract all TCP flows from the data packet capture file. The TCP flow requires that the L4 network layer is the IPv4 protocol and the L5 transport layer is the TCP protocol; the extracted TCP flow includes the following three fields: TCP flow sequence number, packet arrival timestamp, and TCP timestamp;

[0011] Step 2: Extract the above three fields in the TCP flow, filter and delete the data with empty TCP timestamp value and the flow with less than the set number of packets, and then extract the packet arrival timestamp and TCP timestamp of the remaining data according to the same TCP flow sequence number to form different sample sets;

[0012] Step 3: Assume there are N TCP flows, each with Data packets, corresponding to Packet arrival timestamp value and TCP timestamp value , so for the i-th TCP flow, we have The samples were used for quasi-linear regression analysis;

[0013] Arrival timestamp value of the jth packet 、TCP timestamp value After standard normalization, we get ,in, Indicates the normalized packet arrival timestamp value, Indicates the normalized TCP timestamp value;

[0014] For the i-th TCP flow, by determining the slope and the intercept , so that the following residual sum of squares is minimized:

[0015] ;

[0016] Wherein, e is a natural constant, and ln() represents a logarithmic function with the natural constant e as the base;

[0017] Step 4: Obtain N parameter pairs corresponding to a total of N TCP flows from step 3 , as cluster samples, perform unsupervised clustering on these N samples to obtain the number of clusters and the corresponding A category set, where each category represents a terminal device and the number of clusters is the number of terminals; all TCP flows corresponding to each category can also be associated to distinguish all TCP flows of the terminal device.

[0018] Furthermore, the formula for the standard normalization is as follows:

[0019] ;

[0020] in, represents the jth normalized packet arrival timestamp value, represents the j-th normalized TCP timestamp value; is the number of packets, is the arrival timestamp value of the jth packet, is the jth TCP timestamp value ; is the arrival timestamp value of the kth packet, is the kth TCP timestamp value. The subscript min indicates the minimum value of this parameter.

[0021] Furthermore, in step 1, all TCP flows are extracted from the data packet capture file according to five-tuples, where the five-tuples are source IP, destination IP, source port, destination port, and protocol group above L4 layer.

[0022] Furthermore, the set number is 8.

[0023] Furthermore, the slope is determined using the linear regression function and the intercept .

[0024] Furthermore, the unsupervised clustering is specifically a spectral clustering algorithm or a density clustering algorithm.

[0025] The method of the present invention utilizes the sequence characteristics of the timestamp field (TSval) in the TCP protocol, and realizes effective identification of each personal device in the home network under the NAT (network address translation) environment by analyzing and clustering it. It breaks through the limitations of traditional MAC address or IP address-based device identification and is particularly suitable for environments where IPv4 addresses are scarce and NAT is widely used. By analyzing the change pattern of the timestamp field in the TCP connection, a device identification model is established to accurately distinguish and count terminal devices in the network. It improves the accuracy and reliability of device identification in complex network environments, and at the same time provides operators with a new and efficient device identification and counting method, which helps to optimize network resource allocation and management.

[0026] Combined with precise device identification technology, the traffic in the home network is deeply analyzed to obtain user behavior data and further distinguish the TCP flows belonging to different devices. Based on device identification, the refined analysis of user behavior and accurate distinction of TCP flows are achieved, providing strong support for subsequent network management, service optimization and business expansion. By monitoring network traffic and combining device identification results, the user's Internet behavior, device usage preferences, etc. are deeply analyzed, and specific algorithms are used to distinguish the TCP flows of different devices. It provides operators with more in-depth and detailed user behavior data, which helps to formulate more personalized service strategies and improve user experience, and improves the accuracy and efficiency of network traffic analysis, helps to discover potential network problems and take corresponding optimization measures, and provides rich data support and technical reserves for subsequent network research and technological innovation.

[0027] The present invention has the following technical effects:

[0028] 1. Significantly improve the accuracy of home network monitoring and user behavior analysis: By innovatively and accurately identifying the various devices used by users from network traffic, the present invention overcomes the identification difficulties caused by the various types of devices and the surge in number in traditional methods, and realizes the refined processing of home network monitoring and user behavior analysis. This can not only provide operators with more in-depth and detailed user behavior data, but also significantly improve the accuracy and application value of data analysis. It enhances the operator's insight into the home network environment and user habits, helps them to formulate more personalized service strategies, improve user experience, and strengthen network security monitoring to prevent potential risks. The algorithm of this application does not use a machine learning model, but a fixed mode algorithm, which consumes less computing resources and calculates results faster, and is more suitable for real-time terminal device number recognition scenarios; at the same time, because the algorithm itself does not rely on complex machine learning models, it avoids uncertainty and potential overfitting problems in the model training process, further enhances the stability and reliability of the system, reduces false positives or omissions caused by algorithm errors, and provides users with a more accurate and stable terminal device number recognition service.

[0029] 2. Effectively solve the technical problem of device identification in NAT environment:

[0030] In response to the device identification challenges brought about by the scarcity of IPv4 addresses and the widespread use of NAT, the present invention realizes effective identification of personal devices in a NAT environment by analyzing the sequence of the TCP timestamp field TSval. This innovative method breaks the limitations of traditional identification based on MAC addresses or IP addresses and provides new ideas and approaches for device identification. It enables operators to accurately grasp the specific information of each device in the home network in a NAT environment, providing a solid data foundation for subsequent network management, resource allocation and service optimization, and improving the overall network management efficiency.

[0031] 3. Provide a new method for efficient and accurate calculation of the number of terminal devices and TCP stream differentiation: The present invention uses the characteristics of the TCP timestamp field TSval to propose a novel method to calculate the number of devices in multi-terminal mixed traffic, and can distinguish TCP streams belonging to different devices from the original data. This method not only solves the problem of device counting in the NAT environment, but also realizes the refined differentiation of TCP traffic, which facilitates subsequent more advanced traffic analysis (such as business application category identification). It provides operators with an efficient and accurate device counting and traffic analysis tool, which helps them optimize network resource allocation and management and improve the overall quality and efficiency of network services. At the same time, it also provides strong data support and technical reserves for subsequent network research and technological innovation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic diagram of the process of the present invention;

[0033] Figure 2 Schematic diagram of clustering results after quasi-linear regression fitting in an embodiment of the present invention;

[0034] Figure 3 This is an example of a quasi-linear regression fitting graph drawn by randomly selecting a point (slope-intercept) from each of the three categories of clustering results in an embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to better understand the purpose, structure and function of the present invention, the following is a further detailed description of a terminal device identification method in a NATed network based on TCP timestamp sequence clustering in conjunction with the accompanying drawings.

[0036] The present invention is aimed at the scenario of multiple downstream device terminals of a router in a home optical network environment after NAT conversion, and is based on the following principle as the prerequisite of the TCP timestamp TSval value sequence clustering method: the timestamp value sent in TSval is obtained from a virtual clock, called "timestamp clock". Its value must be at least approximately proportional to the real time to measure the actual RTT.

[0037] Based on the above conditions, the present invention uses a quasi-linear regression algorithm to establish an approximate linear mapping relationship model for the packet arrival timestamp and TCP timestamp TSval value of the same TCP flow. The specific method is: first use a linear function to map the real-time time to obtain a linear output, and then the linear output is passed through a single variable quasi-linear (non-linear but approximately linear) function to obtain the final fitting result. The key point of the present invention is to find a suitable quasi-linear function. Inspired by the neural network activation function Softplus (this function is approximately linear in the x>0 part), the present invention designs a quasi-linear function suitable for non-negative approximate linear regression fitting analysis as follows:

[0038]

[0039] Where ln() represents a logarithmic function with the natural constant e as the base, and the parameters a and b correspond to the slope and intercept of the linear regression function. Quasi-linear function without parameters It satisfies f(0)=0, and approximates the pure linear function y=x in the part x>0. This quasi-linear model has two parameters: slope a and intercept b. At this time, a tuple (a, b) corresponding to the TCP flow can be obtained. This method can obtain the tuples (a, b) corresponding to all TCP flows of the original mixed traffic, thereby obtaining samples for clustering (the number of samples is equal to the number of TCP flows, and the number of features is 2). Use spectral clustering or density clustering methods to cluster the samples to determine the number of terminal devices and the corresponding associated TCP flows. Specifically, Figure 1 As shown, the method comprises the following steps:

[0040] Step S1: extracting TCP stream from original multi-terminal mixed traffic;

[0041] A PCAP (Packet Capture) file is formed from the mixed traffic collected simultaneously from multiple terminals under a router. All TCP flows are extracted from the PCAP file according to the five-tuple (source IP, destination IP, source port, destination port, and protocol group above L4 layer). The TCP flow requires that the L4 network layer is the IPv4 protocol and the L5 transport layer is the TCP protocol.

[0042] The extracted TCP stream corresponding fields are prepared for the subsequent clustering method. Each data packet corresponds to a packet arrival timestamp value and at most one TCP timestamp value (a data packet may not have a TCP timestamp value). The extracted TCP stream corresponding fields only need to extract three fields:

[0043] 1) TCP flow sequence number, which provides a basis for sample set division for subsequent quasi-linear regression analysis;

[0044] 2) Packet arrival timestamp, which is used as the input value for subsequent quasi-linear regression analysis;

[0045] 3) TCP timestamp, which is used as the output value of the subsequent quasi-linear regression analysis.

[0046] Specifically, you can use the tshark.exe executable file that comes with Wireshark (a professional software that can capture and analyze network traffic) to collect network traffic data. The function of this file is to quickly and conveniently read the TCP stream of the original mixed traffic data file through the command line in the console (such as Windows PowerShell), extract the TCP stream sequence number (tcp.stream), packet arrival timestamp (frame.time_epoch), TCP timestamp (tcp.options.timestamp.tsval) from the TCP stream and output them to a CSV file. Later, you can read the CSV file to extract the packet arrival timestamp and TCP timestamp of the same TCP stream for quasi-linear regression analysis.

[0047] Step 2: Packet filtering;

[0048] Read the above three fields from the CSV file output in step 1. It is necessary to filter out data with empty TCP timestamp values ​​and flows with too few packets (less than 8). The remaining data is then extracted according to the same TCP flow sequence number to form different sample sets of packet arrival timestamps and TCP timestamps. The quasi-linear regression model is as follows:

[0049]

[0050] Each TCP flow sequence number corresponds to a sample set of a quasi-linear regression model (1 TCP flow corresponds to multiple data packets (no less than 8), each data packet corresponds to a packet arrival timestamp (quasi-linear regression model input value ) and TCP timestamp (output value of the quasi-linear regression model ). The number of samples for the quasi-linear regression analysis of each TCP flow is the number of packets in the flow.

[0051] Step 3: Obtain clustering characteristic values, namely slope and intercept, through quasi-linear regression analysis;

[0052] Assume there are N TCP flows, each with Data packets, corresponding to Packet arrival timestamp value and TCP timestamp value , so for the i-th TCP flow, there is The samples are used for quasi-linear regression analysis. Quasi-linear regression analysis seeks the best function match for the data by minimizing the sum of squares of a linear function error. Its goal is to solve the slope a and intercept b of the linear function so that the sum of squares of the packet arrival timestamp value and the TCP timestamp value is minimized.

[0053] Since the timestamp is the number of seconds since January 1, 1970, which is generally a specific value greater than 1 billion, the two timestamps need to be normalized before quasi-linear regression analysis. Since the approximate linearity of nonlinear functions is only obvious in the non-negative domain, the traditional standard normalization (Z-score) method cannot be used for normalization. Therefore, this patent improves the standard normalization method and proposes a normalization method suitable for normalization to the positive domain, replacing the mean value of the standard normalization with the minimum value, while the calculation method of the denominator variance remains unchanged. For the arrival timestamp value of the jth packet The standard normalization calculation formulas are as follows:

[0054]

[0055] right After normalization calculation, we get ,in, Indicates the normalized packet arrival timestamp value , Indicates the normalized TCP timestamp value; is the arrival timestamp value of the kth packet, is the kth TCP timestamp value. The subscript min indicates the minimum value of this parameter.

[0056] The goal of quasi-linear regression analysis is to determine the slope of the linear function for the i-th TCP flow and the intercept , so that the following residual sum of squares is minimized:

[0057]

[0058] Using the linear regression function of Python's Skit-learn package, you can quickly implement a quasi-linear regression analysis algorithm to determine the slope and the intercept , an example is as follows:

[0059] from sklearn.linear_model import LinearRegression

[0060] model = LinearRegression()

[0061] model.fit(s, t)

[0062] print(model.coef_[0], model.intercept_)

[0063] Where model.coef_[0] and model.intercept_ correspond to the slope a and intercept b respectively.

[0064] Step 4: Parameter clustering based on quasi-linear regression analysis;

[0065] From step 3, we can get TCP flows generate a unique pair of parameters for quasi-linear regression analysis , so a total of N TCP flows generate N parameter pairs , which is equivalent to generating N samples with 2 eigenvalues.

[0066] Perform unsupervised clustering on these N samples to obtain the number of clusters and the corresponding A set of categories, where each category represents a terminal device, and the number of clusters is the number of terminals. Corresponding to a unique TCP flow sequence number, each category (terminal device) can be associated with all corresponding TCP flows, thereby distinguishing all TCP flows of the terminal device.

[0067] Since the clustering algorithm needs to output the number of categories (number of devices) through the algorithm and cannot be specified in advance, only clustering algorithms that do not require the pre-determined number of clusters can be used, such as spectral clustering and density clustering algorithms. The present invention adopts a spectral clustering algorithm, which is a method of converting a data set into a graph structure and using the eigenvalues ​​and eigenvectors of the Laplacian matrix of the graph for dimensionality reduction and clustering. It constructs the edge weights of the graph by calculating the similarity between data points, then uses eigenvectors to re-represent the data points in a low-dimensional space, and finally applies the K-nearest neighbor algorithm to achieve clustering in the reduced-dimensional space to discover clusters of complex shapes and structures.

[0068] Spectral clustering can be quickly implemented by calling the spectral clustering algorithm of Python's Skit-learn package (from sklearn.cluster importSpectralClustering).

[0069] Experimental verification: A router is connected to an optical modem, and three mobile phones are connected to the router via WiFi. The router image is used to collect data for 10 minutes, about 900,000 data packets, a total of 537 TCP flows, and 304 valid TCP flows are retained after the packet filtering rules (only TCP flows with non-empty TCP timestamps and a number of packets of not less than 8 are retained). The clustering results after fitting by the quasi-linear regression analysis method are as follows Figure 2 As shown in , it is divided into 3 categories; in each category, the points (outliers) that are far away from the center point are deleted, and the optimized clustering results are as follows Figure 3 As shown in the figure, a point (slope-intercept) is randomly selected from each of the three categories of spectral clustering to draw a quasi-linear regression fitting diagram. It can be seen from the figure that the method of the present invention can accurately identify three mobile phone devices. In addition, the boundaries of the three categories are relatively large, indicating that the method of the present invention can effectively distinguish the TCP flows of different devices. Select a point at the center of these three clusters to draw a quasi-linear regression analysis result example as follows. The quasi-linear regression analysis shows that the TCP timestamp of the mobile phone is positively correlated with the packet arrival time.

[0070] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.

Claims

1. A method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering, characterized in that: The following steps are involved: Step 1: Collect mixed traffic from a router to form a data packet capture file, and extract all TCP flows from the data packet capture file. The TCP flow requires that the L4 network layer is the IPv4 protocol and the L5 transport layer is the TCP protocol; the extracted TCP flow includes the following three fields: TCP flow sequence number, packet arrival timestamp, and TCP timestamp; Step 2: Extract the above three fields in the TCP flow, filter and delete the data with empty TCP timestamp value and the flow with less than the set number of packets, and then extract the packet arrival timestamp and TCP timestamp of the remaining data according to the same TCP flow sequence number to form different sample sets; Step 3: Assume there are N TCP flows, each with Data packets, corresponding to Packet arrival timestamp value and TCP timestamp value , so for the i-th TCP flow, we have The samples were used for quasi-linear regression analysis; Arrival timestamp value of the jth packet 、TCP timestamp value After standard normalization, we get ,in, Represents the normalized packet arrival timestamp value, Indicates the normalized TCP timestamp value; For the i-th TCP flow, by determining the slope and the intercept , so that the following residual sum of squares is minimized: ; Wherein, e is a natural constant, and ln() represents a logarithmic function with the natural constant e as the base; Step 4: Obtain N parameter pairs corresponding to a total of N TCP flows from step 3 , as cluster samples, perform unsupervised clustering on these N samples to obtain the number of clusters and the corresponding A set of categories, where each category represents a terminal device and the number of clusters is the number of terminals.

2. According to claim 1, a method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering, characterized in that: The formula for the standard normalization is as follows: ; in, represents the jth normalized packet arrival timestamp value, represents the j-th normalized TCP timestamp value; is the number of packets, is the arrival timestamp value of the jth packet, is the jth TCP timestamp value ; is the arrival timestamp value of the kth packet, is the kth TCP timestamp value. The subscript min indicates the minimum value of this parameter.

3. A method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering according to claim 2, characterized in that: In the step 1, all TCP flows are extracted from the data packet capture file according to five tuples, where the five tuples are source IP, destination IP, source port, destination port, and protocol groups above L4 layer.

4. The method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering according to claim 3, characterized in that: The set number is 8.

5. The method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering according to claim 4, characterized in that: Determine the slope using the linear regression function and the intercept .

6. The method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering according to claim 5, characterized in that: The unsupervised clustering is specifically a spectral clustering algorithm or a density clustering algorithm.

7. A method for identifying terminal devices in a NATed network based on TCP timestamp sequence clustering according to claim 6, characterized in that: By associating all corresponding TCP flows with each category, all TCP flows of the terminal device can be distinguished.

Citation Information

Patent Citations

  • Method and device for counting number of Linux hosts after NAT and electronic equipment

    CN117176407A

  • Counting method for system terminal equipment in NATed network based on optical network

    CN119277242A

  • Network node, destination network node, second network node and methods performed therein

    WO2024223732A1