Video streaming user platform identification
The video streaming user platform identification process and apparatus enable ISPs to accurately identify user platforms in real-time, addressing service issues and optimizing network configurations for improved video streaming experiences.
Patent Information
- Application Number
- PCT/AU2024/051267
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-05
AI Technical Summary
Broadband ISPs face challenges in determining user platforms, including device types and software agents, which hinders their ability to identify and address service issues, optimize bandwidth provisioning, and enhance customer support.
A video streaming user platform identification process and apparatus that analyzes TCP/QUIC and TLS handshake packets to generate attributes, which are then used by a trained classifier to identify user platforms and content providers in real-time, enabling ISPs to adjust network settings for improved service.
The solution allows ISPs to accurately identify user platforms with high accuracy (>96%), enabling them to optimize network configurations, improve customer support, and enhance the overall video streaming experience.
Smart Images

Figure AU2024051267_05062025_PF_FP_ABST
Abstract
Description
[0001] VIDEO STREAMING USER PLATFORM IDENTIFICATION
[0002] TECHNICAL FIELD
[0003] The present invention relates to network traffic management and video streaming, and in particular to a video streaming user platform identification process and apparatus.
[0004] BACKGROUND
[0005] Broadband Internet Service Providers (ISPs) can no longer afford to simply provide dumb data pipes, oblivious to specific characteristics of the behaviours and end user platforms of their customers, and their effects on user experience. However, for technical reasons the ability of ISPs to determine such information is extremely limited, making it difficult for them to identify or predict service issues and thus take preemptive or remedial technical actions such as appropriate bandwidth provisioning and other network configuration actions, issuing pre-emptive advisories, and other forms of customer support.
[0006] It is desired to overcome or alleviate one or more difficulties of the prior art, or to at least provide a useful alternative.
[0007] SUMMARY
[0008] In accordance with some embodiments of the present invention, there is provided a video streaming user platform identification process for processing data packets of ISP level network traffic to automatically identify in real-time user platforms and content providers of video streams of the network traffic, the process being for execution at the ISP level by at least one processor of a video streaming user platform identification apparatus of a network service provider, the process including the steps of: receiving data packets of general network traffic of a plurality of network users of the network service provider; processing the received data packets to detect TCP / QUIC and TLS handshake packets of video streams of corresponding ones of the network users; processing the detected TCP / QUIC and TLS handshake packets of video streams to generate, for each of the video streams, a corresponding plurality of values of respective predetermined attributes; for each of the video streams, processing the corresponding attributes values with a trained classifier to classify the video stream into a corresponding one of a plurality of predetermined classes, each of the predetermined classes representing a corresponding combination of a corresponding content provider of the video stream, and a corresponding user platform being used by a corresponding one of the network users to stream the video stream from the corresponding content provider; and generating statistical data representing a statistical distribution of the classes of video streams being streamed via the network service provider; wherein the user platform includes at least one of:
[0009] (i) the user's operating system; and
[0010] (ii) the user's software application being used to stream the video stream.
[0011] In some embodiments, the user platform includes the user's operating system and the user's software application being used to stream the video stream.
[0012] In some embodiments, the process further includes a step of, responsive to the statistical data, changing one or more network settings to effect one or more of the following network reconfigurations: provisioning bandwidth and network capability in accordance with bandwidth demands of the classes of video streams, prioritising traffic flows, and mapping traffic to network slices.
[0013] In some embodiments, at least some of the video streams are streamed using a QUIC transport protocol.
[0014] In some embodiments, the attributes include at least 17 attributes selected from:
[0015] (i) 19 numerical attributes;
[0016] (ii) 31 categorical attributes; and
[0017] (iii) 11 list type attributes.
[0018] In some embodiments, the attributes are generated from IP header, TCP header, QUIC header, and / or TLS handshake fields in packets of video streaming flows.
[0019] In some embodiments, the attributes include attributes generated from IP headers, TCP headers or QUIC headers include a numerical init_packet_size attribute and a numerical ttl attribute.
[0020] In some embodiments, the attributes include attributes generated from mandatory fields in TLS handshake fields, including a numerical handshakejength attribute, a categorical tls_version attribute, a cipher_suites list attribute, and a compression_methods length attribute. In some embodiments, the attributes include attributes generated from optional extensions of TLS handshake fields, including a signed_certificate_timestamp length attribute, a tls_extensions list attribute, a signature_algorithms list attribute, a categorical ec_point_formats attribute, and a supported_versions list attribute.
[0021] In some embodiments, the attributes include attributes generated from QUIC parameters of TLS handshake fields, including a quic_para meters list attribute, a numerical max_udp_payload_size attribute, a numerical max_ack_delay attribute, a grease_quic_bit presence attribute, and a categorical user_agent attribute.
[0022] In some embodiments, the process further includes a step of training the classifier, including: receiving data packets of network traffic of a plurality of known user platforms streaming video from one or more known video content providers, each of the user platforms being one of:
[0023] (i) a corresponding known operating system; and
[0024] (ii) a corresponding known software application used to stream video; and
[0025] (iii) a known combination of a corresponding known software application used to stream video and a corresponding known operating system; processing the received data packets to detect TCP / QUIC and TLS handshake packets of video streams of corresponding ones of the known user platforms streaming video from one or more known video content providers; processing the detected TCP / QUIC and TLS handshake packets of video streams to generate, for each combination of a known video content provider and a known user platform, a corresponding plurality of attribute values; processing the generated attributes values and corresponding combinations to generate training data for the classifier to enable the classifier to classify further video streams of an unknown one of the user platforms streaming video from a corresponding unknown one of the video content providers into a corresponding one of the predetermined classes.
[0026] In some embodiments, each user platform includes the corresponding known operating system and the corresponding known software application. In accordance with some embodiments of the present invention, there is provided a computer-readable storage medium having stored thereon processor-executable instructions that, when executed by at least one processor of a video streaming user platform identification apparatus of a network service provider, cause the video streaming user platform identification apparatus to execute any one of the above processes.
[0027] In accordance with some embodiments of the present invention, there is provided a video streaming user platform identification apparatus for processing data packets of ISP level network traffic to automatically identify in real-time user platforms and content providers of video streams of the network traffic, the apparatus including: random access memory ; and at least one processor configured to execute any one of the above processes.
[0028] In accordance with some embodiments of the present invention, there is provided a video streaming user platform identification apparatus for processing data packets of ISP level network traffic to automatically identify in real-time user platforms and content providers of video streams of the network traffic, the apparatus including: random access memory; at least one processor; at least one network interface to receive data packets of general network traffic of a plurality of network users of the network service provider; a streaming video flow detector to process the received data packets to identity subsets of the received packets containing video streams; a packet processing module to process the received data packets containing video streams to detect TCP / QUIC and TLS handshake packets of video streams of corresponding ones of the network users of the network service provider; a handshake attribute generator to process the detected TCP / QUIC and TLS handshake packets of the video streams to generate, for each of the video streams, a corresponding plurality of values of respective predetermined attributes; a user platform detector to process, for each of the video streams, the corresponding attributes values with a trained classifier to classify the video stream into a corresponding one of a plurality of predetermined classes, each of the predetermined classes representing a corresponding combination of a corresponding content provider of the video stream, and a corresponding user platform being used by a corresponding one of the network users to stream the video stream from the content provider; a streaming video telemetry module to generate, for each of the video streams, corresponding telemetry data representing data metrics of the video stream, and statistical data representing a statistical distribution of the classes of video streams being streamed via the network service provider; wherein the user platform includes at least one of:
[0029] (i) the user's operating system; and
[0030] (ii) the user's software application being used to stream the video stream.
[0031] In some embodiments, the user platform includes the user's operating system and the user's software application being used to stream the video stream.
[0032] In some embodiments, the apparatus further includes one or more network APIs to change, responsive to the statistical data, one or more network settings to effect one or more of the following network reconfigurations: provisioning bandwidth and network capability in accordance with bandwidth demands of the classes of video streams, prioritising traffic flows, and mapping traffic to network slices.
[0033] BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Some embodiments of the present invention are hereinafter described, by way of example only, with reference to the accompanying drawings, in which:
[0035] Figure 1 is a schematic diagram of a training network used to capture network packets representing video streams and generate corresponding PCAP files;
[0036] Figures 2 and 3 are schematic diagrams illustrating the network communications for a video streaming session;
[0037] Figure 4 is a map of the median (normalised) and number of distinct values shown as (x, y) taken by fields in handshake messages for YouTube flows over QUIC across 12 different user platforms;
[0038] Figure 5 is a map of the median (normalised) and number of distinct values shown as (x, y) taken by fields in handshake messages for YouTube flows over TCP across 14 different user platforms;
[0039] Figure 6 is a flow diagram of a video stream user platform identification process in accordance with embodiments of the present invention;
[0040] Figures 7 and 8 are respective sets of charts showing the relative importance of different attributes in classifying user platforms of YouTube flows over QUIC and TCP, respectively; Figure 9 illustrates hyper-parameter tuning of a random forest model for YouTube over QUIC;
[0041] Figures 10 to 12 illustrate the classification accuracies of the tuned random forest model to classify the user platform, operating system, and software agent, respectively;
[0042] Figure 13 is a block diagram of a video streaming user platform identification apparatus in accordance with an embodiment of the present invention;
[0043] Figure 14 is a block diagram showing video streaming user platform identification components of the video streaming identification apparatus of Figure 13;
[0044] Figure 15 is a chart of video watch times for four video content providers across different user operating systems;
[0045] Figure 16 to 19 are respective bar charts showing the video watch times broken down by operating system and software agent for respective different video content providers;
[0046] Figure 20 is a chart showing bandwidth demand for different video content providers and user operating systems;
[0047] Figure 21 to 24 are respective bar charts showing the distributions of bandwidths broken down by operating system and software agent, for respective different video content providers; and
[0048] Figures 25 to 28 are respective charts of data usage as a function of time of day for respective different video content providers.
[0049] DETAILED DESCRIPTION
[0050] Customer support is a significant burden for ISPs, since they are often the first to be blamed when users have, for example, a video freeze or grainy resolution, even if the issue is due to the user platform rather than the network. For example, recent examples of user-platform-specific streaming issues include choppy video playback on Pixel devices running Android 11 Beta in 2020, and YouTube app errors on iOS devices in late 2022. Similarly, a software update to Roku devices in 2021 caused intermittent video freezes, and Hulu has been accused of deliberately lowering resolution on PC browsers in order to force users to download their proprietary app.
[0051] In work leading up to the invention, the inventors determined that, by knowing the user platform, i.e., the device type and / or operating system (e.g., iOS / Android smartphone / tablet, Windows / Mac PC, smart TV, Xbox / PlayStation console, etc.), and the video streaming software agent / application ( / .e., native video streaming app or a specific web browser such as Chrome, Firefox, Safari or Edge) on which a household user is having a poor video streaming experience, customer support staff could rapidly identify or filter out known platform-specific issues, prioritize handling of support tickets based on platform prevalence, and issue pre-emptive advisories to users, all of which can substantially reduce the ISP's support costs. Moreover, knowledge of the platforms used by their customer base, and changes of these platforms over time, could assist an ISP to identify a need to reconfigure network configurations such as provisioning appropriate bandwidths and the like, in response to technical difficulties such as poor streaming performance, and / or pre-emptively to avoid such problems.
[0052] The inventors determined that knowledge of their customers' platforms would allow an ISP to configure their network and their streaming algorithms to meet the data needs of those customers. For example, the same video watched via a content provider's dedicated app can consume significantly higher bandwidth than watching it on a web browser, and these differences can amplify across operating systems. Given that video streaming dominates network traffic, ISP routing configuration, bandwidth provisioning, and other network configuration and management actions need to account for user device type and software agent heterogeneity, which vary widely from one content provider (e.g., YouTube) to the other (e.g., Netflix).
[0053] However, for technical reasons it is non-trivial for an ISP to determine the user platform associated with a streaming session. Much of the traffic to / from a residential home today is generally on a single IPv4 address that is shared among all household devices via network address translation (NAT). IPv6 is expected to partially overcome this issue in the long term, as each device will have a unique address, and once a device has been identified (e.g., via even one unencrypted HTTP interaction that reveals the user platform), all streams from that unique address can be pinned to the specific device. However, IPv6 deployment is immature in many ISPs globally, and may take many more years, if ever, to displace IPv4. Further, determining the software agent, e.g., native app versus browser, would require a different technique, no matter whether the traffic is IPv6 or IPv4, since client-server interactions are now predominantly carried within SSL / TLS encrypted sessions.
[0054] Knowing software agents together with operating systems ("OS"es) is also important for ISPs because superlative video streaming experience requires system / application- level support (e.g., optimized streaming algorithms) from OSes, browsers, and native apps. Prior works have investigated distinguishable patterns of TCP-based handshake fields across device firmware and application types; however, these studies have focussed only on TLS over TCP, and thus cannot be applied to video streams (e.g., YouTube) using the increasingly popular QUIC streaming protocol, which measurements in tier-1 ISPs indicate already accounts for nearly 30% of traffic in EMEA and 16% in North America. By way of background, the QUIC ("Quick UDP Internet Connections") protocol is a low (sub-second) latency transport protocol being developed by the Internet Engineering Task Force (IETF) as a standard for streaming with the HTTP / 3 protocol.
[0055] In view of the above, the inventors identified a need for ISPs to gain visibility into both user device types and software agents per video streaming session using not only all TCP-based TLS handshake fields, but also those only existing in QUIC flows.
[0056] In order to address these technical challenges, in work leading up to the invention the inventors performed extensive research to investigate whether it might be possible to identify the user platform (device type, operating system, and / or software agent) in real-time by analysing high-level characteristics of individual video streams of 30 different user platforms.
[0057] As a result of these investigations, and as described below, the inventors determined that video streaming user platforms can indeed be accurately (> 96%) identified from 61 attributes generated from video flow handshake fields during the establishment of streaming video sessions.
[0058] Based on these findings, the inventors developed the video streaming identification processes and apparatuses described herein to automatically identify the user video streaming user platforms of individual video streaming sessions from the aggregated data flows of an ISP.
[0059] Handshake Characteristics of Video Streaming Sessions
[0060] In the following, the methodology used to identify the attributes characteristic of each user platform is first described so that the reader can implement and apply the methodology to other (e.g., future) user platforms beyond the 30 user platforms described herein.
[0061] Figure 1 is a schematic diagram of the inventors' training network 100 used to collect network traffic traces (i.e., packet captures, or "PCAPs") for video streaming sessions from the four major video content providers 102, namely YouTube, Netflix, Amazon Prime Video and Disney+. The training network includes a variety of different user / client devices 104, consisting of three mobile devices (iPhone, iPad, Android phone) running iOS and Android operating systems, two Windows desktop PCs, two macOS MacBooks, two smart TVs, one with an inbuilt Android TV system and another that connects to an Android TV set-top box, and a PlayStation gaming console. Having a variety of different user device types 104 expedites data collection with multiple users streaming videos concurrently. These devices 104 connect via Wi-Fi to an access gateway 106 for access to the Internet 108, from which PCAP files 110 are collected using the Wireshark tool.
[0062] In work leading up to the invention, the inventors used the training network 100 to capture traffic traces spanning in excess of 100 video sessions across each of the 30 different user platforms, comprising various device types / OSes and software agents. Each session was composed of one or more video flows, for a total of over 10,000 flows. The duration of each session was at least one minute, sufficient to capture all the handshake messages exchanged between the corresponding user device and the server of the corresponding provider. Note that the captured packets do not contain any video content payload. The client devices 104 were running the most recent versions of operating systems, browsers and apps. Some browsers, e.g., Chrome, Firefox and Edge on Windows / Mac PCs, allow users to configure the transport layer protocol as either QUIC or TCP, which can impact the connection establishment messages exchanged. The resulting dataset encompasses all these different scenarios, and has comprehensive coverage across all different configuration options.
[0063] Table 1 below shows the number of video flows in the PCAPs 110 for each "user platform", in this instance being a combination of device type, OS and software application / agent. Some of the devices run the same OS, and thus can be considered as being the same from a user platform perspective. For example, both iPhone and iPad in the training network are using iOS, and are captured under a single (Mobile, iOS) category in Table 1. Accordingly, in this specification, although the term "user platform" usually refers to the user's OS and streaming agent / app in combination, it can also refer to only the user's OS, or the user's streaming agent / app.
[0064] The PCAPs 110 were analysed to determine the typical communication process involved in establishing a video streaming session between each combination of user platform and streaming server. Table 1: Number of video flows per content provider, captured for each combination of device type and software agent in the collected traffic traces. A user platform not supported by the content provider is marked as Anatomy of Video Streaming Sessions
[0065] Communication process
[0066] Based on this analysis, the inventors found that each streaming video session consists essentially of two sequential stages, initialization and playback. A flow diagram depicting the two stages is shown in Figure 2. In the initialization stage, the client device 202 interacts with a streaming service management server 204 specifying the service request, user configurations and connection parameters such as the type of the client device 202, the video streaming software agent 206 used, and the supported network protocols; step © in Figure 2.
[0067] Then, in step (2), the management server 204 responds with the streaming information
[0068] (e.g., URLs of the content servers and video / audio formats) along with control parameters (e.g., media player configurations) to be adapted by the software agent 206. In step (3), the actual video playback process starts after the client 202 requests video and audio data from the streaming video content server 208 located using the URL information acquired in the previous step. This request also specifies video quality metrics such as resolution that can be adjusted at any time during the playback process, either manually or dynamically by the client-side player 206, depending on the network conditions and playback quality. In step @, the video and audio data are streamed to the client device 202. The software agent 206 on it may periodically send playback status information to a management server 204, as indicated in step (5), to help the content provider keep track of service usage, session status, video quality and the like. Steps © to @ were identified in all of video streaming traces collected across the different providers, whereas step ® was only observed in certain video sessions, such as on macOS devices watching YouTube on a Chrome browser.
[0069] Anatomy of network communication
[0070] From the perspective of network communication, the above steps are carried by HTTPS flows over either TCP or QUIC as the transport layer protocol. For example, Figure 3 is a schematic diagram representing YouTube watched via the Safari browser 302 on an iPhone 304 running iOS as an example of the detailed communication process. Video sessions on other device types, software agents and content providers share a similar anatomy. The steps circled in Figure 3 correspond to those shown in Figure 2. The initial client request flow, i.e., step ©, is always sent via a single HTTPS flow to a management server 306, typically youtube.com, netflix.com, primevideo.com, or disneyplus.com for the four content providers, respectively. For YouTube, the flow is either carried by TCP or QUIC depending on the client configuration, whereas the other three providers, i.e., Netflix, Amazon and Disney+ use only TCP. In step (2), the management server 306 sends player configuration and streaming information to the client 304 on the same HTTPS flow. However, if the software agent 302 is configured to pre-load video metadata on a menu page, rather than requesting it dynamically from the server, then this information may be carried over multiple subsequent HTTPS flows.
[0071] During the video playback stages, steps (3) and @, the client 302 / 304 fetches the video from a content server 308; e.g., googlevideo.com, nflxvideo.net, avodmp4s3ww- a. akamaihd.net and media.dssott.com for the four providers respectively, using one or more HTTPS flows over either TCP or QUIC. Three scenarios are observed in this process. One, in which playback is delivered by a single HTTPS flow containing both video and audio data. Two, in which playback is delivered over multiple concurrent HTTPS flows carrying video and audio. Three, in which multiple HTTPS flows are activated in different time slots, each delivering chunks of video and audio data. For example, in some YouTube sessions, there are flows that send several chunks of video and audio data in the first few seconds (e.g., 3 seconds) and go idle, while the remaining video and audio data is streamed by another flow. Since client information and playback data are all encrypted by TLS, network operators only have visibility into the TCP / QUIC and TLS handshake messages (310 in Figure 3) and volumetric information of the payloads.
[0072] Handshake Characteristics Across User Platforms
[0073] The flows that stream video content are first initialized by a series of handshake messages. Analysis of the trace data reveals that the information in these handshake messages is highly correlated with the OS, the software agent, and the provider of the video content. A full list of the relevant fields in handshake messages is shown in Table 2 below. Table 2: Handshake fields of video flows and formalized attributes
[0074] Handshake fields and their categorization
[0075] A video flow delivered by HTTPS has two handshake processes: one for the transport layer protocol ( / .e., TCP or QUIC), and the other for TLS encryption, both of which occur prior to the encrypted video content being streamed.
[0076] Transport layer handshake
[0077] Both TCP and QUIC require a handshake process to establish connections. A TCP three- way handshake contains TCP header flags and options, such as window size, selective acknowledgment and max segment size, which are set by the user device and the software agent. Most of the TCP flags in handshake messages such as SYN, ACK, RST and FIN are ignored because they do not differ across user platforms, the exceptions being TCP CWR and ECE, which are related to congestion control policies used by the device OS or software agent. QUIC is designed to reduce the connection setup latency overhead, and so its handshake via the first flow packet (i.e., QUIC Initial packet) is integrated with the TLS handshake, as discussed below. In addition, the time-to-live field and packet size in the IP header of the initial packet are often correlated with the device type.
[0078] TLS handshake
[0079] A TLS handshake contains a ClientHello followed by ServerHello and subsequent encryption negotiations, and is executed right after the TCP three-way handshake or along with the first QUIC flow packet. The ClientHello contains customized information provided by the user device, and is highly correlated with the device type and software agent combination that plays the video. The TLS handshake fields are categorised into the following three categories:
[0080] (1) mandatory fields’, these fields always appear in the ClientHello of streaming video flows regardless of the underlying transport layer protocol (TCP or QUIC) and specification of user devices, OSes and software agents. These include handshake length, TLS version, cipher suites and compression methods.
[0081] (2) optional extensions’, these fields only appear in video flows as defined by the logic embedded in a device OS and software agent. In the dataset, a given user platform typically uses unique combinations of optional extensions, each set to a specific value. For example, Firefox browsers running on Windows and macOS PCs typically set the value of record_size_limit extension to 16385, whereas other user platforms do not use this extension.
[0082] (3) In addition to the above two categories, there are parameters in the ClientHello that are specifically available for video flows over QUIC. These parameters are contained in the collection quic_transport_para meters with extension code 57 [6, 47], which are set for specific QUIC preferences in connection establishment of certain user platforms. For example, Firefox browsers on Windows desktop PCs use the parameter grease_quic_bit to indicate its deprecation of certain flag bits in QUIC headers.
[0083] Handshake fields across user platforms
[0084] The distribution of values contained in the fields of various TCP / QUIC and TLS handshake messages are now described to demonstrate the similarities and differences in the values contained within these fields across user platforms, which form the basis of the machine learning model described below. Some fields are not numerical but categorical or lists, such as the mandatory fields tls_version and cipher_suites in TLS CHLO and supported_g roups, signature_algorithms in TLS optional extensions. To simplify the analysis, the values contained in such fields are converted to integers by a 1: 1 mapping between the values contained in the fields to a unique number. For instance, in the dataset, the field compress_certificate takes 2 values when carried over QUIC, i.e., zlib and brotli, which are uniquely mapped as 1 and 2 for the purposes of the analysis. Therefore, a video flow containing zlib as the value for compress_certificate is represented as 1 in the dataset. If a field does not appear in a flow, a value of 0 is assigned to it.
[0085] For the purpose of illustration, Figure 4 shows a map where each cell, corresponding to a user platform (indicated at the bottom of the figure), is a two tuple (x, y), where x is the median value of the field, shown via labels on the left-hand side of the figure, and y is the number of distinct values that field takes in the dataset. The map is based on YouTube over QUIC flows across 12 user platforms.
[0086] As shown in Figure 4, there are 7 fields whose median values are all the same (consistently either 0 or 1) across user platforms. These fields are highlighted with bold labels in the figure, and are: tls_version, compression_methods, server_name, ec_point_formats, ALPN, session_ticket and psk_key_exchange_modes. This means they are not useful in differentiating user platforms for YouTube video flows over QUIC. However, as shown in Figure 5, four of these fields, namely ec_point_formats, ALPN, session_ticket and psk_key_exchange_modes, take different values across user platforms for YouTube video flows over TCP, meaning they can serve as useful indicators for identifying specific user platforms.
[0087] Some fields (e.g., compression_methods) are not particularly helpful because their value is the same for both TCP and QUIC flows. The importance of these fields is discussed below.
[0088] CLASSIFYING USER PLATFORMS FOR VIDEO STREAMS IN REAL-TIME
[0089] Packet Processing Pipeline
[0090] Figure 6 is a schematic diagram of a generalized packet processing pipeline used to classify the user platform of each streaming video flow and applied to the four content providers. The pipeline takes raw packet streams 602 as input, parses them 604 and filters them 606 to identify video flows that belong to the four providers using port numbers and service names extracted from unencrypted packet headers (HDR) and ClientHello (CHLO) SNIs. These packets are further separated 608 into handshake packets 610 for classification and payload packets 612 for telemetry, i.e., to obtain session duration, volume and throughput. This forms the preprocessing stage 614.
[0091] Next, the handshake packets 610 are processed 615 to extract attributes 616 that are fed to three machine learning classifiers 620, 622, 624, which predict the user platform, device type and software agent, respectively. In the small minority of cases where the confidence in predicting the user platform is < 80%, the device type ( / .e., OS) and software agent (i.e., native app or specific browser) are predicted individually, to have high confidence in classifying either of them accurately. The predicted user platform for each video flow is then correlated 626 using flow metadata and timestamps with realtime telemetry 628 and stored 630 in association with the corresponding real-time telemetry 628 and the classification confidence value.
[0092] Attributes from Handshake Fields
[0093] For the user platforms and streaming video providers described above, 61 fields of interest are identified, of which 19 are numerical, 31 categorical, and 11 of type list. Relevant details of these attributes, applicable to both QUIC and TCP video flows, are provided in Table 2.
[0094] Creating attributes from fields contained in handshake messages
[0095] For the purpose of feeding the fields of interest into the machine learning models, their non-numerical values (if any) are first converted to numerical attributes. Some fields are inherently numerical, such as handshakejength and extensionsjength, for which no transformation is needed.
[0096] Fields that are categorical take values from a finite category of elements. For example, compress_certificate can take one of zlib and brotli, so a unique positive integer is assigned to them, i.e., 1 and 2 respectively, and thus this field is transformed to a numerical attribute. 0 is assigned to categorical fields that are not present in a flow.
[0097] Fields that are lists can in turn include many categorical fields. For example, the field cipher_suites contains multiple cipher suites from a finite list supported by a client. The order of the categorical items in the field indicates the client's preference. Therefore, to preserve the information provided by the choice of items and its order in a list-type field, a fixed-length vector is used to indicate the placement of each item, with zeropadding for non-existent items. This simplified representation is readily processed by the classifiers. whether they are present or not differs across device types and software agents. Accordingly, a value of 1 is assigned to represent their presence in a video flow, and 0 otherwise. Also, 7 fields, such as initial_source_connction_id in QUIC, contain values that are not useful because they are randomly chosen, but the length of the values in bytes can be helpful, and hence they are treated as length-based attributes, as shown in Table 2.
[0098] Different computational costs are incurred by converting each field to its corresponding attribute (by the handshake attribute generator 615 shown in Figure 6). The numerical fields are directly taken as attributes, and thus can be treated as low cost <r(i). The categorical fields require medium costs (m) where m denotes the number of unique categorical items. The most computationally expensive process, i.e., high cost <r(mn) is for transforming a list type field, where there are n items in the list to be selected from m choices. As shown below, computationally expensive attributes do not necessarily increase the overall predictive capability.
[0099] Importance of attributes
[0100] As described above, the attributes derived 616 from the handshake fields are not equally important in predicting user platforms. Accordingly, the importance of each attribute is systematically benchmarked using the information gain metric described in J. R. Quinlan, 1986. Induction of decision trees, in Machine Learning 1, 1 (March 1986), 81- 106 (https: / / doi.org / 10.1007 / BF00116251). The information gain of an attribute is computed as the difference in the entropy of the ground-truth labels before and after the dataset is sorted by that attribute. Attributes of the highest importance have information gain close to 1, since the predicted labels are well organized after being sorted by the attribute, while an irrelevant attribute has information gain 0. In the described embodiments, the information gains are computed for both TCP and QUIC video flows for each of the four video providers, with prediction objectives being user platform, device type, and software agent. The relative importance of the attributes is described below, using YouTube video flows over QUIC as a representative example.
[0101] Each of Figures 7 and 8 includes three charts showing the relative importance, i.e., normalized information gain, of each attribute for three prediction objectives (user platform, device type / OS, and software agent) and their computational costs for YouTube over QUIC flows (Figure 7) and YouTube over TCP flows (Figure 8). Attributes are denoted by their ordered labels, and a full mapping between labels and attribute names is provided in Table 2.
[0102] To elicit the relative importance of attributes in satisfying the prediction objectives, two thresholds, 0.2 and 0.1, are defined, and the relative importance of each attribute is rated as high, medium or low if its information gain value is > 0.2, between 0.1 and 0.2 or < 0.1, respectively. As shown in Figure 7, for YouTube over QUIC flows, 17 attributes including tl, ml, m3, ol, o3, o4, o&, 08, ol2, ol8, o21, q2, q4, q5, ql2, ql7 and 20 have high relative importance for all three prediction objectives, i.e., user platform, only device type, and only software agent. Eleven attributes, i.e., m2, m4, o2, o5, o7, ol0, ol5, ol7, ol9, o20 and l8, have an information gain < 0.1 for all three prediction objectives, implying that their effectiveness is limited.
[0103] Other attributes have at least one high or medium information gain value, and one medium or low information gain value. For instance, t2 has a normalized importance score of 1 (high) for device type, but only 0.18 (medium) for software agent. This is not surprising since certain handshake fields (e.g., time-to-live) are highly dependent on device type, while others (e.g., version information) exhibit strong variations across software agents, as described above. In addition, an attribute with low importance for QUIC video flows (Figure 7) may have medium or high importance for TCP flows (Figure 8). An example is 015, which has a value near 0 for QUIC, but over 0.1 for TCP. These observations indicate that a select combination of attributes with medium to high importance can also serve as useful predictors.
[0104] Finally, even though certain attributes have low computational costs for encoding, such as the numerical and categorical fields in handshake messages, relying solely on them does not necessarily compromise prediction accuracy. For instance, out of 42 low-cost attributes, three of them (tl, t2, and 08) have high importance for both QUIC and TCP video flows. On the other hand, four out of nine medium-cost and one out of ten high- cost attributes have low importance. In production environments, where a switch typically carries hundreds of Gbps of traffic aggregated across several thousand households, ISPs may not have the necessary computational resources to execute the entire packet processing and classification pipeline in real-time. Under these circumstances, an ISP can carefully select only those low- and / or medium-cost attributes to achieve an acceptable degree of prediction accuracy. This trade-off between the cost of generating attributes and the accuracy of predicting user platforms is discussed below. Machine Learning Classification Models
[0105] The described embodiments of the present invention use machine learning models to predict user platforms for the four video content providers described above, although it will be apparent to those skilled in the art in light of this disclosure that the described models can readily be adapted to predict user platforms for other video content providers, following the methodology described herein. At the time of writing, only YouTube supports QUIC, the other providers supporting only TCP; however, this is expected to change in future. By following the methodology described above, the models can also be updated to account for changes in the traffic characteristics of any of the four providers (and / or to include new or other content providers not described herein).
[0106] Model training, tuning and selection
[0107] For the purpose of comparison, classifiers from three popular machine learning algorithms, namely random forest (decision tree based), MLP (neural network), and KNN (clustering based) were applied, and their hyper-parameters tuned accordingly. For random forest, these parameters include maximum tree depth, number of trees, and number of attributes. The MLP model is tuned for the number of hidden layers, the number of perceptrons per layer and the activation functions. KNN classifiers are tuned for the number of neighbours, weight functions and leaf size.
[0108] Each of these algorithms was trained on the dataset collected from the training network 100 shown in Figure 1. As described above, the dataset consists of attributes extracted from traces spanning over 10,000 video flows across 30 different user platforms. The performance of each of the models was evaluated in terms of the overall accuracy using 10-fold validation, which determined that the random forest model outperforms both the MLP and the KNN classifiers, not only for YouTube, but for all other providers and user platforms as well.
[0109] For example, in classifying the user platform for YouTube flows over QUIC, the random forest model achieves an overall accuracy of 96.4%, whereas the MLP and KNN models achieved accuracies of only 65.1% and 69.1%, respectively, confirming that decision tree based models are better suited for network traffic classification problems. Based on these results, the described embodiments of the video streaming user platform identification process and apparatus use a random forest model. However, it will be appreciated that other classifier types may be used in other embodiments. The following illustrates how to tune the random forest model for best performance. Out of the 61 attributes overall, only 47 are applicable to QUIC. Figure 9 is a table of the overall classification accuracy as a function of the number of attributes (vertical axis) and the maximum tree depth (horizontal axis). The highest accuracy of 96.4% was attained when these two hyperparameters were set to 34 and 20, respectively, and consequently these two hyperparameters were used to configure the random forest model to provide the best classification performance with the YouTube / QUIC dataset.
[0110] With the random forest model thus configured, Figure 10 shows the classification accuracy of the model per prediction class (as a confusion matrix). Of the 12 prediction classes, 5 classes are identified with 100% accuracy, including all browser types on Windows PC, and Chrome and the native YouTube application on Android phone. Misclassified instances are observed within two groups, namely iOS and macOS devices. For example, native YouTube app on iOS has a small chance (< 4%) of being misclassified as a native app on Android. For the iOS native app instances, their device types (i.e., iOS) are classified with 96% accuracy, whereas their software agents (i.e., native app) might be misclassified as Chrome or Safari with a slim chance of less than 6%.
[0111] Delving deeper into the fields for (iOS, Safari) and (iOS, Chrome) platforms (see columns 10 and 11 in Figure 4), a vast majority of them take on similar values. Only a small number of attributes, e.g., handshakejength, extensionsjength, have different values. Moreover, Chrome has started to randomize TLS extension orders since version 110. These variations in attribute values explain the small fraction of misclassifications seen for these user platforms.
[0112] It is also observed that those misclassified instances are with low confidence, i.e., less than 50%, while the correctly classified ones are with high confidence, i.e., over 80%. The performance of the random forest models to classify only the device type and software agent for YouTube QUIC are depicted in Figures 11 and 12, respectively. It is noted that their accuracy in identifying the device type is high, > 97%, for all device types, and it can predict all software agents with > 91% accuracy. The marginal decline in the accuracy of the latter is for the reasons described above. Nonetheless, the overall results demonstrate that the random forest classifiers offer high prediction accuracy with high confidence. Open-set evaluation
[0113] Relying solely on the accuracy reported by 10-fold validation can result in over-fitting, which can mask the true accuracy of a classifier. To overcome this problem, the performance of the random forest model was further evaluated on a dataset collected by one of the inventors from their home network. While the devices in the home are the same as those in the training network of Figure 1, the OS versions as well as that of the software agents are different. These variations could impact the values of the different attributes. The aim of this exercise therefore is to validate the accuracy of the model, trained on the lab trace data, in predicting the user platforms seen in a different environment, i.e., the home.
[0114] This dataset contains over 2000 video flows spread evenly across all user platforms. As shown in Table 3 below, the results are comparable to the ones reported above, confirming the high accuracy of the classifier. User platforms are classified with > 94% accuracy for YouTube, > 96% for Netflix and Disney+ and > 98% for Amazon Prime Video.
[0115] Table 3: Model performance in open-set evaluation for three classification objectives, namely user platform (in this instance being a combination of device type (i.e., OS) and software agent), device type only, and software agent only.
[0116] Models with a subset of attributes
[0117] As described above, not all attributes are equally important from the perspective of information gain, especially those requiring high costs for encoding. An ISP carrying very high data rates, e.g., multiple hundreds of Gbps, may require servers with significant computational capabilities to deploy the end-to-end classification pipeline in real-time. In the absence of such resources, one can choose to omit the high-cost low- importance attributes to reduce the processing load with negligible impact on classification accuracy, as described below.
[0118] For the purposes of illustration, random forest models were trained with three subsets of attributes. Each subset excluded attributes that are deemed to be of low importance (i.e., < 0.1 information gain). Additionally, the first subset excludes attributes that require high encoding costs, the second excludes those with high or medium encoding costs, and the third excludes attributes with high, medium, or low encoding costs.
[0119] The overall accuracy of the random forest classifier for YouTube QUIC flows is shown in Table 4 below. Compared with the accuracy of the model that uses the full (i.e., 47) attribute set, the reduced models result in a slight reduction (« 3%) of accuracy across the three scenarios. For example, the accuracies for classifying user platforms of QUIC YouTube flows is 96.4% with the full attribute set, dropping to 93.3%, 93.0% and 92.8%, with the respective subsets. The performance is similar when predicting only the device type or software agent. These results demonstrate that if an ISP is faced with limited computational resources, then they can safely ignore attributes with low information gain requiring varying degrees of encoding costs, to deploy the end-to-end user platform classification pipeline in real-time.
[0120] Table 4: Accuracies of models for YouTube QUIC video flows with three subsets of attributes, excluding low-importance attributes associated with high cost, high or medium cost, and high, medium or low cost, respectively. Benchmarking against the state-of-the-art
[0121] Finally, the user platform identification method was benchmarked against six state-of- the-art techniques (referenced as techniques A to F) for user platform identification from the published literature of the last five years, using the ground-truth dataset described above. Three qualitative aspects, as specified in the second to fourth columns of Table 6, including inference objective, covered protocol and inference granularity are compared to demonstrate the superior visibility our method can provide. Our method outperforms all alternatives.
[0122] Table 6: Benchmarking the classification accuracy of the user platform identification method against state-of-the-art methods (after necessary methodological adaptations).
[0123] Two of the six techniques, described in references D (Fatemeh Marzani, Fatemeh Ghassemi, Zeynab Sabahi-Kaviani, Thijs Van Ede, and Maarten Van Steen. 2023. Mobile App Fingerprinting through Automata Learning and Machine Learning, in Proc. IFIP Networking, Barcelona, Spain, 1-9) and F (Omar Richardson and Johan Garcia. 2020, A Novel Flow-level Session Descriptor With Application to OS and Browser Identification, in Proc. IEEE / IFIP NOMS, Budapest, Hungary, 1-9), require collecting statistics of all flows from a candidate host, and thus cannot be used to identify user platforms of individual video flows from clients behind NAT. The other four techniques, described in references A (Blake Anderson and David McGrew 2019, TLS Beyond the Browser: Combining End Host and Network Data to Understand Application Behavior, in Proc. ACM IMC, Amsterdam, Netherlands, 379-392), B (Xinlei Fan, Gaopeng Gou, Cuicui Kang, Junzheng Shi, and Gang Xiong, 2019, Identify OS from Encrypted Traffic with TCP / IP Stack Fingerprinting, in Proc. IEEE International Performance Computing and Communications Conference, London, United Kingdom, 1- 7), C (Martin Lastovicka, Stanislav Spacek, Petr Velan, and Pavel Celeda. 2020, Using TLS Fingerprints for OS Identification in Encrypted Traffic, in Proc. IEEE / IFIP NOMS, Budapest, Hungary, 1-6), and E (Qiuning Ren, Chao Yang, and Jianfeng Ma. 2021, App Identification Based on Encrypted Multi-smartphone Sources Traffic Fingerprints, in Computer Networks 201 (Dec. 2021), 108590), either directly offer flow-level granularity, or have a subset of articulated attributes from individual flows, and thus can be adapted for flow-level identification as specified in the fifth column of Table 6.
[0124] First of all, since all six techniques are designed only for TLS over TCP flows, a generic adaptation has to be made for all of these techniques to handle video flows over QUIC, including identifying and decrypting QUIC Initial packets and extracting handshake attributes from TLS CHLO messages over QUIC. In addition, to make the four techniques applicable for user platform identification, the required adaptations specific to each technique are listed in the fifth column of Table 6 such as extracting fine-grained flowlevel telemetry, constructing features from collected statistics, expanding inference objective, developing classification pipeline and models. For example, the technique described in reference A combines TLS handshake fields to generate string-based fingerprints of applications. Therefore, for benchmarking purposes, this technique was adapted by constructing usable features from their fingerprint strings and developing a classification process. Also, the techniques in references B and C classify device types with IP-level attributes. Thus, their attributes were adapted to be extracted for individual video flows and used to classify not only device types but also software agents.
[0125] The quantitative accuracy of the present method in comparison to the other techniques (after necessary adaptations) is reported in the last five columns of Table 6. It is clearly apparent that the present method outperforms all of the other techniques in all five classification scenarios, i.e., different providers and their supported protocols. It is noted that the three techniques (references A, B, and C) that can achieve overall 80% accuracy require significant adaptations from constructing flow-level telemetry to articulating attributes and developing classification models, whereas the technique in reference E uses attributes extracted from flow metadata (e.g., length) and only one TLS field "TLS_message_type" which becomes unavailable / encrypted in QUIC, and thus has only 11.3% accuracy for YouTube flows over QUIC and less than 60% for other scenarios.
[0126] Video streaming user platform identification Process and Apparatus
[0127] Based on the research and insights described above, the inventors have developed a video streaming user platform identification process and apparatus that are able to process general network traffic at the ISP level to identify, in real-time and with high accuracy, individual video streaming sessions within that general network traffic, and for each of those video streaming sessions, the corresponding user platform (including the user's OS and software agent in combination).
[0128] In the described embodiments, the video streaming user platform identification apparatus is a blade server 1300 configured with an 8-core Intel Xeon E5-2620 CPU 1302, as shown in Figure 13, and the video streaming user platform identification processes executed by the video streaming user platform identification apparatus are implemented as executable instructions of one or more software modules 1400, as shown in Figure 14, stored on non-volatile (e.g., hard disk or solid-state drive) storage 1304 associated with the computer system. However, it will be apparent to those skilled in the art that in other embodiments at least parts of the video streaming user platform identification process can alternatively be implemented in other forms, such as configuration data of one or more field programmable gate arrays (FPGAs), or as one or more dedicated hardware components, such as application-specific integrated circuits (ASICs), or as any combination of these various forms.
[0129] In the described embodiments, the video streaming user platform identification apparatus 900 includes 64GB of DDR4 random access memory (RAM) 1306, a bus 1308, and two 10 Gbps network interface connectors (NIC) 1310, which connect the video streaming user platform identification apparatus 1300 to a communications network such as the Internet 1312.
[0130] The video streaming user platform identification apparatus 900 also includes a number of standard software modules 1314 to 1318, including an operating system 1314 such as Linux or Microsoft Windows, and structured query language (SQL) support 1316 such as Postg reSQL, which allows data to be stored in and retrieved from an SQL database 1318. In the described embodiments, the software modules 1400 of the apparatus 1300, as shown in Figure 14, include a packet processing module 1402 written in Golang, built over the open-source Virtual Network Functions Framework NFF-GO 1404, available from h tips : Z / q ith u b . com / a req m / nff -go . In turn, the NFF-GO framework utilises the DPDK (Data Plane Development Kit) 1406, available from https : ZZ www , d k. Q r Z , as its underlying packet processing functions.
[0131] The packet processing module 1402 implements the pipeline shown in Figure 6 as a virtual network function (VNF) system. The DPDK 1406 and NFF-Go 1404 packet processing frameworks are used for the packet pre-processing stage 614 in Figure 6, which parses input packet streams 602 and handles the handshake and data packets of video flows.
[0132] The packet processing module 1402 maintains states for packets and flows that are candidates for video streaming services, based on their Server Name Identifications (SNI) or flow 5-tuple metadata. The packet processing module 1402 also maintains the states of detected video streaming sessions for their active TLS flows over either TCP or QUIC protocols, domain names identified by Server Name Indications (SNI), packet counts, byte counts, and durations.
[0133] A streaming video flow detector 1408 is implemented on top of the stateful packet processing module 1402, and uses the Server Name Indications or flow 5-tuple metadata to detect active streaming video flows and their service domain names.
[0134] A streaming video packet labeler 1410 is implemented on top of the streaming video flow detector 1108 and packet processing module 1402. The labeler 1410 labels packets belonging to an isolated video flows extracted by the module 1402 as either handshake packets or video data packets according to their TLS header metadata.
[0135] A handshake attribute generator 1412 is implemented (in Golang) on top of the packet processing module 1402. The attribute generator 1412 produces formalized attributes from handshake fields in handshake packets of streaming video flows maintained in module 1402.
[0136] A user platform identifier 1414 is implemented on top of the packet processing module 1402 and handshake attribute generator 1412. The user platform identifier 1414 classifies user platform (including device type and software agent) of video streaming flows maintained in module 1402 using the handshake attributes produced by the handshake attribute generator 1412. The classifier banks of the user platform identifier 1414 are based on the Python scikit-learn library described in F. Pedregosa et. al., Scikit-learn: Machine Learning in Python, Journal of Machine Learning Research 12 (2011) 2825-2830.
[0137] A streaming video telemetry module 1416 is implemented (in Golang) on top of the packet processing module 1402, and uses video data packets from the tracked video flows with their identified user platform types to generate real-time telemetry such as bandwidth and packet rate of each video sessions per client IP address. The video session telemetry along with user platform labels are stored in the SQL database 1318.
[0138] Finally, a Network APIs module 1418 is implemented on top of the streaming video telemetry module 1416 for changing network settings to improve streaming network performance of video sessions or client IP addresses and user quality-of-experience, based on the classifications.
[0139] EXAMPLE
[0140] The demonstrate the capabilities of the video streaming user platform identification apparatus 1300, it was deployed in the inventors' university campus network to classify network traffic between the university campus network and the Internet in real-time. As the apparatus blade server has sufficient computational resources to process realtime traffic streams at 20 Gbps peak rate, the classifiers were trained with the full attribute set described above to achieve high accuracy.
[0141] The apparatus 1300 received a copy of the traffic from the university network border router that is connected to the Internet 1312. The traffic is delivered to the two 10 Gbps network interfaces 1310 of the apparatus 1300 for inbound and outbound traffic, respectively.
[0142] As an additional sanity check of the classification methodology, over 1000 video sessions were played using all available user platforms in the inventors' university laboratory, and were captured by the apparatus 1300. The classification accuracy of those groundtruth video sessions was similar to the open-set evaluation described above. Almost all misclassified instances (less than 4% of streams) were with relatively low confidence, i.e., < 50%; therefore all such instances are ignored. The university campus network serves staff, students, visitors, and over 10 residential dormitories. Over the two months from July 7th 2023, 00:00 to September 3rd 2023, 23:59, the apparatus 1300 collected telemetry statistics, i.e., duration, volume and throughput, for every video flow from the four content providers, and tagged each of them by their user platforms. Over 50 million streaming video flows were collected, representing a total of 200,000 hours of watch time. About 20% of the sessions had low classification confidence, and are excluded from the following analysis.
[0143] Video watch time across user platforms
[0144] Figure 15 shows how the level of engagement, measured in terms of total watch time, varied by device type across the four video content providers. Not surprisingly, across the entire campus demographic, YouTube, where content is mostly free, dominated engagement, with an average daily total watch time of 2000 hours, followed by subscription-based providers such as Netflix, Disney+, and Amazon Prime Video. Furthermore, the majority of subscription-based videos were watched on PCs (Windows / Mac), rather than on mobile devices. In contrast, up to 40% of YouTube engagement occurred on mobile devices (iOS / Android). Such insights enable ISPs to prioritize the troubleshooting of issues, as they can expect more support calls relating to PCs than mobiles for Netflix, and the converse for YouTube.
[0145] Figures 16 to 19 are charts quantifying the breakdown of software agents for the video content providers YouTube, Netflix, Disney+, and Amazon Prime Video, respectively. The Chrome browser on Windows PCs was the most popular software agent used to watch YouTube, clocking up 677 hours, as shown in Figure 16. Amongst mobiles, iOS was preferred, with over 90% of watch time on its YouTube native app. The other devices used a relatively diverse set of software agents, as shown.
[0146] The watch time profiles for Netflix, Disney+ and Amazon Prime are shown in Figures 17 to 19, respectively. While Safari on Mac PCs was popular for viewing Netflix and Amazon, the native Disney+ app on iOS dominated engagement of mobile users by over 90%.
[0147] This detailed visibility into engagement patterns helps ISPs to enhance customer satisfaction via rapid troubleshooting and root cause analysis. Moreover, it enables better segmentation of their customer base, allowing them to craft attractive up- sell / cross-sell offers such as innovative video streaming bundles, accessories for different device types, device-specific speed boosting plans, and so on. Bandwidth demand
[0148] Figures 21 to 24 are charts showing the impact of software agents on bandwidth for the video providers Amazon prime, Disney+, Netflix, and YouTube, respectively. Android mobiles, iOS mobiles and TVs only support native Amazon Prime Video apps and, as shown in Figures 21, consume less bandwidth (median < 3 Mbps) than their PC counterparts. All browsers on Windows / Mac PCs exhibit higher median bandwidth and interquartile range spread for Amazon compared to native mobile apps. In addition, Mac PCs generally require higher median throughput than Windows PCs. These insights can help ISPs to better provision capacity as shows / movies on Amazon rise in popularity.
[0149] Interestingly, with the exception of the Safari browser, Netflix streamed to PCs on browsers consumes lower median bandwidth (< 2 Mbps), which suggests that lower resolution is supported via browsers than via the native app. Such information can be used to improve ISPs' network configuration strategies, such as configuration of routing paths, provisioning bandwidth, and mapping of users or video streaming flows to network slices (with high bandwidth and low latency) based on the network demands required by specific user platforms through automatic logics with network APIs, which are currently largely agnostic to user platforms.
[0150] Capacity planning is a key ongoing activity for ISPs as they strive to meet the ever- increasing demands for quality service assurance from their customers. A key aspect to this is bandwidth management. Given the significant load imposed by video streaming, fine-grained insights into content watched by users, broken down by device types and software agents, give ISPs valuable information about bandwidth demand. This can be used to optimize capacity planning decisions such as right-sizing regions to reduce over- or under-provisioning, create custom plans that allow prioritization of video streams on specific user platforms, improve caching mechanisms to reduce associated transit costs, etc.
[0151] Figure 20 shows the distribution of bandwidth consumption, as box plots indicating the median and the quartiles to either side of the median, for the four video streaming providers across different device types. It is apparent that the bandwidth demand imposed by subscription-based videos is higher than that of YouTube, as the interquartile range for these providers is 3 to 9 Mbps higher. Notably, videos streamed from Amazon Prime Video to Mac PCs demand the highest median bandwidth (of 5.7 Mbps), which is 50% higher compared to smart TVs. Having insights into how bandwidth demand varies across device types and providers allows an ISP to better plan the provisioning and management of network resources, such as by fine-tuning the bandwidth on the links between the edge and CDNs so users continue to enjoy a good streaming experience.
[0152] Temporal usage patterns
[0153] Temporal usage patterns of video streaming services can provide valuable information to ISPs. For example, knowing when peak usage occurs and for what kind of content, ISPs can proactively allocate adequate bandwidth to ensure high levels of service assurance. Conversely, during off-peak periods, bandwidth can be allocated differently to save costs. Alternatively, traffic management policies can be implemented that prioritize video streaming during peak hours, and other types of traffic during off-peak hours.
[0154] Figures 25 to 28 show the median traffic volume during each hour of the day in the 2- month deployment consumed by Amazon, Disney+, Netflix and YouTube videos, respectively, on PCs and mobile devices. Overall, Amazon and Disney+ exhibit fairly similar daily usage patterns with a 4-hour peak period from about 7 pm to 11 pm. Also, mobile usage for Amazon is low compared to Disney+.
[0155] Comparing YouTube and Netflix, it is apparent that the former has a long and sustained peak window from about 4 pm to midnight, whereas the latter has a shorter peak, between 8 pm to 10 pm. In terms of mobile usage, YouTube dominates with relatively steady hourly peak usage in the range 17 to 20 GB from about 4 pm to midnight.
[0156] These trends enable ISPs to anticipate the impact of popular content releases (e.g., One Piece on Netflix) or live sports on YouTube. Moreover, it assists in enforcing appropriate management policies and allocating the necessary network resources, especially during peak periods, to mitigate any degradation in QoE when streaming highly anticipated content.
[0157] As described above, the video streaming user platform identification processes and apparatuses described herein are able to automatically and accurately identify the user video streaming user platforms of individual video streaming sessions in real-time from the aggregated data flows of an ISP. Given the significant load imposed by video streaming, fine-grained insights into the bandwidth demand of popular streaming service users, broken down by device types and software agents, will help ISPs improve the fidelity of their bandwidth forecasting models. It also allows network operators to better understand their customer segments, provision bandwidth, and troubleshoot video streaming issues pertinent to device firmware, OS, or software for customer experience and satisfaction. In addition, temporal usage patterns of video streaming services can also provide valuable information to ISPs. For example, by knowing when peak usage occurs and for what kind of content, ISPs can proactively allocate adequate bandwidth to ensure high levels of service assurance. Conversely, during off-peak periods, bandwidth can be allocated differently to save resources. Alternatively, traffic management policies can be implemented that prioritize video streaming during peak hours and other types of traffic during off-peak hours.
[0158] This fine-grained visibility into user platforms of streaming video flows over both TCP and QUIC, allows network operators to automatically reconfigure their networks to satisfy the different network demands required by video streaming sessions with different user platforms, such as by mapping video streaming sessions from a certain identified user platform to network slices, and provisioning sufficient bandwidth and configuring high-priority routing paths for user platforms with high network demands using network APIs through automatic network configuration logics.
[0159] Many modifications will be apparent to those skilled in the art without departing from the scope of the present invention.
Claims
CLAIMS:
1. A video streaming user platform identification process for processing data packets of ISP level network traffic to automatically identify in real-time user platforms and content providers of video streams of the network traffic, the process being for execution at the ISP level by at least one processor of a video streaming user platform identification apparatus of a network service provider, the process including the steps of: receiving data packets of general network traffic of a plurality of network users of the network service provider; processing the received data packets to detect TCP / QUIC and TLS handshake packets of video streams of corresponding ones of the network users; processing the detected TCP / QUIC and TLS handshake packets of video streams to generate, for each of the video streams, a corresponding plurality of values of respective predetermined attributes; for each of the video streams, processing the corresponding attributes values with a trained classifier to classify the video stream into a corresponding one of a plurality of predetermined classes, each of the predetermined classes representing a corresponding combination of a corresponding content provider of the video stream, and a corresponding user platform being used by a corresponding one of the network users to stream the video stream from the corresponding content provider; and generating statistical data representing a statistical distribution of the classes of video streams being streamed via the network service provider; wherein the user platform includes at least one of:(i) the user's operating system; and(ii) the user's software application being used to stream the video stream.
2. The video streaming user platform identification process of claim 1, wherein the user platform includes the user's operating system and the user's software application being used to stream the video stream.
3. The video streaming user platform identification process of claim 1 or 2, further including a step of, responsive to the statistical data, changing one or more network settings to effect one or more of the following network reconfigurations: provisioning bandwidth and network capability in accordance with bandwidth demands of the classes of video streams, prioritising traffic flows, and mapping traffic to network slices.
4. The video streaming user platform identification process of any one of claims 1 to 3, wherein at least some of the video streams are streamed using a QUIC transport protocol.
5. The video streaming user platform identification process of any one of claims 1 to 4, wherein the attributes include at least 17 attributes selected from:(i) 19 numerical attributes;(ii) 31 categorical attributes; and(iii) 11 list type attributes.
6. The video streaming user platform identification process of any one of claims 1 to 5, wherein the attributes are generated from IP header, TCP header, QUIC header, and / or TLS handshake fields in packets of video streaming flows.
7. The video streaming user platform identification process of any one of claims 1 to 6, wherein the attributes include attributes generated from IP headers, TCP headers or QUIC headers include a numerical init_packet_size attribute and a numerical ttl attribute.
8. The video streaming user platform identification process of any one of claims 1 to 7, wherein the attributes include attributes generated from mandatory fields in TLS handshake fields, including a numerical handshakejength attribute, a categorical tls_version attribute, a cipher_suites list attribute, and a compression_methods length attribute.
9. The video streaming user platform identification process of any one of claims 1 to 8, wherein the attributes include attributes generated from optional extensions of TLS handshake fields, including a signed_certificate_timestamp length attribute, a tls_extensions list attribute, a signature_algorithms list attribute, a categorical ec_point_formats attribute, and a supported_versions list attribute.
10. The video streaming user platform identification process of any one of claims 1 to 9, wherein the attributes include attributes generated from QUIC parameters of TLS handshake fields, including a quic_parameters list attribute, a numerical max_udp_payload_size attribute, a numerical max_ack_delay attribute, a grease_quic_bit presence attribute, and a categorical user_agent attribute.
11. The video streaming user platform identification process of any one of claims 1 to 10, further including a step of training the classifier, including: receiving data packets of network traffic of a plurality of known user platforms streaming video from one or more known video content providers, each of the user platforms being one of:(i) a corresponding known operating system; and(ii) a corresponding known software application used to stream video; and(iii) a known combination of a corresponding known software application used to stream video and a corresponding known operating system; processing the received data packets to detect TCP / QUIC and TLS handshake packets of video streams of corresponding ones of the known user platforms streaming video from one or more known video content providers; processing the detected TCP / QUIC and TLS handshake packets of video streams to generate, for each combination of a known video content provider and a known user platform, a corresponding plurality of attribute values; processing the generated attributes values and corresponding combinations to generate training data for the classifier to enable the classifier to classify further video streams of an unknown one of the user platforms streaming video from a corresponding unknown one of the video content providers into a corresponding one of the predetermined classes.
12. The video streaming user platform identification process of claim 11, wherein each user platform includes the known operating system and the known software application used to stream video.
13. The video streaming user platform identification process of any one of claims 1 to 12, wherein the user platform includes only the user's operating system.
14. The video streaming user platform identification process of any one of claims 1 to 12, wherein the user platform includes only the user's software application being used to stream the video stream.
15. The video streaming user platform identification process of any one of claims 1 to 14, wherein the user's software application being used to stream the video stream is one of: a dedicated video streaming software application of the corresponding content provider, and one of a plurality of web browser applications.
16. A computer-readable storage medium having stored thereon processorexecutable instructions that, when executed by at least one processor of a video streaming user platform identification apparatus of a network service provider, cause the video streaming user platform identification apparatus to execute the process of any one of claims 1 to 15.
17. A video streaming user platform identification apparatus for processing data packets of ISP level network traffic to automatically identify in real-time user platforms and content providers of video streams of the network traffic, the apparatus including: random access memory; and at least one processor configured to execute the process of any one of claims 1 to 15.
18. A video streaming user platform identification apparatus for processing data packets of ISP level network traffic to automatically identify in real-time user platforms and content providers of video streams of the network traffic, the apparatus including: random access memory; at least one processor; at least one network interface to receive data packets of general network traffic of a plurality of network users of the network service provider; a streaming video flow detector to process the received data packets to identity subsets of the received packets containing video streams; a packet processing module to process the received data packets containing video streams to detect TCP / QUIC and TLS handshake packets of video streams of corresponding ones of the network users of the network service provider; a handshake attribute generator to process the detected TCP / QUIC and TLS handshake packets of the video streams to generate, for each of the video streams, a corresponding plurality of values of respective predetermined attributes; a user platform detector to process, for each of the video streams, the corresponding attributes values with a trained classifier to classify the video stream into a corresponding one of a plurality of predetermined classes, each of the predeterminedclasses representing a corresponding combination of a corresponding content provider of the video stream, and a corresponding user platform being used by a corresponding one of the network users to stream the video stream from the content provider; a streaming video telemetry module to generate, for each of the video streams, corresponding telemetry data representing data metrics of the video stream, and statistical data representing a statistical distribution of the classes of video streams being streamed via the network service provider; wherein the user platform includes at least one of:(iii) the user's operating system; and(iv) the user's software application being used to stream the video stream.
19. The video streaming user platform identification apparatus of claim 18, wherein the user platform includes the user's operating system and the user's software application being used to stream the video stream.
20. The video streaming user platform identification apparatus of claim 18 or 19, further including one or more network APIs to change, responsive to the statistical data, one or more network settings to effect one or more of the following network reconfigurations: provisioning bandwidth and network capability in accordance with bandwidth demands of the classes of video streams, prioritising traffic flows, and mapping traffic to network slices.
Citation Information
Patent Citations
Encrypted traffic classification method and related equipment
CN113783795A
Encrypted traffic identification and classification method based on deep learning model
CN115378701A
Method to Identify Video Applications from Encrypted Over-the-top (OTT) Data
US20190279113A1
Method and Apparatus for Inferring ABR Video Streaming Behavior from encrypted traffic
US20200153805A1
Method and Apparatus for Identifying Encrypted Data Stream
US20200280584A1
Cited By
Data hierarchical compression storage method, energy storage management system and device and storage medium
CN121217147A