Counteracting mac address randomization and spoofing attempts and identifying wi-fi devices based on user behavior
Patent Information
- Application Number
- JP2022178290
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-28
- Filing Date
- 2022-11-07
- Publication Date
- 2025-10-14
AI Technical Summary
Existing technologies face challenges in identifying and tracking devices in Wi-Fi networks due to MAC address randomization and spoofing, which disrupts network-based solutions and inhibits legitimate functionality such as parental control and user tracking.
A system and method to associate disparate device identifiers, such as MAC addresses, by analyzing operating parameters and network metadata to determine when different identifiers represent the same physical device, using machine learning and behavioral analysis to create continuity in device identification.
Enables accurate tracking and identification of devices despite MAC address changes, maintaining consistent network management and user tracking capabilities.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to networking systems and methods. More specifically, this disclosure relates to countering attempts to randomize or spoof device identifiers such as Media Access Control (MAC) addresses in a networking environment and identifying devices within a Wi-Fi network based on user behavior, such as by stitching MAC addresses.
Background Art
[0002] Wi-Fi networks are deployed to provide network connectivity to various devices (e.g., mobile devices, smartphones, computers, laptops, tablets, TVs, smart TVs, Internet of Things (IoT) devices, media players, etc.). A device's Media Access Control (MAC) address is a unique identifier that not only uniquely identifies the device but also identifies the device vendor, device type, etc. That is, Wi-Fi networks utilize MAC addresses to uniquely track connected devices.
[0003] The MAC address can be used as a network address in communications within a network segment such as the local Wi-Fi section of a network. It can be used in Ethernet®, Wi-Fi, and Bluetooth® communications. Within the Open Systems Interconnection (OSI) network model, the MAC address is used in the media access control protocol sublayer of the data link layer. The MAC address is typically recognized by six groups of two-digit hexadecimal numbers.
[0004] MAC addresses may be assigned by device manufacturers as Ethernet® hardware addresses, such as the manufacturer's Organizational Unique Identifier (OUI). A device's MAC address may be stored in hardware, such as read-only memory (ROM). The address can be either a universally managed address (UAA) or a locally managed address (LAA).
[0005] In some cases, network interfaces may allow these MAC addresses to be changed. For example, on most Unix-like systems, the command utility "ifconfig" can be used to remove and add link address aliases, which can be used to specify which address to activate. Thus, in some configurations, MAC addresses can be randomized at specific points in time, such as during the boot process or before establishing a network connection.
[0006] MAC address spoofing is sometimes performed to exploit security vulnerabilities in computer systems. Some modern operating systems, particularly on mobile devices, such as Apple iOS and Android, are designed to randomize MAC address assignments to network interfaces when scanning wireless access points to evade tracking systems. To avoid tracking cell phone movement, Apple and other vendors often randomize MAC addresses when scanning networks on iOS (and other) devices. MAC address anonymization techniques can also be used to allow users to remain anonymous.
[0007] Apple platforms and similar platforms from other vendors can use randomized MAC addresses when performing Wi-Fi scans that are not associated with a specific Wi-Fi network. These scans can be performed to find and connect to existing Wi-Fi networks. Randomization of Wi-Fi MAC addresses is supported on iPhone® 5 and later models. Apple platforms also use randomized MAC addresses when performing Enhanced Preferred Network Offload (ePNO) scans, and when using location services in geofencing applications such as location-based reminders that determine if a device is near a specific location. In this case, the device's MAC address may change, for example, if it is disconnected from a Wi-Fi network, so even if the device is connected to a cellular network, it cannot be persistently tracked using its MAC address. Additionally, iOS 14, iPad®OS 14, and watchOS 7 introduce a new Wi-Fi privacy feature that allows iPhone®, iPad®, iPod® touch, or Apple Watch to identify themselves with a unique (i.e., randomized) MAC address when connected to a Wi-Fi network. This feature and other privacy features can be disabled by the user, or disabled using a new option in the Wi-Fi payload. Under certain circumstances, the device will fall back to its actual MAC address.
[0008] One problem with MAC address randomization and anonymization technologies is that they may involve one-way functionality of MAC addresses, which can suppress legitimate and useful functions. For example, in parent control functions used in Wi-Fi systems, it is beneficial to provide specific control to certain devices that can be more easily identified using MAC addresses. Also, including tracking systems can be beneficial to users in order to ensure consistency in application functionality. For example, legitimate companies such as Google and Apple track user movements to maintain the identity of the person being tracked, as well as the hardware itself. Therefore, in the field of MAC address processing, it is necessary to prevent spoofing, randomization, and anonymization of device identifiers in order to create continuity throughout the lifespan of a device.
[0009] Furthermore, if the operating systems (OS) of devices such as Apple and Android implement MAC address randomization, network-based solutions may be unable to use MAC addresses to uniquely identify devices from network traffic analysis. Moreover, all OS vendors are increasingly stringent with their privacy measures, obscuring other indicators that could be used to track devices. For example, Wi-Fi (i.e., IEEE 802.11) may use random sequence numbers and / or anonymize probe messages. This trend is likely to continue, making it more difficult for network-based solutions to track devices using protocol fields or other data controlled by protocol specifications or OS vendors. Therefore, a mechanism independent of vendor-owned software and components (e.g., networking discovery protocols) is needed. [Overview of the project]
[0010] This disclosure relates to a system, method, and non-temporary computer-readable medium for determining when different device identifiers are used to represent the same physical device. In this case, embodiments of the disclosure are configured to form links (e.g., associations, stitching, relating, etc.) between sets of information stored in databases that are otherwise considered independent of each other. By analyzing specific operating parameters, network metadata, etc., embodiments are configured to determine when two or more different device identifiers should be linked, stitched, and relating.
[0011] One implementation of the method includes the step of obtaining a first set of operational parameters relating to a first set of devices operating in a section of the network, the first set of operational parameters may include at least a first set of device identifiers representing the first set of devices. The method also includes the step of obtaining a second set of operational parameters relating to a second set of devices operating in that section of the network, the second set of operational parameters may include at least a second set of device identifiers representing the second set of devices. The method also includes the step of comparing a first set of device identifiers with a second set of device identifiers to find mismatched device identifiers. The method then includes the step of determining whether there are mismatched device identifiers in the first and second sets of device identifiers. With respect to mismatched device identifiers, the method includes the step of analyzing the first and second sets of operational parameters to determine whether it is likely that the device identifiers in the first set of device identifiers and the device identifiers in the second set of device identifiers represent the same device.
[0012] The method may include the step of associating device identifiers that are non-matching (e.g., different numbers) but represent the same device. The association may include linking together information of two device identifiers to record that two (or more) device identifiers actually represent the same physical device, even though the device identifiers had changed at some point. The association or linking may include some joining of data in a suitable database. According to some embodiments, the method may further include additional steps. For example, the method may store a first set of device identifiers in a database, a second set of device identifiers in a database, and then associate non-matching device identifiers in the database that are likely to represent the same device. In some cases, the first and second sets of device identifiers may include Media Access Control (MAC) addresses. Also, metrics for the first set of operational parameters may be measured in an earlier time frame, and metrics for the second set of operational parameters may be measured in a subsequent time frame.
[0013] Furthermore, the step of determining whether a device identifier in a first set of device identifiers and a device identifier in a second set of device identifiers are likely to represent the same device may include the steps of calculating a confidence score based on the relationship between the first and second sets of operating parameters, and then determining whether the confidence score exceeds a predetermined threshold. The relationship may include, for example, a) matching device operating factors, b) matching device features, c) uniqueness of matching device features, d) a weighted sum of device features, e) the number of matching device features, f) a machine learning (ML) model of the device matching technique, and / or other types of matching features. The first and second sets of operating parameters may include a) the time when the device identifier is first used, b) the time when the device identifier is no longer used, c) the time when the software or firmware is newly rolled out or upgraded, d) the type of device, e) the operating system of the device, f) the language of the device, g) the destination port or address used, h) the transmission pattern, i) the packet length, j) time information regarding packet transmission, k) one or more applications used by the device, l) device connection information, m) device disconnection information, n) the location of the device in a section of the network, o) supplemental device identification information, p) the carrier service used by the device, and / or other appropriate types of operating parameters.
[0014] In some embodiments, the first and second sets of operating parameters may constitute networking metadata, which may include information obtained via a) Address Resolution Protocol (ARP), b) Logical Link Control (LLC), c) Internet Control Message Protocol (ICMP), d) ICMP version 6 (ICMPv6), e) Bootstrap Protocol (BOOTP), f) Network Time Protocol (NTP), g) Transmit Control Protocol (TCP), h) Transport Layer Security (TLS), i) Dynamic Host Configuration Protocol (DHCP), j) DHCP version 6 (DHCPv6), k) Domain Name System (DNS), l) Multicast DNS (mDNS), m) User Agent, n) Universal Plug and Play (UPNP), o) Shared Serial Data Protocol (SSDP), p) Device Capability Information, q) Port Information, r) Protocol Information, s) 5-tuple Internet Protocol (IP) data, and / or other network protocols.
[0015] According to some embodiments, the network section described above may be a local Wi-Fi network. The step of determining whether a device identifier in a first set of device identifiers and a device identifier in a second set of device identifiers are likely to represent the same device may include setting an end time window around the last occurrence of each mismatched device identifier in the first set of device identifiers, setting an start time window around the first occurrence of each mismatched device identifier in the second set of device identifiers, and narrowing the end time window and the start time window until a single device identifier remains in the first set of device identifiers and a single device identifier remains in the second set of device identifiers.
[0016] The step of determining whether a device identifier in a first set of device identifiers and a device identifier in a second set of device identifiers are likely to represent the same device may also include storing a first set of sequence numbers used to identify packet transmission events associated with the first set of device identifiers, storing a second set of sequence numbers used to identify packet transmission events associated with the second set of device identifiers, and associating a first device identifier in the first set of device identifiers with a second device identifier in the second set of device identifiers if the difference between the end time of the sequence numbers in the first set of sequence numbers associated with the first device identifier and the start time of the sequence numbers in the second set of sequence numbers associated with the second device identifier is less than or equal to a predetermined threshold.
[0017] In some embodiments, the method may also include the step of running an application on one or more devices of a first and second set of devices to individually identify one or more devices. For example, the step of individually identifying one or more devices may include a) using Wi-Fi Protected Access (WPA) Enterprise, b) using installed certificates, c) reading Media Access Control (MAC) addresses, d) obtaining previously installed unique identification codes, e) receiving identifiers provided by the user through a captive portal, f) accessing user profile information, g) receiving user feedback regarding the devices to be correlated, and / or other identification procedures. Furthermore, according to some implementations, the method may also include the step of creating a new identifier for each device determined to be represented by a mismatched device identifier. The method may also include the step of creating a mapping table that connects real device identifiers, randomized device identifiers, and new identifiers.
[0018] In some embodiments, this disclosure further describes systems and methods for identifying user devices based on behavioral information, behavioral patterns, user habits, usage information, user trends, etc. In one implementation, the process for identifying user devices may include the step of monitoring one or more user devices operating on a Wi-Fi network. The process may further include the step of analyzing usage parameters for each of the one or more user devices. The process may also include identifying one or more user devices based on the usage parameters.
[0019] In some embodiments, the process may further include a) obtaining a device identifier associated with each of one or more user devices, and b) associating the device identifier of each of the one or more user devices with an operational identity based on usage parameters. For example, the device identifier associated with each of the one or more user devices may be a Media Access Control (MAC) address. The process may also include a) detecting when a new MAC address has been looked up for an unidentified user device operating on a Wi-Fi network, b) analyzing the current usage parameters of the unidentified user device, and c) comparing the current usage parameters of the unidentified user device with the usage parameters of one or more previously identified user devices. In response to determining that the current usage parameters match those of one of the previously identified user devices, the process may perform the step of stitching the new MAC address with the MAC address of the corresponding previously identified user device. Alternatively, in response to determining that the current usage parameters do not match those of one or more previously identified user devices, the process may perform the step of tagging the unidentified user device as a new device to be monitored on the Wi-Fi network.
[0020] The process may also include a step of analyzing usage parameters for each of one or more user devices over time. Next, based on the usage parameters analyzed over time, the process may include a step of creating one or more behavioral models associated with one or more users, so that each behavioral model represents the usage patterns of each user, depending on how the user uses at least one of the user devices. In some embodiments, the step of analyzing usage parameters over time may include creating one or more behavioral models using machine learning techniques. The process may further include a) assigning one or more unique user identifiers to represent one or more users, and b) associating one or more unique user identifiers with one or more behavioral models. The process may also include a step of retraining one or more behavioral models based on changes in the usage parameters of each corresponding user.
[0021] In additional embodiments, the usage parameters described herein may relate to the identity of one or more apps installed on one or more user devices. The usage parameters may also relate to app usage information, which may include a) the frequency of use of one or more apps, b) the time spent on each of the one or more apps, c) the type of communication associated with app usage, d) the time of day of app usage, and / or other information. Furthermore, the usage parameters may relate to the identity of one or more websites or domains accessed by one or more user devices.
[0022] In some embodiments, the process may include a step of refining the identity of one or more user devices based on weighted values of several metrics. The metrics may include a) the identity of one or more installed apps, b) app usage information, c) browsing patterns, and / or other metrics. The weighted values may relate, for example, to the uniqueness of each metric. User devices referred to herein may include smartphones, computers, laptops, tablets, smart TVs, Internet of Things (IoT) devices, media players, or other suitable devices that communicate with a Wi-Fi network. In some implementations, usage parameters may relate to device-based behaviors such as a) Wi-Fi access point usage, b) Wi-Fi network connection patterns, c) Bluetooth®-related transmissions, d) device port usage, and / or other device-related behaviors. [Brief explanation of the drawing]
[0023] This disclosure is illustrated and described herein with reference to various drawings, where similar reference numerals are used to indicate similar system components / method steps, as appropriate.
[0024] [Figure 1] This is a network diagram of a distributed Wi-Fi system with cloud-based control and management. [Figure 2] Figure 1 is a network diagram illustrating the differences in operation between a distributed Wi-Fi system and conventional single access point systems, Wi-Fi mesh networks, and Wi-Fi repeater networks. [Figure 3] This is a block diagram of a server that can be used in the cloud, on other systems, or as a standalone device. [Figure 4] Figure 1 is a block diagram of a user device, such as a mobile device, that may be used in a distributed Wi-Fi system. [Figure 5]A flowchart showing a process for associating different device identifiers together when it is determined that the device identifiers actually represent the same device. [Figure 6] A diagram showing an example of a screenshot of a user device regarding the type of a mobile application. [Figure 7] A table showing an example of the weights of different parameters of a user device. [Figure 8] A flowchart showing a process for stitching a MAC address based on user behavior.
Mode for Carrying Out the Invention
[0025] The present disclosure relates to a system and method for analyzing device identifiers in a networking environment. Since device identifiers (e.g., media access control (MAC) addresses, etc.) may be changed during the lifetimes of various network devices (e.g., mobile phones), the system and method are configured to determine from various metadata obtained by the network whether a device is represented by two (or more) different device identifiers. To accurately track a device, this embodiment may be configured to form associations or links between all different device identifiers that actually represent the same physical device.
[0026] The system and method may obtain a first set of operational parameters relating to a first set of devices operating in a section of the network. This first set of operational parameters may include at least a first set of device identifiers representing the first set of devices. The system and method may also obtain a second set of operational parameters relating to a second set of devices operating in that section of the network. This second set of operational parameters may include at least a second set of device identifiers representing the second set of devices. The system and method may also compare the first set of device identifiers with the second set of device identifiers to find mismatched device identifiers and determine whether there are mismatched device identifiers in the first and second sets of device identifiers. With respect to mismatched device identifiers, the system and method may analyze the first and second sets of operational parameters to determine whether it is likely that the device identifiers in the first set of device identifiers and the device identifiers in the second set of device identifiers represent the same device.
[0027] (Distributed Wi-Fi system) Figure 1 is a network diagram of a distributed Wi-Fi system 10 controlled via a cloud service 12. The distributed Wi-Fi system 10 may operate in accordance with the IEEE 802.11 protocol and its variations. The distributed Wi-Fi system 10 includes multiple access points 14 (labeled as access points 14A to 14H), which may be distributed throughout a location such as a residence or office. In other words, the distributed Wi-Fi system 10 is intended to operate in any physical location where it would be inefficient or impractical to serve using a single access point, repeater, or mesh system. As described herein, the distributed Wi-Fi system 10 may also be referred to as a network, system, Wi-Fi network, Wi-Fi system, cloud-based system, etc. Access points 14 may be referred to as nodes, access points, Wi-Fi nodes, Wi-Fi access points, etc. The purpose of access points 14 is to provide network connectivity to Wi-Fi client devices 16 (denoted as Wi-Fi client devices 16A to 16E). The Wi-Fi client device 16 may also be referred to as a client device, user device, client, Wi-Fi client, Wi-Fi device, etc.
[0028] In a typical residential deployment, a distributed Wi-Fi system 10 may include 3 to 12 or more access points within the home. The numerous access points 14 (sometimes referred to as nodes in the distributed Wi-Fi system 10) ensure that the distance between any two access points 14 is always small, and similarly, the distance to any Wi-Fi client device 16 requiring Wi-Fi service is also small. In other words, the goal of the distributed Wi-Fi system 10 may be that the distance between access points 14 is similar in size to the distance between a Wi-Fi client device 16 and its associated access point 14. This small distance ensures that the Wi-Fi signal adequately covers every corner of the consumer's home. It also ensures that any given hop in the distributed Wi-Fi system 10 is short and passes through walls very little. As a result, very strong signal strength is obtained at each hop in the distributed Wi-Fi system 10, enabling the use of high data rates and providing robust operation. Those skilled in the art will recognize that the Wi-Fi client device 16 may be a mobile device, tablet, computer, home appliance, home entertainment device, television, IoT device, or any network-enabled device. For external network connectivity, one or more of the access points 14 may be connected to a modem / router 18, which may be a cable modem, a digital subscriber loop (DSL) modem, or any device that provides external network connectivity to a physical location associated with the distributed Wi-Fi system 10.
[0029] While providing excellent coverage, a large number of access points 14 (nodes) present coordination challenges. Centralized control is necessary to properly configure and efficiently communicate with all access points 14. This cloud service 12 can provide control via a server 20, which is reachable via the internet and remotely accessible, such as through an application ("app") running on a user device 22. Thus, the operation of the distributed Wi-Fi system 10 becomes what is commonly known as a "cloud service." The server 20 is configured to receive measurement data via the cloud 12, analyze the measurement data, and configure the access points 14 within the distributed Wi-Fi system 10 based on that data. The server 20 may also be configured to determine which access point 14 each Wi-Fi client device 16 will associate with. In other words, in an exemplary embodiment, the distributed Wi-Fi system 10 includes cloud-based control (a cloud-based controller or cloud service in the cloud) for optimizing, configuring, and monitoring the operation of the access points 14 and Wi-Fi client devices 16. This cloud-based control is in contrast to conventional operation that relies on local settings, such as logging in locally to the access point. In the distributed Wi-Fi system 10, control and optimization do not require a local login to the access point 14, and the user device 22 (or local Wi-Fi client device 16) communicates with the server 20 in the cloud 12 via a heterogeneous network (a network different from the distributed Wi-Fi system 10) (e.g., LTE®, another Wi-Fi network, etc.).
[0030] Access point 14 may include both wireless and wired links for connectivity. In the example in Figure 1, access point 14A has an exemplary Gigabit Ethernet® (GbE) wired connection to modem / router 18. Optionally, access point 14B also has a wired connection to modem / router 18, for redundancy or load balancing, etc. Access points 14A and 14B may also have wireless connections to modem / router 18. Access point 14 may have a wireless link for client connections (referred to as a client link) and a wireless link for backhaul (referred to as a backhaul link). The distributed Wi-Fi system 10 differs from conventional Wi-Fi mesh networks in that the client links and backhaul links do not necessarily share the same Wi-Fi channel, thereby reducing interference. In other words, the access point 14 can support at least two Wi-Fi wireless channels, which can be flexibly used to provide either a client link or a backhaul link, and may have at least one wired port for connecting to the modem / router 18 or to other devices. In the distributed Wi-Fi system 10, only a small subset of the access points 14 require a direct connection to the modem / router 18, while the unconnected access points 14 communicate with the modem / router 18 via a backhaul link and return to the connected access points 14.
[0031] (Distributed Wi-Fi system compared to conventional Wi-Fi system) Figure 2 is a network diagram illustrating the differences in operation of a distributed Wi-Fi system 10 compared to a conventional single access point system 30, a Wi-Fi mesh network 32, and a Wi-Fi repeater network 33. The single access point system 30 relies on a single high-power access point 34, which may be centrally located to serve all Wi-Fi client devices 16 in a location (e.g., a house). Again, as described herein, in a typical house, the single access point system 30 may have several walls, floors, etc., between the access point 34 and the Wi-Fi client devices 16. In addition, the single access point system 30 operates on a single channel, leading to potential interference from neighboring systems. The Wi-Fi mesh network 32 solves some of the problems of the single access point system 30 by having multiple mesh nodes 36 that distribute Wi-Fi coverage. Specifically, the Wi-Fi mesh network 32 operates on the basis that the mesh nodes 36 are fully interconnected with each other and share a channel such as channel X between each mesh node 36 and the Wi-Fi client devices 16. In other words, the Wi-Fi mesh network 32 is a fully interconnected grid that shares the same channel and allows for multiple different paths between the mesh nodes 36 and the Wi-Fi client devices 16. However, because the Wi-Fi mesh network 32 uses the same backhaul channel, each hop between source points divides the network capacity by the number of hops required to deliver the data. For example, if it takes 3 hops to stream video to the Wi-Fi client device 16, the Wi-Fi mesh network 32 will have only 1 / 3 of its capacity remaining. The Wi-Fi repeater network 33 includes access points 34 wirelessly coupled to Wi-Fi repeaters 38. The Wi-Fi repeater network 33 has a star topology, where there is at most one Wi-Fi repeater 38 between the access points 14 and the Wi-Fi client devices 16.From a channel perspective, access point 34 can communicate with Wi-Fi repeater 38 on the first channel Ch.X, and Wi-Fi repeater 38 can communicate with Wi-Fi client device 16 on the second channel Ch.Y.
[0032] The distributed Wi-Fi system 10 solves the problem of Wi-Fi mesh networks 32, which require the same channel for all connections, by using different channels or bands for various hops to prevent a decrease in Wi-Fi speed (note: some hops may use the same channel / band, but this is not required). For example, the distributed Wi-Fi system 10 can use different channels / bands between access points 14 and between Wi-Fi client devices 16 (e.g., Ch.X, Y, Z, A), and the distributed Wi-Fi system 10 does not necessarily need to use all access points 14, based on configuration and optimization by the cloud 12. The distributed Wi-Fi system 10 solves the problem of single access point systems 30 by providing multiple access points 14. The distributed Wi-Fi system 10 is not constrained to a star topology, like a Wi-Fi repeater network 33 that allows a maximum of two wireless hops between Wi-Fi client devices 16 and the gateway. Furthermore, while the distributed Wi-Fi system 10 has a single path between the Wi-Fi client device 16 and the gateway, unlike the Wi-Fi repeater network 33, it forms a tree topology that allows for multiple wireless hops.
[0033] Wi-Fi is a shared simplex protocol, meaning that only one conversation between two devices can occur on the network at any given time; if one device is speaking, the other device must be listening. By using different Wi-Fi channels, multiple simultaneous conversations can occur simultaneously in a distributed Wi-Fi system 10. Interference and congestion can be avoided by selecting different Wi-Fi channels among the access points 14. Server 20 via the cloud 12 automatically configures the access points 14 with an optimized channel hop solution. The distributed Wi-Fi system 10 can select routes and channels to support the ever-changing needs of consumers and their Wi-Fi client devices 16. The approach of the distributed Wi-Fi system 10 is to ensure that the Wi-Fi signal does not need to travel long distances (either in backhaul or client connections). Therefore, by communicating on the same channel as the Wi-Fi mesh network 32 or Wi-Fi repeater, the Wi-Fi signal remains strong and interference can be avoided. In an exemplary embodiment, Server 20 in the cloud 12 is configured to optimize channel selection for the best user experience.
[0034] Of note, this disclosure for identifying MAC addresses is not limited to distributed Wi-Fi systems 10, but is intended for any of the Wi-Fi networks 10, 30, 32, or 33, including monitoring via the cloud 12 and local monitoring.
[0035] (Cloud-based Wi-Fi management) Conventional Wi-Fi systems utilize local management, where users on the Wi-Fi network connect to a designated address (e.g., 192.168.1.1). The distributed Wi-Fi system 10 is configured to perform cloud-based management via a server 20 in the cloud 12. Furthermore, a single access point system 30, a Wi-Fi mesh network 32, and a Wi-Fi repeater network 33 can also support cloud-based management, as described above. For example, AP 34 and / or mesh nodes 36 can be configured to communicate with a server 20 in the cloud 12. This configuration can be achieved via a software agent installed on each device, such as OpenSync. As described herein, cloud-based management includes reporting Wi-Fi-related performance metrics to the cloud 12 and receiving Wi-Fi-related configuration parameters from the cloud 12. The system and method are intended for use with any Wi-Fi system (i.e., a distributed Wi-Fi system 10, a single access point system 30, a Wi-Fi mesh network 32, and a Wi-Fi repeater network 33, etc.), including systems that only support reporting of Wi-Fi-related performance metrics (but do not support cloud-based configurations).
[0036] Cloud computing utilizes cloud computing systems and methods to abstract away physical servers, storage, networking, etc., and instead provides them as on-demand, resilient resources. The National Institute of Standards and Technology (NIST) provides a concise and specific definition, which states that cloud computing is a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned and released with minimal administrative effort or interaction with a service provider. Cloud computing differs from the traditional client-server model in that it provides applications from servers that run and are managed in the client's web browser, etc., without requiring the installation of client versions of the application. Centralization allows cloud service providers to have complete control over the versions of browser-based applications and other applications provided to clients, eliminating the need for version upgrades and license management on individual client computing devices. The term SaaS is sometimes used to describe application programs delivered through cloud computing. The common abbreviation for the cloud computing service provided (or the collection of all existing cloud services) is "cloud."
[0037] (Example server architecture) Figure 3 is a block diagram of a server 200 that may be used in Cloud 12, in other systems, or as a standalone device. From a hardware architecture standpoint, the server 200 may generally be a digital computer including a processor 202, an input / output (I / O) interface 204, a network interface 206, a data store 208, and memory 210. Figure 3 depicts the server 200 in an overly simplified manner, and those skilled in the art will understand that actual embodiments may include additional components and appropriately configured processing logic to support known or conventional operating features not described in detail herein. The components (202, 204, 206, 208, and 210) are connected communicatively via a local interface 212. The local interface 212 may be, but is not limited to, one or more buses or other wired or wireless connections, as known in the art, for example. The local interface 212 may have additional elements omitted for simplification, such as controllers, buffers (caches), drivers, repeaters, and receivers, in order to enable communication. Furthermore, the local interface 212 may include address, control, and / or data connections to enable appropriate communication between the aforementioned components.
[0038] The processor 202 is a hardware device for executing software instructions. The processor 202 may be any custom-made or commercially available processor, a central processing unit (CPU), an auxiliary processor among several processors associated with the server 200, a semiconductor-based microprocessor (in the form of a microchip or chipset), or any device in general for executing software instructions. When the server 200 is operating, the processor 202 is configured to execute software stored in memory 210, communicate data with memory 210, and generally control the operation of the server 200 according to software instructions. The I / O interface 204 may be used to receive user input from one or more devices or components and / or to provide system output to one or more devices or components. User input may be provided, for example, via a keyboard, touchpad, and / or mouse. System output may be provided via a display device and a printer (not shown). The I / O interface 204 may include, for example, a serial port, a parallel port, a Small Computer System Interface (SCSI), Serial ATA (SATA), Fibre Channel, Infiniband, iSCSI, a PCI Express interface (PCI-x), an infrared (IR) interface, a radio frequency (RF) interface, and / or a Universal Serial Bus (USB) interface.
[0039] The network interface 206 may be used to enable the server 200 to communicate over a network such as the Internet. The network interface 206 may include, for example, an Ethernet® card or adapter (e.g., 10BaseT, Fast Ethernet®, Gigabit Ethernet®, 10GbE) or a Wireless Local Area Network (WLAN) card or adapter (e.g., 802.11a / b / g / n / ac). The network interface 206 may include address, control, and / or data connections to enable proper communication over the network. The data store 208 may be used to store data. The data store 208 may include volatile memory elements (e.g., random access memory (RAM) such as DRAM, SRAM, SDRAM), non-volatile memory elements (e.g., ROM, hard drive, tape, CD-ROM, etc.), or a combination thereof. Furthermore, the data store 208 may incorporate electronic, magnetic, optical, and / or other types of storage media. In one example, the datastore 208 may be located inside the server 200, for example, as an internal hard drive connected to a local interface 212 within the server 200. In another embodiment, the datastore 208 may be located outside the server 200, for example, as an external hard drive connected to an I / O interface 204 (e.g., SCSI or USB connection). In yet another embodiment, the datastore 208 may be connected to the server 200 via a network, for example, as a network-attached file server.
[0040] Memory 210 may include volatile memory elements (e.g., random access memory (RAM) such as DRAM, SRAM, SDRAM, etc.), non-volatile memory elements (e.g., ROM, hard drives, tapes, CD-ROMs, etc.), or combinations thereof. Furthermore, memory 210 may incorporate electronic, magnetic, optical, and / or other types of storage media. It should be noted that memory 210 may have a distributed architecture in which various components are located remotely from one another but can be accessed by the processor 202. The software in memory 210 may include one or more software programs, each software program including an ordered list of executable instructions for implementing a logical function. The software in memory 210 includes a suitable operating system (O / S) 214 and one or more programs 216. The operating system 214 essentially controls the execution of other computer programs such as one or more programs 216 and provides scheduling, input / output control, file and data management, memory management, and communication control and related services. One or more programs 216 may be configured to implement various processes, algorithms, methods, techniques, etc., as described herein.
[0041] (Example user device architecture) Figure 4 is a block diagram of a user device 300 that may be used in user device 22, etc. From a hardware architecture standpoint, the user device 300 can generally be a digital device including a processor 302, an input / output (I / O) interface 304, a radio 306, a data store 308, and memory 310. Figure 4 depicts the user device 300 in an overly simplified manner, and those skilled in the art will understand that actual embodiments may include additional components and appropriately configured processing logic to support known or conventional operating features not described in detail herein. The components (302, 304, 306, 308, and 302) are communicatively coupled via a local interface 312. The local interface 312 may be, but is not limited to, one or more buses or other wired or wireless connections, as known in the art, for example. The local interface 312 may have additional elements omitted for simplification, in particular controllers, buffers (caches), drivers, repeaters, and receivers, in order to enable communication. Furthermore, the local interface 312 may include address, control, and / or data connections to enable appropriate communication between the aforementioned components.
[0042] The processor 302 is a hardware device for executing software instructions. The processor 302 may be any custom-made or commercially available processor, a central processing unit (CPU), an auxiliary processor among several processors associated with the user device 300, a semiconductor-based microprocessor (in the form of a microchip or chipset), or any device in general for executing software instructions. When the user device 300 is operating, the processor 302 is configured to execute software stored in the memory 310, communicate data with the memory 310, and generally control the operation of the user device 300 according to software instructions. In one embodiment, the processor 302 may include a mobile-optimized processor, such as one optimized for power consumption and mobile applications. The I / O interface 304 can be used to receive user input and / or provide system output. User input can be provided, for example, via a keypad, touchscreen, scroll ball, scroll bar, buttons, barcode scanner, etc. System output can be provided via a display device such as a liquid crystal display (LCD) or touchscreen. The I / O interface 304 may also include, for example, a serial port, a parallel port, a Small Computer System Interface (SCSI), an infrared (IR) interface, a radio frequency (RF) interface, a Universal Serial Bus (USB) interface, and the like. The I / O interface 304 may also include a graphical user interface (GUI) that allows a user to interact with the user device 300. Furthermore, the I / O interface 304 may further include an imaging device, i.e., a camera, a video camera, and the like.
[0043] The radio 306 enables wireless communication to an external access device or network. Any number of suitable wireless data communication protocols, technologies, or methodologies can be supported by the radio 306. The radio 306 includes, but is not limited to, RF; IrDA (infrared); Bluetooth®; ZigBee® (and other variations of the IEEE 802.15 protocol); IEEE 802.11 (any variation); IEEE 802.16 (WiMAX or other variations); Direct Sequence Spread Spectrum; Frequency Hopping Spread Spectrum; Long Term Evolution (LTE®); Cellular / Wireless / Cordless Telecommunications Protocols (e.g., 3G / 4G / 5G, etc.); Wireless Home Network Communication Protocols; Proprietary Wireless Data Communication Protocols such as variations of Wireless USB; and any other arbitrary protocols for wireless communication. The data store 308 may be used to store data. The datastore 308 may include any of the following: volatile memory elements (e.g., random access memory (RAM) such as DRAM, SRAM, SDRAM, etc.), non-volatile memory elements (e.g., ROM, hard disk, tape, CD-ROM, etc.), or combinations thereof. Furthermore, the datastore 308 may incorporate electronic, magnetic, optical, and / or other types of storage media.
[0044] Memory 310 may include volatile memory elements (e.g., random access memory (RAM) such as DRAM, SRAM, SDRAM), non-volatile memory elements (e.g., ROM, hard drive, etc.), or any combination thereof. Furthermore, memory 310 may incorporate electronic, magnetic, optical, and / or other types of storage media. It should be noted that memory 310 may have a distributed architecture in which various components are located remotely from each other but can be accessed by the processor 302. The software in memory 310 may include one or more software programs, each software program containing an ordered list of executable instructions for implementing logical functions. In the example in Figure 3, the software in memory 310 includes a suitable operating system (O / S) 314 and program 316. The operating system 314 essentially controls the execution of other computer programs and provides scheduling, input / output control, file and data management, memory management, and communication control and related services. Program 316 may include various applications, add-ons, etc., configured to provide end-user functions on the user device 300. For example, exemplary programs 316 may include, but are not limited to, web browsers, social networking applications, streaming media applications, games, mapping and location-based applications, email applications, financial applications, and the like. In a typical example, an end user typically uses one or more programs 316 in conjunction with a network.
[0045] (Association of device identifiers) The server 200's program 216 and / or the user device 300's program 316 may include a “device identifier association program” or other similar program for associating mismatched device identifiers in order to provide continuity of identification information when a device identifier (e.g., MAC address) changes during the device’s lifetime. Thus, the function of associating device identifiers may be performed on the server 200, on the user device 300 itself, or in combination with other systems or devices in the network on which the server 200, user device 300, and / or communication device (e.g., user device) are operating. In some cases, the server 200 may be configured to download one or more applications or programs (e.g., device identifier association programs) to one or more user devices 300 in order to perform at least some of the steps described in this disclosure.
[0046] Device identifier association programs according to various embodiments of this disclosure may be stored in a non-temporary computer-readable storage medium (e.g., memory 210, 310, etc.). The device identifier association program may have computer-readable code configured to program a server 200, a user device 300, etc., to perform specific functions that can be assisted by a processor 202, 302, or other suitable processing device.
[0047] A device identifier association program is configured to associate two different device identifiers when it is determined with at least a reasonable level of certainty that the device identifiers actually represent the same device (e.g., user device 300). Thus, by measuring and utilizing network metadata, the device identifier association program is configured to determine when two different device identifiers (e.g., MAC addresses) obtained at two different points in time appear to represent a situation in which the device identifiers had changed. For example, by analyzing network metadata (e.g., device type, the time when one device identifier disappears and another first appears, etc.), the systems and methods of the disclosure are configured to process this evidence to determine whether a first device identifier is associated with a second device identifier. The process of “correlating” two (or more) different device identifiers (each representing the same device) may include storing links connecting the two (or more) device identifiers in a database (e.g., datastores 208, 308, etc.). For example, the term “associate” may also be referred to in this disclosure as “stitch,” “relationship,” “link,” “related,” “connect,” “integrate,” or “combine,” and includes creating continuity in the records of each device, whether the device identifier of each specific device changes or does not change. The device identifier association program may be configured to examine evidence in network metadata to make reasonable inferences about which device identifiers should be linked.
[0048] According to various embodiments of this disclosure, device identifier association programs 216, 316 can be deployed in various forms on a server 200, one or more user devices 300, a network management system, a control system, a router, a modem, etc. Each version of the software / firmware for detecting device identity (and other functions described throughout this disclosure) may include any functionality suitable for operation in various parts of the network, as described in this disclosure and understood by a person skilled in the art who is knowledgeable about this disclosure. For example, some processing functions may be incorporated into (or deployed) one or more access points (APs) of a Wi-Fi network, a router or modem, a cloud, an internet device (e.g., a server), offline (e.g., for non-real-time operation), online in real-time, etc. Also, some implementations of systems and methods may include a historical approach that can retrieve information retrospectively to "fill in the gaps" with respect to MAC address changes.
[0049] Figure 5 is a flowchart illustrating one embodiment of process 400 for associating, stitching, or linking heterogeneous device identifiers in a database when they relate to the same physical device. Process 400 includes a first step of obtaining a first set of operational parameters related to a first set of devices operating in a network section, as shown in block 402. For example, the first set of operational parameters may include at least a first set of device identifiers representing the first set of devices. Process 400 also includes a step of obtaining a second set of operational parameters related to a second set of devices operating in a network section, as shown in block 404. For example, the second set of operational parameters may include at least a second set of device identifiers representing the second set of devices. Process 400 also includes a step of finding mismatched device identifiers by comparing the first set of device identifiers with the second set of device identifiers, as shown in block 406.
[0050] Next, process 400 includes a step of determining whether there are any mismatched device identifiers in the first and second sets of device identifiers, as shown in condition diamond 408. If there are no mismatched device identifiers (or if there are no more mismatched device identifiers that have not yet been processed), process 400 terminates. Otherwise, if mismatched device identifiers are found in condition diamond 408, process 400 proceeds to block 410. Thus, with respect to mismatched device identifiers, process 400 includes a step of analyzing the first set of operating parameters and the second set of operating parameters to determine whether it is likely that the device identifiers in the first set of device identifiers and the device identifiers in the second set of device identifiers represent the same device. If it is determined in condition diamond 412 that they represent the same device, process 400 proceeds to block 414. Otherwise, process 400 loops back to condition diamond 408 to determine whether there are any more mismatched device identifiers. As shown in block 414, process 400 includes a step of associating device identifiers that represent the same device but are mismatched (e.g., different numbers). The association may involve linking together information from two device identifiers to record that, although the device identifier had changed at some point, two (or more) device identifiers actually represent the same physical device. The association or link may involve some joining of data in the appropriate database (e.g., data stores 208, 308). After the association, process 400 returns to condition diamond 408 to process the remaining mismatched device identifiers.
[0051] According to some embodiments, process 400 may further include additional steps. For example, process 400 may store a first set of device identifiers in a database, a second set of device identifiers in a database, and then associate mismatched device identifiers in the database that are likely to represent the same device. In some cases, the first and second sets of device identifiers may include Media Access Control (MAC) addresses. Also, metrics for the first set of operating parameters may be measured in an earlier time frame, and metrics for the second set of operating parameters may be measured in a subsequent time frame.
[0052] Furthermore, the step of determining whether the device identifiers in a first set of device identifiers and the device identifiers in a second set of device identifiers are likely to represent the same device (e.g., block 410) may include the steps of calculating a confidence score based on the relationship between the first and second sets of operating parameters, and then determining whether the confidence score exceeds a predetermined threshold. The relationship may include, for example, a) matching device operating factors, b) matching device features, c) uniqueness of the matching device features, d) a weighted sum of the device features, e) the number of matching device features, f) a machine learning (ML) model of the device matching technique, and / or other types of matching features. The first and second sets of operating parameters may include a) the time when the device identifier is first used, b) the time when the device identifier is no longer used, c) the time when the software or firmware is newly rolled out or upgraded, d) the type of device, e) the operating system of the device, f) the language of the device, g) the destination port or address used, h) the transmission pattern, i) the packet length, j) time information regarding packet transmission, k) one or more applications used by the device, l) device connection information, m) device disconnection information, n) the location of the device in a section of the network, o) supplemental device identification information, p) the carrier service used by the device, and / or other appropriate types of operating parameters.
[0053] In some embodiments, the first and second sets of operating parameters may constitute networking metadata, which may include information obtained via a) Address Resolution Protocol (ARP), b) Logical Link Control (LLC), c) Internet Control Message Protocol (ICMP), d) ICMP version 6 (ICMPv6), e) Bootstrap Protocol (BOOTP), f) Network Time Protocol (NTP), g) Transmit Control Protocol (TCP), h) Transport Layer Security (TLS), i) Dynamic Host Configuration Protocol (DHCP), j) DHCP version 6 (DHCPv6), k) Domain Name System (DNS), l) Multicast DNS (mDNS), m) User Agent, n) Universal Plug and Play (UPNP), o) Shared Serial Data Protocol (SSDP), p) Device Capability Information, q) Port Information, r) Protocol Information, s) 5-tuple Internet Protocol (IP) data, and / or other network protocols.
[0054] According to some embodiments, the network section described above may be a local Wi-Fi network. The step of determining whether a device identifier in a first set of device identifiers and a device identifier in a second set of device identifiers are likely to represent the same device (e.g., block 410) may include setting an end time window around the last occurrence of each mismatched device identifier in the first set of device identifiers, setting an start time window around the first occurrence of each mismatched device identifier in the second set of device identifiers, and narrowing the end time window and the start time window until a single device identifier remains in the first set of device identifiers and a single device identifier remains in the second set of device identifiers.
[0055] The step of determining whether a device identifier in a first set of device identifiers and a device identifier in a second set of device identifiers are likely to represent the same device (e.g., block 410) may also include storing a first set of sequence numbers used to identify packet transmission events associated with the first set of device identifiers, storing a second set of sequence numbers used to identify packet transmission events associated with the second set of device identifiers, and associating a first device identifier in the first set of device identifiers with a second device identifier in the second set of device identifiers if the difference between the end time of the sequence numbers in the first set of sequence numbers associated with the first device identifier and the start time of the sequence numbers in the second set of sequence numbers associated with the second device identifier is less than or equal to a predetermined threshold.
[0056] In some embodiments, process 400 may include the step of running the application on one or more devices of a first and second set of devices to individually identify one or more devices. For example, the step of individually identifying one or more devices may include a) using Wi-Fi Protected Access (WPA) Enterprise, b) using installed certificates, c) reading Media Access Control (MAC) addresses, d) obtaining previously installed unique identification codes, e) receiving identifiers provided by the user through a captive portal, f) accessing user profile information, g) receiving user feedback regarding the associated device, and / or other identification steps.
[0057] Furthermore, process 400 may include, according to some implementations, a step of creating a new identifier for each device determined to be represented by a mismatched device identifier. Process 400 may also include a step of creating a mapping table that connects the actual device identifier, the randomized device identifier, and the new identifier.
[0058] (Identifying MAC addresses that should be stitched / associated) The systems and methods of this disclosure may utilize temporal factors to determine whether a device identifier used at an earlier point in time represents the same physical device represented by a different device identifier used at a later point in time. For example, device identifier association programs 216, 316 (or other systems and methods described in this disclosure) may be configured to analyze network metadata that can be obtained from a network in operation. From this metadata, device identifier association programs 216, 316 can keep track of the occurrence of MAC addresses (or other device identifiers) being used. Device identifier association programs 216, 316 may record the time when one or more new MAC addresses first appeared and the time when one or more old MAC addresses no longer appear. For example, it may be determined that at a certain time t, a particular MAC address is no longer used on the network, and then a new MAC address first appears shortly thereafter (e.g., at t+x, where x may be a relatively short time). In this case, using temporal information, it can be considered that these two MAC addresses may represent the same device. Naturally, further analysis can be used to confirm (with reasonable certainty) that they actually represent (or are likely to represent) the same device.
[0059] In some embodiments, the time analysis procedure may include setting time windows before and after the last and first occurrences, respectively, of each MAC address that is interrupted or started, under the observation of device identifier association programs 216, 316. These time windows may be initially set to a wide time span (e.g., several days), so that there may be large overlaps where several MAC addresses are interrupted and / or started. The device identifier association programs 216, 316 may then narrow the time windows until a single match exists between the end MAC address and the start MAC address. Stitching operations (e.g., correlation, joining, etc.) may be performed to connect the data of the two MAC addresses in order to store this association in data stores 208, 308.
[0060] In some embodiments, additional processing steps may include determining whether a new software / firmware upgrade has been rolled out to a device in the network. If so, this information may be used to confirm that two heterogeneous MAC addresses (connected by temporal processing) are more likely to represent the same device. This confirmation may be based on the recognition that some companies (e.g., Apple) may randomize or otherwise change the MAC addresses of devices when new software / firmware is rolled out. Thus, the device identifier association programs 216, 316 may be configured to use multiple characteristics, features, parameters, or other appropriate information obtained from network metadata in detecting MAC address changes. In other words, the device identifier association programs 216, 316 can determine the connection between heterogeneous MAC addresses using one analysis (e.g., temporal processing) or multiple analyses (e.g., temporal processing, software / firmware rollout information, and other types of analysis and information processing). Other types of analysis, processing steps, etc., of network metadata may be used as described throughout this disclosure.
[0061] For example, another type of analysis that may be performed by device identifier association programs 216, 316 is the process of considering device identifiers that could be candidates for association based on device type. That is, if it is shown that two association candidates are both used to represent the same type of device, this information can be used as another confirmation that they represent the same device. Otherwise, if it is determined that the two candidates represent two different types of devices, it is possible to determine that these device identifiers do not represent the same device and to eliminate any matching / linking or to withdraw consideration for matching / linking.
[0062] Here too, device identifier association programs 216, 316 can determine the device type using the received network metadata. For example, the network metadata may include any one or more of the following: ARP, LLC, ICMP, ICMPv6, BOOTP, NTP, TCP, TLS client hello, DHCP, DHCPv6, mDNS, user agent, UPNP, SSDP, DNS, ICMP, device capabilities, port, protocol, 5-tuple IP data, etc.
[0063] Another factor that may be used to determine whether two (or more) different MAC addresses represent the same device is the device's behavior. For example, if the behavior of a device represented by one MAC address is similar to or the same as the behavior of a device represented by another MAC address, the device identifier association programs 216, 316 may be configured to use (to some extent) this information to verify that the MAC addresses are associated with the same device. Some examples of device behavior may include application ("app") usage, connection and disconnection patterns (e.g., within a network), time patterns (e.g., when the device is used), location patterns (e.g., where the device is used), etc. For example, location patterns may refer to usage within an area (e.g., a city, a country, etc.) or smaller-scale usage within a home or office, etc. For example, the detection of smaller-scale locations may be based on which access points are used to use the device.
[0064] Furthermore, considering and using device behavior may also relate to detecting the programming language used by the device. This may involve detecting the packet destination port, packet destination address, transmission pattern, packet length, Tx / Rx bytes moved, time between packet transmissions, protocol, etc. Device type and device behavior may be used together to determine whether two device identifiers are likely to be associated with the same device, or alternatively to determine whether device identifiers are likely to represent different devices.
[0065] Furthermore, in the process of determining whether two (or more) MAC addresses should be stitched, associated, related, etc., the device identifier association programs 216, 316 may also obtain and use other parameters that form a unique device ID other than the MAC address. For example, other IDs may be associated with DHCP unique identifiers (DUIDs), DHCPv6 DUIDs, TCP identifiers, mDNS option data, NetBios, ICMPv6 Neighbor Solicitation and Neighbor Advertisement packets, etc.
[0066] Device identifier association programs 216, 316 may utilize sequence numbers commonly used in some communication protocols. In this case, the sequence number may be applied to packets in a sequential manner to identify a particular packet. These sequence numbers may be counted up for each packet and eventually rolled over or reset. A device disconnected from the network may be assumed to be reconnected soon after. If MAC addresses are randomized in this situation, the sequence numbers may be approximately the same value, but other devices operating on the network may find that packets with sequence numbers at very different points in the sequence number counting / rollover process are being communicated. Therefore, device identifier association programs 216, 316 may parse these sequence numbers to determine whether two identifiers represent the same device.
[0067] To determine the MAC address to stitch / link, the device identifier association programs 216, 316 may be configured to consider the operating system (O / S) used by the device, the device type, and / or the firmware version the device has. Here again, the possibility that two identifiers represent the same device may be based on whether these factors are the same or similar, while if these factors are different, this observation may be used to suggest that the identifiers represent different devices.
[0068] Metadata indicating where device identifiers (e.g., MAC addresses) are managed can also be used to determine whether two (or more) device identifiers represent the same device. For example, some devices may be configured to receive locally managed MAC addresses (e.g., within the Wi-Fi network itself), while others may be configured to receive conventionally managed MAC addresses, which may include the management of Organization-Specific Identifiers (OUIs). Furthermore, metadata regarding carrier information (e.g., the cellular service the device is using) may be used to determine whether different MAC addresses represent the same or different devices.
[0069] (Active observation vs. passive observation for information gathering for stitching) Network metadata may be obtained using passive and / or active processes. In a passive system, a monitoring device may be configured to simply observe messages from the device, and device identifier association programs 216, 316 may use these messages to build rules and / or train machine learning (ML) models based on passively observed traffic patterns.
[0070] On the other hand, an active system may be configured to use a monitoring device to send a message to the device requesting information that would help identify the device. In this case, the monitoring device may elicit a response containing information useful to device identifier association programs 216, 316 for identifying the device. With respect to the active system, the monitoring device may obtain fields that the device freely provides but which may be withheld for privacy reasons. Eliciting a response may include, for example, requesting ICMP timestamp information and receiving a response in the form of ICMP message types 13 and 15. Other requests may include DHCPv6 queries, mDNS scans, SSDP scans, TCP scans, UDP-based scans, etc., to identify the device type and device identifier.
[0071] (Obtain user feedback and / or install a new device identifier on the device) According to some implementations, the systems and methods of this disclosure may further include other proactive methods for obtaining information that can be used to determine whether different MAC addresses represent the same device. For example, instead of simply observing readily available metadata from the network, some implementations may include allowing device identifier association programs 216, 316 to request useful information from the user themselves, and this information being used productively to stitch MAC addresses. Also, as mentioned above, certain functions may be installed on new devices to more easily track changes in MAC addresses.
[0072] In some examples, the systems and methods of this disclosure may use Wi-Fi Protected Access (WPA) Enterprise, install certificates, install an app on a device (e.g., newly issued to a user) that presents a unique identity, and install an app on the device that reads the actual MAC address or other identifiers already present on the device. This information can then be reported to device identifier association programs 216, 316 as needed for stitching / linking. In some cases, the user of a new device may need to go through a captive portal, and software may be installed on the new device to require the user to provide a device identifier (e.g., login name). This information can be stored and reported for device identification purposes.
[0073] Users may be required to enter a user profile that includes information about themselves and the device itself. This may be included in apps (e.g., internet access apps to provide parental controls, usage restrictions for different times of day, restrictions on access to specific sites, etc.). Once an app is installed on the device, if the device may undergo MAC address randomization, such as when a new firmware update is released, the user may need to re-enter certain information. In this case, the user data can be used by device identifier association programs 216, 316 to determine that different MAC addresses actually refer to the same device.
[0074] In some embodiments, the app may be configured to suggest which device identifiers to stitch / link and prompt the user for a response. The user can then accept, reject, or refuse to respond to the suggestion. Alternatively, the app may be configured to provide a selection of eligible device identifiers and ask the user to choose which ones correspond to the same device. In some embodiments, the app may be configured to notify the user that two or more device identifiers have been automatically associated, stitched, and linked, and in response, allow the user to accept or reject the stitching if desired, and / or request to reverse the stitching. The app may also be configured to provide the user with an opportunity to manually modify any stitching / linking process.
[0075] Depending on the implementation, an application installed on a user device ("App") may include various settings and / or be associated with other functionalities. For example, embodiments of the present disclosure may be added on to other software / firmware products to enable a device identification process, along with any other combination of software functions for performing other services.
[0076] For example, an app may be configured to transfer configurations applied by the user or system to the device. These configurations may include settings for parental controls, cybersecurity settings (e.g., blacklisted sites, whitelisted sites, etc.), device nicknames, user IDs, user profiles, access control zones, motion detection settings, sensor alerts (e.g., health-related or bio-related settings), content filters (e.g., for teenagers and children), policy settings, internet freeze settings and schedules, QoS (Quality of Service) priority settings (e.g., set by devices, services, applications, etc.), room assignment information, previously obtained captive portal login information, device access restrictions, quarantine or block status of a given device, group assignments, access sharing, screen time settings and status, app usage status, and / or other appropriate setting restrictions, constraints, parameters, etc. If multiple stitching options exist, the app may be configured to transfer any of the settings, restrictions, etc. that include the same option.
[0077] (Create a continuity of data related to the device) If the device identifier association programs 216, 316 determine (with reasonable certainty) that two (or more) device identifiers represent the same physical device, then certain steps may be taken to create a continuity of data relating to those two (or more) device identifiers. In other words, the device identifier association programs 216, 316 may create any appropriate links or stitchings in the data stores 208, 308 to consolidate records relating to the same device.
[0078] In one embodiment, the device identifier association programs 216, 316 may be configured to rewrite the database with newly discovered / determined local MAC addresses. Alternatively, during querying, the device identifier association programs 216, 316 may be configured to combine multiple MAC address records based on an association table listing aliases.
[0079] According to some embodiments, the systems and methods of the present disclosure may be configured to create a new identifier (or new device identifier) that is unique to each device and retained for the lifetime of the device. This can be thought of as a preferred method such as stitching, linking, relating, or joining. In this case, the device identifier association programs 216, 316 may be configured to rewrite entries with one or more databases with the new unique identifier. The systems and methods may also store a mapping table that links real MAC addresses and any randomized MAC addresses found to the new identifier. The mapping table may be configured as a slowly changing dimensional table. The mapping function may be performed on the fly and may include an action to store new data along with the unique identifier as it is acquired and analyzed. Alternatively, data may be stored with the currently used MAC address and converted to the unique identifier when the data is read based on the mapping table.
[0080] (Confidence score related to association / stitching) Another aspect of the various systems and methods of this disclosure is the confidence that device identifier association programs 216, 316 can reasonably infer from metadata that two (or more) different MAC addresses actually refer to a single physical device. The confidence (or level of certainty) of this discovered connection may be characterized by a specific score or value, which may be referred to herein as the “confidence score.” The confidence score may be based on how many elements (at least partially) match. The confidence score may be based (at least partially) on the uniqueness of the matching factors. The confidence score may be based (at least partially) on a weighted sum of the various matching factors and the degree of uniqueness of each of those factors. The confidence score may also be influenced by how many potential candidates (e.g., non-matching device identifier candidates) exist (e.g., those to be analyzed for potential stitching / association).
[0081] Calculating a confidence score may involve fine-grained weighting. For example, it may involve weighting each factor based on device type, firmware version, temporal factors, application usage, device language, location information, etc. For example, any iOS device may be uniquely identified based on various factors (e.g., DHCP, DHCPv6, ICMP, ICMPv6, QUIC, HTTP UA, etc.). However, for example, in iOS 14, DHCPv6 may be given a higher weight compared to other fields (e.g., mDNS, ICMPv6, etc.) that are given a higher weight in iOS 15. Therefore, rules can be built and modified as needed to accommodate different types of devices and operating systems. Furthermore, as new devices and operating systems are deployed to the network, the systems and methods of this disclosure may be updated to specifically characterize the weights of various factors for these new devices. Accordingly, in some implementations, the device identifier association programs 216, 316 may periodically train and modify new ML models for various versions and models of devices using artificial intelligence such as machine learning (ML). ML-based models may be used for specific behaviors in each version and model of the device.
[0082] In some embodiments, the device identifier association programs 216, 316 may be configured to take a multi-layered approach for device identification. For example, instead of applying a flat weighted average, the device identifier association programs 216, 316 may be configured to apply a hierarchical ML clustering approach in the step of determining whether two (or more) MAC addresses should be stitched together. This can be done first to classify devices based on their type, model, operating system, etc. Then, in the next step, the device identifier association programs 216, 316 may be configured to uniquely identify each device in the cluster and ML model obtained in the stitching step.
[0083] The disclosure may also include embodiments that utilize a process for automatically adjusting the time interval for collecting data from the network. For example, this may include data collection intervals that are adjusted based on feedback from an identification engine. Data collection for a particular device may be stopped or paused as soon as the required level of confidence for device identification is achieved. In some cases, stitching / association may be performed only if the confidence score is above a certain threshold (e.g., about 90% or higher). The system and method may use the confidence score to delay stitching / association so that more information can be collected or so that the reappearance of some potential stitching candidates can be observed periodically in the network. The confidence score may also be used to determine when it is necessary to ask for user feedback to more accurately determine whether a MAC address should be linked / stitched.
[0084] (Hostname masking) According to some embodiments, an app for stitching MAC addresses may enable a user to input a nickname for a device on the network. This nickname can effectively replace any unique hostname within the user interface. Device identifier association programs 216, 316 may be configured to input or associate a hostname with a device (e.g., following MAC randomization) based on the MAC association / stitching performed. In some embodiments, the app may generate random hostnames for a given device (e.g., iPhone® Blue) so that different devices are distinguishable within the app. The user can determine which is their device by observing over time. The app may also prompt the user to input a nickname when a device with a masked hostname connects to the network. Device type information at that time can be used to help identify which device is requesting a nickname. The system and method may also use any available device type information (e.g., Apple iPhone® 7 Max) as a hostname.
[0085] (MAC stitching based on user behavior) The server 200's program 216 and / or the user device 300's program 316 may include a “user device identification program” or other similar program for identifying user devices connected to a Wi-Fi network based on usage parameters. Thus, the function of identifying user devices on a Wi-Fi network may be performed by the server 200, the user device 300 itself, or a combination of the server 200, the user device 300, and / or other systems or devices in the network on which the communication device (e.g., the user device) is operating. In some cases, the server 200 may be configured to download one or more applications or programs (e.g., user device identification programs) to one or more user devices 300 to perform at least some of the steps described herein.
[0086] User device identification programs according to various embodiments of this disclosure may be stored in a non-temporary computer-readable storage medium (e.g., memory 210, 310, etc.). The user device identification program may have computer-readable code configured to program a server 200, user device 300, etc., to perform a specific function, which may be assisted by a processor 202, 302, or other suitable processing device.
[0087] In some embodiments, a user device identification program may be configured to associate two different device identifiers when it is determined with at least a reasonable degree of certainty that the device identifiers actually represent the same device (e.g., user device 300). Thus, by analyzing usage parameters related to the use of a user device on a Wi-Fi network, a user device identification program may be configured to determine when two different device identifiers (e.g., MAC addresses) obtained at two different points in time are likely to represent a situation in which the device identifiers were changed. For example, by analyzing usage parameters (e.g., the type or category of apps installed on the user device, app usage, ports open on the user device, user browsing patterns, Bluetooth® communication information, Wi-Fi access points used, etc.), the systems and methods of the present disclosure are configured to process these usage parameters to determine the identifier of a user device and to determine whether a first user device identifier can be associated with a second user device identifier.
[0088] According to various embodiments of this disclosure, user device identification programs (e.g., programs 216, 316) may be deployed in various forms on a server 200, one or more user devices 300, a network management system, a control system, a router, a modem, etc. Each version of the software / firmware for detecting user device identity (and other functions described throughout this disclosure) may include any functionality suitable for operation in various parts of the network, as described in this disclosure and understood by a person skilled in the art who is knowledgeable about this disclosure. For example, some processing functions may be incorporated into (or deployed) one or more access points (APs) of a Wi-Fi network, a router or modem, a cloud, an internet device (e.g., a server), offline (e.g., for non-real-time operation), online in real-time, etc.
[0089] According to some embodiments, the systems and methods of the Disclosure may be configured to perform a process for stitching MAC addresses (or other user device identifiers) based on the user behavior of one or more user devices. In one generalized embodiment, the systems and methods of the Disclosure may include the step of monitoring one or more user devices operating on a Wi-Fi network. The systems and methods may also include the step of analyzing usage parameters for each of the one or more user devices. The systems and methods may then identify one or more user devices based on these usage parameters.
[0090] In some embodiments, the system and method may include obtaining a device identifier associated with each of one or more user devices, and then associating each device identifier of one or more user devices with an operational identity based on usage parameters. For example, the device identifier associated with each of one or more user devices may be a Media Access Control (MAC) address. The system and method may also include a) the step of detecting when a new MAC address has been retrieved for an unidentified user device operating on a Wi-Fi network; b) the step of analyzing the current usage parameters of the unidentified user device; and c) the step of comparing the current usage parameters of the unidentified user device with the usage parameters of one or more previously identified user devices. In response to determining that the current usage parameters match those of one or more previously identified user devices, the process may include the step of stitching the new MAC address with the MAC address of the corresponding previously identified user device. In response to determining that the current usage parameters do not match those of one or more previously identified user devices, the process may include the step of tagging the unidentified user device as a new device to be monitored on the Wi-Fi network.
[0091] Furthermore, the systems and methods of this disclosure may also include the steps of a) analyzing usage parameters for each of one or more user devices over time, and b) creating one or more behavioral models associated with one or more users based on the usage parameters analyzed over time. Each behavioral model may represent the usage patterns of each user, depending on how the user uses the user devices. Alternatively, if a user uses multiple devices, the behavioral models may be associated with multiple devices to represent a single user using common usage behaviors on each device. The step of analyzing usage parameters over time may include utilizing machine learning techniques to create one or more behavioral models. Systems and methods according to some embodiments may further include the steps of a) assigning one or more unique user identifiers to represent one or more users, and b) associating one or more unique user identifiers with one or more behavioral models. The systems and methods may also include the step of retraining one or more behavioral models based on changes in the usage parameters of each corresponding user.
[0092] For example, the “Usage Parameters” described in this disclosure may include any type or category of user-based behavioral patterns and / or device-based behavioral patterns. Usage parameters may relate to the identity of one or more applications (“Apps”) installed on one or more user devices. Usage parameters may also relate to app usage information. For example, app usage information may include one or more of the following: a) how often each app is used, b) the amount of time the user spends on each app, c) the types of communications associated with app usage, and d) the time periods when the user uses the apps. Usage parameters may also relate to the identity of one or more websites or domains accessed by one or more user devices.
[0093] According to some embodiments, the systems and methods of the Disclosure may perform a step of refining the identity of one or more user devices based on weighted values of several metrics, such as the metrics described above. For example, the weighted metrics may be associated with a) the identity of one or more installed apps, b) app usage information, and / or c) browsing patterns. The weights may be associated with the uniqueness factors of each metric. As described in the Disclosure, user devices may include smartphones, computers, laptops, tablets, smart TVs, Internet of Things (IoT) devices, and / or media players. Usage parameters may be device-based behaviors, which may include a) Wi-Fi access point usage, b) Wi-Fi network connection patterns, c) Bluetooth®-related transmissions, d) device port usage, and / or other behaviors.
[0094] Identifying a user device may be based on its MAC address, user behavior, and other important unique indicators. The systems and methods of this disclosure may be configured to leverage and utilize user behavior of a device to uniquely identify the device. Note that user behavior is independent of the device vendor and is an aspect that cannot be controlled or altered by the device vendor. Thus, this disclosure is configured to identify a user device by leveraging behavioral patterns. Note that one user may use multiple devices, and / or a single device may be used by one or more users. Embodiments of this disclosure are configured to create a model of this user behavior and associate the behavior with one or more devices for identification. As described above, each user device may have a MAC address that can be stitched to the MAC address of another user device, based on the knowledge that MAC addresses may be changed or randomized, and / or that user behavior may be associated with one or more devices.
[0095] In some embodiments, user behavior may be categorized into the following behavioral buckets (or categories). These parameters may be key indicators of each of the user's areas of interest (or work) and may be unique to each user. It should also be noted that analyzing user behavior over time can lead to finer adjustments to each user's behavior and create a better understanding of their habits or patterns. These may distinguish users from other users who have different interests, patterns, behaviors, habits, etc. It may also be possible to identify when one or more visitors use the Wi-Fi network. These visitors may not typically be members of a family or work group. This disclosure may also store records of these visitors' usage behavior.
[0096] User behavior buckets (or categories) may include the following: 1) Application (App) Installation Information: This information may include apps installed on the user's device. For example, this information may include the category of the app (e.g., shopping apps, game apps, news apps, podcast apps, activity / hobby apps, business-related apps such as LinkedIn, health or medical apps, etc.). This information may also include a list of specific apps that have been installed. Additional information may include whether these apps were pre-installed on the user's device or downloaded and installed after being listed on the Market. 2) App Usage / Activity Information: This information may include classifications based on a) frequency of use, b) duration of use for each app, c) type of communication used by the app, and d) time of day when the app was used. This allows devices to be categorized based on the user's app usage patterns. 3) Browsing Patterns: This category may include a) a list of websites and / or domains actually accessed, and b) the type or category of websites and / or domains accessed. Examples of website types or categories may include finance / stocks, school / academics, entertainment, etc. This allows devices to be classified based on the user's interests or areas of interest. 4) Wi-Fi access point used: This information indicates the Wi-Fi access point to which the user device connects, and the patterns of connection and disconnection. For example, the systems and methods of this disclosure may be configured to detect in most cases that the user device is using a particular access point in the home and to indicate where in the home the user is likely to be located within the coverage area of this access point. This categorizes the device based on its physical location and movement within the home. 5) Open ports on the device: This information indicates the ports that are open and used accordingly for communication along specific routes to access various services (e.g., websites). 6) Bluetooth®-based: This information indicates communication and transmission patterns related to the Bluetooth® classification. This may include clock skew information, Bluetooth® Low Energy (BLE) status, piconet, traffic patterns, universally unique identifier (UUID) information, local name, etc. Please note that items 1-3 relate to specific (and / or intentional) "user-based" behaviors. Items 4-6 relate to specific "device-based" behaviors related to relevant protocols and settings associated with the device.
[0097] With respect to “1) App installation information,” this disclosure can identify the “category” of installed applications (apps). The advantage of this is that this method may be less vulnerable than checking a precise, detailed list of apps. Also, knowing the category of an app may be sufficient to identify a device in many embodiments. For example, a student might switch from one particular homework app to another, but this disclosure may still be able to identify the phone as the student's device. Embodiments described herein may be configured to identify users and devices by a unique collection of app categories present on the device (installed), rather than a specific list of apps themselves. For example, different app categories may include finance, school, shopping, games, entertainment, fitness, security, privacy, VPN, social networking, healthcare, productivity, enterprise (confluence), work-related, lifestyle, mobile apps, news apps, etc.
[0098] In some embodiments, the Disclosure may be configured to create a location category map. For example, this may be a category-based device map, which may be useful for a home and may also use machine learning. This category map may be supported by cloud-based global intelligence using analytics on a backend cloud server. Installed apps may be defined based on their uniqueness. The advantage of this approach is that if an app is replaced by another app in the same category, the systems and methods of the Disclosure may be configured to continue fingerprinting the device. It should be noted that app-based stitching can often be fragile. For example, if a home has only one device in a category and MAC addresses are randomized, a new MAC is likely to be the only device with that app category.
[0099] In some embodiments, the disclosure may include a category building system that uses analytics. For example, a cloud-based server may maintain a list of categories and apps belonging to those categories. These categories may be continuously updated and enhanced based on backend research, analytics, algorithms, etc., running on the backend cloud server. The system of the disclosure may use telemetry data collected from current customer households associated with installed apps. The category classification technology may look for the purpose of the app and the category of the app as specified by the app store. The system may build its own in-house intelligence or partner with an external vendor to assign categories to individual mobile apps.
[0100] In some cases, the processing of app category characteristics may be configured to make each app category more specific, tagging it with geolocation, household size, number of devices, device type, product license type, and other demographics that can be inferred from router visibility. Data analysis and algorithm modules may be configured to run on a backend cloud system. In some cases, processing app category characteristics may be configured so that app categorization also takes app security assessment into account. This may involve working with a third-party app security scoring database of internal intelligence.
[0101] Each app may have associated weights, which may be based on how unique it is to a device. These weights may be location / household specific and / or global trends learned from backend analytics processes. Another characteristic to consider is whether the app is a standard pre-installed app on a particular device. This is similar to using the device type as a stitching parameter, but may be based on pre-installed default apps.
[0102] Figure 6 shows an example of a screenshot 600 of a user device relating to the types of mobile apps. In this example, as shown in the "Category" section of screenshot 600, apps on the user device may be categorized into specific categories such as "Health & Fitness".
[0103] 2) App Usage / Activity Information may be monitored to monitor the usage of each app on the user's device. Usage information may include a) frequency of use, b) time spent on the app, c) type of communication, d) time of day, etc. This information may be used to classify user devices based on the user's app usage patterns. This may differ from general traffic analysis information, and as a result, traffic analysis may be in the context of a specific application.
[0104] Each app may have a weight associated with it based on how unique it is to the device itself or to other devices in a set (e.g., devices operating on a Wi-Fi network, devices across a region, state, or country). The systems and methods of this disclosure may be configured to utilize global intelligence for this weighting. This information may be based on data collected in the backend regarding the relationship between apps and device types. For example, a particular app may be found on only one device in a home, or across multiple homes in a certain area. The uniqueness of these apps may be related to school apps (e.g., Google Classroom, Schoology, etc.).
[0105] Regarding "3) Browsing Patterns," we analyze classifications based on browsing patterns. This may include the order in which domains / websites are accessed, the types of content categories accessed, the number of times a website is revisited, and the rating of accessed websites (e.g., graylist, blacklist, etc.). Embodiments described herein may use any of the many existing website classification tools (e.g., WebPulse, Brightcloud, Akamai, etc.). The following are some examples of thematic categories: a) Real Estate, b) Finance, c) Shopping, d) Travel, e) Adult, f) Dating, g) Religion, h) Law, i) Health, j) Automobiles, etc. Note that some of the browsing pattern information may be device-based rather than user-based. For example, the device-based aspect of a browsing pattern may include one or more protocols that can be used by a device. For example, a fitness tracking app may rely heavily on BLE technology.
[0106] With respect to "4) Wi-Fi access points used," this disclosure may include analysis to detect access point scanning patterns, radio signal data (e.g., radio transmission patterns), fingerprinting based on firmware clock skew, sequence numbers associated with Wi-Fi packets, etc. In some cases, there may be no standards around the request / response process of active probing. Timing between probe requests may be constant for the driver. Detected timing patterns can be used to fingerprint the device. Also, unique characteristics may be detected from transient signals at the start of packet transmission by the systems and methods described herein. These characteristics may be specific to a particular vendor, device, model, etc. Furthermore, the systems and methods may use machine learning techniques to measure Tx and Rx clock drift and fingerprint the device. The systems may also utilize TCP timestamps. In some embodiments, the systems and methods may analyze clock skew specific to the stack.
[0107] The sequence number, in this implementation, might appear as a huge counter that may roll over from time to time. Over time, different devices may reside in different regions of the overall sequence number space. Even if a device is out of reach (e.g., the user leaves the house) and sending Wi-Fi traffic, over a span of time, the device may remain in a range of the sequence number space that distinguishes it from other devices in the home. The concept of sequence numbers may apply to other communication methods, including Bluetooth®, Ethernet®, or perhaps IP packets with sequence numbers, although those sequence numbers are specific to a conversation, perhaps they come from a single counter or resume from where they left off, etc. It should be noted that the sequence number may be reset after a reboot, which may happen with firmware updates that may cause a change in the MAC address. However, for something like daily MAC randomization, the system would not normally reboot to cause this.
[0108] With respect to "5) Open ports on the device," the system and methods of this disclosure may observe the open ports and services running on the device. The system may determine vulnerabilities found on the device, the security level and security policies on the device, violations detected on the device, etc. For example, some services may include FTP, SSH, HTTP, etc. Ports may include, for example, 8890, 20, 21, etc.
[0109] With respect to "6) Bluetooth®-based," the system and method may analyze Bluetooth® pattern-based classification. This may include firmware clock skew-based fingerprinting, BLE advertisements, device piconets, traffic patterns, etc. Each set of devices (e.g., clients) may have a unique clock skew that remains constant. However, if a device leaves the Wi-Fi system's premises and then returns, the device may end up with a different slot or skew.
[0110] Bluetooth® processing may include periodic packets from the device, a fully local name that uniquely identifies the device, and fields such as a UUID. The advertisement interval is unique to the device and can be used for identification, but this may change when the device leaves the premises and returns. Bluetooth® connections may also be based on information created by the device at startup or when the MAC address changes. This category may also include connection / disconnection information, conversation patterns, time-based patterns, size-based patterns, etc.
[0111] In some embodiments, the systems and methods of the Disclosure (e.g., using user device identification programs 216, 316) may be configured to calculate behavioral patterns based on weighted values associated with each of the above parameters. The weighted values may be based on the level of uniqueness with respect to other devices. For example, the weighting may be high if the parameter value is unique to a device (e.g., within a household or across a larger population) and low if two or more devices have the same value. In one example, it may be determined that a small percentage of the population (e.g., only one person in a household) may install a particular app on a user device, use a particular app regularly, use the app at a specific time of day, and use the app while accessing a Wi-Fi network via a specific access point in the household. Compound parameters may be used to learn unique characteristics.
[0112] Figure 7 is Table 700, which shows an example of the weights of different parameters for a user device. Note that Table 700 shows the uniqueness of certain aspects of an app and the lack of uniqueness of other aspects of the app. Therefore, aspects with higher weights can identify the user device with greater specificity.
[0113] The systems and methods of this disclosure may be configured to learn specific usage patterns or characteristics of different user devices operating on a section of a network (e.g., a Wi-Fi network) using the above classification parameters in conjunction with machine learning techniques. The systems and methods may also be configured to create a behavioral model for each user. Each model may be assigned a unique user ID that identifies the user of the device. A device ID may also be assigned to identify the user device. The device ID may be distinct from and stored separately from the MAC address associated with the user device. However, if it is determined that they are related to the same device, the disclosure may be configured to stitch these two identification fields together. Thus, the disclosure may be configured to create and maintain the MAC address and device ID and map the two together with other stitched connections.
[0114] When a device first comes online on a Wi-Fi system with a new MAC address, its behavior may be recorded for a sufficient period (e.g., 3-5 minutes) to determine specific usage patterns. This current behavior can be compared to previously learned models and rules for each of the other user devices on the network. If the behavior matches an existing model, the MAC address-device ID mapping may be updated. Additionally, the old and new MAC addresses may be stitched together, if necessary, to maintain a consistent identity for a single device (or multiple devices used by the same user). If the behavior does not match any existing model, the device may be tagged as a new device. The learning algorithm may then be started for this new device.
[0115] One of the advantages of this methodology can be seen in the following example: Whenever a user device's MAC address is randomized, its behavioral indicators may remain essentially the same. Embodiments of this disclosure are configured to take advantage of this fact to quickly identify a user device and map the randomized MAC address to a previously used MAC address (along with other device-related information).
[0116] It should be noted that users may use multiple devices. In this case, these devices may exhibit the same user behavior patterns. In this overlapping behavior, this disclosure may be configured to use device types (e.g., iPhone®, iPad®, etc.) to distinguish devices from each other and enable correct mapping.
[0117] The systems and methods of this disclosure may be configured to utilize machine learning techniques for analyzing usage behavior over time in order to train one or more behavioral models. These models can be retrained using continuous learning to update the models as needed. The machine learning models may then be used for analysis to identify specific user devices based on usage patterns. The systems described herein can continue to learn (using machine learning techniques) and adapt to any changes in user behavior. For example, a user may delete an app and no longer use it, or no longer visit a particular domain. In another example, a user (e.g., a sixth grader) may exhibit behavior regarding accessing sixth-grade level assignments. However, in the following year, the user's behavior may change (e.g., accessing seventh-grade level assignments), and this can be detected by the systems and methods of this disclosure.
[0118] For example, one advantage of this disclosure is that a vendor (who may own specific software and protocol stacks on a device) can control the visibility of protocols and network traffic by encrypting one or more elements of network traffic. A vendor seeking to protect user privacy may make it extremely complex and difficult for adversaries to identify MAC addresses. However, this may result in a decrease in the performance of reputable network analysis systems that can rely on network protocol characteristics. Nevertheless, this disclosure is configured to overcome this encryption strategy by identifying based on user behavior, usage patterns, user habits, user trends, user models, etc. Thus, this disclosure leverages behavioral aspects of users that are typically unique and specific to a given user. Vendors will no longer be able to obscure or hide user behavior on a network (e.g., LAN, Wi-Fi network, etc.). The techniques or algorithms of this disclosure may be self-adaptive to allow the system to continuously learn new behavioral indicators and evolve to respond to any changes in user behavior.
[0119] Therefore, one key concept of this disclosure may be to use the usage parameters listed above, create a behavioral model around these usage parameters, and leverage this model to uniquely tag users. The systems and methods described herein may use user profiles to uniquely identify user devices owned or used by that user. Consequently, this disclosure may be configured to identify devices using the owner's behavioral profile.
[0120] Figure 8 is a flowchart illustrating one embodiment of a process 800 for identifying a user device based on usage behavior information. Based on device identification, the system and method may further be configured to stitch MAC addresses based on this usage behavior. As shown, process 800 includes the step of monitoring one or more user devices operating on a Wi-Fi network, as shown in block 802. Process 800 further includes the step of analyzing one or more usage parameters and operation parameters for each of the one or more user devices, as shown in block 804. Process 800 also includes the step of identifying one or more user devices based on one or more usage parameters and operation parameters in response to the randomization of Media Access Control (MAC) addresses of one or more user devices, as shown in block 806.
[0121] In some embodiments, process 800 may further include the steps of: a) obtaining a device identifier associated with each of one or more user devices; and b) associating the device identifier of each of the one or more user devices with an operational identity based on usage parameters. For example, the device identifier associated with each of the one or more user devices may be a Media Access Control (MAC) address. Process 800 may also include the steps of: a) detecting when a new MAC address has been retrieved for an unidentified user device operating on a Wi-Fi network; b) analyzing the current usage parameters of the unidentified user device; and c) comparing the current usage parameters of the unidentified user device with the usage parameters of one or more previously identified user devices. In response to determining that the current usage parameters match those of one of the previously identified user devices, process 800 may perform the step of stitching the new MAC address with the corresponding MAC address of the previously identified user device. Alternatively, in response to determining that the current usage parameters do not match those of one or more previously identified user devices, process 800 may perform the step of tagging the unidentified user device as a new device to be monitored on the Wi-Fi network.
[0122] Process 800 may include a step of analyzing usage parameters for each of one or more user devices over time. Next, based on the usage parameters analyzed over time, Process 80 may include a step of creating one or more behavioral models associated with one or more users, so that each behavioral model represents the usage patterns of each user, depending on how the user uses at least one of the user devices. In some embodiments, the step of analyzing usage parameters over time may include utilizing machine learning techniques to create one or more behavioral models. Process 800 may further include a) assigning one or more unique user identifiers to represent one or more users, and b) associating one or more unique user identifiers with one or more behavioral models. In some embodiments, the step of analyzing usage parameters over time may include utilizing machine learning techniques to create one or more behavioral models.
[0123] In additional embodiments, the usage parameters described herein may relate to the identity of one or more apps installed on one or more user devices. The usage parameters may also relate to app usage information, which may include a) the frequency of use of one or more apps, b) the time spent on each of the one or more apps, c) the type of communication associated with app usage, d) the time of day of app usage, and / or other information. Furthermore, the usage parameters may relate to the identity of one or more websites or domains accessed by one or more user devices.
[0124] In some embodiments, process 800 may include a step of refining the identity of one or more user devices based on weighted values of several metrics. The metrics may include a) the identity of one or more installed apps, b) app usage information, c) browsing patterns, and / or other metrics. The weighted values may, for example, relate to the uniqueness of each metric. User devices referred to herein may include smartphones, computers, laptops, tablets, smart TVs, Internet of Things (IoT) devices, media players, or other suitable devices that communicate with a Wi-Fi network. In some implementations, usage parameters may relate to device-based behaviors such as a) Wi-Fi access point usage, b) Wi-Fi network connection patterns, c) Bluetooth®-related transmissions, d) device port usage, and / or other device-related behaviors.
[0125] (Conclusion) It will be understood that some embodiments described herein may include, in combination with certain non-processor circuits, one or more general-purpose or specialized processors ("one or more processors") such as microprocessors; a central processing unit (CPU); customized processors such as digital signal processors (DSPs), network processors (NPs) or network processing units (NPUs), graphics processing units (GPUs); field-programmable gate arrays (FPGAs), etc., along with specific stored program instructions (including both software and firmware) for their control, in order to carry out some, most, or all of the functions of the methods and / or systems described herein. Alternatively, some or all of the functions may be implemented by a state machine that does not store program instructions, or in one or more application-specific integrated circuits (ASICs) in which each function or certain combinations of functions are implemented as custom logic or circuitry. Of course, combinations of the approaches described above may also be used. For some of the embodiments described herein, the corresponding devices, consisting of hardware and optionally software, firmware, and combinations thereof, may also be referred to as “circuits configured or adapted to perform a set of operations, steps, methods, processes, algorithms, functions, techniques, etc.” for digital and / or analog signals, as described herein for various embodiments.
[0126] Furthermore, some embodiments may include a non-temporary computer-readable storage medium having computer-readable code stored thereon for programming computers, servers, appliances, devices, processors, circuits, etc., each of which may include a processor, to perform functions as described and claimed herein. Examples of such computer-readable storage media include, but are not limited to, hard disks, optical storage devices, magnetic storage devices, ROM (read-only memory), PROM (programmable read-only memory), EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), flash memory, etc. When stored on a non-temporary computer-readable medium, software may include instructions that can be executed by a processor or device (e.g., any type of programmable circuit or logic), and in response to such execution, the processor or device can be caused to perform a set of operations, steps, methods, processes, algorithms, functions, techniques, etc., described herein for various embodiments.
[0127] While this disclosure has been illustrated and described herein with reference to preferred embodiments and specific examples thereof, it will be readily apparent to those skilled in the art that other embodiments and examples may perform similar functions and / or achieve similar results. All such equivalent embodiments and examples are in the spirit and scope of this disclosure, contemplated thereby, and intended to be covered by the following claims.
Claims
1. Monitoring (802) one or more user devices (16) operating on a Wi-Fi network (10, 30, 32, 33); analyzing (804) one or more usage and operational parameters for each of the one or more user devices (16); In response to the media access control (MAC) address randomization (806) of the one or more user devices (16), identifying the one or more user devices (16) based on the one or more of usage and operational parameters; The method (800) includes:
2. retrieving a device identifier associated with each of the one or more user devices (16); associating the device identifier of each of the one or more user devices (16) with an operational identity based on the one or more of a usage parameter and an operational parameter; The method (800) of claim 1, further comprising:
3. Detecting when a new MAC address is found for an unidentified user device (16) operating on said Wi-Fi network (10, 30, 32, 33); analyzing one or more of the current usage and operational parameters of the unidentified user device (16); comparing the one or more current usage and operational parameters of the unconfirmed user device (16) with one or more current usage and operational parameters of the one or more previously identified user devices (16); The method (800) of claim 2, further comprising:
4. analyzing the one or more usage and operational parameters for each of the one or more user devices (16) over time; creating one or more behavioral models associated with one or more users based on the one or more usage and operational parameters analyzed over time, each behavioral model representing a respective user's pattern according to how the respective user uses at least one of the one or more user devices (16); The method (800) of any one of claims 1 to 3, further comprising:
5. 5. The method (800) of claim 4, wherein analyzing the one or more of usage and operational parameters over time comprises utilizing machine learning techniques to create the one or more behavioral models.
6. The parameters used are: the identities of one or more apps installed on the one or more user devices (16); one or more categories of apps installed on the one or more user devices (16); app usage information, the app usage information including one or more of a frequency of use of one or more apps, a time spent on each of the one or more apps, a type of communication associated with app usage, and a time period of app usage; the identity of one or more websites or domains accessed by said one or more user devices (16); categories of the websites or domains accessed by the one or more user devices (16); and one or more of open ports, network services, security levels, security policies, and potential security vulnerabilities of the one or more user devices (16); The method (800) of any one of claims 1 to 3, wherein the method (800) relates to one or more of:
7. refining the identity of each of the one or more user devices (16) based on weighted values of a plurality of metrics including one or more of the identity of one or more installed apps, app usage information, and browsing patterns, the weighted values relating to the uniqueness of each of the metrics; The method (800) of any one of claims 1 to 3, further comprising:
8. refining the identity of each of the one or more user devices (16) based on both usage parameters of the device and the determined device type of the device. The method (800) of any one of claims 1 to 3, further comprising:
9. The method (800) of any one of claims 1 to 3, wherein the one or more user devices comprise one or more smartphones, computers, laptops, tablets, smart TVs, Internet of Things (IoT) devices, and / or media players.
10. 4. The method (800) of claim 1, wherein the usage parameters are related to device-based behavior, the device-based behavior including one or more of Wi-Fi access point usage, Wi-Fi network connection patterns, Bluetooth-related transmissions, and device port usage.
11. calculating a confidence score based on a relationship between the current one or more of the usage parameters and the operating parameters and the stored one or more of the usage parameters and the operating parameters; determining whether the confidence score exceeds a predetermined threshold; The method (800) of any one of claims 1 to 3, further comprising:
12. 12. The method (800) of claim 11, wherein the relationship comprises one or more of: matching device behavior factors, matching device features, uniqueness of matching device features, weighted sum of device features, number of matching device features, and machine learning (ML) model of device matching techniques.
13. 4. The method (800) of claim 1, wherein the operational parameters are networking metadata including information obtained via one or more of Address Resolution Protocol (ARP), Logical Link Control (LLC), Internet Control Message Protocol (ICMP), ICMP Version 6 (ICMPv6), Bootstrap Protocol (BOOTP), Network Time Protocol (NTP), Transmission Control Protocol (TCP), Transport Layer Security (TLS), Dynamic Host Configuration Protocol (DHCP), DHCP Version 6 (DHCPv6), Domain Name System (DNS), Multicast DNS (mDNS), User Agent, Universal Plug and Play (UPNP), Shared Serial Data Protocol (SSDP), device capability information, port information, protocol information, and 5-tuple Internet Protocol (IP) data.
14. transferring settings and controls to a given device (16) based on said identifying. The method (800) of any one of claims 1 to 3, further comprising:
15. one or more processors (202); a memory (210) storing instructions that, when executed, cause the one or more processors to perform a method (800) according to any one of claims 1 to 3; A server (200) comprising: