IoT device identification using a machine learning model of packet flow behavior

A data appliance with an IoT server and module uses machine learning to classify IoT device behavior and enforce policies, addressing the inadequacies of existing malware detection in diverse networks and enhancing security by managing IoT devices.

JP2026001101APending Publication Date: 2026-01-06PALO ALTO NETWORKS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025160064
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-26
Filing Date
2025-09-26
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

Existing malware detection techniques are inadequate for diverse computing environments and are easily evaded by malware authors, posing a continuous threat to computer systems, especially in networks with unmanaged IoT devices that lack endpoint protection.

Method used

Implementing a data appliance with an IoT server and module to passively monitor network traffic, identify IoT devices, and provide AAA support, using machine learning models to classify device behavior and enforce policies, thereby enhancing security in heterogeneous networks.

Benefits of technology

Effectively detects and mitigates malicious activity in networks with unmanaged IoT devices, improving security and reducing potential harm by dynamically managing network access and enforcing fine-grained policies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026001101000001_ABST
    Figure 2026001101000001_ABST
Patent Text Reader

Abstract

Identifying internet of things (IoT) devices with packet flow behavior is disclosed, including by using a machine learning model.SOLUTION: Information associated with network communications of an IoT device is received. A determination is made whether the IoT device has been previously classified. In response to determining that the IoT device has not been previously classified, a determination is made that the probability match of the IoT device to the behavior signature exceeds a threshold. Based at least in part on the probabilistic match, a classification of the IoT device is provided to a security appliance configured to apply a policy to the IoT device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Malicious individuals attempt to compromise computer systems in a variety of ways. As one example, such individuals may embed or otherwise contain malicious software (“malware”) in email attachments and send them to unsuspecting users. When executed, the malware compromises the victim's computer and can perform additional illicit tasks (e.g., extracting sensitive data, propagating to other systems, etc.). Various approaches can be used to harden computers against exposure to such and other risks. Unfortunately, existing approaches to protecting computers are not necessarily suitable for all computing environments. Furthermore, malware authors continually adapt their techniques to evade detection. Thus, there is a continuing need for improved techniques for detecting malware and preventing its harm in a variety of situations. [Brief explanation of the drawings]

[0002] Various embodiments of the present invention are disclosed in the following detailed description and accompanying drawings. [Figure 1] FIG. 1 illustrates an example of an environment in which malicious activity may be detected and its harm mitigated. [Figure 2A] FIG. 2A illustrates one embodiment of a data appliance. [Figure 2B] FIG. 2B is a functional diagram of the logical components of one embodiment of a data appliance. [Figure 2C] FIG. 2C shows one exemplary event path between the IoT Server and the IoT Module. [Figure 2D] FIG. 2D shows an example of a device discovery event. [Figure 2E]FIG. 2E shows an example of a session event. [Figure 2F] FIG. 2F illustrates one embodiment of an IoT module. [Figure 2G] FIG. 2G illustrates one exemplary method for implementing IoT device analytics. [Figure 3] FIG. 3 illustrates one embodiment of a process for passively providing AAA support to IoT devices in a network. [Figure 4A] FIG. 4A illustrates an example of a RADIUS message sent by an IoT server to an AAA server on behalf of an IoT device, in accordance with various embodiments. [Figure 4B] FIG. 4B illustrates an example of a RADIUS message sent by an IoT server to an AAA server on behalf of an IoT device, in accordance with various embodiments. [Figure 4C] FIG. 4C illustrates an example of a RADIUS message sent by an IoT server to an AAA server on behalf of an IoT device, in accordance with various embodiments. [Figure 5] FIG. 5 shows one embodiment of an IoT module. [Figure 6] FIG. 6 shows an example of a process for classifying IoT devices. [Figure 7A] FIG. 7A shows an example firewall rule. [Figure 7B] FIG. 7B shows an example firewall rule. [Figure 8] FIG. 8 shows a portion of an exemplary interface. [Figure 9] FIG. 9 shows a portion of an example interface. [Figure 10] FIG. 10 shows a portion of an exemplary interface. [Figure 11]FIG. 11 illustrates one example of a process for generating policies to apply to communications involving IoT devices. [Figure 12] FIG. 12 shows a sequence of packet inter-arrival times. [Figure 13A] FIG. 13A shows the training data. [Figure 13B] FIG. 13B shows the results of training a model using the dataset shown in FIG. 13A. [Figure 14A] FIG. 14A shows two examples of devices that match the device behavior signature. [Figure 14B] FIG. 14B shows two examples of devices that do not match the device behavior signature. [Figure 15] FIG. 15 illustrates a process for classifying an exemplary IoT device. DETAILED DESCRIPTION OF THE INVENTION

[0003] The present invention can be implemented in numerous ways, including as a process, an apparatus, a system, a composition of matter, a computer program product embodied on a computer-readable storage medium, and / or a processor, such as instructions stored on a memory and / or a processor configured to execute instructions stored and / or provided by a memory coupled to the processor. These implementations, or any other form the present invention may take, may be referred to herein as techniques. In general, the order of steps in disclosed processes may be varied within the scope of the present invention. Unless otherwise specified, components, such as a processor or memory, described as configured to perform a task may be implemented as general-purpose components temporarily configured to perform the task at a given time, or as specific components manufactured to perform the task. As used herein, the term “processor” refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.

[0004] A detailed description of one or more embodiments of the present invention is provided below along with accompanying figures that illustrate the principles of the invention. While the present invention will be described in connection with such embodiments, the present invention is not limited to any embodiment. The scope of the present invention is limited only by the claims, and the present invention encompasses numerous alternatives, modifications, and equivalents. Numerous specific details are set forth in the following description to provide a thorough understanding of the present invention. These details are provided for the purpose of example, and the present invention may be practiced according to the claims without some or all of these specific details. For the purposes of clarity, technical material known in the art related to the present invention has not been described in detail so as not to unnecessarily obscure the present invention.

[0005] I. Overview

[0006] A firewall generally allows authorized communications to pass through the firewall while protecting the network from unauthorized access. A firewall is typically a device, set of devices, or software running on devices that provides firewall functionality for network access. For example, a firewall can be integrated into the operating system of a device (e.g., a computer, smartphone, or other type of network-enabled device). Firewalls can also be integrated or run as software applications on various types of devices or security devices, such as computer servers, gateways, network / routing devices (e.g., network routers), or data appliances (e.g., security appliances or other types of special-purpose devices), and in some implementations, certain operations can be implemented in special-purpose hardware, such as an ASIC or FPGA. do.

[0007] Firewalls typically deny or allow network transmissions based on a set of rules. These sets of rules are often referred to as policies (e.g., network policies or network security policies). For example, a firewall can filter inbound traffic by applying a set of rules or policies to prevent unwanted external traffic from reaching a protected device. A firewall can also filter outbound traffic by applying a set of rules or policies (e.g., allow, block, monitor, notify, or log, and / or other actions that may be specified in a firewall rule or firewall policy, which may be triggered based on various criteria, as described herein). A firewall can also filter local network (e.g., intranet) traffic by similarly applying a set of rules or policies.

[0008] Security devices (e.g., security appliances, security gateways, security services, and / or other security devices) may perform various security operations (e.g., firewalls, anti-malware, intrusion prevention / detection, proxies, and / or other security functions), network functions (e.g., routing, quality of service (QoS), workload balancing of network-related resources, and / or other network functions), and / or other security and / or network-related functions. For example, routing may be performed based on source information (e.g., IP addresses and ports), destination information (e.g., IP addresses and ports), and protocol information.

[0009] Basic packet filtering firewalls filter network communication traffic by inspecting individual packets sent over the network (e.g., stateless packet filtering firewalls, or first-generation firewalls). Stateless packet filtering firewalls typically inspect the individual packets themselves and then apply rules based on the inspected packets (e.g., using a combination of the packet's source and destination address information, protocol information, and port numbers).

[0010] Application firewalls can also perform application-layer filtering (e.g., using an application-layer filtering firewall or a second-generation firewall that functions at the application level of the TCP / IP stack). Application-layer filtering firewalls or application firewalls can generally identify certain applications and protocols (e.g., web browsing using the Hypertext Transfer Protocol (HTTP), Domain Name System (DNS) requests, file transfers using the File Transfer Protocol (FTP), and various other types of applications and other protocols, such as Telnet, DHCP, TCP, UDP, and TFTP (GSS)). For example, an application firewall can block unauthorized protocols that attempt to communicate on standard ports (e.g., unauthorized / out-of-policy protocols that attempt to sneak through by using a non-standard port for that protocol can generally be identified using an application firewall).

[0011] Stateful firewalls can also perform stateful-based packet inspection, where each packet is inspected within the context of the set of packets associated with its network outbound packet flow. This firewall technique is commonly referred to as stateful packet inspection because it keeps a record of all connections passing through the firewall and can determine whether a packet is the start of a new connection, part of an existing connection, or an invalid packet. For example, the state of a connection can itself be one of the criteria that triggers a rule in a policy.

[0012] Advanced or next-generation firewalls can perform stateless and stateful packet filtering and application layer filtering, as described above. Next-generation firewalls can also implement additional firewall technologies. For example, certain newer firewalls, often referred to as advanced or next-generation firewalls, can also identify users and content. In particular, certain next-generation firewalls have expanded the list of applications that they can automatically identify to thousands of applications. Examples of such next-generation firewalls are commercially available from Palo Alto Networks (e.g., the Palo Alto Networks PA Series firewalls). For example, Palo Alto Networks next-generation firewalls use various identification technologies to enable enterprises and service providers to identify and control applications, users, and content—not just ports, IP addresses, and packets. Various identification technologies include Application ID (App-ID) for precise application identification, User ID (User-ID) for user identification (e.g., User ID), Content ID (Content-ID) for real-time content scanning (e.g., to control web surfing and restrict data and file transfers), and Device ID (e.g., for identifying IoT device types). These identification technologies allow enterprises to securely enable application usage using business-relevant concepts instead of following the traditional approach provided by traditional port-blocking firewalls.Additionally, special-purpose hardware for next-generation firewalls (e.g., implemented as dedicated devices) generally provides higher performance levels for application inspection than software running on general-purpose hardware (e.g., security appliances from Palo Alto Networks, Inc., which utilize dedicated, function-specific processing that is tightly integrated with a single-pass software engine to minimize latency while maximizing network throughput, as in the case of Palo Alto Networks' PA Series Next-Generation Firewalls).

[0013] Advanced or next-generation firewalls can also be implemented using virtualized firewalls. An example of such a next-generation firewall is commercially available from Palo Alto Networks (Palo Alto Networks firewalls are deployed on VMware® ESXi). TM and NSX TM , Citrix® Netscaler SDX TM It supports a variety of commercial virtualization environments, including KVM / OpenStack (Centos / RHEL, Ubuntu®), and Amazon Web Services (AWS). For example, the virtualized firewall can support similar or identical next-generation firewall and advanced threat prevention capabilities available in physical form factor devices, allowing enterprises to safely enable the influx of applications to private, public, and hybrid cloud computing environments. Automation capabilities such as VM monitoring, dynamic address groups, and REST-based APIs enable enterprises to dynamically monitor VM changes and update security policies with that context, thereby eliminating policy lag that can occur when VMs change.

[0014] II. Example Environment

[0015] Figure 1 illustrates an example of an environment in which malicious activity is detected and its harm mitigated. In the example shown in Figure 1, client devices 104-108 are a laptop computer, a desktop computer, and a tablet (respectively) that reside on an enterprise network 110 of a hospital (also referred to as "Acme Hospital"). Data appliance 102 is configured to enforce policies regarding communications between client devices, such as client devices 104 and 106, and nodes outside enterprise network 110 (e.g., reachable via external network 118).

[0016] Examples of such policies include those governing traffic shaping, quality of service, and traffic routing. Other examples of policies include security policies, such as those requiring scanning for threats in incoming (and / or outgoing) email attachments, website content, files exchanged via instant messaging programs, and / or other file transfers. In some embodiments, data appliance 102 is also configured to enforce policies with respect to traffic remaining within enterprise network 110.

[0017] Network 110 also includes a directory service 154 and an authentication, authorization, and accounting (AAA) server 156. In the example shown in FIG. 1, directory service 154 (also called an identity provider or domain controller) utilizes the Lightweight Directory Access Protocol (LDAP) or other suitable protocol. Directory service 154 is configured to manage user identity and credential information. One example of directory service 154 is Microsoft's Active Directory server. Instead of an Active Directory server, other types of systems, such as a Kerberos-based system, can also be used, and the techniques described herein would be adapted accordingly. In the example shown in FIG. 1, AAA server 156 is a network admission control (NAC) server. AAA server 156 is configured to authenticate wired, wireless, and VPN users and devices to the network, evaluate and remediate devices for policy compliance before allowing access to the network, differentiate access based on role, and then audit and report on who is on the network. One example of AAA server 156 is Cisco's Identity Services Engine (ISE) server, which utilizes Remote Authentication Dial-In User Service (RADIUS). Other types of AAA servers, including those using protocols other than RADIUS, can be used with the techniques described herein.

[0018] In various embodiments, the data appliance 102 is configured to listen for communications (e.g., passively monitor messages) to / from the directory service 154 and / or the AAA server 156. In various embodiments, the data appliance 102 is configured to communicate (i.e., actively communicate messages) with the directory service 154 and / or the AAA server 156. In various embodiments, the data appliance 102 is configured to communicate with an orchestrator (not shown), which communicates (e.g., actively communicates messages) with various network elements, such as the directory service 154 and / or the AAA server 156. Other types of servers may also be included in the network 110 and may communicate with the data appliance 102 where applicable, and the directory service 154 and / or the AAA server 156 may also be omitted from the network 110 in various embodiments.

[0019] 1 is shown as having a single data appliance 102, a given network environment (e.g., network 110) may include multiple embodiments of data appliances, whether operating individually or in concert. Similarly, while the term "network" is generally referred to herein in the singular (e.g., as "network 110") for simplicity, the techniques described herein may be deployed in a variety of network environments of different sizes and topologies, including various mixes of networking technologies (e.g., virtual and physical), using various networking protocols (e.g., TCP and UDP) and infrastructures (e.g., switches and routers) across various network layers, where applicable.

[0020] Data appliance 102 may be configured to operate in cooperation with remote security platform 140. Security platform 140 may provide various services, including performing static and dynamic analysis on malware samples (e.g., via sample analysis module 124) and providing a list of signatures of known malicious files, domains, etc. to a data appliance, such as data appliance 102, as part of a subscription. As described in more detail below, security platform 140 may also provide information associated with the discovery, classification, management, etc. of IoT devices present in a network, such as network 110 (e.g., via IoT module 138). In various embodiments, the signatures, analysis results, and / or additional information (e.g., regarding samples, applications, domains, etc.) are stored in database 160. In various embodiments, security platform 140 includes one or more dedicated, commercially available hardware servers (e.g., having multi-core processors, 32G+ RAM, Gigabit network interface adapters, and hard drives) running a typical server-class operating system (e.g., Linux). Security platform 140 may be implemented across a scalable infrastructure including multiple such servers, solid-state drives, or other storage 158, and / or other applicable high-performance hardware. Security platform 140 may include several distributed components, including components provided by one or more third parties. For example, some or all of security platform 140 may be implemented using Amazon's Elastic Compute Cloud (EC2) and / or Amazon's Simple Storage Service (S3).Additionally, similar to data appliance 102, whenever security platform 140 is referred to as performing a task, such as storing data or processing data, it should be understood that a subcomponent or subcomponents of security platform 140 may cooperate (individually or in cooperation with third-party components) to perform that task. For example, security platform 140 may cooperate with one or more virtual machine (VM) servers to perform static / dynamic analysis (e.g., via sample analysis module 124) and / or IoT device functionality (e.g., via IoT module 138). One example of a virtual machine server is a physical machine including commercially available server-class hardware (e.g., multi-core processors, 32+ gigabytes of RAM, and one or more gigabit network interface adapters) running commercially available virtualization software such as VMware ESXi, Citrix XenServer, or Microsoft Hyper-V. In some embodiments, the virtual machine server is omitted. Furthermore, the virtual machine server may be under the control of the same entity that manages security platform 140, but may also be provided by a third party. As one example, the virtual machine servers may rely on EC2, with the remainder of security platform 140 provided by dedicated hardware owned by and under the control of the operator of security platform 140.

[0021] One embodiment of a data appliance is shown in FIG. 2A. The illustrated example is a representation of the physical components included in a data appliance 102, in various embodiments. Specifically, the data appliance 102 includes a high-performance multi-core central processing unit (CPU) 202 and random access memory (RAM) 204. The data appliance 102 also includes storage 210 (such as one or more hard disks or solid-state units). In various embodiments, the data appliance 102 stores (either in RAM 204, storage 210, and / or other suitable locations) information used to monitor the enterprise network 110 and implement the disclosed techniques. Examples of such information include application identifiers, content identifiers, user identifiers, requested URLs, IP address mappings, policy and other configuration information, signatures, hostname / URL categorization information, malware profiles, machine learning models, IoT device classification information, etc. The data appliance 102 may also include one or more optional hardware accelerators. For example, the data appliance 102 may include a cryptographic engine 206 configured to perform encryption and decryption operations, and one or more field programmable gate arrays 208 configured to perform matching, act as a network processor, and / or perform other tasks.

[0022] The functionality described herein as being performed by data appliance 102 may be provided / implemented in various ways. For example, data appliance 102 may be a dedicated device or set of devices. A given network environment may include multiple data appliances, each configured to service a particular portion of the network, which may cooperate to service a particular portion of the network, etc. The functionality provided by data appliance 102 may be integrated or executed as software on a general-purpose computer, computer server, gateway, and / or network / routing device. In some embodiments, at least some functionality described as being provided by data appliance 102 is instead (or in addition) provided to a client device (e.g., client device 104 or client device 106) by software executing on the client device. Functionality described herein as being performed by data appliance 102 may also be performed, at least in part, by or in cooperation with security platform 140, and / or functionality described herein as being performed by security platform 140 may also be performed, where applicable, at least in part, by or in cooperation with data appliance 102. As one example, various functions described as being performed by IoT module 138 may be performed by an embodiment of IoT server 134.

[0023] Whenever the data appliance 102 is described as performing a task, a single component, a subset of components, or all components of the data appliance 102 may cooperate to perform the task. Similarly, whenever a component of the data appliance 102 is described as performing a task, a subcomponent may perform the task and / or the component may perform the task in conjunction with other components. In various embodiments, portions of the data appliance 102 are provided by one or more third parties. Depending on factors such as the amount of computing resources available to the data appliance 102, various logical components and / or features of the data appliance 102 may be omitted, and the techniques described herein may be adapted accordingly. Similarly, additional logical components / features may be included in embodiments of the data appliance 102, as applicable. One example of a component included in the data appliance 102 in various embodiments is an application identification engine configured to identify applications (e.g., using various application signatures to identify applications based on packet flow analysis). For example, the application identification engine may determine the type of traffic a session involves, such as web browsing-social networking, web browsing-news, SSH, etc. Another example of a component included in data appliance 102 in various embodiments is IoT server 134, which is described in more detail below. IoT server 134 can take various forms, including as a standalone server (or set of servers), whether physical or virtualized, and may also be co-located / embedded with data appliance 102 (e.g., as shown in FIG. 1 ), where applicable.

[0024] 2B is a functional diagram of the logical components of one embodiment of a data appliance. The example shown is a representation of the logical components that may be included in data appliance 102 in various embodiments. Unless otherwise specified, the various logical components of data appliance 102 may generally be implemented in a variety of ways, including as a set of one or more scripts (e.g., written in Java, Python, etc., as applicable).

[0025] As shown, data appliance 102 includes a firewall and includes a management plane 212 and a data plane 214. The management plane is responsible for managing user interaction, such as by providing a user interface for setting policies and displaying log data, and the data plane is responsible for data management, such as by performing packet processing and session handling.

[0026] The network processor 216 is configured to receive packets from client devices, such as the client device 108, and provide them to the data plane 214 for processing. The flow module 218 creates a new session flow whenever it identifies a packet as part of a new session. Subsequent packets are identified as belonging to the session based on the flow lookup. If applicable, SSL decryption is applied by the SSL decryption engine 220. Otherwise, processing by the SSL decryption engine 220 is skipped. The decryption engine 220 helps the data appliance 102 inspect and control SSL / TLS and SSH encrypted traffic and, therefore, helps stop threats that might otherwise remain hidden within the encrypted traffic. The decryption engine 240 can also help prevent sensitive content from leaving the enterprise network 110. Decryption can be selectively controlled (e.g., enabled or disabled) based on parameters such as URL category, traffic source, traffic destination, user, user group, and port. In addition to decryption policies (e.g., specifying which sessions to decrypt), decryption profiles can be assigned to control various options of the session controlled by the policy, for example, the use of specific cipher suites and encryption protocol versions can be required.

[0027] The application identification (APP-ID) engine 222 is configured to determine the type of traffic a session involves. As one example, the application identification engine 222 may recognize a GET request in the received data and conclude that the session requires an HTTP decoder. In some cases, such as during a web browsing session, the identified application may change, and such changes may be noted by the data appliance 102. For example, a user may first browse a company wiki (categorized as "Web Browsing-Productivity" based on the URLs visited) and then browse a social networking site (categorized as "Web Browsing-Social Networking" based on the URLs visited). Different types of protocols have corresponding decoders.

[0028] Based on the determination made by the application identification engine 222, the packets are sent by the threat engine 224 to the appropriate decoder, which is configured to assemble the packets (which may be received out of order) into the correct order and extract the information. The threat engine 224 also performs signature matching to determine what should happen to the packets. If necessary, the SSL encryption engine 226 can re-encrypt the decrypted data. The packets are forwarded using the forwarding module 228 for forwarding (e.g., to a destination).

[0029] 2B, policies 232 are also received and stored in the management plane 212. The policies can include one or more rules, which can be specified using domain names and / or host / server names, and the rules can apply one or more signatures or other matching criteria or heuristics, such as for enforcing security policies on subscriber / IP flows, based on various extracted parameters / information from the monitored session traffic flows. An interface (I / F) communicator 230 is provided for management communications (e.g., via (REST) ​​APIs, messages, or network protocol communications, or other communication mechanisms). The policies 232 can also include policies for managing communications involving IoT devices.

[0030] III. IoT Device Discovery and Identification

[0031] Returning to FIG. 1 , assume that a malicious individual (e.g., using system 120) has created malware 130. The malicious individual wants vulnerable client devices to run copies of malware 130, compromising the client devices and causing them to become bots in a botnet. The compromised client devices can then be instructed to perform tasks (e.g., cryptocurrency mining, participating in denial-of-service attacks, and propagating to other vulnerable client devices), and to report information to or otherwise extract data from an external entity (e.g., command and control (C&C) server 150), as well as receive instructions from C&C server 150, if applicable.

[0032] Some client devices shown in FIG. 1 are commodity computing devices typically used within enterprise organizations. For example, client devices 104, 106, and 108 each run a typical operating system (e.g., macOS, Windows, Linux, Android, etc.). Such commodity computing devices are often provisioned and maintained by administrators (e.g., as corporate-issued laptops, desktops, and tablets, respectively) and often operate with user accounts (e.g., managed by a directory service provider (also called a domain controller) configured with user identity and certificate information). As one example, employee Alice is issued laptop 104, which she uses to access her ACME-related email and perform various ACME-related tasks. Other types of client devices (generally referred to herein as Internet of Things or IoT devices) are also increasingly present within networks and are often “unmanaged” by IT departments. Some such devices (e.g., teleconferencing devices) may be found across various different types of enterprises (e.g., as IoT whiteboards 144 and 146). Such devices may also be vertical specific. For example, infusion pumps and computed tomography scanners (e.g., CT scanner 112) are examples of IoT devices that may be found within a healthcare enterprise network (e.g., network 110), and a robotic arm is an example of a device that may be found within a manufacturing enterprise network. Additionally, consumer IoT devices (e.g., cameras) may also reside within enterprise networks. Similar to commodity computing devices, IoT devices residing within a network may communicate with resources both inside or outside (or both, if applicable) such networks.

[0033] Like commodity computing devices, IoT devices are targets for malicious individuals. Unfortunately, the presence of IoT devices in a network can present some unique security / management challenges. IoT devices are often low-power or dedicated devices and are often deployed without the knowledge of network administrators. Even with the knowledge of such administrators, it may not be possible to install endpoint protection software or agents on IoT devices. IoT devices may be managed by and communicate solely / directly with third-party cloud infrastructure (e.g., with industrial thermometer 152 communicating directly with cloud infrastructure 126) using proprietary (or otherwise non-standard) protocols. This can confound attempts to monitor network traffic in and out of such devices to determine when a threat or attack is occurring against the device. Furthermore, some IoT devices (e.g., in healthcare environments) are mission-critical (e.g., network-connected surgical systems). Unfortunately, the compromise of an IoT device (e.g., by malware 130) or the misapplication of security policies to traffic associated with an IoT device can have potentially devastating effects. The techniques described herein can be used to improve the security of heterogeneous networks including IoT devices and reduce harm caused to such networks.

[0034] In various embodiments, the data appliance 102 includes an IoT server 134. The IoT server 134, in some embodiments, is configured to cooperate with an IoT module 138 of the security platform 140 to identify IoT devices within a network (e.g., the network 110). Such identification may be used, for example, by the data appliance 102 to help create and enforce policies regarding traffic associated with the IoT devices and to enhance the functionality of other elements of the network 110 (e.g., by providing contextual information to the AAA 156). In various embodiments, the IoT server 134 incorporates one or more network sensors configured to passively sniff / monitor traffic. One exemplary way of providing such network sensor functionality is as a tap interface or a switch mirror port. Other approaches to monitoring traffic may also be used (additionally or alternatively) where applicable.

[0035] In various embodiments, the IoT server 134 is configured to provide logs or other data (e.g., passively collected from the monitoring network 110) to the IoT module 138 (e.g., via the front end 142). FIG. 2C shows one exemplary event path between the IoT server and the IoT module. The IoT server 134 sends device discovery events and session events to the IoT module 138. Exemplary discovery and session events are shown in FIGS. 2D and 2E, respectively. In various embodiments, a discovery event is sent by the IoT server 110 whenever it observes a packet that can uniquely identify or confirm the device's identity (e.g., whenever a DHCP, UPNP, or SMB packet is observed). Each session that a device has (with another node, either inside or outside the device's network) is described in a session event that summarizes information about the session (e.g., source / destination information, number of packets received / sent, etc.). If applicable, multiple session events may be batched together by the IoT server 134 before sending to the IoT module 138. In the example shown in Figure 2E, two sessions are included. The IoT module 138 provides device classification information to the IoT server 134 via a device verdict event (234).

[0036] One exemplary way to implement IoT module 138 is using a microservices-based architecture. IoT module 138 may also be implemented, where applicable, using different programming languages, databases, hardware, and software environments, and / or as services that are messaging-enabled, context-bounded, autonomously developed, independently deployable, distributed, and built and released using automated processes. One task performed by IoT module 138 is to identify IoT devices in data provided by IoT server 134 (and by other embodiments of data appliances, such as data appliances 136 and 148) and provide additional contextual information about those devices (e.g., back to the respective data appliances).

[0037] FIG. 2F illustrates one embodiment of an IoT module. Area 295 illustrates a set of Spark Applications that run at intervals (e.g., every 5 minutes, every hour, and every day) across all tenant data. Area 297 illustrates a Kafka message bus. Session event messages received by IoT module 138 (e.g., from IoT server 134) bundle multiple events together as they are observed at IoT server 134 (e.g., to conserve bandwidth). Transformation module 236 is configured to flatten received session events into individual events and publish them at 250. The flattened events are aggregated by aggregation module 238 using a variety of different aggregation rules. One example rule is "aggregate all event data for a specific device and each application (APP-ID) it uses for a time interval (e.g., 5 minutes)." Another example rule is "aggregate all event data for a specific device communicating with a specific destination IP address for a time interval (e.g., 1 hour)." For each rule, the aggregation engine 238 tracks a list of attributes that need to be aggregated (e.g., a list of applications used by a device or a list of destination IP addresses). A feature extraction module 240 extracts features (252) from the attributes. An analysis module 242 uses the extracted features to perform device classification (e.g., using supervised and unsupervised learning), and the results (254) are used to power other types of analysis (e.g., via an operational intelligence module 244, a threat analysis module 246, and anomaly detection module 248). The operational intelligence module 244 provides analytics related to the OT framework and operational or business intelligence (e.g., how the device is being used). Alerts (256) can be generated based on the results of the analysis.In various embodiments, MongoDB 258 is used to store aggregated data and feature values. Background services 262 receive data aggregated by Spark applications and write the data to MongoDB 258. API server 260 pulls and merges data from MongoDB 258 to fulfill requests received from frontend 142.

[0038] FIG. 2G illustrates one exemplary method for implementing IoT device identification analysis (e.g., in IoT module 138 as one embodiment of analytics module 242 and related elements). Discovery events and session events (e.g., as shown in FIGS. 2D and 2E, respectively) are received as Kafka topics on the message bus and as raw data 264 on the message bus (and also stored in storage 158). Features are extracted by feature engine 276 (which may be implemented, for example, using Spark / MapReducer). The raw data is enriched 266 by security platform 140 with additional contextual information, such as geolocation information (e.g., of source / destination addresses). During metadata feature extraction 268, features such as the number of packets sent from an IP address within a time interval, the number of applications used by a particular device during that time interval, and the number of IP addresses contacted by the device during that time interval are constructed. The features are both passed in real time (e.g., in JSON format) to the inline analytics engine 272 (e.g., over a message bus) and stored (e.g., in a suitable format such as Apache Parquet / Data Frame in the feature database 270) for subsequent querying (e.g., during offline modeling 299).

[0039] In addition to features constructed from metadata, a second type of feature, referred to herein as an analytical feature, may be constructed 274 by the IoT module 138. One exemplary analytical feature is one constructed over time based on time series data using aggregated data. The analytical features are similarly passed to the analytics engine 272 in real time and stored in the feature database 270.

[0040] The inline analytics engine 272 receives features over the message bus via a message handler. One task it performs is activity classification (278), which attempts to identify activity associated with a session (such as a file download, a login / authentication process, or a disk backup activity) based on the received feature values / session information and attaches any applicable tags. One way to implement activity classification 278 is via a neural network-based multi-layer perceptron combined with a convolutional network.

[0041] Assume that the activity classification determines that a particular device is engaged in printing activity (i.e., using a printing protocol) and also periodically contacts resources owned by HP (e.g., to check for updates by calling an HP URL and using it to report status information). In various embodiments, the classification information is passed to both a clustering process (unsupervised) and a prediction process (supervised). If either process is successful in classifying the device, the classification is stored in device database 286.

[0042] Devices may be clustered into multiple clusters (e.g., behaves like a printer, behaves like an HP device, etc.) by a stage one clustering engine 280 based on their attributes and other behavioral patterns. One way to implement the clustering engine 280 is to use an extreme gradient boosting framework (e.g., XGB). The stage 1 classifier may be useful for classifying devices that have not been seen before but are similar to existing, known devices (e.g., a new thermostat vendor begins selling a thermostat device that behaves similarly to a known thermostat).

[0043] As shown in FIG. 2G, the activity classification information is also provided to a set of classifiers 282, which perform predictions based on the provided device characteristics. Two possibilities can arise. In the first scenario, it is determined that there is a high probability (i.e., a high confidence score) that the device matches a known device profile. If so, information about the device is provided to a stage two classifier (284), which makes a final verdict on the device's identity (e.g., using the provided information and any additional applicable contextual information) and updates the device database 286 accordingly. One way to implement the stage two classifier is to use a gradient boosting framework. In the second scenario, assume a low confidence score (e.g., the device matches both an HP printer and an HP laptop with 50% confidence). In this scenario, the information determined by the classifiers 282 can be provided to the clustering engine 280 as additional information usable in clustering.

[0044] Also shown in FIG. 2G is an offline modeling module 299. The offline modeling module 299 is not time-constrained, as opposed to the inline analysis engine 272, which attempts to provide device classification information (e.g., as messages 234) in real time. Periodically (e.g., once a day or once a week), the offline modeling module 299 (e.g., implemented using Python) rebuilds the models used by the inline analysis module 272. The activity modeling engine 288 builds models for activity classification 278, which are also used for device type models (296), which are used by the classifier for device identification during inline analysis. The baseline modeling engine 290 builds models of the baseline behavior of the device model, which are also used when modeling specific types of device anomalies (292), such as kill chains, and specific types of threats (294). The generated models are stored in a model database 298 in various embodiments.

[0045] IV. Network Entity Identity AAA

[0046] As mentioned above, assume that Alice has been issued laptop 104 by ACME. Various components of network 110 cooperate to authenticate Alice's laptop as she uses it to access various resources. As one example, when Alice connects laptop 104 to a wireless access point (not shown) located within network 110, the wireless access point can communicate (either directly or indirectly) with AAA server 156 while provisioning network access. As another example, when Alice uses laptop 104 to access her ACME email, laptop 104 can communicate (either directly or indirectly) with directory service 154 while fetching her inbox, etc. As a commodity laptop running a commodity operating system, laptop 104 can generate appropriate AAA messages (e.g., RADIUS client messages) that help laptop 104 gain access to the appropriate resources it needs.

[0047] As mentioned above, one problem posed by IoT devices (e.g., device 146) in networks such as 110 is that such devices are often “unmanaged” (e.g., not configured, provisioned, managed, etc., by a network administrator), do not support protocols such as RADIUS, and therefore cannot be integrated with AAA services like other devices, such as laptop 104. Various approaches can be adopted to provide IoT devices with network access within network 110, each of which has drawbacks. One option is for ACME to restrict IoT devices to using a guest network (e.g., via a pre-shared key). Unfortunately, this can limit the usefulness of IoT devices if they are unable to communicate with other nodes in network 110 to which they should legitimately have access. Another option is to grant IoT devices unrestricted access to network 110, mitigating the security benefits of having a segmented network. Yet another option is for ACME to manually specify rules governing how a given IoT device should be able to access resources within network 110. This approach is generally unacceptable / impractical for a variety of reasons. As one example, administrators may often be uninvolved in the deployment of IoT devices and therefore may not know what policies for such devices should be included (e.g., in the data appliance 102). Even if an administrator could, for example, manually configure policies for a particular IoT device within the appliance 102 (e.g., for a device such as device 112), keeping such policies up to date is error-prone and generally unacceptable given the vast number of IoT devices that may be present in the network 110.Moreover, such policies are likely to be simple (e.g., assigning a CT scanner 112 to a particular network by IP address and / or MAC address) and do not allow for more fine-grained control over connections / policies involving the CT scanner 112 (e.g., dynamically including policies applicable to surgical devices versus point-of-sale (POS) terminals). Furthermore, even if the CT scanner 112 is manually included in the data appliance 102, as mentioned above, IoT devices generally do not support technologies such as RADIUS, and the benefits of having such an AAA server manage the networking access of the CT scanner 112 would be limited compared to other types of devices (e.g., laptop 104) that more fully support such technologies. As described in more detail below, in various embodiments, the data appliance 102 (e.g., via the IoT server 134) is configured to provide support for AAA functions for IoT devices present in the network 110 in a passive manner.

[0048] In the following description, assume that Alice's department at ACME recently purchased an interactive whiteboard 144 so that Alice can collaborate with other ACME employees as well as individuals outside of ACME (e.g., Bob, a researcher at Beta University who has his own network 114, data appliance 136, and whiteboard 146). As part of the initial setup of the whiteboard 146, Alice connects the whiteboard to a power source and provides it with either a wired connection (e.g., to an outlet in a conference room) or wireless credentials (e.g., credentials for use by visitors to the conference room). When the whiteboard 146 provisions its network connection, the IoT server 134 recognizes the whiteboard 146 as a new device in the network 110 (e.g., via mechanisms such as network sensors, as described above). One action taken in response to this detection is to communicate with security platform 140 (e.g., create a new record for whiteboard 146 in database 160 and retrieve any currently available contextual information associated with whiteboard 146 (e.g., obtain the manufacturer of whiteboard 146, the model of whiteboard 146, etc.)). Any contextual information provided by security platform 140 may be provided to (and stored in) data appliance 102, which in turn may provide it to directory service 154 and / or AAA server 156, if applicable. If applicable, IoT module 138 may provide updated contextual information about whiteboard 146 to data appliance 102 as it becomes available. Data appliance 102 (e.g., via IoT server 134) may then provide ongoing information about whiteboard 146 to security platform 140.Examples of such information include observations about the behavior of whiteboard 146 on network 110 (e.g., statistics about the connections it makes), which may be used by security platform 140 to build a behavioral profile for a device such as whiteboard 146. Similar behavioral profiles may be built by security platform 140 for other devices (e.g., whiteboard 144). Such profiles may be used for a variety of purposes, including detecting anomalous behavior. As one example, data appliance 148 may use information provided by security platform 140 to detect whether thermometer 152 is behaving abnormally by comparing it to historical observations of thermometer 152 and / or to other thermometers of similar model or manufacturer (not shown), or more generally, including thermometers present in other networks. If anomalous behavior is detected (e.g., by data appliance 148), appropriate corrective action may be automatically taken, such as restricting thermometer 152's access to other nodes on network 116, generating an alert, etc.

[0049] 3 illustrates one embodiment of a process for passively providing AAA support for IoT devices in a network. In various embodiments, process 300 is performed by the IoT server 134. The process begins at 302 when a set of packets transmitted by an IoT device is obtained. As one example, such packets may be passively received by the IoT server 134 at 302 when a whiteboard 146 is initially provisioned on the network 110. Packets may also be received at 302 during subsequent use of the whiteboard 144 (e.g., when Alice has a whiteboard session with Bob via the whiteboard 146). At least one packet in the set of data packets is analyzed at 304. As one example of processing performed at 304, the IoT server 134 determines that the packets received at 302 were transmitted by the whiteboard 146. One action the IoT server 134 can take is to identify the whiteboard 146 as a new IoT device on the network 110 and, if available, obtain context information from the IoT module 138. At 306, the IoT server 134 sends an AAA message on behalf of the IoT device that includes information associated with the IoT device. An example of such a message is shown in FIG. 4A. As mentioned above, the whiteboard 146 does not support the RADIUS protocol. However, the IoT server 134 can generate a message such as that shown in FIG. 4A on behalf of the whiteboard 146 (e.g., using the information received at 302 and, if applicable, from the security platform 140). As mentioned above, when the IoT server 134 provides information about the whiteboard 146 to the IoT module 138, the IoT module 138 can take various actions.4A , RADIUS messages generated by the IoT server 134 on behalf of the whiteboard 146 may contain limited information. As additional contextual information about the whiteboard 146 is collected by the security platform 140, the profile may be updated and propagated to the data appliance 102. When the whiteboard 146 is initially provisioned in the network 110, the additional contextual information may not be available (e.g., the security platform 140 may not have such additional information, or providing such information to the IoT server 134 by the security platform 140 may not be instantaneous). Therefore, and as shown in FIG. 4A , RADIUS messages generated by the IoT server 134 on behalf of the whiteboard 146 may contain limited information. As additional contextual information is received (e.g., by the IoT server 134 from the IoT module 138), subsequent RADIUS messages sent by the IoT server 134 on behalf of the whiteboard 146 may be enriched with such additional information. Examples of such subsequent messages are shown in Figures 4B and 4C. Figure 4B shows an example of a RADIUS message that the IoT server 134 can send on behalf of the whiteboard 146 once the context information about the whiteboard 146 (e.g., including a database of context information about various IoT devices) has been provided by the IoT module 138. In the example shown in Figure 4B, context information such as the whiteboard's manufacturer (Panasonic) and the nature of the device (e.g., that it is an interactive whiteboard) is included. Such context information can be used by an AAA server, such as AAA server 156, to provide AAA services for the whiteboard 146 (without the need to modify the whiteboard 146).4C , additional context information is included in the RADIUS message by the IoT server 134 on behalf of the whiteboard 146. The additional context information includes additional attribute information such as the device model, operating system, and operating version. When the whiteboard 146 is initially provisioned in the network 110, all of the context information shown in FIG. 4C is likely not available. As the whiteboard 146 is used within the network 110 over time, additional contextual information may be collected (e.g., as the IoT server 134 continues to passively observe packets from the whiteboard 146 and provide information to the security platform 140). This additional information may be leveraged (e.g., by the data appliance 102) to enforce fine-grained policies. As one example, as shown in FIG. 4C, the whiteboard 146 runs a particular operating system that is Linux-based and has version 3.16. Frequently, IoT devices run operating system versions that are not upgradeable / patchable.Such devices may pose security risks when exploits are developed for their operating systems. The data appliance 102 can implement security policies based on the context information, such as by isolating IoT devices with outdated operating systems from other nodes in the network 110 (or otherwise restricting their access), while allowing less restrictive network access to IoT devices with current operating systems.

[0050] 4A-4C show example RADIUS Access-Request messages. Where applicable, the IoT server 134 can generate various types of RADIUS messages on behalf of the whiteboard 146. As one example, a RADIUS Accounting Start message can be triggered when traffic from the whiteboard 146 is first observed. Periodic RADIUS Accounting Interim Update messages can be sent while the whiteboard is in use, and a RADIUS Accounting Stop message can be sent when the whiteboard 146 goes offline.

[0051] V. IoT Device Discovery and Identification

[0052] As mentioned above, one task performed by the security platform 140 (e.g., via the IoT module 138) is IoT device classification. As one example, when the IoT server 134 sends a device discovery message to the IoT module 138, the IoT module 138 attempts to determine a classification for the device and respond (e.g., via decision 234 shown in FIG. 2C ). If applicable, the device is associated with a unique identifier by the IoT module 138 so that subsequent classification of the device does not need to be performed (or, if applicable, is performed less frequently than would otherwise be performed). Also, as mentioned above, the determined classification can be used (e.g., by the data appliance 102) to enforce policies on traffic to / from the device.

[0053] Various approaches can be used to classify devices. The first approach is to perform the classification based on a set of rules / heuristics that leverage static attributes of the device, such as its organizationally unique identifier (OUI), the type of application it runs, etc. The second approach is to perform the classification using machine learning techniques that leverage dynamic but predefined attributes of the device extracted from its network traffic (e.g., the number of packets sent per day). Unfortunately, both of these approaches have weaknesses.

[0054] Rule-based approaches generally require that a separate rule be manually created for each type of IoT device (describes which attributes / values ​​should be used as signatures for each type of device signature). One challenge presented by this approach is determining which signatures are both relevant to identifying the device and unique among other device signatures. Furthermore, rule-based approaches are limited in the number of static attributes available that can be easily obtained from traffic (e.g., user agent, OUI, URL destination, etc.). Attributes generally need to be simple enough to appear in patterns that a regular expression can match. Another challenge is identifying new static attributes that may be present / identifying as new devices enter the market (e.g., a new brand or model of CT scanner is offered). Another challenge is that all matching attributes in network traffic must be collected for a rule to be triggered. Fewer attributes may result in a verdict. As an example, a signature may require that a specific device with a specific OUI connects to a specific URL. While having the OUI itself may already be a sufficient indicator of a device's identity, the signature does not trigger until the URL is also observed. This causes additional delays in determining the device's identity. Another challenge is maintaining and updating the signature as a static attribute of the device that changes over time (e.g., due to updates made to the device or the services used by the device). As an example, a particular device may have initially been manufactured using one type of network card, but over time the manufacturer may have switched to a different network card (which exhibits a different OUI). If the rule-based system is unaware of the change, false positives may result. Yet another challenge lies in scaling signature generation / verification as the number of new IoT devices brought online each day approaches millions of new device instances.As a result, newly created rules may conflict with existing rules and cause false positives in classification.

[0055] Machine learning-based approaches typically involve creating a training model based on static and / or dynamic features extracted from network traffic. Prediction results for network data from new IoT devices are based on a pre-trained model, which provides device identification information with associated accuracy. Examples of problems with machine learning approaches include: When predictions are made for each new device or for devices that do not have a constant / unique ID (e.g., MAC address), the computation time required to reach the desired accuracy may be unacceptable. Because there are thousands or tens of thousands of features that need to be generated, and these features may vary over a given time window, it can take a significant amount of time before a sufficient number of features are available to make valid predictions (which may defeat the purpose of policy enforcement). Furthermore, when minimizing latency in predictions is the goal, building and maintaining large data pipelines for streaming network data can be costly. Yet another problem is that noise introduced by irrelevant features specific to a given deployment environment can reduce prediction accuracy. And, when the number of device types reaches tens of thousands or more, maintaining and updating models presents challenges.

[0056] In various embodiments, security platform 140 addresses the issues of each of the two aforementioned approaches by using a hybrid approach to classification. In an exemplary hybrid approach, a network behavior pattern identifier (also referred to herein as a pattern ID) is generated for each type of device. In various embodiments, the pattern ID is a list of attributes or sequence features combined with their respective probabilities (as importance scores for the features or behavior categories), which form a distinct network behavior description and can be used to identify the type of IoT device. The pattern ID can be stored (e.g., in a database) and used to identify / verify the identity of the device.

[0057] When training on a set of attributes, certain approaches, such as extreme gradient boosting frameworks (e.g., XGB), can provide a top list of important features (whether static attributes, dynamic attributes, and / or aggregated / transformed values). Pattern IDs can be used to uniquely identify device types once established. If certain features are dominant for a device (e.g., if a particular static feature (such as contacting a very specific URL at boot time) identifies the device with 98% confidence), they can be used to automatically generate rules. Even if no dominant features exist, a representation of the top features can still be used as the pattern ID (e.g., if multiple sets of features are concatenated into a pattern). By training on a dataset that includes all known models (and all known IoT devices), potential conflicts between models / uniquely identifying features can be avoided. Furthermore, pattern IDs do not need to be human-readable (but can be stored, shared, and / or reused for identification purposes). Significant time savings can also be realized with this approach, resulting in near-real-time classification. Classification of a particular device can occur as soon as a dominant feature is observed (instead of having to wait until a large number of features occur).

[0058] One example of data that may be used to create a Pattern ID for a “Teem Room Display iPad®” device may include the following (with the complete list generated automatically through training a multivariate model or training multiple binary models): *Apple devices (100%) *Special iPad (>98.5%) *Teamroom app (>95%) *Meets volume pattern VPM-17 (>95%) *Servers in the cloud (>80%)

[0059] One exemplary method for implementing a hybrid approach is as follows: A neural network-based machine learning system may be used for automated pattern ID training and generation. Example features that may be used to train a neural network model include both static features extracted from network traffic (e.g., OUI, hostname, TLS fingerprint, matched L7 payload signature, etc.) and sequence features extracted from network traffic but not specific to the environment (e.g., application, L7 attributes of the application, volume ranges converted to categorical features, etc.). A lightweight data pipeline may be used to stream selected network data for feature generation in real time. A prediction engine may be used to import the model and provide caching to minimize prediction latency. Short (e.g., minute-based) aggregation may be used in prediction to stabilize selected sequence features. Customized data normalization, enrichment, aggregation, and transformation techniques may be used to manipulate sequence features. For better accuracy, longer aggregation windows may be used in training. Using merged and aggregated features over time, prediction accuracy may be improved. A back-end feedback engine can be used to route the results of a "slow path" prediction system (e.g., a machine learning-based approach including a device type modeling subsystem and a device group modeling subsystem) that helps expand the attributes used for pattern ID prediction. A device group model can be trained to compensate for issues with the device type model when sufficient samples or features are not available and improve accuracy beyond an acceptable threshold (e.g., assigning prediction results based on a set of predefined types, some of which are associated with another subsystem to cluster unlabeled devices of similar types).Finally, a decision module may be used to publish results from the real-time prediction engine.

[0060] Exemplary advantages of the hybrid approach to classification as described herein are: First, fast convergence occurs, allowing a given device to potentially be identified within minutes or seconds. Second, the invention addresses the individual problems of rule-based and machine learning-based systems. Third, it provides stability and consistency in prediction results. Fourth, it has the scalability to support tens of thousands (or more) of different types of IoT devices. Prediction is typically only required for new devices (even if a given device lacks a unique ID assignment, such as L3 network traffic-based identification).

[0061] An embodiment of module 138 is shown in Figure 5. One exemplary way to implement IoT module 138 is to use a microservices-based architecture where services are fine-grained and protocols are lightweight. Services can also be implemented using relatively small services that are context-bounded, autonomously developed, independently deployable, distributed, and built and released using automated processes, spanning different programming languages, databases, hardware and software environments, and / or messaging environments, where applicable.

[0062] As mentioned above, in various embodiments, security platform 140 periodically receives information (e.g., from data appliance 102) about IoT devices on a network (e.g., network 110). In some cases, the IoT device has been previously classified by security platform 140 (e.g., a CT scanner installed on network 110 last year). In other cases, the IoT device is newly seen by security platform 140 (e.g., whiteboard 146 is being installed for the first time). Assume that a given device has not been previously classified by security platform 140 (e.g., there is no entry for the device in database 286, which stores a set of unique device identifiers and associated device information). As shown in FIG. 5, information about a new device can be provided to two different processing pipelines for classification. Pipeline 504 represents a “fast path” classification pipeline (corresponding to a pattern ID-based approach), and pipeline 502 represents a “slow path” classification pipeline (corresponding to a machine learning-based approach).

[0063] In pipeline 504, fast path feature engineering is performed (508) to identify applicable static and sequence features of the device. Fast path prediction is performed (510) using the pattern ID or a previously built model (e.g., a model built based on top significant features and using offline processing pipeline 506). A confidence score for devices matching a particular pattern is determined (512). If the confidence score for the device meets a pre-trained threshold (e.g., based on the overall prediction accuracy of module 138 or its components, such as 0.9), a classification can be assigned or updated to the device (in device database 516), if applicable. Initially, the confidence score is based on near real-time fast path processing. The advantage of this approach is that data appliance 102 can begin applying policies to the device's traffic very quickly (e.g., within minutes of module 138 identifying the device as new / unclassified). The appliance 102 may be configured to either fail-safe (e.g., reduce / limit the device's ability to access various network resources) or fail-safe (e.g., allow broad access for the device) pending a classification decision from the system 140. As additional information becomes available (e.g., via slow-path processing), the confidence score can be based on that additional information, if applicable (e.g., increasing the confidence score or correcting / amending mistakes made during fast-path classification).

[0064] Examples of features (eg, static attributes and array features) that can be used include: The pattern ID can be any combination of these attributes with logical conditions, including: *OUI in mac address *Hostname string from decoded protocol *User agent strings from HTTP and other clear-text protocols *System name string from the decoded SNMP response *OS, hostname, domain, and username from decrypted LDAP protocol *URL from decoded DNS protocol *SMB version, command, error from decoded SMB protocol *TCP flags *Decoded option string from the DHCP protocol *Strings from decoded IoT protocols such as Digital Imaging and Communications in Medicine (DICOM) * List of inbound applications from the local network *List of inbound applications from the internet * List of outbound applications to the local network *List of outbound applications to the internet *List of inbound server ports from the local network *List of inbound server ports from the internet *List of outbound server ports to the local network *List of outbound server ports to the internet *List of inbound IPs from your local network *List of inbound URLs from the internet *List of outbound IPs to your local network *List of outbound URLs to the internet

[0065] In some cases, the confidence score determined in 512 may be very low. One reason this may occur is because the device is a new type (e.g., a new type of IoT toy or other type of product not previously analyzed by security platform 140) and there is no corresponding pattern ID available for the device on security platform 140. In such a scenario, information about the device and the classification results may be provided to offline processing pipeline 506, which may perform clustering (514) on behaviors exhibited by the device and other applicable information (e.g., to determine that the device is a wireless device, acts like a printer, uses the DICOM protocol, etc.). The clustering information is applied as a label and, if applicable, flagged for further investigation 518, so that any similar devices subsequently seen are automatically grouped together. If, as a result of the investigation, additional information about a given device is determined (e.g., identified as corresponding to a new type of consumer IoT meat thermometer), the device (and all other devices with similar characteristics) can be relabeled accordingly (e.g., as a Brand XYZ Meat Thermometer), and an associated pattern ID can be generated and made available by pipelines 502 / 504, if applicable (e.g., after the models have been rebuilt). In various embodiments, offline modeling 520 is a process that runs daily to train and update various models 522 used for IoT device identification. In various embodiments, the models are refreshed daily to cover newly labeled devices and (for the slow-path pipeline 502) rebuilt weekly to reflect behavioral changes and adapt to new features and data insights added during the week.Note that when adding a new type of device to security platform 140 (i.e., creating a new device pattern), multiple existing device patterns may be affected, requiring either the list of features or their importance scores to be updated. This process can be performed automatically (and is a major advantage compared to rule-based solutions).

[0066] For fast path modeling, neural network-based models (e.g., FNN) and general classification models (e.g., XGB) are widely used for multivariate machine learning models. To improve results and help provide input for clustering, a binary model is also built for the selected profile. The binary model provides a yes / no answer to the device's identity or a given device behavior. For example, a binary model can be used to determine whether a device is an IP phone type or is unlikely to be an IP phone. A multivariate model has many outputs normalized to a probability of 1, each corresponding to a device type. Binary models are generally faster, but require the device to pass through many of them in the prediction to find the correct "yes" answer. A multivariate model can achieve this in one step.

[0067] The slowpath pipeline 502 is similar to pipeline 504 in that features are extracted (524). However, the features used by pipeline 502 typically take a period of time to build. As one example, a feature such as "bytes transmitted per day" requires one day to collect. As another example, a given usage pattern may take a period of time to develop / be observed (e.g., a CT scanner is used hourly to perform scans (first behavior), backs up data daily (second behavior), and checks the manufacturer's website weekly for updates (third behavior)). The slowpath pipeline 502 invokes a multivariate classifier (526) in an attempt to classify new device instances against the complete set of features. The features used are not limited to static or sequence features, but also include volume and time-series-based features. This is generally referred to as stage 1 prediction. If the results of stage 1 prediction are suboptimal (low confidence), for a particular profile, stage 2 prediction is used to improve the results. The slow path pipeline 502 invokes a set of decision tree classifiers (528) supported by additional imported device context to classify new device instances. The additional device context is imported from an external source. As one example, a URL connected to by a device may have been given a category and a risk-based reputation, which may be included as a feature. As another example, an application used by the device may have been given a category and a risk-based score, which may be included as a feature. By combining the results from the stage 1 prediction 526 and the stage 2 prediction 528, a final slow path classification decision may be reached using a derived confidence score.

[0068] There are generally two stages included in the slow path pipeline 502. In the slow path pipeline, in some embodiments, a Stage 1 model is built using a multivariate classifier based on neural network techniques. Stage 2 of the slow path pipeline is generally a set of decision-based models with additional logic to handle probability-related exceptions from Stage 1. In prediction, Stage 2 integrates inputs from Stage 1, applies rules and context to validate the Stage 1 output, and generates the final output of the slow path. The final output includes a device identification, an overall confidence score, a pattern ID that can be used in future fast path pipelines 504, and an explanation list. The confidence score is based on the reliability and accuracy of the model (the model also has a confidence score) and the probability as part of the classification. The explanation list includes a list of features that contribute to the result. As mentioned above, an investigation can be triggered if the result deviates from a known pattern ID.

[0069] In some embodiments, two types of models are built for slow path modeling: one for individual identification and one for group identification. For example, telling the difference between two printers from different vendors or with different models is often more difficult than distinguishing a printer from a thermometer (e.g., because printers tend to exhibit network behavior, speak similar protocols, etc.). In various embodiments, various printers from different vendors are included in a group, and a “printer” model is trained for group classification. This group classification result can provide better accuracy than a specific model for a specific printer and can be used to update the device's confidence score, if applicable, or to provide reference and validation for the identity-based classification of individual profiles.

[0070] FIG. 6 shows an example process for classifying an IoT device. In various embodiments, process 600 is performed by security platform 140. Process 600 may also be performed by other applicable systems (e.g., systems collocated on-premise with IoT devices). Process 600 begins at 602 when information associated with network communications of an IoT device is received. As one example, such information is received by security platform 220 when data appliance 102 sends a device discovery event for a given IoT device to security platform 140. At 604, a determination is made that the device is unclassified (or, if applicable, reclassification should be performed). As one example, platform 140 can query database 286 to determine whether the device is classified. At 606, a two-part classification is performed. As one example, at 606, a two-part classification is performed by the platform 140, providing information about the device to both the fast-path information provision pipeline 504 and the slow-path classification pipeline 502. Finally, at 608, the results of the classification process performed at 606, along with the network behavior aggregated from the baseline modeling (290), are provided to a security appliance configured to enforce policies on the IoT device. Examples of such network behavior include most used applications, URLs, and other attributes that can be “extracted” from the machine-learned baseline for the IoT device profile to help form security appliance policies. As discussed above, this allows highly granular security policies to be implemented in potentially mission-critical environments with minimal administrative effort.

[0071] In a first example of implementing process 600, assume that an Xbox One game console is connected to network 110. During classification, a determination may be made that the device has the following key features: a "vendor=Microsoft" feature with 100% confidence, a "communicates with Microsoft cloud server" feature with 89.7% confidence, and a "game console" match feature with 78.5% confidence. These three features / confidence scores may be collectively matched against a set of profile IDs (a process driven by neural network-based prediction) to identify the device as an Xbox One game console (i.e., a profile ID match that meets a threshold is found at 512). In a second example, assume that an AudioCodes IP phone is connected to network 110. During classification, it may be determined that the device matches the "vendor=AudioCodes" feature with 100% confidence, the "is IP audio device" feature with 98.5% confidence, and the "act as local server" feature with 66.5% confidence. These three features / confidence scores are also matched against a set of profile IDs, although in this scenario, it is assumed that no existing profile IDs match with sufficient confidence. Information about the device may then be provided to a clustering process 514, and, if applicable, a new profile ID may eventually be generated and associated with the device (and used to classify future devices).

[0072] If applicable, security platform 140 may recommend specific policies based on the determined classification, which will be described in more detail below. The following are examples of policies that may be implemented: * Deny internet traffic for all infusion pumps (regardless of vendor). * Deny internet traffic for all GE ECG machines except from / to specified GE hosts. *Allow only internal traffic to Picture Archiving and Communication System (PACS) servers for all CT scanners (regardless of vendor).

[0073] VI. IoT Security Policy in Firewalls

[0074] As mentioned above, IoT devices are often purpose-built devices (as opposed to general-purpose computing devices such as laptops) with predefined behaviors that can be observed on a network. As one example, a CT scanner, regardless of manufacturer (e.g., GE or Fujitsu), functions similarly / behaves similarly on a network as other CT scanners, using one or more specific protocols to transmit captured patient images to a networked image server (e.g., via an interface to the server) for review by medical staff. Other types of systems (e.g., heating, ventilation, and air conditioning (HVAC) systems) exhibit their own set of similar, typical, predefined behaviors (e.g., reporting temperature values ​​to a server once per minute via a specific protocol).

[0075] As described above, analysis of these behaviors (e.g., by security platform 140) from observed traffic (e.g., by data appliance 102) can identify specific IoT devices (including by identifying specific device instances, device models, device manufacturers, device types, etc.). Furthermore, a by-product of device identification will be a device baseline model trained (e.g., by baseline modeling engine 290) for classification purposes. If applicable, anomaly detection module 248 can be used to filter out known anomalous behaviors when creating a baseline for a device (or group of devices). This deep machine learning model captures the network behaviors described above. This baseline model can be used not only in device identification prediction, but also to generate a common list of behavior aggregations ranked by how popular the network behaviors are on device profiles. One approach to behavior aggregation is to use an ML algorithm, such as XGB, to extract and rank the top contributing features (used in device identification) from the baseline model during the training process. Other approaches can also be used or combined (e.g., heuristic approaches). The top contributing features (according to a reliability / confidence threshold) essentially highlight what are the most common network behaviors that a type of device exhibits from the thousands of attributes or features used in training, and can be used in recommendations (e.g., whitelisting / blacklisting specific URLs, protocols, etc.). Aggregated behaviors can include which applications are used, which connections are made to given network domains, what payloads are carried in the applications, the volume, time, and frequency of communication, etc.Each attribute is assigned a frequency category, such as “rare,” “often,” or “regular.” Each attribute can also be assigned a range category, such as “less than 1 MB per hour.” Anomalies (e.g., compromised or malfunctioning / misconfigured IoT devices) can be detected (e.g., by the data appliance 102 working in conjunction with the anomaly detection module 248) as deviations from the baseline. These attributes (and known vulnerabilities to specific attacks) can be used as a blueprint to automatically create recommended firewall policies to constrain network activity associated with specific IoT devices, device types, etc. For example, regularly used applications and URLs can be used to build an “allow” firewall policy. In another example, applications that are not part of the baseline behavior can be used to build a “deny” firewall policy. Users can adjust policies based on the frequency of network behavior summarized from hundreds of thousands of similar devices. And any known vulnerabilities (e.g., a particular device's vulnerability to a particular attack), if applicable, can be modeled separately and incorporated into the recommended policy. An example of a top feature for a given device type is that the device checks for updates approximately once a day at a particular URL (e.g., www.siemens.com / updates). If a threshold number of devices sharing a device type exhibit similar baseline behavior, the feature can be selected as a recommended whitelist item for the device profile associated with that device type.

[0076] 7A illustrates a first approach for implementing a set of policies regarding CT scanners / image servers that Acme may deploy within network 110. In particular, assume that Acme has deployed two types of CT scanners (manufactured by GE and Fujitsu). An administrator of network 110 (hereafter referred to as Charlie) can interact with an interface (e.g., provided by data appliance 102 and / or security platform 140, where applicable) and manually specify, for each CT scanner and image server in network 110, the protocols, ports, and IP addresses over which they are allowed to communicate. Unfortunately, this approach is time-consuming and error-prone. As one example, when a new CT scanner is added to the environment, Charlie must manually add to the rules shown in FIG. 7A and also potentially modify / remove some of the rules (e.g., if the new CT scanner replaces an existing one and / or network information changes). If the total number of IoT devices in an environment is small and the IoT devices are assigned static IP addresses, manual maintenance of rules as shown in Figure 7A may be feasible. In practice, however, a given environment may have hundreds or thousands (or more) of IoT devices and / or use DHCP, and manual maintenance of rules may be infeasible.

[0077] An alternative approach is to abstract the application (e.g., “DICOM-App,” which indicates the particular protocol / port / etc. corresponding to the network traffic used to communicate medical image information) and device type (e.g., GE-Xray-Device) according to the techniques described herein. An abstraction of the rules shown in FIG. 7A is shown in FIG. 7B. Notably, Charlie does not need to provide the IP addresses, ports, or protocols associated with the IoT policy; rather, he can use the abstracted application and device types. A policy such as that shown in FIG. 7B can be compiled and used by the data appliance 102 at runtime. During compilation, the abstracted element (e.g., GE-Xray-Device) is replaced (e.g., with the IP address of each IoT device that matches that device identification) based on information stored in the data appliance 102 (such as APP-ID information, IP information, and / or a dictionary of device types).

[0078] Charlie can choose to manually write IoT device rules (e.g., using the abstractions described above, if desired), but policy recommendations can also be provided by the security platform 140. Recommendations, in various embodiments, are based on device profiles (including device type or other information) for sets of devices sharing various characteristics and baseline / typical behavior (across many different customer environments / deployments). If Charlie accepts the recommended policies, appropriate rules (e.g., rules such as those shown in FIG. 7B ) can be automatically created (e.g., by the security platform 140) and imported into the security device 102 for enforcement. The security device 102 learns the device profiles of IoT devices on its network and matches applicable policies to devices as source or destination. If applicable, policies can be translated (e.g., by the security platform 140) into a format usable by other types of infrastructure than the security device 102, such as a network access controller.

[0079] In the following discussion, assume that Acme recently purchased a set of building automation devices (e.g., a set of badge readers), installed them inside various Acme facilities, and brought the devices online with a portion of network 110. Using the device identification / classification techniques described above, security platform 140 (in cooperation with data appliance 102) identifies that Acme has added 28 new badge reader devices to its network and learns the various behaviors that those particular badge reader devices exhibit as they operate within Acme's network environment (e.g., during an initial observation period of one week or one month). A portion of the management interface provided by security platform 140 is shown in FIG. 8. Interface 800 shows that Acme currently has a total of 65 different IoT devices (with corresponding profiles) operating within its environment. The newly added badge reader is shown in row 802.

[0080] If Charlie clicks on link 804, he is directed to the interface shown in FIG. 9. Area 902 shows that security platform 140 has identified 28 new devices as matching the Siemens Building Technology Device profile with a high degree of confidence. Behavioral information collected about the 28 devices while operating within the Acme environment is also shown and summarized in area 904. Collectively, the 28 devices run eight applications in the Acme environment, communicate with 23 destinations (22 within Acme and 1 external), and currently have a risk score of 56. A count of how many of the 28 devices are using each application in the Acme environment is shown in area 906, and whether the destination is internal or external is shown in area 908. Area 910 contains the number of applications used by Siemens Building Technology devices in the Acme environment compared to how these devices (sharing the same Siemens Building Technology device profile) behave across environments of other customers of security platform 140. As shown in area 910, a typical customer deployment of Siemens Building Technology devices uses between three and five applications (912), placing Acme's deployment outside the typical range (914). When Charlie hovers his cursor over area 912, he is presented with a box providing additional information about the comparison, such as the following: "In this profile, eight different applications were used by the device. Based on data from all IoT security customers, the minimum number of applications used was three, the average was three, and the maximum was five. Application usage by Siemens Building Technology devices was higher than normal. Review the application list."

[0081] Charlie can review the application list by scrolling further down interface 900. As shown in FIG. 10 , after such scrolling, Charlie is reviewing the badge reader device's usage of the “dhcp” and “bacnet” applications. The “usage” designation (1002) indicates the network usage pattern (device profile + application + frequency of URL (e.g., “www.siemens.com / update”) and / or frequency of destination profile (e.g., “PACS server”)) for IoT devices sharing the profile. In various embodiments, the usage of each application is generated based on the first month of collected traffic. Charlie can use the usage information in deciding whether to allow or block certain behaviors. For example, he can learn (e.g., based on his knowledge of Acme's environment) that if bacnet is used infrequently, it should be allowed only within internal domains or only to predefined external domains.

[0082] If Charlie clicks on area 916 of interface 900, he is presented with two options for creating a set of policies that can be applied to the badge reader device. As previously described, Charlie can manually create his own policy set for the badge reader device (e.g., by interacting with various elements of an embodiment of interface 900). Charlie can also choose to load a recommended policy set that security platform 140 has generated using baseline / other information obtained from the environments of other customers of security platform 140. Once Charlie clicks on area 916 and selects to use a recommended policy set (if available), security platform 140 lists any available recommended policy sets, and Charlie can download / apply them to the Acme environment with the ability to refine / adjust the applicable policies (e.g., by interacting with various features provided by interface 900).

[0083] FIG. 11 illustrates an example process for generating policies to apply to communications involving IoT devices. In various embodiments, process 1100 is performed by security platform 140. Process 1100 may also be performed by other applicable systems (e.g., systems collocated on-premise with IoT devices). Process 1100 begins at 1102 when information associated with network communications of an IoT device is received. As one example, such information is received by security platform 140 when data appliance 102 sends a device discovery event for a given IoT device (e.g., a badge reader device). At 1104, the received information is used to determine a device profile to associate with the IoT device. As one example, a determination is made that the IoT device is a Siemens SIMATEC RF10000 device, having a particular serial / MAC address, a particular IP address, etc. In this example, the determined “device type” may be “Siemens Building Technology Device.” In this section, device types (e.g., badge reader devices) and device profiles (e.g., Siemens Building Technology devices) are generally referred to interchangeably. However, multiple profiles may be created for a given device type (e.g., Siemens Building Technology devices located in Acme's research area vs. Acme's retail area), and a given profile may include multiple device types, if applicable (e.g., a Siemens Building Technology device profile may include a badge reader device and a motion trigger sensor). Finally, at 1106, a recommended policy to be applied to the IoT device by the security appliance is generated.9, security device 140 can recommend a policy set that allows only the three most commonly used badge reader applications (or the five most commonly used badge reader applications) that correspond to the information shown in area 910. Once the recommended policies are downloaded and applied, if Charlie needs to make adjustments to the recommended set (e.g., whitelist bacnet), he can make the adjustments (e.g., by interacting with an “edit” option provided by interface 800).

[0084] VII. IOT Device Identification Using Machine Learning Models of Packet Flow Behavior

[0085] A. Introduction

[0086] As mentioned above, packet inspection, such as inspecting packet headers and / or inspecting payloads to perform content-based pattern or signature matches, can provide useful information when attempting to identify / classify devices, such as IoT devices. Such packet and content information often includes items such as organizationally unique identifiers (OUIs), source / destination ports, hostnames, application IDs, special destinations, user agent strings, DHCP fingerprints, and DHCP vendor class identifiers. Unfortunately, some devices can be difficult to identify through packet analysis. Below are six examples of these types of devices: *Endpoint devices for which OUI information is not available. *Multiple types of devices sharing a single OUI (for example, an X-ray device and an ultrasound machine all using wireless cards with the same OUI). *Devices whose identities have been discovered but for which it is difficult to define corresponding classification rules (e.g., X-ray devices discovered because other devices in the same subnet also have similar traffic patterns or similar hostnames, or X-ray devices discovered because they communicate with a server with a name that suggests the client is an X-ray device). *Devices with encrypted traffic (which can make signature-based classification difficult). *Devices behind network equipment (e.g., router or switch) where the device MAC appears to be the MAC of the router / switch interface, rather than the MAC of the device. *Devices that do not have a hostname (existing identification rules often use a combination of hostname and OUI as the identification / classification rule).

[0087] In various embodiments, the device identification / classification techniques described above are supplemented / complemented by deploying one or more machine learning models that utilize packet behavior information. Examples of such packet behavior information (described in more detail below) include packet length sequences (SPLN), packet inter-arrival time sequences (SPIT), and transport layer security (TLS) information, all of which may be used as features. In various embodiments, logistic regression is used to train the dataset, and the coefficients in the logistic regression may be used as behavioral signatures for the devices (which may then be used to identify / help identify the devices). Returning to FIG. 2F , in various embodiments, the training phase is performed by training module 241 using labeled events provided to aggregation module 238 for feature extraction and training. The trained model may be stored for later use, for example, by analysis module 242. During production / operation, events (e.g., as initially provided by data appliance 102 and processed by various components of IoT module 138) similarly pass through aggregation module 238 and into analytics module 242, which uses the trained model to perform device identification on the aggregated events (e.g., during in-line analysis by analytics engine 272, as an example of binary / multi-class classifier 282). Where applicable (e.g., as new information is received about new types of devices), training module 241 can retrain existing models / create new models.

[0088] The techniques described herein are particularly well suited to environments such as those shown in Figure 1, where one or more data appliances (e.g., data appliances 102, 136, and 148) are present and used to help enforce security policies. This is in part because such data appliances can maintain a session control block for each session, and the session control block contains the metadata / other information used by the machine learning models described herein. Furthermore, if an existing data appliance does not support particular feature extraction (e.g., packet length sequences, packet inter-arrival time sequences, and / or redundant TLS information), adding such functionality can be done efficiently while keeping overhead to a minimum.

[0089] B. Example of packet behavior information

[0090] The following are example features (related to packet behavior parameters) that may be used to train one or more machine learning models and used individually or in combination with other models (e.g., as described above) to help identify / classify devices, such as IoT devices: The following example features may be combined with other features (if available), such as metadata information associated with the packet (e.g., OUI, destination port, application ID, protocol, outgoing packet count, and outgoing packet byte count).

[0091] In one exemplary embodiment, logistic regression is used (based on linear regression), where the result of the logistic regression is a probability that can be used as a confidence level for device identification. One exemplary formula for logistic regression is:

number

[0092] Other types of classifiers may also be used in accordance with the techniques described herein (including by using the features described herein), such as Gaussian Naive Bayes, K-Nearest Neighbors, decision trees, random forests, and / or vector machines.

[0093] 1. Packet Length Sequence (SPLN)

[0094] A sequence of packet lengths (SPLN) may be obtained for a given flow by, for example, determining the lengths of the first 10 (or other suitable number) non-zero packet length packets (payloads) of the flow. If a given packet in the sequence has a zero-length payload, that packet may be skipped and the next packet length may be examined (until the required 10 or other number of non-zero packet lengths is considered). If a given flow has fewer than the total number of non-zero packet lengths (e.g., the entire flow has only 7 non-zero length packets), the SPLN may be padded with zeros at the end to reach the configured number of packet lengths (e.g., 10).

[0095] 2. Packet Inter-Arrival Time Sequence (SPIT)

[0096] FIG. 12 illustrates a sequence of packet inter-arrival times. As shown, "Time 1" represents the time between when the first packet (Packet 0) and the second packet (Packet 1) are received. "Time 2" represents the time between when the second packet (Packet 2) and the third packet (Packet 2) are received, and so on. As with the SPLN described above, in some embodiments, SPIT is determined only for non-zero-length packets. For example, if "Packet 2" shown in FIG. 12 has a zero-length payload, the delta between the time "Packet 3" and "Packet 1" are received is used (and a total of 10 or other suitable number of packet inter-arrival times is determined between the first 11 or other suitable number of non-zero-length packets).

[0097] 3. Transport Layer Security (TLS)

[0098] When a device (e.g., an IoT device) engages in an encrypted traffic session, the first part of the session is a handshake in the clear. Information can be captured from the handshake, and information about the handshake can be extracted. Examples of such information include the TLS version, the ordered list of offered cipher suites (e.g., in the Client Hello message), the list of supported TLS extensions (e.g., in the Client Hello message), the selected cipher suite (e.g., in the Server Hello Response message), the selected TLS extension (e.g., also in the Server Hello Response message), and the public key length. Similar to the packet metadata, SPLN, and SPIT information described above, TLS information can be used as a feature in determining packet behavior signatures.

[0099] C. Example - SampleCo Smart Meter

[0100] Below is a simplified example of using the techniques described herein to train (e.g., via training module 241) a model that recognizes a SampleCo brand smart meter (e.g., a water meter that collects data about how much water is used and when it is used, and then transmits the collected data to a network for billing, e.g., by a water company). The SampleCo smart meter exhibits various characteristics, e.g., by communicating via NTP, TCP, and FTP both internally (over an intranet) and externally (e.g., to an external server).

[0101] Positive and negative training data (pertaining to SampleCo smart meters) is collected and provided (e.g., by researchers) to the training module 241. An example of such data is shown in FIG. 13A. In the illustrated embodiment, a total of 10 training samples are provided (each numbered 0-9). Column 1302 indicates whether a given sample is a positive sample (indicated by a "1") or a negative sample (indicated by a "0"). Thus, the first five devices are SampleCo smart meters, and the last five are not. The remaining nine columns (1304) each correspond to a feature. An example includes OUI (1306), protocol (protocol, prot) (1308), destination port (dport) (1310), packet length (plen) (1312), and packet timing information (t) (1314). In various embodiments, such sample data may be obtained using one or more Python scripts configured to parse traffic flow logs. Examining the training data for Device 0 (a positive example of a SampleCo smart meter), it has an OUI of 6935 and is communicating using Protocol 6 (e.g., FTP) over port 21. The first three non-zero packets for Device 0 were observed to be 21 bytes, 26 bytes, and 22 bytes, respectively. In this example data set, the first packet time (t0) is always zero, and the delta between each time (e.g., t1-t0, etc.) is determined by the training module 241. Other forms of data may also be used (e.g., where deltas are determined by the Python script and appear in place of each of the times listed in field 1314).

[0102] Figure 13B shows the results of training a model using the dataset shown in Figure 13A. In particular, the coefficients and intercepts for the model for determining whether traffic corresponds to a SampleCo smart meter are shown. Collectively, the coefficients / intercepts shown in Figure 13B are an example of a behavioral signature of a SampleCo smart meter device and may be stored (e.g., in database 160 or other suitable location) for use in subsequent identification of other SampleCo smart meter devices (e.g., as they are added to network 110 and their traffic is received by data appliance 102).

[0103] Figure 14A shows two examples of devices that match a device behavior signature. The two rows shown in area 1402 correspond to data collected from two respective devices, namely, device 0 and device 1. Both of these devices are SampleCo smart meter devices. Area 1404 shows the probability that each device (for device 0 in the top row and device 1 in the bottom row) is not a SampleCo smart meter. Area 1406 shows the probability that each device (for device 0 in the top row and device 1 in the bottom row) is a SampleCo smart meter. Both devices match the SampleCo smart meter device behavior profile with a probability of over 99%.

[0104] Figure 14B shows two examples of devices that do not match the device behavior signature. The two rows shown in area 1452 correspond to data collected from two respective devices, namely, device 0 and device 1. Neither of these devices are SampleCo smart meter devices. Area 1454 shows the probability that each device (for device 0 in the top row and device 1 in the bottom row) is not a SampleCo smart meter. Area 1456 shows the probability that each device (for device 0 in the top row and device 1 in the bottom row) is a SampleCo smart meter. Both devices have less than a 1% probability of being a SampleCo smart meter device based on their behavior profiles.

[0105] D. Process for Performing Device Identification

[0106] Models trained using the techniques described herein can be used for a variety of purposes. As mentioned above, one such purpose is to efficiently perform inline device identification / classification (e.g., classification performed in real time as traffic associated with a device is observed on the network). Another such purpose is to perform offline analysis (e.g., on pcap or other previously captured network traffic files). In both inline and offline device classification, device datasets and other traffic are used to train a model (e.g., using logistic regression) and generate a set of coefficients for each device. The set of coefficients for each device can be used as a device behavior signature. The set of coefficients can then be fed into a model for use in device identification. Features extracted (from live traffic or from pcap files by a security appliance, such as a firewall) are provided to a device identification engine (e.g., implemented as analysis engine 242 and related elements shown in FIG. 2G). The device identification engine generates as output a probability that a given device matches a considered device behavior signature. Each possible device behavior signature may be looped through and the one with the highest probability (eg, subject to a quality threshold) may be assigned as the device's classification.

[0107] FIG. 15 illustrates a process for classifying an exemplary IoT device. In various embodiments, process 1500 is performed by security platform 140. Process 1500 may also be performed by other systems (e.g., systems co-located on-premises with IoT devices) where applicable. Process 1500 begins at 1502 when information associated with an IoT device's network communications is received. As one example, such information is received by security platform 140 when data appliance 102 sends a device discovery event for a given IoT device. At 1504, a determination is made that the device is unclassified (or, if applicable, reclassification should be performed). As one example, platform 140 can query database 286 to determine whether the device is classified. At 1506, a comparison against one or more behavioral signatures is performed. As one example, as described above, a variety of different machine learning and rule / heuristic-based models may be used by inline analytics engine 272. Models may be applied individually or (more typically) collectively, with different types of models being better at detecting specific kinds of devices. As mentioned above, in some situations, machine learning models that utilize packet behavior information may be effective in classifying devices. In various embodiments, at 1506, one or more such models are used to help identify a given IoT device. In particular, at 1506, a probability match to such models is determined. Then, if a given device has a probability match above a threshold for a given packet behavior signature, the device may be classified (e.g., by platform 140) as being of a particular type (e.g., a device type corresponding to the matched profile). Finally, at 1508, the classification of the IoT device (e.g., performed at 1506) is provided to a security appliance configured to apply policies to the IoT device.As mentioned above, this allows highly fine-grained security policies to be implemented in potentially mission-critical environments with minimal administrative effort.

[0108] Although the foregoing embodiments have been described in some detail for purposes of clarity of understanding, the invention is not limited to the details provided. There are many alternative ways of implementing the invention. The disclosed embodiments are illustrative and not restrictive.

Claims

1. A system comprising a processor and a memory, The processor: receiving information associated with network communications of an Internet of Things (IoT) device; determining whether the IoT device has previously been classified as a particular type of IoT device; and in response to determining that the IoT device has not previously been classified as an IoT device of the particular type, determining that a probability match of the IoT device to behavioral signatures exceeds a threshold, the behavioral signatures being generated at least in part by using machine learning models trained on features extracted from exemplary IoT devices of the particular type, wherein determining the probability match includes using at least one machine learning model trained using packet behavior information, and determining that the probability match exceeds a threshold includes determining that multiple signatures match above the threshold and selecting the resulting highest ranked match; providing a classification of the IoT device based at least in part on the probability match to a security appliance configured to apply a policy to the IoT device; It is structured as follows: The memory includes: coupled to the processor and configured to provide instructions to the processor; system.

2. the received information includes packet length sequence information; The system of claim 1 .

3. the received information includes packet inter-arrival time sequence information; The system of claim 1 .

4. the received information includes Transport Layer Security (TLS) information; The system of claim 1 .

5. An Organizationally Unique Identifier (OUI) for the IoT device is not obtained; The system of claim 1 .

6. The OUI for the IoT device corresponds to a network card; and The IoT device is not a network card. The system of claim 1 .

7. The OUI for the IoT device corresponds to a network appliance, and The IoT device is not a network device. The system of claim 1 .

8. At least a portion of the network communications are encrypted. The system of claim 1 .

9. A hostname for the IoT device is not obtained; The system of claim 1 .

10. The behavioral signature includes a set of coefficients. The system of claim 1 .

11. 1. A method comprising: receiving, by a security platform, information associated with network communications of an Internet of Things (IoT) device; determining, by the security platform, whether the IoT device has previously been classified as a particular type of IoT device; and in response to determining that the IoT device has not previously been classified as an IoT device of the particular type, determining, by the security platform, a probability match of the IoT device to behavioral signatures that exceeds a threshold, the behavioral signatures being generated at least in part by using machine learning models trained on features extracted from exemplary IoT devices of the particular type, the determining of the probability match including using at least one machine learning model trained using packet behavior information, and determining that the probability match exceeds a threshold includes determining that multiple signatures match above the threshold and selecting the resulting highest ranked match; providing, by the security platform, a classification of the IoT device based at least in part on the probability match to a security appliance configured to apply a policy to the IoT device; A method comprising:

12. A computer program stored on a non-transitory computer-readable medium, the computer program comprising computer instructions; When executed, the instructions cause the computer to: receiving information associated with network communications of an Internet of Things (IoT) device; determining whether the IoT device has previously been classified as a particular type of IoT device; and in response to determining that the IoT device has not previously been classified as an IoT device of the particular type, determining that a probability match of the IoT device to behavioral signatures exceeds a threshold, the behavioral signatures being generated at least in part by using machine learning models trained on features extracted from exemplary IoT devices of the particular type, wherein determining the probability match includes using at least one machine learning model trained using packet behavior information, and determining that the probability match exceeds a threshold includes determining that multiple signatures match above the threshold and selecting the resulting highest ranked match; providing a classification of the IoT device based at least in part on the probability match to a security appliance configured to apply a policy to the IoT device; A computer program that performs the above.