Iot device discovery and identification
By working together with data devices and IoT servers, and utilizing rule-based and machine learning classification processes, the challenge of detecting malware in IoT devices has been solved, enabling effective monitoring and management of IoT devices, improving network security, and reducing the harm caused by attacks.
Patent Information
- Application Number
- CN202180032361.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-23
- Filing Date
- 2021-06-01
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2041-06-01
AI Technical Summary
Existing technologies struggle to effectively detect and prevent malware attacks on IoT devices, especially unmanaged ones, which can lead to serious security threats and potentially catastrophic consequences.
By working together with data devices and IoT servers, and utilizing a two-part classification process based on rules and machine learning, the network behavior patterns of IoT devices are identified, policy applications are provided to enhance network security, and the discovery and management of IoT devices are achieved by combining a security platform and an AAA server.
It improves the security of IoT device networks, reduces the harm of malware attacks, enhances the monitoring and policy enforcement capabilities of IoT devices, and ensures the stable operation of the network.
Smart Images

Figure CN115486105B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 033,004, filed June 1, 2020, entitled “IOT DEVICE DISCOVERY AND IDENTIFICATION,” which is incorporated by reference herein for all purposes. BACKGROUND
[0003] Evil individuals attempt to harm computer systems in various ways. As one example, such individuals can embed or otherwise include malicious programs (“malware”) in email attachments and transmit the malware to, or cause it to be transmitted to, unsuspecting users. When executed, the malware harms the victim’s computer and can perform additional nefarious tasks (e.g., leak sensitive data, propagate to other systems, etc.). Various methods can be used to harden computers against such or other harm. Unfortunately, existing methods of protecting computers do not necessarily apply to all computing environments. Moreover, malware authors continually adapt their techniques to evade detection, and there is a continuing need for improved techniques to detect malware and prevent it from doing harm in various situations. BRIEF DESCRIPTION OF DRAWINGS
[0004] Various embodiments of the application are disclosed in the following detailed description and the accompanying drawings.
[0005] Figure 1 An example environment in which malicious activity is detected and its harm reduced is illustrated.
[0006] Figure 2A An embodiment of a data device is illustrated.
[0007] Figure 2B A functional diagram of a logical component of a data device embodiment.
[0008] Figure 2C An example event path between an IoT server and an IoT module is illustrated.
[0009] Figure 2D A device discovery event is illustrated.
[0010] Figure 2E A session event is illustrated.
[0011] Figure 2F An embodiment of an IoT module is illustrated.
[0012] Figure 2G An example way in which IoT device analysis is implemented is illustrated.
[0013] Figure 3 An embodiment of a process to passively provide AAA support for IoT devices in a network is illustrated.
[0014] Figures 4A-4C Examples of RADIUS messages sent by an IoT server on behalf of an IoT device to an AAA server in various embodiments are illustrated.
[0015] Figure 5 An embodiment of an IoT module is illustrated.
[0016] Figure 6 An example of a process to classify IoT devices is illustrated. DETAILED DESCRIPTION
[0017] The application can be implemented in numerous ways, including as a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer readable storage medium; and / or a processor, with a configured to perform instructions stored on and / or provided by a memory coupled to the processor. In this specification, these implementations, or any other form that the application can take, can be referred to as techniques. In general, the order of the steps of disclosed processes can be altered, unless such alteration would materially affect the implementation. Unless otherwise specified, components described as being configured to perform a task can be implemented as a general component that is temporarily configured to perform the task at a given time, or a specific component that is manufactured to perform the task. As used herein, the term 'processor' refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.
[0018] A detailed description of one or more embodiments of the application follows, with reference to which one or more implementations of the application are illustrated. The application is described in connection with such embodiments, but the application is not limited to any embodiment. The scope of the application is limited only by the claims and the application encompasses numerous alternatives, modifications and equivalents. For the purpose of providing a clear and concise description of the application, numerous specific details are set forth in the following description. These details are provided for embodiments of the application, and so as not to unduly complicate the application. One skilled in the art will recognize that the application can be practiced without these specific details. In other instances, specific details are not provided in order not to unnecessarily obscure the application. In the following description, numerous details are set forth to provide thorough description of the application. These details are provided for the purpose of example and the application can be practiced without some or all of these specific details. For the purpose of clarity, technical material that is known in the technical fields related to the application has not been described in detail so that the application is not unnecessarily obscured.
[0019] A system and method are described. Information associated with network communications of an Internet of Things (IoT) device is received. In some aspects, the information is received from a data appliance configured to monitor the IoT device. In some aspects, the received information includes network traffic metadata. In response to determining that the IoT device has not been classified, a two-part classification process is performed. The first part of the classification process includes an online classification, and the second part of the classification process includes a subsequent verification of the online classification. In some aspects, the first part of the classification includes a rule-based classification. In some aspects, the second part of the classification process includes a machine learning-based classification. In some aspects, performing the classification process includes determining whether any features of the IoT device are dominant. In some aspects, performing the two-part classification process includes determining a confidence that one or more features of the IoT device match a network behavior pattern identifier. The network behavior pattern identifier can be generated by the system / method / computer program product. Results of the classification process are provided to a security appliance configured to apply a policy to the IoT device. In some aspects, the results include a group classification. In some aspects, the results include a profile classification.
[0020] I. SUMMARY
[0021] Firewalls generally protect networks from unauthorized access while allowing authorized communications to pass through the firewall. Firewalls are generally a device, collection of devices, or software executing on the devices that provides firewall functionality for network access. For example, a firewall can be integrated into an operating system of a device (e.g., a computer, a smart phone, or other type of device with network communication capabilities). Firewalls can also be integrated into or executed as one or more software applications on various types of devices, such as computer servers, gateways, network / routing devices (e.g., network routers), and data appliances (e.g., security appliances or other types of specialized devices), and certain operations can be implemented in specialized hardware, such as ASICs or FPGAs, in various implementations.
[0022] Firewalls generally deny or permit network transmissions based on a set of rules. These sets of rules are generally referred to as policies (e.g., network policies or network security policies). For example, a firewall can filter inbound traffic by applying a set of rules or policies to prevent unwanted external traffic from reaching a protected device. A firewall can also filter outbound traffic by applying a set of rules or policies (e.g., allow, block, monitor, notify, or log, and / or other actions that can be specified in a firewall rule or firewall policy that can be triggered based on various criteria, such as described herein). A firewall can also filter local network (e.g., intranet) traffic by similarly applying a set of rules or policies.
[0023] Security devices (e.g., security appliances, security gateways, security services, and / or other security devices) can include various security functionality (e.g., firewall, anti-malware, intrusion prevention / detection, data loss prevention (DLP), and / or other security functionality), networking functionality (e.g., routing, quality of service (QoS), workload balancing of network-related resources, and / or other networking functionality), and / or other functionality. For example, routing functionality can be based on source information (e.g., IP address and port), destination information (e.g., IP address and port), and protocol information.
[0024] Basic packet filtering firewalls filter network communication traffic by inspecting individual packets transmitted on a network (e.g., packet filtering firewalls or first generation firewalls, which are stateless packet filtering firewalls). Stateless packet filtering firewalls typically inspect individual packets themselves and apply rules based on the inspected packets (e.g., using a combination of source and destination address information, protocol information, and port numbers of the packets).
[0025] Application firewalls can also perform application layer filtering (e.g., application layer filtering firewalls or second generation firewalls, which operate at the application layer of the TCP / IP stack). Application layer filtering firewalls or application firewalls can typically identify certain applications and protocols (e.g., web browsing using hypertext transfer protocol (HTTP), domain name system (DNS) requests, file transfer using file transfer protocol (FTP), and various other types of applications and other protocols, such as Telnet, DHCP, TCP, UDP, and TFTP (GSS)). For example, application firewalls can block unauthorized protocols that attempt to communicate through standard ports (e.g., unauthorized / policy- noncompliant protocols that attempt to sneak through using non-standard ports for the protocol can typically be identified using an application firewall).
[0026] Stateful firewalls can also perform state-based packet inspection, in which each packet is examined in the context of a series of packets associated with a stream of packets transmitted on a network. Such firewall technology is typically referred to as stateful packet inspection because it maintains a record of all connections passing through the firewall and is able to determine whether a packet is the beginning of a new connection, part of an existing connection, or an invalid packet. For example, connection state itself can be one of the criteria that triggers rules within a policy.
[0027] As noted above, advanced or next-generation firewalls can perform stateless and stateful packet filtering as well as application-layer filtering. Next-generation firewalls can also perform additional firewall techniques. For example, certain newer firewalls, sometimes referred to as advanced or next-generation firewalls, can also identify users and content (e.g., next-generation firewalls). In particular, certain next-generation firewalls are expanding the list of applications that these firewalls can automatically identify to thousands of applications. Examples of such next-generation firewalls are available from Palo Alto Networks, Inc. (e.g., Palo Alto Networks PA Series firewalls). For example, Palo Alto Networks next-generation firewalls enable enterprises to identify and control applications, users, and content using a variety of identification technologies, such as APP-ID for accurate application identification, user ID for identifying users (e.g., by user or user group), and content ID for real-time content scanning (e.g., controlling web surfing and limiting data and file transfers), rather than just ports, IP addresses, and packets. These identification technologies allow enterprises to safely enable applications using business-relevant concepts, rather than following the traditional approach provided by traditional port-blocking firewalls. Moreover, the specialized hardware used for next-generation firewalls (e.g., implemented as specialized appliances) generally provides higher levels of application review performance than software executing on general-purpose hardware (e.g., security appliances such as those provided by Palo Alto Networks, Inc. that use specialized, function-specific processing that is tightly integrated with single-pass software engines to maximize network throughput while minimizing latency).
[0028] Advanced or next-generation firewalls can also be implemented using virtualized firewalls. Examples of such next-generation firewalls are available from Palo Alto Networks, Inc. (e.g., Palo Alto Networks VM Series firewalls, which support a variety of commercial virtualization environments, including, for example, VMware® ESXi TM and NSX TM , Citrix® Netscaler SDX TMKVM / OpenStack (Centos / RHEL, Ubuntu®) and Amazon Web Services (AWS). For example, virtualized firewalls can support advanced threat prevention features available in similar or identical next-generation firewalls and physical form factors, allowing enterprises to securely allow applications to flow into and through their private, public, and hybrid cloud environments. Automation features such as VM monitoring, dynamic address groups, and REST-based APIs allow enterprises to proactively monitor VM changes and dynamically feed that context back into security policies, eliminating policy lag that can occur when VMs change.
[0029] II. Example Environment
[0030] Figure 1 The illustration shows an example of an environment in which malicious activity was detected and its harm mitigated. Figure 1 In the example shown, client devices 104-108 are laptops, desktop computers, and tablets located in the corporate network 110 of the hospital (also known as "Acme Hospital"). Data device 102 is configured to implement policies regarding communication between client devices (such as client devices 104 and 106) and nodes outside the corporate network 110 (e.g., accessible via external network 118).
[0031] Examples of such policies include those controlling traffic shaping, quality of service, and traffic routing. Other examples of policies include security policies, such as those requiring scanning for threats in incoming (and / or outgoing) email attachments, website content, files exchanged via instant messaging programs, and / or other file transfers. In some embodiments, data device 102 is also configured to enforce policies regarding traffic residing within corporate network 110.
[0032] Network 110 also includes directory services 154 and authentication, authorization, and accounting (AAA) servers 156. Figure 1 In the example shown, directory service 154 (also known as an identity provider or domain controller) utilizes Lightweight Directory Access Protocol (LDAP) or other suitable protocols. Directory service 154 is configured to manage user identity and credential information. An example of directory service 154 is a Microsoft Active Directory server. Other types of systems, such as Kerberos-based systems, can also be used instead of an Active Directory server, and the techniques described herein are adapted accordingly. Figure 1In the illustrated example, the AAA server 156 is a network access control (NAC) server. The AAA server 156 is configured to authenticate wired, wireless, and VPN users and devices to the network, assess and remediate devices for policy compliance before granting access to the network, differentiate access based on roles, and then audit and report who is on the network. One example of the AAA server 156 is a Cisco Identity Services Engine (ISE) server that utilizes the Remote Authentication Dial-In User Service (RADIUS). Other types of AAA servers can be used in conjunction with the techniques described herein, including servers that use protocols other than RADIUS.
[0033] In various embodiments, the data appliance 102 is configured to listen to communications to / from the directory service 154 and / or the AAA server 156 (e.g., passively monitor messages). In various embodiments, the data appliance 102 is configured to communicate with the directory service 154 and / or the AAA server 156 (i.e., actively pass messages with the directory service 154 and / or the AAA server 156). In various embodiments, the data appliance 102 is configured to communicate with a director (not shown) that communicates with various network elements such as the directory service 154 and / or the AAA server 156 (e.g., actively pass messages with the various network elements). Other types of servers can also be included in the network 110 as appropriate, and can communicate with the data appliance 102, and the directory service 154 and / or the AAA server 156 can also be omitted from the network 110 in various embodiments.
[0034] Although depicted in Figure 1 as having a single data appliance 102, a given network environment (e.g., the network 110) can include multiple embodiments of data appliances, whether operating individually or in concert. Similarly, although the term "network" is generally referred to in the singular for simplicity herein (e.g., referred to as "the network 110"), the techniques described herein can be deployed in various network environments of various sizes and topologies, including various mixtures of networking technologies (e.g., virtual and physical), using various networking protocols (e.g., TCP and UDP) and infrastructure (e.g., switches and routers) across various network layers, as appropriate.
[0035] The data appliance 102 can be configured to work in cooperation with a remote security platform 140. The security platform 140 can provide various services, including performing static and dynamic analysis on malware samples (e.g., via a sample analysis module 124), and providing lists of signatures for known malicious files, domains, etc. as part of a subscription to data appliances, such as the data appliance 102. As will be described in greater detail below, the security platform 140 can also provide information associated with discovery, classification, management, etc. of IoT devices that exist within a network, such as the network 110 (e.g., via an IoT module 138). In various embodiments, signatures, analysis results, and / or additional information (e.g., about samples, applications, domains, etc.) are stored in a database 160. In various embodiments, the security platform 140 comprises one or more commercially available dedicated hardware servers (e.g., with multi-core processor(s), 32G+ of RAM, gigabit network interface adapter(s), and hard drive(s)) running a typical server-level operating system (e.g., Linux). The security platform 140 can be implemented across a scalable infrastructure, including multiple such servers, solid state drives or other storage devices 158, and / or other applicable high-performance hardware. The security platform 140 can comprise several distributed components, including components provided by one or more third parties. For example, portions or all of the security platform 140 can be implemented using Amazon Elastic Compute Cloud (EC2) and / or Amazon Simple Storage Service (S3). Furthermore, just as with the data appliance 102, whenever the security platform 140 is referred to as performing a task such as storing data or processing data, it should be understood that one or more sub-components of the security platform 140 (whether alone or in cooperation with third party components) can cooperate to perform the task. As an example, the security platform 140 can perform static / dynamic analysis (e.g., via the sample analysis module 124) and / or IoT device functionality (e.g., via the IoT module 138) in cooperation with one or more virtual machine (VM) servers. One example of a virtual machine server is a physical machine comprising commercially available server-level hardware (e.g., multi-core processor, 32+ gigabytes of RAM, and one or more gigabit network interface adapters) running commercially available virtualization software such as VMware ESXi, Citrix XenServer, or Microsoft Hyper-V. In some embodiments, the virtual machine server is omitted. Furthermore, the virtual machine server can be under the control of the same entity that manages the security platform 140, but can also be provided by a third party.As one example, the virtual machine server can rely on EC2, with the rest of the security platform 140 being provided by dedicated hardware owned and controlled by the operator of the security platform 140.
[0036] Figure 2A One embodiment of a data appliance is shown. The example shown is a representation of the physical components included in the data appliance 102 in various embodiments. Specifically, the data appliance 102 includes a high-performance multi-core central processing unit (CPU) 202 and random access memory (RAM) 204. The data appliance 102 also includes storage 210, such as one or more hard disks or solid state storage units. In various embodiments, the data appliance 102 stores information for monitoring the enterprise network 110 and implementing the disclosed technology, whether in the RAM 204, the storage 210, and / or other appropriate locations. Examples of such information include application identifiers, content identifiers, user identifiers, requested URLs, IP address mappings, policies and other configuration information, signatures, hostname / URL classification information, malware profiles, machine learning models, IoT device categorization information, and the like. The data appliance 102 can also include one or more optional hardware accelerators. For example, the data appliance 102 can include a cryptographic engine 206 configured to perform encryption and decryption operations, and one or more field programmable gate arrays (FPGAs) 208 configured to perform matching, act as a network processor, and / or perform other tasks.
[0037] The functions described herein as being performed by the data appliance 102 can be provided / implemented in a variety of ways. For example, the data appliance 102 can be a dedicated device or set of devices. A given network environment can include multiple data appliances, each of which can be configured to provide services to one or more particular portions of the network, can cooperate to provide services to one or more particular portions of the network, and the like. The functions provided by the data appliance 102 can also be integrated into or performed as software on a general-purpose computer, a computer server, a gateway, and / or a network / routing device. In some embodiments, at least some of the functions described as being provided by the data appliance 102 are instead (or in addition) provided to a client device (e.g., the client device 104 or the client device 106) by software executing on the client device. The functions described herein as being performed by the data appliance 102 can also be performed at least in part by or in cooperation with the security platform 140, and / or the functions described herein as being performed by the security platform 140 can also be performed at least in part by or in cooperation with the data appliance 102, as appropriate. As one example, various functions described as being performed by the IoT module 138 can be performed by an embodiment of the IoT server 134.
[0038] Whenever the data appliance 102 is described as performing a task, individual components, a subset of components, or all components of the data appliance 102 can cooperate to perform the task. Similarly, whenever a component of the data appliance 102 is described as performing a task, a sub-component can perform the task and / or the component can perform the task in conjunction with other components. In various embodiments, portions of the data appliance 102 are provided by one or more third parties. Depending on factors such as the amount of computing resources available to the data appliance 102, various logical components and / or features of the data appliance 102 can be omitted, and the techniques described herein adapted accordingly. Similarly, additional logical components / features can be included in embodiments of the data appliance 102 as circumstances warrant. In various embodiments, one example of a component included in the data appliance 102 is an application identification engine configured to identify applications (e.g., using various application signatures, for identifying applications based on packet flow analysis). For example, the application identification engine can determine what type of traffic a session involves, such as web browse - social; web browse - news; SSH; and so on. In various embodiments, another example of a component included in the data appliance 102 is the IoT server 134 described in greater detail below. The IoT server 134 can take various forms, including as a standalone server (or set of servers), whether physical or virtual, and as circumstances warrant also collocated / merged into the data appliance 102 (e.g., as shown in Figure 1
[0039] Figure 2B is a functional diagram of logical components of an embodiment of a data appliance. The example shown is representative of logical components that can be included in the data appliance 102 in various embodiments. Unless otherwise specified, various logical components of the data appliance 102 can generally be implementable in various ways, including as a set of one or more scripts (e.g., written in Java, python, etc., as circumstances warrant).
[0040] As shown, the data appliance 102 includes a firewall, and includes a management plane 212 and a data plane 214. The management plane is responsible for managing user interactions, such as by providing a user interface for configuring policies and viewing log data. The data plane is responsible for managing data, such as by performing packet processing and session handling.
[0041] The network processor 216 is configured to receive packets from a client device, such as the client device 108, and provide them to the data plane 214 for processing. Whenever the flow module 218 identifies a packet as part of a new session, it creates a new session flow. Based on the flow lookup, subsequent packets will be identified as belonging to the session. Depending on the actual situation, SSL decryption is applied by the SSL decryption engine 220. Otherwise, the processing by the SSL decryption engine 220 is omitted. The decryption engine 220 can help the data appliance 102 review and control SSL / TLS and SSH encrypted traffic, and thereby help stop threats that might otherwise be hidden in encrypted traffic. The decryption engine 220 can also help prevent sensitive content from leaving the enterprise network 110. Decryption can be selectively controlled (e.g., enabled or disabled) based on parameters such as URL category, traffic source, traffic destination, user, user group, and port. In addition to a decryption policy (e.g., specifying which sessions to decrypt), a decryption profile can be assigned to control various options for sessions controlled by the policy. For example, a particular cipher suite and encryption protocol version can be required to be used.
[0042] The application identification (APP-ID) engine 222 is configured to determine what type of traffic a session involves. As one example, the application identification engine 222 can recognize a GET request in received data and infer that the session requires an HTTP decoder. In some cases, such as in the case of a web browsing session, the identified application can change, and such a change will be noted by the data appliance 102. For example, a user can initially browse to a company Wiki (classified as "Web Browsing - Productivity" based on the URL accessed), and then subsequently browse to a social networking site (classified as "Web Browsing - Social" based on the URL accessed). Different types of protocols have corresponding decoders.
[0043] Based on the determination made by the application identification engine 222, the packets are sent to the appropriate decoder by the threat engine 224, which is configured to assemble the packets (which can be received out of order) into the correct order, perform tokenization, and extract out information. The threat engine 224 also performs signature matching to determine what action should be taken on the packet. As needed, the SSL encryption engine 226 can re-encrypt the decrypted data. The packets are forwarded for transmission (e.g., to a destination) using the forwarding module 228.
[0044] Also as Figure 2BAs shown, policies 232 are received and stored in management plane 212. The policies can include one or more rules that can be specified using domain names and / or host / server names, and the rules can apply one or more signatures or other matching criteria or heuristics, such as for enforcing security policies on subscribers / IP flows based on various parameters / information extracted from monitored session traffic flows. Interface (I / F) communicator 230 is provided for managing communications (e.g., via (REST) API, messaging or network protocol communications or other communication mechanisms). Policies 232 can also include policies for managing communications involving IoT devices.
[0045] III. IOT DEVICE DISCOVERY AND IDENTIFICATION
[0046] Returning to Figure 1 , assume that a malicious individual (e.g., using system 120) has created malware 130. The malicious individual wants vulnerable client devices to execute a copy of malware 130, compromise the client devices, and make the client devices zombies in a botnet. The compromised client devices can then be instructed to perform tasks (e.g., cryptocurrency mining, participate in denial of service attacks, and spread to other vulnerable client devices) and report information or otherwise leak data to external entities (e.g., command and control (C&C) servers 150), as well as receive instructions from C&C servers 150, as appropriate.
[0047] Figure 1Some of the client devices depicted are commercial computing devices that are typically used within an enterprise organization. For example, client devices 104, 106, and 108 each execute a typical operating system (e.g., macOS, Windows, Linux, Android, etc.). Such commercial computing devices are typically provisioned and maintained by an administrator (e.g., as company-issued laptops, desktops, and tablets, respectively), and are typically operated in conjunction with a user account (e.g., managed by a directory service provider (also known as a domain controller) configured with user identity and credential information). As one example, an employee Alice can be issued a laptop 104 that she uses to access her ACME-related email and perform various ACME-related tasks. Other types of client devices (generally referred to herein as Internet of Things, or IoT, devices) are also increasingly present in networks, and are typically not "managed" by an IT department. Some such devices (e.g., teleconference devices) can be found across a variety of different types of enterprises (e.g., as IoT whiteboards 144 and 146). Such devices can also be vertical-specific. For example, infusion pumps and computerized tomography scanners (e.g., CT scanner 112) are examples of IoT devices that can be found in a healthcare enterprise network (e.g., network 110), and robotic arms are examples of devices that can be found in a manufacturing enterprise network. Further, consumer-oriented IoT devices (e.g., cameras) can also be present in enterprise networks. Like commercial computing devices, IoT devices present within a network can communicate with resources both internal and external to such a network (or both, as the case can be).
[0048] Like commercial computing devices, IoT devices are targets of nefarious individuals. Unfortunately, the presence of IoT devices in a network can present several unique security / management challenges. IoT devices are often low-power devices or special-purpose devices, and are often deployed without the knowledge of a network administrator. Even when such an administrator is aware, it can not be possible to install end-point protection software or agents on the IoT devices. IoT devices can be managed by third-party cloud infrastructure and communicate directly / individually with the third-party cloud infrastructure using proprietary (or other non-standard) protocols (e.g., industrial thermometer 152 communicates directly with cloud infrastructure 126). This can obfuscate attempts to monitor network traffic to / from such devices to make decisions about when threats or attacks against the devices are occurring. In addition, some IoT devices (e.g., in a healthcare environment) are mission-critical (e.g., network-connected surgical systems). Unfortunately, compromise of an IoT device (e.g., by malware 130) or misuse of security policies against traffic associated with an IoT device can have potentially catastrophic impact. Using the techniques described herein, the security of a heterogeneous network including IoT devices can be improved, and the damage caused to such a network can be reduced.
[0049] In various embodiments, data appliance 102 includes IoT server 134. In some embodiments, IoT server 134 is configured to identify IoT devices within a network (e.g., network 110) in cooperation with IoT module 138 of security platform 140. For example, such identification can be used by data appliance 102 to assist in formulating and enforcing policies regarding traffic associated with IoT devices, and to enhance the functionality of other elements of network 110 (e.g., to provide context information to AAA 156). In various embodiments, IoT server 134 incorporates one or more network sensors configured to passively sniff / monitor traffic. One example way to provide such network sensor functionality is as a tap interface or switch mirror port. Other methods of monitoring traffic can also be used, as appropriate (in addition or instead).
[0050] In various embodiments, IoT server 134 is configured to provide logs or other data (e.g., collected from passively monitoring network 110) to IoT module 138 (e.g., via front end 142). Figure 2C An example event path between IoT server and IoT module is illustrated. IoT server 134 sends device discovery events and session events to IoT module 138. Figure 2D and 2EAn example discovery event and session event are illustrated in FIG. 23. In various embodiments, IoT server 134 sends a discovery event each time it observes a packet that can uniquely identify or confirm a device identity (e.g., each time a DHCP, UPNP, or SMB packet is observed). Each session that a device has (with other nodes inside or outside the device's network) is described within a session event, which summarizes information about the session (e.g., source / destination information, number of packets received / sent, etc.). Depending on practical considerations, multiple session events can be batched together by IoT server 134 before being sent to IoT module 138. In Figure 2E In the example shown, two sessions are included. IoT module 138 provides device classification information to IoT server 134 via a device verdict event (234).
[0051] One example way to implement IoT module 138 is to use a microservices-based architecture. IoT module 138 can also be implemented using different programming languages, databases, hardware and software environments, as appropriate, and / or as services that enable implementation of message passing, context constraints, autonomous development, independent deployability, decentralization, and building and publishing with automated processes. One task performed by IoT module 138 is to identify IoT devices in the data provided by IoT server 134 (and by other embodiments of data appliances such as data appliances 136 and 148) and to provide additional contextual information about those devices (e.g., back to the respective data appliance).
[0052] Figure 2FAn embodiment of an IoT module is illustrated. Region 294 depicts a Spark application set that runs across all tenants' data intervals (e.g., every five minutes, every hour, and every day). Region 296 depicts a Kafka message bus. Session event messages received by IoT module 138 (e.g., from IoT server 134) bundle multiple events observed at IoT server 134 together (e.g., to save bandwidth). Transform module 236 is configured to flatten the received session events into individual events and publish them at 250. Flattened events are aggregated by aggregation module 238 using various different aggregation rules. One example rule is "aggregate all event data for a particular device and each (APP-ID) application it uses for a time interval (e.g., 5 minutes)." Another example rule is "aggregate all event data for a particular device that communicates with a particular destination IP address for a time interval (e.g., 1 hour)." For each rule, aggregation engine 238 tracks a list of attributes that need to be aggregated (e.g., a list of applications used by a device or a list of destination IP addresses). Feature extraction module 240 extracts features from the attributes (252). Analysis module 242 uses the extracted features to perform device classification (e.g., using supervised and unsupervised learning), the results (254) of which are used to drive other types of analysis (e.g., via operations intelligence module 244, threat analysis module 246, and anomaly detection module 248). Operations intelligence module 244 provides analysis related to OT frameworks and operations or business intelligence (e.g., how to use a device). Alerts (256) can be generated based on the results of the analysis. In various embodiments, MongoDB 258 is used to store aggregated data and feature values. Background service 262 receives data aggregated by the Spark application and writes the data to MongoDB 258. API server 260 pulls and fuses data from MongoDB 258 to service requests received from frontend 142.
[0053] Figure 2G An example way to implement IoT device identification analysis is illustrated (e.g., within IoT module 138 as an embodiment of analysis module 242 and related elements). Discovery events and session events (e.g., as described in Figure 2D and 2EThe features are extracted by a feature engine 276 (e.g., can be implemented using Spark / MapReducer). The security platform 140 enriches 266 the raw data with additional contextual information, such as geo-location information (e.g., of source / destination addresses). During meta-data feature extraction 268, features such as the number of packets sent from an IP address in a time interval, the number of applications used by a particular device in that time interval, and the number of IP addresses contacted by that device in that time interval are constructed. The features are delivered (e.g., on a message bus) to an online analytics engine 272 (e.g., in JSON format) in real-time and stored (e.g., in a feature database 270 in an appropriate format such as Apache Parquet / DataFrame) for subsequent querying (e.g., during offline modeling 299).
[0054] In addition to features constructed from meta-data, a second type of feature can be constructed by the IoT module 138 (274), referred to herein as analytic features. An example analytic feature is a feature constructed over time using time-series data based on the temporal evolution of the data. Similarly, analytic features are delivered to the analytics engine 272 in real-time and stored in the feature database 270.
[0055] The online analytics engine 272 receives features on the message bus via a message processor. One task performed is activity classification 278, which attempts to identify the activity associated with a session (such as a file download, login / authentication process, or disk backup activity) based on the received feature values / session information and attach any applicable labels. One way to implement activity classification 278 is via a neural network-based multilayer perceptron combined with a convolutional neural network.
[0056] Suppose, as a result of activity classification, it is determined that a particular device is conducting a print activity (i.e., using a print protocol) and also periodically contacts a resource owned by HP (e.g., by calling an HP URL and using it to report status information to check for updates). In various embodiments, the classification information is delivered to both a clustering process (unsupervised) and a prediction process (supervised). If either process results in a successful classification of the device, the classification is stored in a device database 286.
[0057] Devices can be clustered into multiple clusters (e.g., behaving like printers, behaving like HP devices, etc.) by the phase one clustering engine 280 based on their attributes and other behavioral patterns. One way to implement the clustering engine 280 is to use the extreme gradient boosting framework (e.g., XGB). Phase one classifiers can be useful for classifying devices that have not been seen before but are similar to existing known devices (e.g., a new vendor of thermostats starts selling thermostat devices that behave similarly to known thermostats).
[0058] As shown in Figure 2G classification information is also provided to the classifier set 282 and a prediction is performed based on the features provided for the device. Two possibilities can arise. In a first scenario, it is determined that the probability of the device matching a known device profile is high (i.e., a high confidence score). If so, information about the device is provided to the phase two classifier (284) which makes a final determination of the identity of the device (e.g., using the information provided to it and any additional applicable contextual information) and updates the device database 286 accordingly. One way to implement the phase two classifier is to use the gradient boosting framework. In a second scenario, assume the confidence score is low (e.g., the device matches an HP printer and an HP laptop with 50% confidence). In this scenario, the information determined by the classifier 282 can be provided to the clustering engine 280 as additional information that can be used for clustering.
[0059] Figure 2G Also shown in FIG. 3 is an offline modeling module 299. The offline modeling module 299 is in contrast to the online analysis engine 272 in that it is not time constrained (whereas the online analysis engine 272 attempts to provide device classification information (e.g., as messages 234) in real time). Periodically (e.g., once a day or once a week), the offline modeling module 299 (e.g., implemented using Python) re-creates the models used by the online analysis module 272. The activity modeling engine 288 builds models for the activity classifier 278, which are also used for the device type model (296) used by the classifier for device identification during online analysis. The baseline modeling engine 290 builds models of baseline behavior of devices, which are also used when modeling anomalies (292) for a particular type of device and threats (294) for a particular type of device, such as kill chains. In various embodiments, the models generated are stored in the model database 298.
[0060] IV. Network Entity ID AAA
[0061] As previously mentioned, assume that ACME issues a laptop 104 to Alice. As Alice uses the laptop to access various resources, various components of the network 110 will cooperate to authenticate the laptop of Alice. As one example, when Alice connects the laptop 104 to a wireless access point (not shown) located within the network 110, the wireless access point can communicate with the AAA server 156 (whether directly or indirectly) while providing network access. As another example, when Alice uses the laptop 104 to access her ACME email, the laptop 104 can communicate with the directory service 154 (whether directly or indirectly) while obtaining her inbox, etc. As a commercial laptop running a commercial operating system, the laptop 104 is capable of generating appropriate AAA messages (e.g., RADIUS client messages) that will help the laptop 104 to obtain access to the appropriate resources that it needs.
[0062] As noted previously, one problem caused by IoT devices (e.g., device 146) in a network such as 110 is that such devices are typically "unmanaged" (e.g., not configured, provisioned, managed, etc. by a network administrator), do not support protocols such as RADIUS, and thus cannot integrate with AAA services such as other devices (such as laptop 104). Various approaches can be employed to provide network access within network 110 for IoT devices, each with drawbacks. One option is for ACME to restrict IoT devices to use a guest network (e.g., via a pre-shared key). Unfortunately, this can limit the utility of the IoT devices if they cannot communicate with other nodes within network 110 to which they should have legitimate access. Another option is to allow IoT devices unrestricted access to network 110, mitigating the security benefits of having a segmented network. Yet another option is for ACME to manually specify rules that govern how a given IoT device should be able to access resources in network 110. For various reasons, this approach is typically untenable / feasible. As one example, an administrator can often not be involved in the deployment of IoT devices, and thus will not know what policies should be included (e.g., in data appliance 102) for such devices. Even in cases where an administrator can manually configure policies for a particular IoT device (e.g., for a device such as 112) in appliance 102, keeping such policies up-to-date is error prone, and typically untenable given the sheer number of IoT devices that can exist in a given network 110. Moreover, such policies will likely be overly simplistic (e.g., assign CT scanner 112 to a particular network by IP address and / or MAC address), and not allow for more granular control over connections / policies involving CT scanner 112 (e.g., dynamically include with policies applicable to surgical devices versus point-of-sale devices). Moreover, even in cases where CT scanner 112 is manually included in data appliance 102, as noted previously, IoT devices will typically not support technologies such as RADIUS, and the benefits of having such an AAA server manage network access for CT scanner 112 will be limited as compared to other types of devices (e.g., laptop 104) that more fully support such technologies. As will be described in greater detail below, in various embodiments, data appliance 102 (e.g., via IoT server 134) is configured to provide support for AAA functionality to IoT devices present in network 110 in a passive manner.
[0063] In the following discussion, assume that Alice's department at ACME recently purchased interactive whiteboard 146, enabling Alice to collaborate with other ACME employees as well as individuals outside of ACME (e.g., Bob, a researcher at Beta University, with his own network 114, data device 136, and whiteboard 144). As part of the initial setup of whiteboard 146, Alice connects it to a power source and provides it with a wired connection (e.g., to a socket in the conference room) or wireless credentials (e.g., credentials for use by occupants of the conference room). When whiteboard 146 provides a network connection, IoT server 134 (e.g., via a mechanism such as the network sensor described above) recognizes whiteboard 146 as a new device within network 110. One action taken in response to this detection is to communicate with security platform 140 (e.g., create a new record for whiteboard 146 in database 160, and retrieve any currently available contextual information associated with whiteboard 146 (e.g., obtain the manufacturer of whiteboard 146, the model of whiteboard 146, etc.)). Any contextual information provided by security platform 140 can be provided to (and stored at) data device 102, which in turn can provide it to directory service 154 and / or AAA server 156. Depending on the actual situation, IoT module 138 can provide data device 102 with updated contextual information about whiteboard 146 as it becomes available. Also, data device 102 can similarly provide (e.g., via IoT server 134) security platform 140 with ongoing information about whiteboard 146. Examples of such information include observations about whiteboard 146's behavior on network 110 (e.g., statistics about connections it establishes), which security platform 140 can use to build a behavioral profile for devices such as whiteboard 146. Similar behavioral profiles can be built by security platform 140 for other devices (e.g., whiteboard 144). Such profiles can be used for various purposes, including detecting anomalous behavior. As one example, data device 148 can use information provided by security platform 140 to detect whether thermometer 152 is operating abnormally compared to historical observations of thermometer 152, and / or compared to other thermometers (not shown) of similar model, manufacturer, or more generally, including thermometers that exist in other networks. If anomalous behavior is detected (e.g., by data device 148), appropriate remedial measures can be taken automatically, such as restricting thermometer 152's access to other nodes on network 116, generating an alert, etc.
[0064] Figure 3An embodiment of a process to passively provide AAA support for an IoT device in a network is illustrated. In various embodiments, the process 300 is performed by the IoT server 134. The process begins at 302, where a set of packets transmitted by an IoT device is obtained. As one example, when the whiteboard 146 is first provisioned on the network 110, such packets can be passively received by the IoT server 134 at 302. Packets can also be received at 302 during subsequent use of the whiteboard 146 (e.g., when Alice conducts a whiteboard session with Bob via the whiteboard 144). At 304, at least one packet included in the set of data packets is analyzed. As one example of processing performed at 304, the IoT server 134 determines that the packet received at 302 is being transmitted by the whiteboard 146. One action that the IoT server 134 can take is to identify the whiteboard 146 as a new IoT device on the network 110, and to obtain context information from the IoT module 138 (if available). At 306, the IoT server 134 transmits an AAA message on behalf of the IoT device that includes information associated with the IoT device. One example of such information is Figure 4A illustrated. As mentioned previously, the whiteboard 146 does not support the RADIUS protocol. However, the IoT server 134 can generate a message such as Figure 4A depicted in Figure 4A on behalf of the whiteboard 146 (e.g., using information received at 302, and also information received from the security platform 140, as applicable). As mentioned previously, when the IoT server 134 provides information about the whiteboard 146 to the IoT module 138, the IoT module 138 can take various actions, such as creating a record for the whiteboard 146 in the database 160, and populating the record with context information about the whiteboard 146 (e.g., determining its manufacturer, model number, etc.). As additional context information about the whiteboard 146 is collected by the security platform 140, its profile can be updated and propagated to the data appliance 102. When the whiteboard 146 is initially provisioned within the network 110, no additional context information can be available (e.g., the security platform 140 can not have such additional information, or providing such information by the security platform 140 to the IoT server 134 can not be immediate). Thus, and as depicted in Figure 4A , the RADIUS message generated by the IoT server 134 on behalf of the whiteboard 146 can include limited information. When additional context information is received (e.g., by the IoT server 134 from the IoT module 138), subsequent RADIUS messages sent by the IoT server 134 on behalf of the whiteboard 146 can be enriched with such additional information. Examples of such subsequent messages are illustrated in Figure 4B and 4C . Figure 4BAn example of a RADIUS message that IoT server 134 can send on behalf of whiteboard 146 is illustrated once context information about whiteboard 146 is provided by IoT module 138 (e.g., it contains a database of context information about a wide variety of IoT devices). In Figure 4B In the illustrated example, context information is included, such as the manufacturer of the whiteboard (Panasonic) and the nature of the device (e.g., it is an interactive whiteboard). Such context information can be used by an AAA server, such as AAA server 156, to provide AAA services to whiteboard 146 (without having to modify whiteboard 146), such as by automatically providing it on a subnetwork dedicated to teleconference equipment. Other types of IoT devices can also be automatically grouped based on attributes such as device type, use, etc. (e.g., automatically providing critical surgical equipment on a subnetwork dedicated to such equipment, and thus isolated from other devices on the network). Such context information can be used to enforce policies, such as traffic shaping policies, such as a policy that gives whiteboard 146 packets priority over social networking packets (e.g., as determined using APP-ID). Fine-grained policies can similarly be applied to communications with critical surgical equipment (e.g., preventing any device that communicates with such equipment from having an out-of-date operating system, etc.). In Figure 4C In the illustrated example, IoT server 134 includes more additional context information in the RADIUS message on behalf of whiteboard 146. Such additional context information includes additional attribute information, such as device model, operating system, and operating version. When whiteboard 146 was initially provisioned in network 110, Figure 4C All of the context information depicted in can not be available at the time. As electronic whiteboard 146 is used within network 110 over time, additional context information can be collected (e.g., as IoT server 134 continues to passively observe packets from electronic whiteboard 146 and provide information to security platform 140). This additional information can be leveraged (e.g., by data appliance 102) to enforce fine-grained policies. As one example, as shown in Figure 4C In the illustrated example, whiteboard 146 runs a Linux-based and has a specific operating system with a 3.16 version. IoT devices will often run non-upgradeable / non-patchable operating system versions. When vulnerabilities are developed against these operating systems, such devices can pose a security risk. Data appliance 102 can implement security policies based on context information, such as by isolating IoT devices with out-of-date operating systems from other nodes in network 110 (or otherwise limiting their access), while permitting less restrictive network access to devices with current operating systems, etc.
[0065] Figures 4A-4CAn example of a RADIUS access request message is depicted. Depending on the actual situation, the IoT server 134 can generate various types of RADIUS messages on behalf of the whiteboard 146. As one example, a RADIUS accounting start message can be triggered when traffic from the whiteboard 146 is first observed. Periodic RADIUS accounting interim update messages can be sent while the whiteboard is in use, and a RADIUS accounting stop message can be sent when the whiteboard 146 goes offline.
[0066] V. IoT Device Discovery and Identification
[0067] As discussed above, one task performed by the security platform 140 (e.g., via the IoT module 138) is IoT device classification. For example, when the IoT server 134 transmits a device discovery message to the IoT module 138, the IoT module 138 attempts to determine the classification of the device and responds accordingly (e.g., with a message as shown in the decision 234 of FIG. 2). The device is associated with a unique identifier by the IoT module 138, such that subsequent classification of the device need not be performed (or is performed less frequently than otherwise) depending on the actual situation. As also discussed above, the determined classification can be used to enforce policies (e.g., by the data appliance 102) for traffic to / from the device. Figure 2C
[0068] A variety of methods can be used to classify devices. A first method is to perform classification based on a rule set / heuristic that utilizes static attributes of the device, such as the organization unique identifier (OUI), the type of application it executes, etc. A second method is to perform classification using machine learning techniques that utilize dynamic but pre-defined attributes of the device extracted from the device's network traffic (e.g., the number of packets sent per day). Unfortunately, both of these methods have weaknesses.
[0069] Rule-based approaches typically require manual creation of separate rules for each type of IoT device (describing which attributes / values should be used as signatures of the type of device signature). One challenge presented by this approach is determining which signatures are both relevant to identifying the device and unique among other device signatures. Further, using a rule-based approach, a limited number of static attributes (e.g., user agent, OUI, URL destination, etc.) can be readily extracted from traffic. Attributes typically need to be simple enough that they exhibit a pattern that can be matched by a regular expression. Another challenge is that new static attributes that can be identified / are present can be identified when new devices enter the market (e.g., a new brand or new model of CT scanner is provided). Another challenge is that in order to trigger a rule, all matching attributes in the network traffic need to be collected. Fewer attributes cannot lead to a decision. For example, a signature can require a particular device with a particular OUI to connect to a particular URL. Having the OUI alone can already be a sufficient indicator of the device identity, but the signature will not trigger until the URL is also observed. This will result in further delay in determining the device identity. Another challenge is maintaining and updating signatures as static attributes that change over time (e.g., due to updates to the device or services used by the device). For example, a particular device can have been manufactured using a first type of network card initially, but over time, the manufacturer can have switched to a different network card (which will exhibit a different OUI). If the rule-based system is not aware of such changes, false positives can result. Yet another challenge is scaling signature generation / verification when the number of new IoT devices coming online each day approaches millions of new device instances. Thus, newly created rules can conflict with existing rules and result in false positives in classification.
[0070] Machine learning-based approaches typically involve creating a trained model based on static and / or dynamic features extracted from network traffic. The prediction result on network data from a new IoT device is based on a pre-trained model that provides the identity of the device with an associated accuracy. An example of a problem with machine learning approaches is as follows. In the case of prediction on each new device or devices that do not have a constant / unique ID (e.g., MAC address), the computation time required to reach a desired precision can be unacceptable. Thousands or tens of thousands of features can need to be generated, and these features can change over a pre-defined time window, taking a significant amount of time before a sufficient number of features are available for effective prediction (which can defeat the purpose of policy enforcement). Further, if the goal is to minimize latency in prediction, the cost of building and maintaining a large data pipeline for network data streams can be high. Yet another problem is that noise from irrelevant features specific to a given deployment environment can degrade the accuracy of the prediction. Further, when the number of device types reaches tens of thousands or higher, there are challenges in maintaining and updating the model.
[0071] In various embodiments, the security platform 140 addresses the problems of each of the above two approaches by using a hybrid approach to classification. In an example hybrid approach, a network behavior pattern identifier (also referred to herein as a pattern ID) is generated for each type of device. In various embodiments, the pattern ID is a list of combined attribute or sequence features, with their corresponding probabilities (as importance scores of the features or behavior classes), forming a distinct network behavior description, and can be used to identify the type of IoT device. The pattern ID can be stored (e.g., in a database) and used to identify / verify the identity of a device.
[0072] When training on a set of attributes, certain methods such as the extreme gradient boosting framework (e.g., XGB) can provide a top list of important features (whether static attributes, dynamic attributes, and / or aggregated / transformed values). Once established, the pattern ID can be used to uniquely identify the device type. If certain features are dominant for a device (e.g., a particular static feature such as contacting a highly specific URL at boot time identifies the device with 98% confidence), they can be used to automatically generate rules. Even in the absence of dominant features, the representation of the top features can still be used as the pattern ID (e.g., in the case where a collection of multiple features are joined into a pattern). By training on a dataset that includes all known models (and all known IoT devices), potential conflicts between models / unique identifying features can be avoided. Furthermore, the pattern ID need not be human-readable (but can be stored, shared, and / or reused for identification purposes). Significant time savings can also be achieved with this approach, such that it can be used for near real-time classification. Once a dominant feature is observed, a particular device can be classified (rather than having to wait for a large number of features to occur).
[0073] An example of data that can be used to create a pattern ID for a "Teem Room Display iPad" device can include the following (a complete list automatically generated by training a multivariate model or training multiple binary models):
[0074] * Apple device (100%)
[0075] * Special iPad (> 98.5%)
[0076] * Teem Room App (> 95%)
[0077] * Meeting volume pattern VPM-17 (> 95%)
[0078] * Server in the cloud (> 80%).
[0079] An example way to implement the hybrid approach is as follows. A neural network based machine learning system can be used for automatic pattern ID training and generation. Examples of features that can be used to train the neural network model include static features extracted from network traffic (e.g., OUI, hostname, TLS fingerprint, matching L7 payload signatures, etc.) and sequential features extracted from network traffic but not specific to the environment (e.g., application, L7 attributes of the application, volume ranges converted to categorical features, etc.). A lightweight data pipeline can be used to stream selected network data for real-time generation of features. A prediction engine can be used to import models and provide caching to minimize latency in predictions. In predictions, a short (e.g., minute-based) aggregation can be used to stabilize selected sequential features. Custom data normalization, enrichment, aggregation, and transformation techniques can be used to design sequential features. For better accuracy, longer aggregation windows can be used in training. The accuracy of predictions can be improved by merging and aggregating features over time. A backend feedback engine can be used to route results of a "slow path" prediction system (e.g., a machine learning based approach including a device type modeling subsystem and a device group modeling subsystem), which helps to expand the attributes used for pattern ID prediction. When there are not enough samples or features available, a device group model can be trained to compensate for issues with a device type model to improve accuracy above an acceptable threshold (e.g., assign a prediction result based on a pre-defined type set, some of which are used with another subsystem to cluster unlabeled similar types of devices). Finally, a decision module can be used to publish results from the real-time prediction engine.
[0080] Example advantages of a hybrid classification approach such as described herein are as follows. First, fast convergence can occur, allowing a given device to be potentially identified within minutes or seconds. Second, it addresses individual issues of rule-based and machine learning based systems. Third, it provides stability and consistency of prediction results. Fourth, it has scalability to support tens of thousands (or more) of different types of IoT devices. Prediction is typically only needed on new devices (even if a given device lacks a unique ID assignment, such as identification based on L3 network traffic).
[0081] Figure 5 One embodiment of module 138 is shown in FIG. 1. One example way to implement IoT module 138 is to use a microservice based architecture, where services are fine-grained and protocols are lightweight. Services can also be implemented using different programming languages, databases, hardware and software environments (as practical), and / or relatively small services that enable message passing, context constraints, autonomous development, independent deployability, decentralization, and leverage automated processes for building and releasing.
[0082] As mentioned previously, in various embodiments, the security platform 140 periodically receives information about IoT devices on a network (e.g., network 110) (e.g., from data appliance 102). In some cases, the IoT device has been previously classified by the security platform 140 (e.g., a CT scanner installed on network 110 last year). In other cases, the IoT device will be newly seen by the security platform 140 (e.g., whiteboard 146 is installed for the first time). Assume that the given device has not been previously classified by the security platform 140 (e.g., there is no entry for the device in database 286 that stores a set of unique device identifiers and associated device information). As Figure 5 As illustrated in FIG. 5, information about the new device can be provided to two different processing pipelines for classification. Pipeline 504 represents a "fast path" classification pipeline (corresponding to a pattern ID based approach), and pipeline 502 represents a "slow path" classification pipeline (corresponding to a machine learning based approach).
[0083] In pipeline 504, a fast path feature design is performed (508) to identify applicable static and sequence features of the device. A fast path prediction is performed (510) using a pattern ID or a previously built model (e.g., a model built based on the most important features and using offline processing pipeline 506). A confidence score is determined (512) for the device matching a particular pattern. If the confidence score for the device satisfies a pre-trained threshold (e.g., based on the overall prediction accuracy of module 138 or its components, such as 0.9), a classification can be assigned to the device (in device database 516) or updated as appropriate. Initially, the confidence score will be based on near real-time fast path processing. The advantage of this approach is that data appliance 102 can very quickly begin applying policies to the device's traffic (e.g., within minutes of module 138 identifying the device as new / unclassified). Device 102 can be configured to be fail-safe (e.g., reduce / limit the device's ability to access various network resources) or fail-dangerous (e.g., allow the device to have broad access) pending a classification decision from system 140. As additional information becomes available (e.g., via slow path processing), the confidence score can be based on this additional information as appropriate (e.g., increase the confidence score or correct / rectify mistakes made during fast path classification).
[0084] Examples of features (e.g., static attributes and sequence features) that can be used include the following. A pattern ID can be any combination of these attributes that include logical conditions as follows:
[0085] * OUI in mac address
[0086] * hostname string from decoded protocol
[0087] * User agent string from HTTP and other cleartext protocols
[0088] * System name string from decoded SNMP responses
[0089] * OS, hostname, domain, and username from decoded LDAP protocol
[0090] * URL from decoded DNS protocol
[0091] * SMB version, command, errors from decoded SMB protocol
[0092] * TCP flags
[0093] * Option strings from decoded DHCP protocol
[0094] * Strings from decoded IoT protocols, such as Digital Imaging and Communications in Medicine (DICOM)
[0095] * Inbound application list from local network
[0096] * Inbound application list from the internet
[0097] * Outbound application list to local network
[0098] * Outbound application list to the internet
[0099] * Inbound server port list from local network
[0100] * Inbound server port list from the internet
[0101] * Outbound server port list to local network
[0102] * Outbound server port list to the internet
[0103] * Inbound IP list from local network
[0104] * Inbound URL list from the internet
[0105] * Outbound IP list to local network
[0106] * Outbound URL list to the internet.
[0107] In some cases, the confidence score determined at 512 can be very low. One reason this can occur is because the device is a new type (e.g., a new type of IoT toy or other type of product that the security platform 140 has not previously analyzed), and there is no corresponding pattern ID on the security platform 140 for the device. In such scenarios, information about the device and the classification result can be provided to the offline processing pipeline 506, which, for example, can perform clustering (514) on the behavior exhibited by the device and other applicable information (e.g., to determine that the device is a wireless device, behaves like a printer, uses DICOM protocol, etc.). The clustering information can be applied as a label and flagged for additional study 518, as appropriate, where any subsequently seen similar devices are automatically grouped together. If as a result of the study, additional information is determined about the given device (e.g., it is identified as corresponding to a new type of consumer-facing IoT meat thermometer), then the device (and all other devices with similar characteristics) can be relabeled accordingly (e.g., as XYZ brand meat thermometer), and an associated pattern ID is generated and made available by the pipeline 502 / 504, as appropriate (e.g., after the model is rebuilt). In various embodiments, the offline modeling 520 is a process that runs daily to train and update various models 522 for IoT device identification. In various embodiments, the models are refreshed daily to cover new labeled devices, and weekly to reflect behavior changes (for the slow path pipeline 502), and to adapt to new features and data insights added during that period. Note that when a new type of device is added to the security platform 140 (i.e., a new device pattern is created), multiple existing device patterns can be affected, which requires updating the feature list or their importance scores. This process can be performed automatically (a major advantage compared to rule-based solutions).
[0108] For fast path modeling, neural network based models (e.g., FNN) and general machine learning models (e.g., XGB) are widely used for multivariate classification models. Binary models are also built for selected profiles to help improve results and provide input for clustering. A binary model gives a yes / no answer to the identity of a device or certain behavior of a device. For example, a binary model can be used to determine whether a device is a type of IP phone or is unlikely to be an IP phone. A multivariate model will have many outputs normalized with a probability of 1. Each output corresponds to a type of device. Although binary models are generally faster, they require a device to traverse many models in a prediction to find the correct “yes” answer. This can be accomplished in one step by a multivariate model.
[0109] The slow path pipeline 502 is similar to the pipeline 504 in that features are extracted (524). However, the features used by the pipeline 502 often take a period of time to build. As one example, the "bytes sent per day" feature will take a day to collect. As another example, certain usage patterns can take a period of time to occur / be observed (e.g., where a CT scanner is used to perform scans every hour (a first behavior), data is backed up every day (a second behavior), and a manufacturer's website is checked for updates every week (a third behavior). The slow path pipeline 502 invokes a multivariate classifier (526) to attempt to classify the new device instance according to the full set of features. The features used are not limited to static or sequence features, but also include quantity and time series based features. This is often referred to as phase one prediction. For certain profiles, when the phase one prediction results are not optimal (with a lower confidence), a phase two prediction is used to attempt to improve the results. The slow path pipeline 502 invokes a set of decision tree classifiers (528) supported by additional imported device context to classify the new device instance. The additional device context is imported from external sources. As an example, URLs that the device has connected to can have been given a category and risk-based reputation, which can be included as a feature. As another example, applications used by the device can have been given a category and risk-based score, which can be included as a feature. By combining the results from the phase one prediction 526 and the phase two prediction 528, a final determination of the slow path classification can be made with an exported confidence score.
[0110] Two phases are typically included in the slow path pipeline 502. In the slow path pipeline, in some embodiments, a phase one model is built with a multivariate classifier based on neural network technology. The phase two of the slow path pipeline is typically a set of decision based models with additional logic to handle phase one's probabilistic related anomalies. In prediction, phase two will incorporate inputs from phase one, apply rules and context to validate phase one's output, and generate the final output of the slow path. The final output will include the identity of the device, an overall confidence score, a pattern ID that can be used for future fast path pipeline 504, and an explanation list. The confidence score is based on the reliability and accuracy of the model (models also have a confidence score), and the probability as part of the classification. The explanation list will include a list of features that contributed to the result. As mentioned above, if the result deviates from a known pattern ID, an investigation can be triggered.
[0111] In some embodiments, for the slow path modeling, two types of models are built, one for individual identity and one for group identity. For example, it is often more difficult to tell the difference between two printers from different vendors or with different models (e.g., because printers tend to exhibit network behavior, speak similar protocols, etc.) than to distinguish between a printer and a thermometer. In various embodiments, various printers from various vendors are included in a group, and a "printer" model is trained for group classification. This group classification result can provide better accuracy than a specific model for a specific printer, and can be used to update the confidence score of the device, or, as the case can be, provide a reference and verification for classification based on individual profile identity.
[0112] Figure 6 An example of a process for classifying IoT devices is illustrated. In various embodiments, process 600 is performed by security platform 140. Process 600 can also be performed by other systems, as the case can be, such as systems collocated in the field with IoT devices. At 602, process 600 begins when information associated with a network communication of an IoT device is received. As one example, such information is received by security platform 140 when data appliance 102 transmits a device discovery event for a given IoT device to security platform 140. At 604, it is determined that the device has not been classified (or, as the case can be, that reclassification should be performed). As one example, platform 140 can query database 286 to determine whether the device has already been classified. At 606, a two-part classification is performed. As an example, at 606, a two-part classification is performed by platform 140, providing information about the device to both fast path classification pipeline 504 and slow path classification pipeline 502. Finally, at 608, the results of the classification process performed at 606 are provided to a security appliance configured to apply a policy to the IoT device. As mentioned above, this allows for highly granular security policies to be implemented in potentially mission-critical environments with minimal administrative effort.
[0113] In a first example of performing process 600, assume that an Xbox One game console has connected to network 110. During classification, it can be determined that the device has the following dominant features: a "Vendor = Microsoft" feature with 100% confidence, a "Communicates with Microsoft cloud servers" feature with 89.7% confidence, and a matching "Game console" feature with 78.5% confidence. These three features / confidence scores can collectively match against the profile ID set (a process accomplished by the neural network-based predictions) to identify the device as an Xbox One game console (i.e., find a profile ID match at 512 that satisfies the threshold). In a second example, assume that an AudioCodes IP phone has connected to network 110. During classification, it can be determined that the device matches a "Vendor = AudioCodes" feature with 100% confidence, a "Is an IP audio device" feature with 98.5% confidence, and a "Behaves like a local server" feature with 66.5% confidence. These three features / confidence scores also match against the profile ID set, but in this scenario, assume that no existing profile ID matches with sufficient confidence. The information about the device can then be provided to clustering process 514, and depending on the actual situation, a new profile ID can ultimately be generated and associated with the device (and used to classify future devices).
[0114] Depending on the actual situation, security platform 140 can recommend a particular policy based on the determined classification. The following are examples of policies that can be implemented:
[0115] * Deny all internet traffic for infusion pumps (regardless of vendor)
[0116] * Deny all internet traffic for GE ECG machines, except traffic to / from certain GE hosts
[0117] * For all CT scanners (regardless of vendor), only allow internal traffic to Picture Archiving and Communication System (PACS) servers.
[0118] While the foregoing embodiments have been described in some detail for purposes of clarity and the known alternative ways of implementing the application, the teachings of this application are not limited to those details. Numerous alternatives will be readily apparent to those skilled in the art. The disclosed embodiments are illustrative rather than restrictive.
Claims
1. A system for IoT device discovery and identification, comprising: a processor configured to: receive information associated with network communications of an IoT device; determine whether the IoT device has been classified; in response to determining that the IoT device has not been classified, perform a two-part classification process, wherein a first part of the classification process comprises online classification of the IoT device, wherein the online classification comprises performing activity classification based on the received information, and wherein a second part of the classification process comprises subsequent verification of the online classification of the IoT device; and provide a result of the classification process to a security appliance configured to apply a policy to the IoT device; and a memory coupled to the processor and configured to provide instructions to the processor, wherein the first part of the classification process comprises rule-based classification and the second part of the classification process comprises machine learning-based classification.
2. The system of claim 1, wherein, The information is received from a data appliance configured to monitor the IoT device.
3. The system of claim 1, wherein, The received information comprises network traffic metadata.
4. The system of claim 1, wherein performing the classification process comprises determining whether any features of the IoT device are dominant.
5. The system of claim 1, wherein performing the two-part classification process comprises determining a confidence that one or more features of the IoT device match a network behavior pattern identifier.
6. The system of claim 5, wherein the processor is further configured to generate a network behavior pattern identifier.
7. The system of claim 1, wherein the result of the classification process comprises a group classification.
8. The system of claim 1, wherein the result of the classification process comprises a profile classification.
9. A method for IoT device discovery and identification, comprising: receiving information associated with network communications of an IoT device; determining whether the IoT device has been classified; in response to determining that the IoT device has not been classified, performing a two-part classification process, wherein a first part of the classification process comprises online classification of the IoT device, wherein the online classification comprises performing activity classification based on the received information, and wherein a second part of the classification process comprises subsequent verification of the online classification of the IoT device; and providing a result of the classification process to a security appliance configured to apply a policy to the IoT device, wherein the first part of the classification process comprises rule-based classification and the second part of the classification process comprises machine learning-based classification.
10. A computer program product, embodied in a tangible computer-readable storage medium, the computer program product comprising computer instructions for: receiving information associated with network communications of an IoT device; determining whether the IoT device has been classified; in response to determining that the IoT device has not been classified, performing a two-part classification process, wherein a first part of the classification process comprises online classification of the IoT device, wherein the online classification comprises performing activity classification based on received information, and wherein a second part of the classification process comprises subsequent verification of the online classification of the IoT device; and providing a result of the classification process to a security appliance configured to apply a policy to the IoT device, wherein the first part of the classification process comprises rule-based classification, and the second part of the classification process comprises machine learning-based classification.
Citation Information
Patent Citations
System, device, and method of adaptive network protection for managed internet-of-things services
US20180375887A1
Active labeling of unknown devices in a network
US20200162425A1