Collecting and processing context attributes on the host
The novel architecture on hosts captures and processes context attributes to enhance middlebox services, addressing the inefficiencies in existing solutions by enabling effective utilization of contextual data for service rule execution in software-defined networking and network virtualization.
Patent Information
- Application Number
- JP2023120131
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-10-30
- Filing Date
- 2023-07-24
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2037-12-10
Smart Images

Figure 0007681068000001 
Figure 0007681068000002 
Figure 0007681068000003
Abstract
Description
[Technical field]
[0001] Concerning the collection and processing of context attributes on the host. [Background technology]
[0002] Middlebox services have historically been hardware appliances implemented at one or more points in the network topology within an enterprise or data center. With the advent of software-defined networking (SDN) and network virtualization, traditional hardware appliances do not take advantage of the flexibility and control offered by SDN and network virtualization. Thus, in recent years, various methods have been proposed to provide middlebox services on hosts. However, most of these middlebox solutions do not take advantage of the rich contextual data that may be captured for each data message flow on the host. One reason for this is that existing technologies do not provide an efficient distributed scheme for filtering thousands of captured contextual attributes in order to efficiently process service rules defined in terms of a much smaller set of contextual attributes. Summary of the Invention
[0003] Some embodiments of the present invention provide a novel architecture for capturing context attributes on a host computer running one or more machines, which in some embodiments are virtual machines (VMs), in other embodiments are containers, and in yet other embodiments are a mix of VMs and containers.
[0004] Some embodiments run a guest-introspection (GI) agent on each machine for which context attributes need to be captured. In addition to running one or more machines on each host computer, these embodiments also run a context engine and one or more attribute-based service engines on each host computer. In some embodiments, via a machine's GI agent on a host, the context engine of that host collects context attributes related to network and / or process events on the machine. As described further below, the context engine then provides the context attributes to a service engine, which uses these context attributes to identify service rules that specify context-based services to execute on processes executing on the machine and / or data message flows sent or received by the machine.
[0005] In some embodiments, the context engine of a host collects context attributes from the GI agents of machines on that host through a variety of different methods. For example, in some embodiments, the GI agent on a machine registers hooks (e.g., callbacks) with one or more modules (e.g., kernel space modules or user space modules) in the machine's operating system for all new network connection events and all new process events.
[0006] When a new network connection event occurs, the GI agent receives a callback from the OS and, based on the callback, provides a network event identifier to the context engine. The network event identifier provides a set of attributes related to the network event. These network event attributes in some embodiments include a 5-tuple identifier of the requested network connection (i.e., source port and IP address, destination port and IP address, and protocol), a process identifier of the process requesting the network connection, a user identifier associated with the requesting process, and a group identifier (e.g., an Activity Directory (AD) identifier) associated with the requesting process.
[0007] In some embodiments, the context engine instructs the GI agent to collect additional process parameters from the OS module that are associated with the process identifier (ID) received with the network event. These additional process parameters in some embodiments include a process name, a process hash, a process path with command line parameters, a process network connection, a process loaded module, and one or more process consumption parameters that specify the process consumption of one or more resources of the machine (e.g., central processing unit consumption, network consumption, and memory consumption). Instead of using the process identifier to query the GI agent for additional process parameters associated with the network event, the context engine of other embodiments receives all process parameters associated with the network event in one shot when the GI agent reports the network event to the context engine.
[0008] In some embodiments, the OS on the machine puts a hold on new network events (i.e., does not begin sending data messages for the network events) until the GI agent on the machine instructs it to continue processing the network events. In some of these embodiments, the GI agent allows the OS to proceed with processing the network event only after the context engine has gathered all attributes required for the event (e.g., after receiving a message from the context engine specifying that the context engine has received all process or network attributes required for the new network event).
[0009] In some embodiments, the context engine uses the process hash received from the GI agent to identify the name and version of the application (i.e., software product) to which the process belongs. To do this, in some embodiments, the context engine stores the process hash and the associated application name / version, compares the process hash received from the GI agent to the stored process hash to identify a matching hash, and then uses the application name / version of the matching hash as the application name and version of the process associated with the event.
[0010] In some embodiments, the context engine obtains the process hash and application name / version from one or more network or compute managers, which may be running on another device or computer. In other embodiments, the context engine provides a hash associated with the process identifier to the network or compute manager, which then matches this hash to its process hash records and provides the application name / version of the associated process to the context engine. Once the context engine obtains the application name / version associated with the network event, the context engine can provide the name and version attributes to the attribute-based service engine, which can use this information (i.e., application name and / or version) to identify service rules to implement.
[0011] When a process event occurs, the GI agent receives a callback from the OS and, based on the callback, provides a process event identifier to the context engine. The process event identifier provides a set of attributes related to the process event. In some embodiments, the set of attributes includes a process identifier. In some embodiments, the settings also include a user identifier and / or a group identifier (e.g., an Activity Directory (AD) identifier).
[0012] In some embodiments, when the GI agent reports a process event to the context engine, it provides all process parameters associated with the process event to the context engine (e.g., process identifier, user ID, group ID, process name, process hash, loaded module identifiers, consumption parameters, etc.). In other embodiments, the context engine instructs the GI agent to collect additional process parameters from the OS module associated with the process identifier received by the context engine along with the process event. These additional process parameters in some embodiments are the same as the process parameters described above for the reported network event (e.g., process name, process hash, loaded module identifiers, consumption parameters, etc.).
[0013] The context engine of some embodiments augments the context attributes received from the GI agent with context attributes received from other modules executing on the host. For example, in some embodiments, a deep packet inspection (DPI) module executes on the host. The context engine or another module (e.g., a firewall engine) instructs the DPI engine to inspect data messages of a data message flow associated with a process ID to identify the type of traffic being sent in these data messages by the application associated with the process ID.
[0014] The identified traffic type identity is commonly referred to today as AppID. And currently there are many DPI modules that analyze messages of data message flow to generate AppID. In some embodiments, the context engine combines the obtained AppID for a network event with other identifying context attributes for this event (e.g., by using the 5-tuple identifier of the event to associate the AppID with the collected context attributes) to generate a very rich set of attributes that the service engine can use to execute services. This rich set of attributes provides true application identity (i.e., application name, application version, application traffic type, etc.) based on which the service engine can execute their services.
[0015] Also, in some embodiments, a threat detection module runs on the host computer together with the context engine. When the context engine obtains a set of process parameters that specify that a process has started on the machine or is sending a data message on the machine, in some embodiments, the context engine provides one or more process parameter hashes, application name, application version, AppID, other process parameters, etc., to the threat detection module. The threat detection module then generates a threat level indicator (e.g., low, medium, high, etc.) for the identified process and provides the threat level indicator to the context engine. The context engine then provides the threat level indicator to one or more service engines as another context attribute for performing services on the data message of the new process event or new network event, and the service engine can use the threat level indicator as another attribute to identify a service rule to be implemented.
[0016] The context engine distributes collected context attributes to the service engines in some embodiments using a push model and in other embodiments using a pull model. In still other embodiments, the content engine uses a push model for some service engines and a pull model for other service engines. In the push model, the context engine delivers to the service engines the context attributes it collects for process or network events having an identifier of a process and / or a flow identifier of the network event (e.g., a 5-tuple identifier of the flow). In some embodiments, the context engine delivers to the service engines only those context attributes that are relevant to the service rules of that service engine.
[0017] In the pull model, the context engine receives a query from a service engine for context attributes that it has collected for a particular process or network connection. In some embodiments, the context engine receives a process ID or flow identifier (e.g., a 5-tuple identifier) from the service engine along with the query and uses the received identifier to identify a set of attributes that it must provide to the service engine. In some embodiments, the context engine generates a service token (also called a service tag) for the collection of attributes related to the service engine and provides this service token to another module (e.g., a GI agent or another module on the host) to pass it to the service engine (e.g., in an encapsulation header of a data message). The service engine then extracts the service token and provides this service token to the context engine to identify the context attributes that it must provide to the service engine.
[0018] The context engine in some embodiments provides the context attributes to a number of context-based service engines on its host computer. In some embodiments, the context engine and the service engine are all kernel space components of a hypervisor on which multiple VMs or containers run. In other embodiments, the context engine and / or one or more service engines are user space processes. For example, one or more service engines in some embodiments are Service VMs (SVMs).
[0019] Different embodiments use different types of context-based service engines. For example, in some embodiments, an attribute-based service engine includes: (1) a firewall engine that performs context-based firewall operations on data messages sent by or received for the machine, (2) a process control engine that implements context-based process control operations (e.g., process evaluation and termination operations) on processes started on the machine, (3) a load balancing engine that performs context-based load balancing operations to distribute data message flows from the machine to different destinations or service nodes in one or more destination / service node clusters, and (4) an encryption engine that performs context-based encryption or decryption operations to encrypt data messages from the machine or decrypt data messages received for the machine.
[0020] Another context-based service engine in some embodiments is a discovery service engine. In some embodiments, the discovery engine captures new process events and new network events from the context engine along with the context attributes that the context engine collects about these process and network events. The discovery service engine then relays these events and their associated context attributes to one or more network managers (e.g., servers) that provide a management layer that allows network administrators to visualize events in the data center and specify policies for computing and network resources in the data center.
[0021] In relaying these events and attributes to the network management layer, the discovery module of some embodiments performs some pre-processing of these events and attributes. For example, in some embodiments, the discovery module filters some of the network or process events while aggregating some or all of these events and their attributes. Also, in some embodiments, the discovery engine instructs the context engine to collect additional context attributes for process or network events via a GI agent or other module (e.g., a DPI engine or a threat detection engine), or to capture other types of events, such as file events and system events.
[0022] For example, in some embodiments, the discovery engine instructs the context engine to build an inventory of applications installed on the machine and periodically refresh this inventory. The discovery engine may do so upon request of the management plane or based on operational configurations that the management or control plane specifies for the discovery engine. In response to a request from the discovery engine, in some embodiments, the context engine causes each GI agent on each of its host's machines to discover all installed processes and all running processes and services on the machine.
[0023] After building an inventory of installed applications and running processes / services, the discovery engines of the host computers in the data center provide this information to the network / computing manager in the management plane. In some embodiments, the management plane collects context attributes from sources other than the host computer discovery engine and the context engine. For example, in some embodiments, the management plane collects compute context from one or more servers (e.g., cloud context from a cloud vendor or compute virtualization context from data center virtualization software), identity context from a directory service server, mobility context from a mobility management server, endpoint context from a DNS (Domain Name Server) and application inventory server, network context from a network virtualization server (e.g., virtualized network context).
[0024] By collecting context information (e.g., from the discovery engine and context engine, and / or from other context sources), the management plane can provide a user interface to network / compute administrators to visualize the compute and network resources in the data center. Additionally, the collected context attributes enable the management plane to provide control through this user interface for these administrators to specify context-based service rules and / or policies. These service rules / policies are then distributed to host computers so that service engines on these computers can perform context-based service operations.
[0025] The preceding summary is intended to be provided as a brief introduction to some embodiments of the present invention. It is not intended to be an introduction or overview of all of the inventive subject matter disclosed in this document. The following detailed description and the figures referenced in the detailed description will further describe other embodiments and the embodiments described in the summary. Therefore, a thorough review of the summary, detailed description, figures and claims is necessary to understand all of the embodiments described by this document. Moreover, the subject matter described in the claims should not be limited by the illustrative details in the summary, detailed description and figures. [Brief description of the drawings]
[0026] The novel features of the invention will become apparent from the appended claims, however, for purposes of illustration, certain embodiments of the invention are set forth in the following figures.
[0027] [Figure 1] FIG. 1 illustrates a host computer that employs the context engine and context-based service engine of some embodiments of the present invention.
[0028] [Diagram 2]FIG. 2 illustrates a more detailed example of a host computer that, in some embodiments, is used to establish a distributed architecture for configuring and running context-rich, attribute-based services in a data center.
[0029] [Diagram 3] FIG. 3 illustrates a process performed by the context engine of some embodiments.
[0030] [Figure 4] FIG. 4 illustrates example load balancing rules of some embodiments.
[0031] [Diagram 5] FIG. 5 illustrates a load balancer of several embodiments of untrusted web server traffic between several application servers.
[0032] [Figure 6] FIG. 6 illustrates a process performed by a load balancer in some embodiments.
[0033] [Figure 7] FIG. 7 shows some examples of such firewall rules.
[0034] [Figure 8] FIG. 8 shows some more detailed examples of context-based firewall rules of some embodiments.
[0035] [Figure 9] , [Figure 10] , [Figure 11] , [Figure 12] 9-12 show various examples illustrating the enforcement of various context-based firewall rules by the firewall engine.
[0036] [Figure 13] FIG. 13 illustrates the process the context engine performs to collect user and group identifiers each time it receives a new network connection event from a GI agent.
[0037] [Figure 14] FIG. 14 illustrates the process performed by the firewall engine in some embodiments.
[0038] [Figure 15] FIG. 15 shows an example of such a context-based encryption rule.
[0039] [Figure 16] FIG. 16 illustrates the process that an encryptor of some embodiments performs to encrypt a data message sent by a VM on a host.
[0040] [Figure 17] FIG. 17 illustrates the process that the encryption engine performs to decrypt an encrypted data message that a forwarding element port receives on a destination host executing the destination VM of the data message.
[0041] [Figure 18] FIG. 18 illustrates the process that the encryption engine performs to decrypt an encrypted data message that contains a key identifier in its header.
[0042] [Figure 19] FIG. 19 shows some examples of process control rules.
[0043] [Figure 20] FIG. 20 illustrates a process performed by a process control engine in some embodiments.
[0044] [Figure 21]FIG. 21 shows an example of how service engines are managed in some embodiments.
[0045] [Figure 22] FIG. 22 conceptually illustrates a computer system upon which some embodiments of the present invention may be implemented. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0046] In the following detailed description of the invention, numerous details, examples, and embodiments of the invention are explained and described. However, it will be apparent and clear to one skilled in the art that the invention is not limited to the described embodiments, and that the invention may be practiced without some of the specific details and examples explained.
[0047] Some embodiments of the present invention provide a novel architecture for capturing context attributes on a host computer running one or more machines and executing services on the host computer using the captured context attributes. Some embodiments execute a guest-introspection (GI) agent on each machine for which context attributes need to be captured. In addition to executing one or more machines on each host computer, these embodiments also execute a context engine and one or more attribute-based service engines on each host computer. In some embodiments, through a machine's GI agent on a host, the context engine of that host collects context attributes related to network events and / or process events on the machine. The context engine then provides the context attributes to a service engine, which uses these context attributes to identify service rules that specify context-based services to execute on processes running on the machine and / or data message flows sent or received by the machine.
[0048] As used herein, a data message refers to a collection of bits in a particular format that is transmitted over a network. Those skilled in the art will appreciate that the term data message may be used herein to refer to various formatted collections of bits that may be transmitted over a network, such as an Ethernet frame, an IP packet, a TCP segment, a UDP datagram, etc. Also, as used herein, references to the L2, L3, L4, and L7 layers (or layer 2, layer 3, layer 4, layer 7) are references to the second data link layer, the third network layer, the fourth transport layer, and the seventh application layer, respectively, of the Open System Interconnection (OSI) layer model.
[0049] 1 illustrates a host computer 100 that employs the context engine and context-based service engine of some embodiments of the present invention. As illustrated, the host computer 100 includes a number of data computing nodes 105, a context engine 110, a number of context-based service engines 130, a threat detector 132, and a deep packet inspection (DPI) module 135. The context-based service engines include a discovery engine 120, a process control engine 122, an encryption engine 124, a load balancer 126, and a firewall engine 128. The host computer 100 also includes a context-based service rule store 140, and an attribute store 145.
[0050] A DCN is an endpoint machine that runs on a host computer 100. In some embodiments, the DCN is a virtual machine (VM), in other embodiments, a container, and in yet other embodiments, a mix of VMs and containers. On each DCN, a GI agent 150 runs to collect context attributes for the context engine 110. In some embodiments, the context engine 110 collects context attributes from the GI agent 150 of the DCN on its host through a variety of different methods. For example, in some embodiments, the GI agent on the DCN registers hooks (e.g., callbacks) with one or more modules (e.g., kernel space modules or user space modules) in the DCN's operating system for every new network connection event and every new process event.
[0051] When a new network connection event occurs, the GI agent 150 receives a callback from the DCN's OS and, based on the callback, provides a network event identifier to the context engine 110. The network event identifier provides a set of attributes related to the network event. These network event attributes in some embodiments include a 5-tuple identifier of the requested network connection (i.e., source port and IP address, destination port and IP address, and protocol), a process identifier of the process requesting the network connection, a user identifier associated with the requesting process, and a group identifier (e.g., an Activity Directory (AD) identifier) associated with the requesting process.
[0052] In some embodiments, the context engine instructs the GI agent 150 to collect additional process parameters from the OS module that are associated with the process identifier (ID) received with the network event. These additional process parameters in some embodiments include a process name, a process hash, a process path with command line parameters, a process network connection, a process loaded module, and one or more process consumption parameters that specify the process consumption of one or more resources of the machine (e.g., central processing unit consumption, network consumption, and memory consumption). Instead of using the process identifier to query the GI agent 150 for additional process parameters associated with the network event, the context engine 110 in other embodiments receives all process parameters associated with the network event in one shot when the GI agent reports the network event to the context engine.
[0053] The OS of a DCN in some embodiments suspends a new network event (i.e., does not begin sending data messages for the network event) until a GI agent 150 on that DCN instructs the DCN to proceed with processing the network event. In some of these embodiments, the GI agent 150 allows the OS to proceed with processing the network event only after the context engine 110 has collected all attributes required for the event (e.g., after receiving a message from the context engine specifying that the context engine has received all process or network attributes required for the new network event).
[0054] In some embodiments, context engine 110 uses the process hash received from GI agent 150 to identify the name and version of the application (i.e., software product) to which the process belongs. To do this, in some embodiments, context engine 110 stores the process hash and the associated application name / version, compares the process hash received from the GI agent to the stored process hash to identify a matching hash, and then uses the application name / version of the matching hash as the application name and version of the process associated with the event.
[0055] In some embodiments, the context engine 110 obtains the process hash and application name / version from one or more network or compute managers, which may be running on another device or computer. In other embodiments, the context engine provides a hash associated with the process identifier to the network or compute manager, which then matches this hash to its process hash record and provides the application name / version of the associated process to the context engine. Once the context engine 110 obtains the application name / version associated with a network event, it can provide the name and version attributes to the attribute-based service engine, which can use this information (i.e., the application name and / or version) to identify the service rules to be implemented.
[0056] When a process event occurs on the DCN 105, the DCN GI Agent 150, in some embodiments, receives a callback from the DCN OS and, based on the callback, provides a process event identifier to the context engine 110. The process event identifier provides a set of attributes related to the process event. In some embodiments, the set of attributes includes a process identifier. In some embodiments, the set also includes a user identifier and / or a group identifier (e.g., an Activity Directory (AD) identifier).
[0057] In some embodiments, when the GI agent reports a process event to the context engine, it provides all process parameters associated with the process event to the context engine (e.g., process identifier, user ID, group ID, process name, process hash, loaded module identifiers, consumption parameters, etc.). In other embodiments, the context engine instructs the GI agent to collect additional process parameters from the OS module associated with the process identifier received by the context engine along with the process event. These additional process parameters in some embodiments are the same as the process parameters described above for the reported network event (e.g., process name, process hash, loaded module identifiers, consumption parameters, etc.).
[0058] The context engine 110 of some embodiments augments the context attributes it receives from the GI agent 150 with context attributes it receives from other modules executing on the host. The DPI module 135 (also referred to as a deep packet inspector) and the threat detector 132 (also referred to as a threat inspection module) are two such modules that provide context attributes to augment the attributes the context engine collects from the GI agent 150. In some embodiments, the DPI module is instructed by the context engine 110 or another module (e.g., the firewall engine 128) to inspect data messages of a data message flow associated with a process ID to identify the type of traffic being sent in these data messages by the application associated with the process ID.
[0059] The identified traffic type identity is commonly referred to today as AppID. Also, there are currently many DPI modules that analyze the messages of a data message flow to generate the AppID of the data message flow. In some embodiments, the context engine combines the AppID obtained for a network event with other context attributes it identifies for this event to generate a very rich set of attributes that the service engine can use to perform its services. This rich set of attributes provides the true application identity (i.e., application name, application version, application traffic type, etc.) based on which the service engine can perform their services. In some embodiments, the context engine 110 uses the 5-tuple identifier of the network event to associate the AppID of the data message flow of this event with the context attributes that the context engine collects from the GI agent of the DCN associated with the data message flow (e.g., of the DCN where the data message flow originates).
[0060] The threat detector 132 provides a threat level indicator that specifies a threat level associated with a particular application running on the DCN. Once the context engine obtains a set of process parameters that specify that a process has started on a machine or is sending a data message on a machine, in some embodiments, the context engine provides one or more of the process parameter hashes (such as an application name, an application version, an AppID, other process parameters, etc.) to the threat detection module.
[0061] The threat detection module then generates a threat level indicator (e.g., low, medium, high, etc.) for the identified process and provides the threat level indicator to the context engine. In some embodiments, the threat detector assigns a threat score to applications running on the DCN based on various application behavioral events such as (1) whether they perform poor input validation, (2) whether they pass authentication credentials over unencrypted network links, (3) whether they use weak password and account policies, (4) whether they store configuration secrets in plain text, (5) whether they can transfer files, (6) whether the application is known to propagate malware, (7) whether the application is intentionally evasive, and (8) whether the application has known vulnerabilities. In some embodiments, the threat detector is a third-party whitelisting application such as Bit9.
[0062] In some embodiments, the context engine provides the threat level indicator generated by the threat detector 132 to one or more service engines as another context attribute for performing services on the data message of a new process event or a new network event, and the service engine can use the threat level indicator as another attribute to identify a service rule to be implemented.
[0063] The context engine 110 stores the collected context attributes for network and process events in the attribute store 145. In some embodiments, the context engine stores each set of context attributes with one or more network event identifiers and / or process identifiers. For example, in some embodiments, the context engine 110 stores the collected context attributes for a new process event with or with a reference to the process identifier. The context engine then provides the collected context attributes using the process identifier to a service engine (e.g., the process control engine 122) that performs a service for the process event.
[0064] The context engine, in some embodiments, uses the 5-tuple identifier of the network connection event or a reference to this 5-tuple identifier to store the collected context attributes for the new network connection event. In some of these embodiments, the context engine provides the context attributes of the network event to the service engine along with the 5-tuple identifier of the event. The data message of this network event uses this 5-tuple identifier, so the service engine can use the provided 5-tuple identifier to identify the context attributes associated with the data message flow.
[0065] The context engine distributes collected context attributes to the service engines 130 in some embodiments using a push model, while in other embodiments the engine distributes these attributes to the service engines 130 using a pull model. In still other embodiments, the context engine uses a push model for some service engines and a pull model for other service engines. In a push model, the context engine in some embodiments delivers to the service engines the context attributes it collects for a process event or network event with an identifier of the process and / or a flow identifier of the network event (e.g., a 5-tuple identifier of the flow).
[0066] In some embodiments, the context engine delivers to the service engine only those context attributes that are relevant to the service rules of that service engine. In other words, in these embodiments, the context engine compares each collected attribute in a set of collected attributes (e.g., for a network event or a process event) with a list of attributes used by the service rules of the service engine, and discards each collected attribute that is not used by the service rule. The context engine then provides to the service engine only the subset of collected attributes (within the set of collected attributes) that are used by the engine's service rules. In other embodiments, the service engine performs this filtering operation to discard context attributes that are not needed.
[0067] In the pull model, the context engine receives queries from the service engine for context attributes it has collected for a particular process or network connection. In some embodiments, the context engine receives a process ID or flow identifier (e.g., a 5-tuple identifier) from the service engine along with the query and uses the received identifier to identify the set of attributes that it must provide to the service engine.
[0068] In some embodiments, the context engine generates a service token (also called a service tag) to collect attributes related to the service engine and provides this service token to another module (e.g., a GI agent or another module on the host) to pass to the service engine (e.g., in an encapsulation header of a data message). The service engine then extracts the service token and provides it to the context engine to identify the context attributes that the context engine must provide to the service engine.
[0069] In some embodiments, the context engine 110 and the service engine 130 are all kernel space components of a hypervisor on which multiple VMs or containers execute, as described further below with reference to FIG. 2. In other embodiments, the context engine and / or one or more service engines are user space processes. For example, one or more service engines in some embodiments are service VMs (SVMs). In some embodiments, the one or more service engines are in the ingress and / or egress data paths of the DCN to receive access to data message flows to and from the DCN and perform services on these data message flows. In other embodiments, one or more other modules on the host 100 intercept data messages from the ingress / egress data paths and forward these messages to one or more service engines for these engines to perform services on the data messages. One such approach is described below with reference to FIG. 2.
[0070] Different embodiments use different types of context-based service engines. In the example shown in Figure 1, the service engines 130 include a discovery engine 120, a process control engine 122, an encryption engine 124, a load balancer 126, and a firewall engine 128. Each of these service engines 130 has an attribute-based service rule store. Figure 1 collectively represents all the context-based service rule stores of these service engines together with a context-based service rule store 140 to simplify the illustration shown in this figure.
[0071] In some embodiments, each service rule in the context-based service rule store 140 has a rule identifier for matching with a process or flow identifier to identify the rule to be executed for a process or network event. In some embodiments, the context-based service rule store 140 is defined in a hierarchical manner to ensure that a rule check matches a higher priority rule before matching a lower priority rule. Also, in some embodiments, the context-based service rule store 140 includes a default rule that specifies a default action for any rule check, as described further below.
[0072] The firewall engine 128 performs firewall actions on data messages sent by or received for the DCN 105. These firewall actions are based on firewall rules in the context-based service rule store 140. Some of the firewall rules are defined purely in terms of layer 2-layer 4 attributes, e.g., in terms of 5-tuple identifiers. Other firewall rules are defined in terms of context attributes, which may include one or more of the collected context attributes, such as application name, application version, AppID, resource consumption, threat level, user ID, group ID, etc. Still other firewall rules in some embodiments are defined in terms of both L2-L4 parameters and context attributes. Because the firewall engine 128 can resolve firewall rules defined by referencing context attributes, this firewall engine is referred to as a context-based firewall engine.
[0073] In some embodiments, the context-based firewall engine 128 can allow, block, or reroute data message flows based on any number of context attributes, since its firewall rules can be identified with respect to any combination of collected context attributes. For example, the firewall engine can block all email traffic from crome.exe when a user is part of a Nurse user group and one firewall rule specifies that data messages should be blocked when the flow is associated with the Nurse group ID, the AppID identifies the traffic type as email, and the application name is Chrome. Similarly, a context-based firewall rule can block data message flows related to video conferencing, online video viewing, or using older versions of software. Examples of such rules are blocking all Skype traffic, blocking all YouTube video traffic, blocking all HipChat audio / video conferencing when the application version number is older than a certain version number, blocking data message flows for any application with a high threat score, etc.
[0074] The load balancing engine 126 performs load balancing operations on data messages transmitted by the DCN 105 to distribute data message flows to different destinations or service nodes in one or more destination / service node clusters. These load balancing operations are based on load balancing rules in the context-based service rule storage 140. In some of these embodiments, each load balancing rule can specify one or more load balancing criteria (e.g., round robin criteria, weighted round robin criteria, etc.) for distributing traffic and can limit each criterion to a particular time range. In some embodiments, the load balancing operation includes replacing a destination network address (e.g., destination IP address, destination MAC address, etc.) of a data message flow with another destination network address.
[0075] Some of the load balancing rules are defined purely in terms of L2-L4 attributes, e.g., in terms of a 5-tuple identifier. Other load balancing rules are defined in terms of context attributes, which may include one or more of the collected context attributes, such as application name, application version, AppID, resource consumption, threat level, user ID, group ID, etc. Still other load balancing rules in some embodiments are defined in terms of both L2-L4 parameters and context attributes. Because the load balancing engine 126 can resolve the load balancing rules defined by referencing the context attributes, the load balancing engine is referred to as a context-based load balancer.
[0076] In some embodiments, the context-based load balancer 126 can distribute data message flows based on any number of context attributes because its load balancing rules can be identified with respect to any combination of the collected context attributes. For example, the data distribution of the load balancer 126 can be based on any combination of user data and application data. Examples of such load balancing operations include (1) distributing data message flows related to the Finance department over all load balancing pools, (2) redirecting all Finance department traffic to another pool when the department's primary pool is down to make the department's traffic highly available, and (3) making all traffic related to the Doctor's user group highly available. In some embodiments, load balancing rules can also be defined with respect to the collected resource consumption to distribute traffic to provide more or less resources to applications that consume a lot of resources on the DCN.
[0077] The encryption engine 124 performs encryption / decryption operations (collectively referred to as cryptographic operations) on data messages sent and received by the DCN 105. These encryption operations are based on encryption rules in the context-based service rule store 140. In some embodiments, each of these rules includes an encryption / decryption key identifier that the encryption engine can use to retrieve an encryption / decryption key from a key manager operating on the host or external to the host. Each encryption rule also, in some embodiments, specifies the type of encryption / decryption operation that the encryption module must perform.
[0078] Each encryption rule also has a rule identifier. In some encryption rules, the rule identifier is defined purely in terms of L2-L4 attributes, e.g., in terms of a 5-tuple identifier. Other encryption rules are defined in terms of context attributes, which may include one or more of the collected context attributes, such as application name, application version, AppID, resource consumption, threat level, user ID, group ID, etc. Still other encryption rules in some embodiments are defined in terms of both L2-L4 parameters and context attributes. Because the encryption engine 124 can resolve encryption rules defined by referencing context attributes, this encryption engine is referred to as a context-based encryption engine.
[0079] In some embodiments, the context-based encryption module 124 can identify its encryption rules with respect to any combination of collected context attributes, and can therefore encrypt or decrypt data message flows based on any number of context attributes. For example, the encryption / decryption operations of the encryption engine 124 can be based on any combination of user data and application data. Examples of such encryption operations include (1) encrypting all traffic from Outlook (initiated on any machine) to an Exchange server, (2) encrypting all communication between applications in a three-tier web server, application server, and database server, (3) encrypting all traffic originating from the Administrators Active Directory group, etc.
[0080] The process control engine 122 performs context-based process control actions (e.g., process evaluation and termination actions) on processes initiated on the DCN 105. In some embodiments, whenever the context engine 110 receives a new process event from the GI agent 150 of the DCN, it provides the process parameters associated with the process event to the process control engine 122. The engine then uses the received set of process parameters to look up its context-based service rule store 140 to identify a matching context-based process control rule.
[0081] The process control engine 122 can instruct the context engine to instruct the DCN's GI agent to perform process control actions on the process. Examples of such process control actions include (1) terminating a videoconferencing application having a particular version number, (2) terminating a browser displaying YouTube traffic, and (3) terminating an application having a high threat level score.
[0082] Discovery engine 120 is another context-based service engine. In some embodiments, discovery engine 120 captures new process events and new network events from the context engine along with the context attributes that the context engine collects about these process and network events. As described further below, the discovery service engine then relays these events and their associated context attributes to one or more network managers (e.g., servers) that provide a management layer that allows network administrators to visualize events in the data center and specify policies for computing and network resources in the data center.
[0083] In relaying these events and attributes to the network management layer, the discovery module of some embodiments performs some pre-processing of these events and attributes. For example, in some embodiments, the discovery module filters some of the network or process events while aggregating some or all of these events and their attributes. Also, in some embodiments, the discovery engine 120 instructs the context engine 110 to collect additional context attributes for process or network events via the GI agent 150 or other modules (e.g., a DPI engine or a threat detection engine), or to capture other types of events, such as file events and system events.
[0084] For example, in some embodiments, the discovery engine instructs the context engine to build an inventory of applications installed on the machine and periodically refresh this inventory. The discovery engine may do so upon request of the management plane or based on operational configurations that the management or control plane specifies for the discovery engine. In response to a request from the discovery engine, in some embodiments, the context engine causes each GI agent on each of its host's machines to discover all installed processes and all running processes and services on the machine.
[0085] After building an inventory of installed applications and running processes / services, the discovery engines of the host computers in the data center provide this information to the network / compute manager in the management plane. In some embodiments, the management plane collects context attributes from sources other than the host computer discovery engine and the context engine. For example, in some embodiments, the management plane collects compute context from one or more servers (e.g., cloud context from a cloud vendor or compute virtualization context from data center virtualization software), identity context from a directory service server, mobility context from a mobility management server, endpoint context from DNS (Domain Name Server) and application inventory servers, network context (e.g., virtualized network context from a network virtualization server).
[0086] By collecting context information (e.g., from the discovery engine and context engine, and / or from other context sources), the management plane can provide a user interface to network / compute administrators to visualize the compute and network resources in the data center. Additionally, the collected context attributes enable the management plane to provide control through this user interface for these administrators to specify context-based service rules and / or policies. These service rules / policies are then distributed to host computers so that service engines on these computers can perform context-based service operations.
[0087] In some embodiments described above, the same service engine 130 (e.g., the same firewall engine 128) performs the same type of service (e.g., firewall service) based on service rules that may be defined with respect to message flow identifiers (e.g., 5-tuple identifiers) or with respect to collected context attributes associated with data message flows (e.g., AppID, threat level, user identifier, group identifier, application name / version, etc.). However, in other embodiments, different service engines provide the same type of service based on message flow identifiers (e.g., 5-tuple identifiers) and based on collected context attributes of data message flows. For example, some embodiments use one flow-based firewall engine that performs firewall operations based on rules defined with respect to flow identifiers and another context-based firewall engine that performs firewall operations based on rules defined with respect to context attributes (e.g., AppID, threat level, user identifier, group identifier, application name / version, etc.).
[0088] 2 illustrates a more detailed example of a host computer 200 used in some embodiments to establish a distributed architecture for configuring and running context-rich, attribute-based services in a data center. The host computer 200 includes many of the same components as the host computer 100, such as the context engine 110, the service engine 130, the threat detector 132, the DPI module 135, the context-based service rule store 140, and the context attribute store 145. As in FIG. 1, the service engine 130 of FIG. 2 includes a discovery engine 120, a process control engine 122, an encryption engine 124, a load balancer 126, and a firewall engine 128.
[0089] In Fig. 2, the DCN is a VM 205 running on a hypervisor. Also in Fig. 2, the host computer 200 includes a software forwarding element 210, an attribute mapping storage 223, a connection state cache storage 225, a MUX (multiplexer) 227, and a context engine policy storage 143. In some embodiments, the context engine 110, the software forwarding element 210, the service engine 130, the context-based service rule storage 140, the connection state cache storage 225, the context engine policy storage 143, and the MUX 227 run in the kernel space of the hypervisor, while the VM 205 runs in the user space of the hypervisor. In other embodiments, one or more of the service engines are user space modules (e.g., service VMs).
[0090] In some embodiments, VMs 205 serve as data endpoints within a data center. Examples of such machines include web servers, application servers, database servers, etc. In some cases, all VMs belong to one entity, such as the entity that operates the host. In other cases, host 200 operates in a multi-tenant environment (e.g., a multi-tenant data center), and different VMs 205 can belong to one tenant or multiple tenants.
[0091] Each VM 205 includes a GI agent 250 that interacts with the context engine 110 to provide it with a set of context attributes and to receive instructions and queries from the engine. The interaction between the GI agent 250 and the context engine 110 is similar to the interaction between the GI agent 150 and the context engine 110 described above. However, as shown in FIG. 2, in some embodiments, all communication between the context engine 110 and the GI agent 250 is relayed through a MUX 227. One example of such a MUX is the MUX used by VMware, Inc.'s ESX hypervisor endpoint security (EPSec) platform.
[0092] In some embodiments, the GI agent communicates with MUX 227 through a high-speed communication channel (such as the VMCI channel of ESX). In some embodiments, this communication channel is a shared memory channel. As mentioned above, the attributes collected by the context engine 110 from the GI agent 250 in some embodiments include a rich group of parameters (e.g., layer 7 parameters, process identifier, user identifier, group identifier, process name, process hash, loaded module identifier, consumption parameters, etc.).
[0093] As shown, each VM 205, in some embodiments, also includes a virtualized network interface card (VNIC) 255. Each VNIC is responsible for exchanging messages between that VM and a software forwarding element (SFE) 210. Each VNIC connects to a specific port 260 of the SFE 210. The SFE 210 also connects to a physical network interface card (NIC) (not shown) of the host. In some embodiments, a VNIC is a software abstraction created by the hypervisor of one or more physical NICs (PNICs) of the host.
[0094] In some embodiments, SFE 210 maintains a single port 260 for each VNIC of each VM. SFE 210 connects to the host PNIC (through a NIC driver (not shown)) to send outgoing messages and receive incoming messages. In some embodiments, SFE 210 is defined to include a port 265 that connects to the PNIC's driver to send and receive messages to and from the PNIC. SFE 210 performs message processing operations to forward messages received on one of its ports to another one of its ports. For example, in some embodiments, SFE uses data in the message (e.g., data in the message header) to match the message to flow-based rules, and upon finding a match, attempts to perform the action specified by the matching rule (e.g., pass the message to one of its ports 260 or 265, send the message to be delivered to the destination VM or PNIC).
[0095] In some embodiments, SFE 210 is a software switch, and in other embodiments, a software router or combined software switch / router. SFE 210 implements one or more logical forwarding elements (e.g., logical switches or logical routers) in some embodiments with SFEs running on other hosts in a multi-host environment. The logical forwarding elements in some embodiments can span multiple hosts to connect VMs running on different hosts but belonging to one logical network.
[0096] Different logical forwarding elements can be defined to specify different logical networks for different users, and each logical forwarding element can be defined by multiple software forwarding elements on multiple hosts. Each logical forwarding element separates traffic of VMs of one logical network from VMs of another logical network served by another logical forwarding element. Logical forwarding elements can connect VMs running on the same host and / or different hosts. In some embodiments, the SFE extracts a logical network identifier (e.g., VNI) and a MAC address from the data message. The SFE in these embodiments uses the extracted VNI to identify a logical port group and then uses the MAC address to identify a port within the port group.
[0097] A software switch (e.g., a hypervisor software switch) may be referred to as a virtual switch because it runs in software and provides VMs with shared access to the host's PNIC. However, for the purposes of this specification, a software switch is referred to as a physical switch because it is an item in the physical world. This term also distinguishes a software switch from a logical switch, which is an abstraction of the type of connectivity provided by a software switch. There are various mechanisms for creating a logical switch from a software switch. VXLAN provides one way of creating such a logical switch. The VXLAN standard is described in the IETF's "VXLAN: A Framework for Overlaying Virtualized Layer 2 Networks over Layer 3 Networks," by Mahalingam, Mallik; Dutt, Dinesh G., et al. (2013 / 05 / 08).
[0098] In some embodiments, a port of SFE 210 includes one or more function calls to one or more modules that perform special input / output (I / O) operations on incoming and outgoing messages received at the port. Examples of I / O operations performed by port 260 include ARP broadcast suppression and DHCP broadcast suppression operations described in U.S. Pat. No. 9,548,965. In some embodiments of the invention, other I / O operations (such as firewall operations, load balancing operations, network address translation operations, etc.) may be so implemented. In some embodiments, by implementing a stack of such function calls, a port may implement a sequence of I / O operations on incoming and / or outgoing messages. Also, in some embodiments, other modules in the data path (such as VNIC 255, port 265, etc.) perform I / O function call operations on behalf of or in conjunction with port 260.
[0099] In some embodiments, one or more function calls of the SFE port 260 can be to one or more service engines 130 that process context-based service rules in the context-based service rule store 140. Each service engine 130 in some embodiments has its own context-based service rule store 140, attribute mapping store 223, and connection state cache store 225. FIG. 2 shows only one context-based service rule store 140, attribute mapping store 223, and connection state cache store 225 for all service engines so as not to obscure the view of this figure with unnecessary details. Also, in some embodiments, each VM has its own instance of each service engine 130 (e.g., its own instance of the discovery engine 120, the process control engine 122, the encryption engine 124, the load balancer 126, and the firewall engine 128). In other embodiments, one service engine can service data message flows for multiple VMs on a host (e.g., VMs of the same logical network).
[0100] To perform its service operation on a data message flow, in some embodiments, the service engine 130 attempts to match a flow identifier (e.g., a 5-tuple identifier) and / or a set of associated context attributes of the flow to a rule identifier of the service rule in its context-based service rule store 140. Specifically, for the service engine 130 to perform a service check operation on a data message flow, the SFE port 260 that invokes the service engine provides a set of attributes of the message that the port receives. In some embodiments, the set of attributes is a message identifier, such as a traditional 5-tuple identifier. In some embodiments, one or more of the identifier values may be logical values defined for a logical network (e.g., IP addresses defined in a logical address space). In other embodiments, all of the identifier values are defined in a physical domain. In still other embodiments, some of the identifier values are defined in the logical domain and other identifier values are defined in the physical domain.
[0101] The service engine in some embodiments then uses the attribute set of the received message (e.g., the 5-tuple identifier of the message) to identify the context attribute set that the service engine has stored in the attribute mapping store 223 for this flow. As described above, the context engine 110 in some embodiments provides the context attributes for the new flow (i.e., the new network connection event) and the new process to the service engine 130 along with the flow identifier (e.g., the 5-tuple identifier) or the process identifier. The context engine policy store 143 includes rules that control the operation of the context engine 110. In some embodiments, these policies instruct the context engine to generate rules for the service engine or to the service engine to generate rules (e.g., when a high-threat application runs on a VM, instruct the encryption engines for all other VMs on the same host to encrypt their data message traffic). The service engine 130 in these embodiments saves the context attributes received from the context engine in the attribute mapping store 223.
[0102] In some embodiments, the service engine 130 saves the context attributes set for each new flow or new process with its flow identifier (e.g., 5-tuple identifier) or its process identifier in the attribute mapping store. In this way, the service engine can identify the context attribute set for each new flow it receives from the SFE port 260 by searching its attribute mapping store 223 for a context record with a matching flow identifier. The context record with the matching flow identifier contains the context attributes set for this flow. Similarly, to identify the context attribute set for a process event, in some embodiments, the service engine searches its attribute mapping store 223 for a context record with a matching process identifier.
[0103] As described above, some or all of the service engines in some embodiments pull the context attribute set for a new flow or new process from the context engine. For example, in some embodiments, the service engine provides the 5-tuple identifier of a new flow that it receives from the SFE port 260 to the context engine 110. This engine 110 then consults its attribute storage 145 to identify the set of attributes stored for this 5-tuple identifier, and then provides this attribute set (or a subset of the attribute set obtained by filtering the attribute set identified for the service engine) to the service engine.
[0104] As described above, some embodiments implement the pull model by encoding a set of attributes for a new message flow using a service token. When notified of a new network connection event, in some embodiments, the context engine 110 (1) collects a set of context attributes for the new event, (2) filters this set to discard attributes that are not relevant to performing one or more services on the flow, (3) saves the remaining filtered attribute subset in the attribute storage 145 along with the service token, and (4) provides the service token to the GI agent 250. The GI agent 250 then passes this token to the service engine either in-band (e.g., in the header of a data message that the agent's VM sends to a destination) or out-of-band (i.e., separate from the data message that the agent's VM sends to a destination).
[0105] When a service engine gets a new flow through the SFE port 260, it provides the flow's service token to the context engine, which uses the service token to identify context attributes in its attribute storage 145 to provide to the service engine. In embodiments where the SFE port does not provide the service token to the service engine, the service engine must first identify the service token by searching its data store using the flow's identifier before providing it to the context engine.
[0106] After identifying the context attribute set for the data message flow, in some embodiments, the service engine 130 performs the service operation based on the service rules stored in the context-based service rule storage 140. To perform the service operation, the service engine 130 matches the received attribute subset with a corresponding attribute set stored for the service rule. In some embodiments, each service rule in the context-based service rule storage 140 has a rule identifier and a set of action parameters.
[0107] As discussed above, the rule identifier of a service rule in some embodiments may be defined in terms of one or more context attributes that are not L2-L4 header parameters (e.g., L7 parameters, process identifiers, user identifiers, group identifiers, process names, process hashes, loaded module identifiers, consumption parameters, etc.). In some embodiments, the rule identifier may also include L2-L4 header parameters. Also, in some embodiments, one or more parameters in the rule identifier may be specified in terms of individual values or wildcard values. Also, in some embodiments, the rule identifier may include an individual set of values or a group identifier, such as a security group identifier, a compute construct identifier, a network construct identifier, etc.
[0108] To match the received attribute set with a rule, the service engine compares the received attribute set with associated identifiers of service rules stored in the context-based service rule storage 140. Upon identifying a matching rule, the service engine 130 performs a service action (e.g., a firewall action, a load balancing action, an encryption action, other middlebox action, etc.) based on a set of action parameters of the matching rule (e.g., based on allow / drop parameters, load balancing criteria, encryption parameters, etc.).
[0109] In some embodiments, the context-based service rule store 140 is defined in a hierarchical manner to ensure that when a message's attribute subset matches multiple rules, the message rule check matches a higher priority rule before matching a lower priority rule. Also, in some embodiments, the context-based service rule store 140 includes a default rule that specifies a default action for any message rule check that cannot identify other service rules, which in some embodiments is a match for a subset of all available attributes, ensuring that the service rule engine returns an action for the subset of all received attributes. In some embodiments, the default rule does not specify a service.
[0110] For example, multiple messages may have the same message identifier attribute set if the messages are part of one flow related to one communication session between two machines. Thus, after matching the data message with a service rule in the context-based service rule storage 140 based on the identified context attribute set of the message, the service engine of some embodiments stores the service rule (or a reference to the service rule) in the connection state cache storage 225 so that the service engine can later use this service rule for subsequent data messages of the same flow.
[0111] In some embodiments, the connection state cache storage 225 stores service rules, or references to service rules, that the service engine 130 identifies for different message identifier settings (e.g., for different 5-tuple identifiers that identify different data message flows). In some embodiments, the connection state cache storage 225 stores each service rule, or a reference to a service rule, along with an identifier generated from a set of matching message identifiers (e.g., a 5-tuple identifier of the flow and / or a hash value of the 5-tuple identifier of the flow).
[0112] Before checking the context-based service rule storage 140 for a particular message, the service rule engine 130 of some embodiments checks the connection state cache storage 225 to determine whether it has previously identified a service rule for this message's flow. If not, the service engine 130 identifies the message flow's context attribute set and then checks the context-based service rule storage 140 for a service rule that matches the message's identified attribute set and / or its 5-tuple identifier. When the connection state data storage has an entry for a particular message, the service engine executes its service action based on the action parameter setting of this service rule.
[0113] In the service architecture of FIG. 2, the DPI module 135 performs deep packet inspection on data message flows at the direction of the firewall engine 128. Specifically, when the firewall engine 128 receives a new data message that is part of a new data message flow, in some embodiments, the firewall engine instructs the DPI module to inspect the new data message and one or more of the next few data messages in the same flow. Based on this inspection, the DPI engine identifies the type of traffic (i.e., the application on the wire) being transmitted in this data message flow, generates an AppID for this traffic type, and saves this AppID in the attribute store 145. In some embodiments, the context attribute set is stored in the attribute store based on a flow identifier and / or a process identifier. Thus, in some embodiments, the DPI engine 135 stores the AppID of the new data message flow in the attribute store 145 based on the 5-tuple identifier of that flow.
[0114] In some embodiments, the context engine 110 pushes the AppID of a new data message flow to the service engine 130 when the DPI engine saves the AppID in attribute storage 145. In other embodiments, the context engine 110 pulls the AppID from attribute storage 145 each time it is queried for a context attribute of a data message flow by a service engine. In some embodiments, the context engine 110 uses the 5-tuple identifier of the flow to identify a record in attribute storage 145 that has a matching record identifier and AppID.
[0115] 3 illustrates a process 300 that the context engine executes (110) in some embodiments whenever it is notified of a new process or network connection event. The process 300 first receives (at 305) a notification about a new process or network connection event from the GI agent 250 of the VM 205. Next, at 310, the process 300 collects all desired context attributes for the notified event.
[0116] As described above, the context engine 110, in some embodiments, interacts (at 310) with the reporting GI agent 250 to gather additional information about the reported event. In some embodiments, the GI agent interacts with the network stack and / or processing subsystem in the OS kernel space of the VM to gather context attributes about the process or network event. In some embodiments, the GI agent also gathers this information from a user space module (e.g., a user mode dynamic link library, DLL) that runs in a user space process (e.g., VMtool.exe) to gather context attributes. In a VM using Microsoft Windows, the GI agent in some embodiments registers hooks with the Windows Filtering Platform (WFP) to get network events and with the Windows process subsystem to gather process-related attributes. In some embodiments, the GI agent hooks are in the Application Layer Enforcement (ALE) layer of the WFP, so that the GI agent hooks can capture all socket connection requests from application processes on the VM.
[0117] In some embodiments, the context engine 110 interacts with a management or control plane to gather context attributes and / or receives records that can be examined to identify context attributes for identified network or process events. In some of these embodiments, the context engine interacts with a management or control plane proxy (running on the host) to obtain data from a management or control plane server running external to the host. In some of these embodiments, the context engine runs in kernel space.
[0118] After collecting the context attributes at 310, the process uses (at 315) the attributes of the received event or the collected context attributes for the received event to identify one or more policies in the context engine policy store 143. At 315, the process identifies any policies that have a policy identifier that matches the collected attributes and the event attributes.
[0119] Next, at 320, the process generates context attribute mapping records for one or more service engines based on the policies identified at 315. One or more of the identified policies may specify that for a particular process or network event, a particular set of service engines should be notified about the event (e.g., about a new data message flow), and that each service engine receives a subset of context attributes that are relevant to that service engine in order to perform its processing for that event. This operation in some embodiments includes the context engine not including attributes that are not relevant to each particular service engine in the subset of context attributes that the context engine provides to that particular service engine.
[0120] In some embodiments, some events may require that new service rules be created for one or more service engines. For example, when a high-threat application is identified on one VM, a policy may specify that other VMs on that host must begin encrypting their data message traffic. In some such embodiments, the context engine policy store 143 includes policies that instruct the context engine to generate service rules for a service engine under certain circumstances, or instruct the service engine to generate such service rules. For such embodiments, the process (320) generates service rules for a service engine under certain circumstances, or instructs the service engine to generate such service rules, as needed.
[0121] At 325, process 300 distributes the mapping records and / or generated service rules / instructions to one or more service engines. As discussed above, the context engine can distribute such records and / or rules / instructions using a push model or a pull model. When using a pull model, process 300, in some embodiments, performs operation 325 in response to a query from a service engine, as well as performing some or all of operation 320 in response to the query. After 325, the process ends.
[0122] Load balancing engine 126 is a context-based load balancer that performs its load balancing operations based on load balancing rules that can be specified not only with respect to L2-L4 parameters, but also with respect to context attributes. FIG. 4 shows an example of such context-based load balancing rules. These rules are used independently by different load balancing engines 126 on different hosts 500 to distribute data messages from different web server VMs 505 on different hosts to different application server VMs 510, as shown in FIG. 5. In some embodiments, host 500 is similar to host 200, and web server VMs 505 and application server VMs 510 are similar to VMs 205 of FIG. 2. To avoid obscuring the figure with unnecessary details, the only components of host 500 shown in FIG. 5 are load balancer 126 and web server VMs 505, with the load balancer appearing as a service module in the egress path of web server VMs 505.
[0123] In the example shown in FIG. 5, the load balancers 126 span multiple hosts and collectively form a distributed load balancer (i.e., a single conceptual logical load balancer) that distributes web server VM traffic uniformly across application server VMs. As described further below, the load balancing criteria of the load balancing rules for the different load balancers 126 in some embodiments are the same to ensure uniform handling of web server VM traffic. As shown by host 500b, multiple web server VMs 505 can run on the same host 500, each of which is served by a different load balancer 126. In other embodiments, one load balancer can serve multiple VMs on the same host (e.g., multiple VMs for the same tenant on the same host). Also, in some embodiments, multiple application server VMs can run on the same host as each other and / or the web server VMs.
[0124] In the example shown in Figures 4-5, the load balancer 126 performs a destination network address translation (DNAT) that translates the virtual IP (VIP) address of an application server to the IP address of a particular application VM, referred to as the destination IP address or DIP. In other embodiments, the DNAT operation can translate other network addresses, for example, translating MAC addresses to achieve MAC redirection. Also, while the load balancer in Figures 4-5 distributes web server traffic among application servers of an application server group (cluster), the load balancer can be used to distribute traffic from any set of one or more services or compute nodes of any service / compute node cluster.
[0125] 4 illustrates a load balancing (LB) rule storage 140 that stores several LB rules, each associated with one load-balanced compute or service cluster. Each load balancing rule includes (1) a rule identifier 405, (2) several IP addresses 410 of several nodes in a load balancing node group, and (3) a weight value 415 for each IP address. Each rule identifier 405 specifies one or more data tuples that can be used to identify a rule that matches a data message flow.
[0126] In some embodiments, the rule identifier may include a VIP address of the rule's associated DCN group (such as the VIP address of the application server group in FIG. 5). As shown, the rule identifier may include any L2-L4 parameters (e.g., source IP address, source port, destination port, protocol, etc.) in some embodiments. Also, as shown, the rule identifier for a LB rule may include context attributes such as AppID, application name, application version, user ID, group ID, threat level, resource consumption, etc. In some embodiments, the load balancer searches the LB data store by comparing one or more message attributes (e.g., 5-tuple header value, context attributes) with the rule identifier 405 to identify the highest priority rule with a matching rule identifier.
[0127] The IP address 410 of each LB rule is the IP address of a compute or service node that is a member of the load-balanced compute or service group associated with the rule (e.g., of the load-balanced group associated with the VIP address specified in the rule's identifier 405). As described further below, in some embodiments, the destinations of these nodes are served by a set of controllers that make up the load balancer. In some embodiments, the load balancer translates the destination IP address of a received data message into one of the IP addresses 410 (e.g., translates a VIP included in the data message into a DIP).
[0128] In some embodiments, the load balancer performs its load balancing operations based on one or more load balancing criteria. The LB rules of some embodiments include such LB criteria within each rule to specify how the load balancer should distribute traffic across the nodes of the load-balanced group when a message flow matches that rule. In the example shown in FIG. 4, the load balancing criteria is a weighted round-robin scheme defined by the weight value 415 of each rule. Specifically, the weight value 415 of each IP address of each LB rule provides the criteria for the load balancer to distribute traffic to the nodes of the LB node group.
[0129] For example, the application server group in Fig. 5 has five application servers. Assume that the load balancing rules for distributing web server traffic among these five application servers specify weight values of 1, 3, 1, 3, and 2 for the five IP addresses of the five application servers. Based on these values, the load balancer distributes data messages that are part of the ten new flows in the following order: the first flow to the first IP address, the second flow to the second IP address, the fifth flow to the third IP address, the sixth through eighth flows to the fourth IP address, the ninth and tenth flows to the fifth IP address, and so on. According to this scheme, the allocation of the next batch of new flows loops back through these IP addresses (e.g., the eleventh flow is assigned to the first IP address, and so on).
[0130] In some embodiments, the weight values of the LB rules are generated and adjusted by the controller configuration based on (1) LB statistics that the load balancer collects regarding data message flows that it distributes among the nodes of the LB node group and (2) provides to the controller configuration. The LB rules in some embodiments also specify time periods for the different load balancing criteria of the LB rules that are valid for different time periods in order to appropriately switch between different load balancing criteria. Thus, in some embodiments, a load balancing rule can specify multiple different configurations of destination network addresses with different time values that specify when each configuration of the address is valid. Each of these sets in some embodiments can include its own set of LB criteria (e.g., its own set of weight values).
[0131] 6 illustrates a process 600 that load balancer 126 executes in some embodiments. As illustrated, process 600 begins when the load balancer receives (at 605) a data message from its corresponding SFE port 260. When the port receives a data message from a VM or a VM, it relays the message. In some embodiments, the port relays the data message by passing the load balancer a reference to the data message or a header value of the data message (e.g., a handle that identifies a location in memory to store the data message).
[0132] The process determines (at 610) whether the connection state cache 225 stores a record identifying a destination to which the data message should be forwarded. As described above, each time the load balancer uses the LB rule to direct a new data message flow to a node of the load-balanced node group, the load balancer, in some embodiments, creates a record in the connection state cache 225 to store the physical IP address of the selected node, so that when the load balancer receives another data message in the same flow (i.e., with the same 5-tuple identifier), it can forward the data message to the same node used for the previous data message in the same flow. The use of the connection state cache 225 allows the load balancer 126 to process the data message flow more quickly. In some embodiments, each cached record in the connection state cache 225 has a record identifier defined in terms of a data message identifier (e.g., 5-tuple identifier). In these embodiments, the process compares the identifier (e.g., 5-tuple identifier) of the received data message to the record identifiers of the cached records to identify any records having a record identifier that matches the identifier of the received data message.
[0133] Once process 600 identifies (at 610) a record in connection state cache 225 for the flow of the received data message, the process replaces (at 615) the destination address of the message (e.g., a VIP address) with the destination address (e.g., a DIP) stored in the record in connection state cache 225. At 615, the process sends the address-translated data message along its data path. In some embodiments, this action involves communicating back to the SFE port 260 (which called the load balancer to initiate process 600) to indicate that the load balancer has completed processing the VM data message. The SFE port 260 can then hand off the data message to an SFE or call another service engine in the I / O chain operator to perform another service operation on the data message. At 615, process 600 also updates statistics it maintains for its data message flow process in some embodiments. After 615, process 600 ends.
[0134] If the process 600 determines (at 610) that the connection state cache 225 does not store a record for the flow of the received data message, the process 600 identifies (at 620) one or more context attributes for this data message flow. As discussed above, the service engines of different embodiments perform this operation differently. For example, in some embodiments, the load balancer 126 checks the attribute mapping store 223 for a record having a record identifier that matches a header value (e.g., its 5-tuple identifier) of the received data message. The context attribute set of this matching record is then used (620) as the context attribute set for the received data message flow.
[0135] In another embodiment, the load balancer 126 queries the context engine to obtain a set of context attributes for a received data message. With this query, the load balancer provides the flow identifier (e.g., a 5-tuple identifier) of the received message or its associated service token. The context engine then uses the flow identifier of the message or its associated service token to identify a set of context attributes in its context attribute store 145, and then provides the identified set of context attributes to the load balancer 126, as described above.
[0136] Once the process 600 has obtained a context attribute set for the received data message, it uses the attribute set along with other identifiers of the message to identify (at 625) a LB rule in the LB rule store 140 for the received data message at 605. For example, in some embodiments, a LB rule has a rule identifier 405 defined in terms of one or more of the 5-tuple attributes along with one or more context attributes such as application name, application version, user ID, group ID, AppID, threat level, resource consumption level, etc. To identify a LB rule in the LB rule store 140, the process in some embodiments compares the context attributes and / or other attributes (e.g., 5-tuple identifier) of the received data message with the rule identifiers (e.g., rule identifier 405) of the LB rules to identify a highest priority rule having an identifier that matches the attribute set of the message.
[0137] In some embodiments, the process uses different sets of message attributes to perform this comparison operation. For example, in some embodiments, the message attribute set includes the destination IP address of the message (e.g., the VIP of the addressed group of nodes) along with one or more context attributes. In other embodiments, the message attribute set includes other attributes, such as one or more of other 5-tuple identifiers (e.g., one or more of source IP, source port, destination port, and protocol). In some embodiments, the message attribute set includes a logical network identifier, such as a virtual network identifier (VNI), a virtual distributed router identifier (VDRI), a logical MAC address, a logical IP address, etc.
[0138] As discussed above, each LB rule in some embodiments includes two or more destination addresses (e.g., IP addresses 410) that are destination addresses (e.g., IP addresses) of nodes that are members of a load balancing node group. When the process identifies (at 630) a LB rule, it selects one of the destination addresses (e.g., IP addresses) of the rule to replace a virtual address (e.g., a VIP address) in the message. Also, as discussed above, each LB rule in some embodiments stores a set of load balancing criteria to facilitate the process's selection of one of the destination addresses of the LB rule to replace the virtual destination identifier of the message. In some embodiments, the stored criteria are weights and / or time values as discussed above with reference to Figures 4 and 5. Thus, in some embodiments, the process 600 selects one of the destination addresses of the matching rule based on the selection criteria stored in the rule and changes the destination address of the message to the selected destination address.
[0139] After modifying the destination address of the data message, the process sends (at 630) the data message along its data path. Again, in some embodiments, this action involves communicating back to SFE port 260 (which called the load balancer to initiate process 600) to indicate that the load balancer has completed processing the VM data message. SFE port 260 can then hand off the data message to an SFE or VM, or call another service engine in the I / O chain operator to perform another service operation on the data message.
[0140] After 630, the process transitions to 635 where in the connection state cache storage 225, the process creates a record identifying the computing or service node in the load balanced group to use for forwarding data messages that are part of the same flow as the data message received in 605 (i.e., identifying the node destination identifier). At 635, the process 600 also updates statistics maintained for the node whose message was addressed by the process 600. This update reflects the transmission of the new data message to this node. After 635, the process ends.
[0141] Because context attributes are used to define rule identifiers for LB rules, the context-based load balancer 126 can distribute data message flows based on any number of context attributes. As described above, examples of such load balancing operations include (1) distributing data message flows associated with a finance department to all load balancing pools, (2) redirecting all finance department traffic to another pool if the first pool for this department goes down to make this department's traffic highly available, (3) making all traffic associated with a Doctor's user group highly available, and (4) distributing data message flows for finance applications among service nodes of a low latency service node group. In some embodiments, load balancing rules can also be defined with respect to the aggregated resource consumption to distribute traffic to provide more or less resources to resource-intensive applications on the DCN.
[0142] In some embodiments, the management plane takes an inventory of all processes and services running on VMs on hosts in the data center. In some embodiments, the discovery engine 120 of the host 200 assists in collecting this data from the VMs running on that host. In some embodiments, the inventoried processes / services are referred to as inventoried applications, which include all client processes, services or daemons that utilize network input / output, and all server processes that have registered to listen (i.e., get messages) on a particular network connection. In some embodiments, the discovery engine uses the GI agent 250 and MUX 227 to collect this data.
[0143] Based on the data collected by all discovery engines on all hosts, the management server (e.g., network manager and / or compute manager) builds an inventory of running applications. In some embodiments, each application is identified by comparing its file hash obtained from the VM 205 with the hash of the application file stored in the management plane's application data store. The management plane in some embodiments causes the discovery engines to update their data collection so that the management plane can refresh its inventory on a schedule.
[0144] The management plane in some embodiments then provides a rule creation interface to allow an administrator to create context-based LB rules and / or policies for the LB engine 126 (as well as service rules for other service engines 130). The rule creation interface allows an administrator to define high-level LB policies (as well as other service policies) based on data collected by the discovery engine 120 and applications inventoried via context attributes collected by the context engine 110 and by the management plane's interfaces with other management server clusters.
[0145] Once high-level LB policies (and other service policies) are defined in the management plane, the management plane feeds some or all of these policies directly to management proxies (not shown) on the host 200 and / or feeds some or all of these policies indirectly to these proxies via a set of controllers (e.g., network controllers). In some embodiments, the management proxies publish the received policies as rules to the context-based service rule store 140. In some embodiments, the proxies translate these policies before publishing them to the context-based service rule store 140. For example, in some embodiments, the policies are published with an AppliedTo tuple that identifies the associated service node and / or logical network. In some of these embodiments, the management proxies on the host remove the AppliedTo tuple from each service policy before pushing the policy to the service rule store 140 as a service rule. Also, as described above, the context engine 110 on the host 200 in some embodiments resolves policies based on the collected context attributes to generate rules for the service engines.
[0146] Firewall engine 128 is a context-based firewall engine that performs its firewall operations based on firewall rules that can be specified not only in terms of L2-L4 parameters, but also in terms of context attributes. Figure 7 illustrates some examples of such firewall rules. The figure illustrates the firewall rule data store 140 of some embodiments. As shown, each firewall rule includes a rule identifier 705 and firewall action parameters 710.
[0147] In some embodiments, the firewall action parameters 710 may specify any one of the traditional firewall actions, such as allow, drop, reroute, etc. Each rule identifier 705 specifies one or more data tuples that can be used to identify a rule that matches a data message flow. As shown, the rule identifier, in some embodiments, may include any L2-L4 parameters (e.g., source IP address, source port, destination port, protocol, etc.). One or more of these parameters may be virtual parameters (e.g., VIP of the destination cluster) or logical identifiers (e.g., logical network identifiers).
[0148] In some embodiments, the rule identifier may also include context attributes such as AppID, application name, application version, user ID, group ID, threat level, and resource consumption. In some embodiments, the firewall engine searches the firewall data store by comparing one or more message attributes (e.g., 5-tuple header value, context attributes) to the rule identifier 705 to identify the highest priority rule with a matching rule identifier.
[0149] In some embodiments, different firewall engines 128 on different hosts enforce the same set of firewall rules. For example, in some embodiments, different firewall engines 128 process the same firewall rules on different hosts for VMs of one logical network to provide a level of security on data messages sent or received by those VMs. For this logical network, these firewall engines 128 collectively form a distributed firewall engine (i.e., a single conceptual logical firewall engine) that spans multiple hosts.
[0150] 8 shows some more detailed examples of context-based firewall rules of some embodiments. In these examples, the rule identifier 705 of each rule is expressed in terms of a 5-tuple identifier and one or more context attributes. Each rule has one or more attributes in its 5-tuple identifier that are wildcard values designated by an asterisk to specify that the values of these attributes do not matter (i.e., a data message flow can have any value for these attributes without failing to match the rule).
[0151] The first rule 835 specifies that all data message flows from Skype version 1024 should be dropped. The rule identifier for this rule is expressed solely in terms of the context attributes of the data message flow. As mentioned above and further described below, each time firewall engine 128 identifies a new data message flow, it identifies a record that specifies the context attributes of the flow's 5-tuple identifier by interacting with the context engine or by examining records in its attribute mapping store 223 to identify the context attributes of the flow.
[0152] A second rule 830 specifies that all data message flows with Group ID equal to Nurses and App ID equal to YouTube traffic are dropped. By enforcing this rule, the firewall engine 128 can ensure that Nurses logged into that VM 205 cannot see YouTube traffic. Again, the rule identifier for this rule is expressed solely in terms of the context attributes of the data message flow. In this example, the context attributes are Group ID and App ID.
[0153] FIG. 9 illustrates an example illustrating the enforcement of a second rule 830 by the firewall engine 128. Specifically, the firewall engine 128 is shown allowing a first data message 905 from a browser 910 to pass while blocking a second data message 915 from the browser. As illustrated, both of these data messages are associated with an operation performed on the browser by a nurse 920. The first data message flow is allowed to pass because it is associated with an email that the nurse is sending through the browser. The firewall engine 128 allows the message to pass because it does not match any firewall rules that require the message to be blocked. However, the firewall engine 128 blocks the second data message because it is associated with the nurse attempting to watch a YouTube video, and this type of data message flow is prohibited by rule 830.
[0154] 8 specifies that all data message flows associated with a high threat level indicator should be blocked if they are destined for a particular destination IP address A. The rule identifier for this rule is defined in terms of a context attribute (i.e., the high threat level indicator) and one attribute (the destination IP address) within the 5-tuple identifier of the data message.
[0155] FIG. 10 shows an example illustrating the enforcement of this rule 825 by the firewall engine 128 in two stages. Specifically, in a first stage 1002, the firewall engine 128 allows a data message 1010 to be passed from the VM 205 to another VM 1020 (outside the host) with a specific destination IP address A, and in a second stage 1004, an application 1005 is shown to be installed on the VM 205. This application is designated as a high threat application by the threat detector 132. Whenever a new data message flow starts on the VM, the context engine associates this data message flow with a high threat level tag. Thus, in the second stage 1004, the firewall engine 128 blocks a data message 1050 from the VM 205 to the other VM 1020 because this data message is associated with a high threat level, and the rule 815 prohibits such a data message from being sent to the IP address A of the VM 1020.
[0156] The fourth and fifth rules 820 and 815 of FIG. 8 specify that data messages related to the Doctor and Nurses group can access VMs related to VIP address A, but data messages related to the Accountants group cannot access these VMs. The rule identifier for the fourth rule is defined in terms of two context attributes (i.e., Doctor and Nurses group identifier) and one attribute (VIP destination address A) in the 5-tuple identifier of the data message. The rule identifier for the fifth rule is defined in terms of one context attribute (i.e., Accountants group identifier) and one attribute (VIP destination address A) in the 5-tuple identifier of the data message. In some embodiments, the VIP address is a destination for a cluster of VMs performing the same function, and the load balancer translates the VIP address to the IP address of one of the VMs in the cluster.
[0157] FIG. 11 illustrates an example illustrating the enforcement of the fourth and fifth rules 820 and 815 by the firewall engine 128. Specifically, it shows two users logged in simultaneously to a VM operating as a terminal server. One of these users is a nurse X, and the other is an accountant Y. FIG. 11 further shows that the firewall engine 128 allows a first data message 1105 from the nurse X's session to pass to a VM in a VM cluster 1150 identified by a destination IP address VIP A. It also shows that the firewall engine blocks a second data message 1110 from the accountant Y's session from reaching any of the VMs in the VM cluster because the fifth rule 815 blocks a data message associated with the accountant group ID from reaching the destination IP address VIP A.
[0158] In FIG. 11, the two data messages are for two different actual users logged into the VM at the same time. In other cases, only one user may actually be logged into the VM, but an administrative process may run on the VM along with processes executed for applications started by the logged-in user. The administrative process may be a service / daemon that runs on the VM in a different user context than the logged-in user. Services generally run in an admin / root context and not in the logged-in user context. This is a potential security hole as it may allow any application running in a non-logged-in user context to access network resources. Thus, even if only a single user is logged into the VM, it may be desirable to specify firewall rules that treat data messages related to background administrative processes differently from data messages related to processes executed for applications started by the logged-in user.
[0159] As an example, the sixth rule 810 allows data messages of processes related to applications operated by individuals in the high security group to access other VMs that have sensitive data (in this example, these VMs are part of a cluster of VMs associated with IP address VIP B), while the seventh rule 805 blocks data messages related to background management processes from accessing such VMs. This is useful to ensure that IT personnel or hackers cannot create a backdoor to access sensitive data by installing a management process that accesses a high security VM that piggybacks off the login session of a user with appropriate clearance.
[0160] FIG. 12 depicts an example illustrating the enforcement of the sixth and seventh rules 810 and 805 by firewall engine 128. Specifically, it illustrates two data message flows originating simultaneously from one VM. One data message flow is associated with the CEO and another data message flow is associated with a background IT utility process called Utility Q. FIG. 12 further illustrates firewall engine 128 allowing a data message 1205 from the CEO's session to pass to a VM in high-security cluster 1250 identified by VIP address B. Because the seventh rule 805 prevents data messages associated with a management process (such as the Utility Q process) from reaching VIP address B, it also illustrates the firewall engine blocking a second data message 1210 from Utility Q's session from reaching any of the VMs in the high-security cluster.
[0161] The firewall engine can distinguish between data message flows for two different processes running simultaneously on a VM for two different login / administration credentials because the context engine collects the user and group identifiers for each data message flow when each flow starts and associates each flow with its user and group identifiers. Figure 13 shows a process 1300 that the context engine executes to collect user and group identifiers each time it receives a new network connection event from the GI agent.
[0162] The process 1300 begins by receiving 1305 a notification of a new network connection event from the GI agent 250 on the VM 205. As described above, the GI agent in some embodiments provides the following information in the new network connection notification: a 5-tuple identifier of the connection, an identifier of the process requesting the network connection, a user identifier associated with the requesting process, and a group identifier associated with the requesting process.
[0163] Next, at 1310, the process 1300 queries the GI agent to gather other context attributes required for the new network connection event. Examples of such additional parameters include additional parameters related to the process requesting the network connection. At 1315, the process 1300 publishes one or more context attribute records to one or more service engines. In some embodiments, each context attribute record to each service engine includes a 5-tuple identifier of the connection and a set of one or more context attributes including a user identifier and / or a group identifier. The service engine then stores the provided context attribute records in its attribute mapping storage 223 so that the service engine can use these records to identify context attribute sets associated with different data message flows that the service engine processes. In the case of a service engine with different service rules for different processes running simultaneously on a VM for different user accounts, the context attribute set includes a user identifier or a group identifier, allowing these service engines to associate different data message flows from the VM with different user / group identifiers.
[0164] In some embodiments, the context engine does not include in the service engine's context attribute record context attributes that are not required by the service engine. Also, in some embodiments, the context engine provides different context attribute records to different context engines for the same network connection event because different service engines require different sets of context attributes. As mentioned above, the context engine 110 in some embodiments does not push the context attribute set for a new network connection to some or all of the service engines, but rather has those service engines pull these attribute sets.
[0165] In some embodiments, the context engine can associate a data message flow on a source host with the context attributes of the source VM (i.e., the VM from which the data message flow originates) and the destination VM on the same host or a different destination host. The firewall engine 128 in these embodiments can then use such destination-based context attributes to resolve firewall rules. For example, the firewall engine can drop all data messages addressed to a particular type of server (e.g., a Sharepoint server). To support such destination-based rules, the context engine of some embodiments instructs the GI agent to identify the process that registers for notifications on a particular port and uses this information along with the process identifier and hash to identify the application that serves as the destination of the data message flow. The information collected by the context engines on different hosts is collected by a management plane (e.g., by a management server running on a separate computer or on the same host that runs the VM), which aggregates this data and distributes the aggregated data to other context engines. The context engines on the hosts can then use the distributed information to resolve context policies on these hosts and feed context-based rules to context-based service engines on these hosts.
[0166] 14 illustrates a process 1400 that firewall engine 128 executes in some embodiments. As shown, process 1400 begins when the firewall engine receives (at 1405) a data message from its corresponding SFE port 260. When the port receives a data message from a VM or a VM, it relays the message. In some embodiments, the port relays the data message by passing the firewall engine a reference to the data message or a header value of the data message (e.g., a handle that identifies a location in memory to store the data message).
[0167] The process determines (at 1410) whether the connection state cache 225 stores a record that identifies a firewall action for the message flow of the received data message. As described above, each time the firewall engine uses a firewall rule to process a new data message, the firewall engine, in some embodiments, creates a record in the connection state cache 225 to store the firewall action that was performed, so that when the firewall engine receives another data message in the same flow (i.e., with the same 5-tuple identifier), it can perform the same firewall action that was performed on a previous data message in the same flow. The use of the connection state cache 225 allows the firewall engine 128 to process the data message flow more quickly. In some embodiments, each cached record in the connection state cache 225 has a record identifier that is defined in terms of the data message identifier (e.g., the 5-tuple identifier). In these embodiments, the process compares the identifier (e.g., the 5-tuple identifier) of the received data message to the record identifiers of the cached records to identify any records that have a record identifier that matches the identifier of the received data message.
[0168] Once process 1400 identifies (at 1410) a record in connection state cache 225 for the flow of the received data message, the process performs (at 1415) the firewall action (e.g., allow, drop, reroute, etc.) specified in this record. Assuming the firewall action does not require dropping the data message, process 1400 sends the processed data message along its data path. In some embodiments, this action involves communicating back to SFE port 260 (which called the firewall engine to initiate process 1400) to indicate that the firewall engine has completed processing the VM data message. SFE port 260 can then hand off the data message to an SFE or VM, or call another service engine in the I / O chain operator to perform another service operation on the data message.
[0169] If the firewall action performed at 1415 results in the data message being dropped, then process 1400 notifies (1415) SFE port 260 of this action. Also, if the firewall action performed at 1415 requires the data message to be rerouted, then process 1400 performs (1415) network address translation on the data message to effect this rerouting, and then returns (1415) the data message to the SFE port so that it can be sent along its data path. After 1415, process 1400 ends.
[0170] If the process 1400 determines (at 1410) that the connection state cache 225 does not store a record for the flow of the received data message, the process 1400 identifies (at 1420) one or more context attributes for this data message flow. As discussed above, the service engines of different embodiments perform this operation differently. For example, in some embodiments, the firewall engine 128 checks the attribute mapping store 223 for a record having a record identifier that matches a header value (e.g., its 5-tuple identifier) of the received data message. The context attribute set of this matching record is then used (1420) as the context attribute set for the received data message flow.
[0171] In another embodiment, the firewall engine 128 issues a query to the context engine to obtain a set of context attributes for a received data message. With this query, the firewall engine provides the flow identifier (e.g., a 5-tuple identifier) of the received message or its associated service token. The context engine then uses the flow identifier of the message or its associated service token to identify a set of context attributes in its context attribute store 145, as described above.
[0172] Once the process 1400 has obtained a context attribute set for the received data message, it uses the attribute set along with other identifiers of the message to identify (at 1425) a firewall rule in the firewall rule data store 140 for the data message received at 1405. For example, in some embodiments, a firewall rule has a rule identifier 705 defined in terms of one or more of the 5-tuple attributes along with one or more context attributes such as application name, application version, user ID, group ID, AppID, threat level, resource consumption level, etc. To identify a firewall rule in the firewall rule data store 140, the process in some embodiments compares the context attributes and / or other attributes (e.g., the 5-tuple identifier) of the received data message to the rule identifiers (e.g., the rule identifier 705) of the firewall rules to identify the highest priority rule having an identifier that matches the attribute set of the message.
[0173] In some embodiments, the process uses different sets of message attributes to perform this comparison operation. For example, in some embodiments, the message attribute set includes one or more of other 5-tuple identifiers (e.g., one or more of source IP, source port, destination port, and protocol) along with one or more context attributes. In some embodiments, the message attribute set includes a logical network identifier, such as a virtual network identifier (VNI), a virtual distributed router identifier (VDRI), a logical MAC address, a logical IP address, etc.
[0174] After the process identifies (at 1425) a firewall rule, it performs (at 1430) the firewall action of that rule (e.g., allow, drop, reroute, etc.) on the received data message. Assuming the firewall action does not require dropping the data message, process 1400 sends (1430) the processed data message along its data path. In some embodiments, this action involves communicating back to SFE port 260 (which invoked the firewall engine to initiate process 1400) to indicate that the firewall engine has completed processing the VM data message. SFE port 260 can then hand off the data message to an SFE or VM, or call another service engine in the I / O chain operator to perform another service operation on the data message.
[0175] If the firewall action performed at 1430 results in the data message being dropped, then process 1400 notifies (1430) SFE port 260 of this action. Also, if the firewall action performed at 1430 requires that the data message be rerouted, then process 1400 performs (1430) network address translation on the data message to effect this rerouting, and then returns (1430) the data message to the SFE port so that it can be sent along its data path.
[0176] After performing the firewall action at 1430, the process creates 1435 a record in the connection state cache storage 225. The record identifies the firewall action for the received flow of data messages. In some embodiments, the record has a record identifier that is defined by referencing the identifier of the data message flow (e.g., its 5-tuple identifier). After 1435, the process ends.
[0177] As mentioned above, the management server in some embodiments interacts with discovery engine 120 running on host 200 in the data center to obtain and refresh an inventory of all processes and services running on the VMs on the host. In some embodiments, the management server (also referred to above and below as the management plane) then provides a rule creation interface to allow an administrator to create context-based firewall rules and / or policies for firewall engine 128 (as well as service rules for other service engines 130). The rule creation interface allows an administrator to define high-level firewall policies (as well as other service policies) based on data collected by discovery engine 120 and applications inventoried via context attributes collected by context engine 110 and by the management plane's interfaces with other management server clusters.
[0178] Once high-level firewall policies (and other service policies) are defined in the management plane, the management plane provides some or all of these policies directly to management proxies (not shown) on the host 200 and / or provides some or all of these policies indirectly to these proxies via a controller configuration (e.g., a network controller). In some embodiments, the management proxies publish the received policies as rules to the context-based service rule store 140. In some embodiments, the proxies translate these policies before publishing them to the context-based service rule store 140. For example, in some embodiments, the policies are published with an AppliedTo tuple that identifies the associated service node and / or logical network. In some of these embodiments, the management proxies on the host remove the AppliedTo tuple from each service policy before pushing the policy to the service rule store 140 as a service rule. Also, as described above, the context engine 110 on the host 200 in some embodiments resolves policies based on the collected context attributes to generate rules for the service engine.
[0179] The encryption engine 124 is a context-based encryptor that performs its encryption / decryption operations based on encryption rules that can be specified not only in terms of L2-L4 parameters, but also in terms of context attributes. FIG. 15 shows an example of such context-based encryption rules. These rules are used independently by different encryption engines 124 on different hosts 200 to encrypt / decrypt data messages sent / received by VMs 205 on these hosts. In this way, encryption engines 124 that implement the same encryption rules on different hosts (e.g., one tenant or one logical network) collectively form a distributed encryption engine (i.e., a single conceptual logical encryption engine) across multiple hosts, uniformly performing the set of desired encryption and decryption operations. In some embodiments, each VM 205 has its own encryption engine, while in other embodiments, one encryption engine 124 can service multiple VMs 205 on the same host (e.g., multiple VMs for the same tenant on the same host).
[0180] FIG. 15 illustrates an encryption rule data store 140 that stores several encryption rules. Each encryption rule includes (1) a rule identifier 1505, (2) an encryption type identifier 1510, and (3) a key identifier 1515. Each rule identifier 1505 specifies one or more data tuples that can be used to identify a rule that matches a data message flow. In some embodiments, the rule identifier can include any L2-L4 parameters (e.g., source IP address, source port, destination port, destination IP, protocol, etc.). In some embodiments, these L2-L4 parameters can be defined in the physical domain or the logical domain. Also, as shown, the rule identifier can include context attributes such as AppID, application name, application version, user ID, group ID, threat level, and resource consumption.
[0181] In some embodiments, the encryptor 124 searches the encryption rules data store 140 by comparing one or more message attributes (e.g., 5-tuple header value, context attribute) to the rule identifier 1505 to identify the highest priority rule with a matching rule identifier. Also, in some embodiments, the encryption rules data store 140 has a default rule that is used when no other rule matches the data message flow. In some embodiments, the default rule does not specify an encryption key because no rule exists for encrypting the data message flow. Also, in some embodiments, when the default rule is returned to the encryptor 124, the encryptor 124 does not encrypt the data message flow for which it is performing the check.
[0182] Each encryption rule's encryption type 1510 specifies the type of encryption / decryption to use, and each rule's key identifier 1515 identifies the key to use for encryption / decryption. In some embodiments, an encryption rule specifies only the key identifier 1515 and not the encryption type 1510, because the key identifier identifies both the key and the encryption / decryption type, or these types are otherwise specified (e.g., pre-configured) in the encryptor.
[0183] The encryption and decryption operations of encryption engine 124 will now be described with reference to Figures 16-18. In some embodiments, these embodiments use a symmetric encryption scheme in which the same key is used to encrypt and decrypt a message, or a permuted version of the same key is used to encrypt and decrypt a message, such that the encryption and decryption operations use the same key or a permuted version of the same key. Other embodiments use an asymmetric encryption scheme (e.g., a source encryptor using its private key and a destination encryptor using the source encryptor's public key).
[0184] 16 illustrates a process 1600 that the encryptor 124 executes to encrypt a data message sent by a VM 205 on a host 200. As illustrated, the process 1600 begins when the encryptor 124 receives (at 1605) a data message from its corresponding SFE port 260. When the port receives the data message from the VM, it relays the message. In some embodiments, the port relays the data message by passing the encryptor a reference to the data message or a header value of the data message (e.g., a handle that identifies a location in memory to store the data message).
[0185] Next, at 1610, the encryptor determines whether its connection state cache 225 stores a cached encryption record for the received data message. In some embodiments, each time the encryptor finds an encryption rule for a VM data message, the encryptor creates a cached encryption record that stores the encryption rule, a criteria to the encryption rule, an encryption key, and / or an identifier for the encryption key in the connection state cache 225.
[0186] The encryptor creates this cache record so that when it receives another data message of the same data message flow, the encryptor does not have to search the encryption rules data store 140 to identify encryption rules for data messages subsequently received in the same flow. In some embodiments, each cached record in the connection state cache 225 has a record identifier defined in terms of a data message identifier (e.g., a 5-tuple identifier). In these embodiments, the process 1600 compares the identifier (e.g., the 5-tuple identifier) of the received data message with the record identifiers of the cached records to identify any records having a record identifier that matches the identifier of the received data message.
[0187] When the process 1600 identifies (at 1610) a cached encryption record for the received data message in the connection state cache 225, the process encrypts (at 1615) the received data message using the key identified by the identified encryption record. In embodiments where the cached record includes an encryption rule or a reference to an encryption rule, the process 1600 retrieves a key identifier from the stored or referenced rule and uses this identifier to retrieve a key from a key data store stored on the host 200, or from a key manager on the host, or from a key data store operating external to the host. Similarly, in embodiments where the cached record includes a key identifier, the process 1600 retrieves the key identifier from the cached record and uses this identifier to retrieve a key from a local or remote key data store or key manager.
[0188] In some embodiments, the process encrypts (at 1615) the payload of the data message (e.g., the L2 payload) by using the identified encryption key, while generating an integrity check value (ICV) hash of the payload, and some or all of the header values (e.g., the physical L3 and L4 header values, and / or the logical L2 or L3 header values), so that the destination of the message must (1) decrypt the encrypted portions of the data message and (2) verify the authenticity and integrity of the payload and header values used in the ICV calculation.
[0189] For some or all of the data message, the encryption process 1600, in some embodiments, also encrypts (at 1615) a portion of the data message header. For data messages exchanged between machines associated with a logical network, some embodiments encrypt all of the physical header values of the data message. Some of these embodiments perform an ICV operation on the logical network identifier (e.g., VNI) and payload so that a decryptor at the destination host can verify the authenticity and integrity of the encrypted data message.
[0190] After encrypting the data message, the process sends (at 1615) the encrypted data message along its data path. In some embodiments, this action involves communicating back to the SFE port 260 (which called the encryptor to begin process 1600) to inform the port that the encryptor has completed processing the data message. The SFE port 260 can then hand off the data message to the SFE 210 or call another I / O chain operator to perform another operation on the data message. After 1615, process 1600 ends.
[0191] If the process 1600 determines (at 1610) that the connection state cache 225 does not store a cached record that matches the received data message, the process 1600 identifies (at 1620) one or more context attributes for this data message flow. As discussed above, the service engines of different embodiments perform this operation differently. For example, in some embodiments, the encryption engine 124 checks the attribute mapping store 223 for a record having a record identifier that matches a header value (e.g., its 5-tuple identifier) of the received data message. The context attribute set of this matching record is then used (1620) as the context attribute set for the received data message flow.
[0192] In another embodiment, the cryptographic engine 124 issues a query to the context engine 110 to obtain a set of context attributes for a received data message. This query causes the cryptographic engine to provide the flow identifier (e.g., a 5-tuple identifier) of the received message or its associated service token, which the context engine then uses to identify a set of context attributes in its context attribute store 145, as described above.
[0193] Once the process 600 has obtained the context attribute set of the received data message, it uses the attribute set by itself or together with other identifiers (e.g., 5-tuple identifier) of the message to identify (at 1625) an encryption rule in the encryption rule data store 140 that matches the attributes of the received data message. For example, in some embodiments, the encryption rule has a rule identifier defined in terms of one or more of the non-context attributes (e.g., 5-tuple attributes, logical attributes, etc.) and / or one or more context attributes such as application name, application version, user ID, group ID, AppID, threat level, resource consumption level, etc. To identify an encryption rule in the encryption rule data store 140, the process in some embodiments compares the context attributes and / or other attributes (e.g., 5-tuple identifier) of the received data message with the rule identifiers (e.g., rule identifier 1505) of the encryption rules to identify the highest priority rule having an identifier that matches the attribute set of the message.
[0194] After 1625, the process determines (at 1630) whether it has identified an encryption rule that specifies that the received data message should be encrypted. As described above, the encryption rule data store 140 has a default encryption rule that matches all data messages and is returned when no other encryption rule matches the received data message. The default encryption rule in some embodiments specifies that the received data message should not be encrypted (e.g., specifies a default key identifier that corresponds to a no encryption operation).
[0195] If the process 1600 determines (at 1630) that the data message should not be encrypted, the process sends (at 1635) the unencrypted message along the message data path. This operation 1630 involves notifying its SFE port 260 that it has completed processing the data message. After 1635, the process transitions to 1645, where it creates a record in the connection state cache storage 225 indicating that encryption should not be performed for the received data message flow. In some embodiments, this record is addressed in the connection state cache 225 based on the 5-tuple identifier of the flow. After 1645, the process ends.
[0196] If the process determines (at 1630) that the data message should be encrypted, the process retrieves (at 1640) a key identifier from the identified rule and uses this identifier to retrieve a key from a local or remote key data store or manager, as described above, and encrypts the received data message with the retrieved key. This encryption (1640) of the data message is identical to the encryption operation 1615 described above. For example, as described above, the process 1600 encrypts the payload (e.g., the L2 payload) of the data message by using the identified encryption key, while performing an ICV operation on some or all of the payload and header values (e.g., the physical L3 and L4 header values, the logical L2 or L3 header values, and / or the logical network identifiers such as VNI and VDRI). For some or all of the data message, the encryption process 1600, in some embodiments, also encrypts (at 1640) some or all of the headers of the data message.
[0197] After encrypting the data message, the process sends (at 1635) the encrypted data message along its data path. Again, in some embodiments, this action involves communicating back to the SFE port 260 to inform the port that the encryptor has completed processing the data message. The SFE port 260 can then hand off the data message to the SFE 210 or call another I / O chain operator to perform another operation on the data message.
[0198] If the encryption rule identified at 1630 is a dynamically created rule after dynamically detecting an event, the encryptor must ensure (at 1640) that the key identifier of the key used to encrypt the data message is included in the data message header before it is sent. Process 1600 achieves this goal in different ways in different embodiments. In some embodiments, process 1600 passes (1640) the key identifier to the (invoked) SFE port 260 so that the invoking port or I / O chain operator can insert the key identifier in the data message header. For example, in some embodiments, one service engine (e.g., another I / O chain operator) encapsulates the data message in a tunnel header that is used to establish an overlay logical network. In some of these embodiments, SFE port 260 passes the key identifier received from process 1600 to this service engine so that this key identifier can be included in its header.
[0199] In other embodiments, the process 1600 does not pass the key identifier to the SFE port, which does not have another service engine encapsulate the key identifier in the overlay network tunnel header. In some of these embodiments, the SFE port 260 has a service engine that simply includes an indication in the overlay network tunnel header that the data message is encrypted. In these embodiments, a decryptor (e.g., encryption engine 124) running on a host with a destination DCN can identify the correct key to use to decrypt the data message based on preconfigured information (e.g., a white-box solution that allows the decryptor to pick the correct key based on a previous key specified for communication with the source DCN, or based on header values in the data message flow), or based on out-of-band communication with a controller or module on the source host regarding the appropriate key to use.
[0200] After 1640, the process transitions to 1645 to create a record in the connection state cache storage 225 that stores the encryption rule identified at 1625, a reference to the encryption rule, the key identifier specified in the encryption rule, and / or the retrieved key specified by the key identifier. As mentioned above, this cached record has a record identifier that, in some embodiments, includes an identifier (e.g., a 5-tuple identifier) of the received data message. After 1645, the process ends.
[0201] 17 illustrates a process 1700 that an SFE port 260 or 265 performs for the encryption engine 124 to decrypt an encrypted data message that it receives on a destination host executing the destination VM of the data message. This encryption engine is referred to below as a decryptor. In some embodiments, the decryptor performs this operation when its corresponding SFE port calls the decryptor to check whether a received data message is encrypted and, if so, to decrypt the message.
[0202] In some embodiments, the SFE port invokes the encryption engine 124 when the SFE port determines that a received data message is encrypted. For example, in some embodiments, the data message is sent to a destination host along a tunnel whose header has an identifier that specifies that the data message is encrypted. In some embodiments, the decoder performs this process only if the header value of the received data message does not specify a key identifier that identifies a key for decrypting the data message. If the header value specifies such a key identifier, the decoder uses the decryption process 1800 of FIG. 18, described below.
[0203] Also, in some embodiments, the received data message has a value (e.g., a bit) that specifies whether the message is encrypted. By analyzing this value, the decoder knows whether the message is encrypted. If this value specifies that the message is not encrypted, the decoder does not invoke either process 1700 or 1800 to decrypt the encrypted message. Instead, the decoder notifies the SFE port that it can send a data message along its data path.
[0204] As shown, process 1700 first identifies (at 1705) a set of message attributes that the process uses to identify encryption rules applicable to a received data message. In different embodiments, the set of message attributes used to derive the rules may be different. For example, in some embodiments, this message attribute set includes a 5-tuple identifier of the received data message. In some embodiments, this message attribute set also includes a logical network identifier associated with the received data message.
[0205] After 1705, the decryptor determines (at 1710) whether its encryption state cache 225 stores a cached encryption rule for the message attribute set identified at 1705. Similar to the encryptor, in some embodiments, each time the encryptor finds an encryption rule for a data message, it stores the encryption rule, a reference to the encryption rule, a key identifier for this rule, or a key for this rule in the connection state cache 225, so that when the decryptor receives another data message having the same identified message attribute set (e.g., when it receives another data message that is part of the same data flow as the original data message), the decryptor does not need to search the encryption rule data store to identify the encryption rule for the subsequently received data message. As mentioned above, the connection state cache 225, in some embodiments, stores encryption rules based on the 5-tuple identifier of the data message. Thus, before searching the encryption rule data store 140, in some embodiments, the decryptor first determines whether the connection state cache 225 stores a matching cached record for the received data message.
[0206] Once the process 1700 identifies (at 1710) a matching cache record for the received data message in the connection state cache 225, the process uses (at 1715) this record to identify a key, and then uses this key to decrypt the encrypted portion of the received data message. In some embodiments, the cached record contains the key, while in other embodiments, the record contains a key identifier or a rule or reference to a rule that contains a key identifier. In the latter embodiment, the process uses the key identifier in the cached record or in a stored or referenced rule to retrieve the key from a local or remote key data store or manager, and then uses the retrieved key to decrypt the encrypted portion of the received data message.
[0207] In some embodiments, part of the decryption operation (at 1715) is to authenticate the ICV-generated hash of the data message header and payload. Specifically, once a portion of the received data message (e.g., its physical (e.g., L3 or L4) header value, or its logical (e.g., VNI) header value) hashed with the payload via an ICV operation by the encryptor, the decryption operation verifies this portion to verify the authenticity and integrity of the encrypted data message.
[0208] After decoding (at 1715) the data message, the process sends (at 1715) the decoded data message along its data path. In some embodiments, this action involves communicating back to the SFE port (which called the decoder to begin process 1700) to inform the port that the decoder has completed processing the data message. The SFE port can then allow the data message to reach its destination VM, or call another I / O chain operator to perform another operation on the data message. After 1715, process 1700 ends.
[0209] If process 1700 determines (at 1710) that the connection state cache 225 does not store an encryption rule for the attribute set identified at 1705, process 1700 searches (at 1720) the encryption rule data store 140 to identify an encryption rule for the received data message. In some embodiments, the destination host receives an out-of-band communication from the source host (directly or via a controller configuration) that provides data that enables the destination host to identify a key identifier, or an encryption rule having the key identifier, to decrypt the encrypted data message. In some of these embodiments, the out-of-band communication includes an identifier for the data message (e.g., a 5-tuple identifier).
[0210] In other embodiments, the encryption engine 124 on the destination host identifies the correct key to use to decrypt the data message based on preconfigured information. For example, in some embodiments, the encryption engines 124 on the source and destination hosts use a white-box solution that (1) steps through encryption keys according to a preconfigured scheme or (2) selects an encryption key based on an attribute of the data message (e.g., a 5-tuple identifier). By having the source and destination encryption engines follow the same scheme to step through or select an encryption key, the white-box scheme ensures that the encryption engine at the destination host's encryptor 124 can select the same encryption key to decrypt the received data message as the source host's encryptor 124 that was used to encrypt the data message.
[0211] If process 1700 cannot find an encryption rule that identifies the key, process 1700, in some embodiments, initiates an error handling process to resolve the unavailability of a decryption key to decrypt the encrypted message. In some embodiments, this error handling process queries a network agent to determine whether it stores an encryption rule for the message attribute set identified in 1705. If the agent has such an encryption rule, the agent provides it to the process (1720). However, in other embodiments, the error handling process does not contact the network agent to obtain the key. Instead, it simply flags the issue for an administrator to resolve.
[0212] Once the process identifies (at 1720) a key identifier or an encryption rule with a key identifier, the process uses (at 1725) the key identifier to retrieve a key from a local or remote key data store or manager and decrypts the received data message with the retrieved key. In some embodiments, part of the decryption operation (at 1725) is to authenticate an ICV-generated hash of the data message header and payload. Specifically, when a portion of the received data message (e.g., its physical (e.g., L3 or L4) header value, or its logical (e.g., VNI) header value) is hashed by the encryptor along with the payload via an ICV operation, the decryption operation validates this portion to verify the authenticity and integrity of the encrypted data message.
[0213] After decoding the data message, the process sends (at 1725) the decoded data message along its data path. In some embodiments, this action involves communicating back to the SFE port (which called the decoder to begin process 1700) to inform the port that the decoder has completed processing the data message. The SFE port can then allow the data message to reach its destination VM, or it can call another I / O chain operator to perform another operation on the data message.
[0214] After 1725, the process transitions to 1730 to create a record in the connection state cache storage 225 that specifies the decryption key, or an identifier for this key, to be used to decrypt data messages having a message attribute set similar to the set identified in 1705 (e.g., to decrypt data messages that are part of the same flow as the received data message). In some embodiments, this record is addressed in the connection state cache 225 based on the 5-tuple identifier of the received data message. After 1730, the process ends.
[0215] In some embodiments, a header value (e.g., a tunnel header) of a received encrypted data message stores a key identifier that identifies a key for decrypting the data message. The encryption engine 124 on the host device then performs its decryption operation by using the key identified by the key identifier. FIG. 18 shows a process 1800 that the encryption engine 124 performs to decrypt an encrypted data message that includes a key identifier in its header. In some embodiments, the decoder in the encryption engine 124 performs this operation when its corresponding SFE port 260 or 265 calls the decoder to check whether the received data message is encrypted and, if so, to decrypt the message. In some embodiments, the decoder performs this process only if the header value of the received data message specifies a key identifier that identifies a key for decrypting the data message.
[0216] As shown, process 1800 first extracts (1805) a key identifier from a received data message. The process then uses (at 1810) the key identifier to retrieve a key from a local key data store / manager on the destination host or a remote key data store / manager not on the destination host, and then uses (at 1815) this key to decrypt the received data message. As described above, part of the decryption operation (at 1815) is to authenticate the ICV-generated hash (of the header and payload of the received data message) that was encrypted with the payload of the data message. In some embodiments, process 1800 stores the key in cache data store 225 so that it does not need to identify this key for other data messages in the same data message flow as the received data message.
[0217] After decoding the data message, the process sends (at 1820) the decoded data message along its data path. In some embodiments, this action involves communicating back to the SFE port (which called the decoder to begin process 1800) to inform the port that the decoder has completed processing the data message. The SFE port can then allow the data message to reach its destination VM, or call another I / O chain operator to perform another operation on the data message. After 1820, process 1800 ends.
[0218] Because context attributes are used to define rule identifiers for encryption rules, the context-based encryptor 124 can distribute data message flows based on any number of context attributes. As discussed above, examples of such encryption operations include (1) encrypting all traffic from Outlook (initiated on any machine) to an Exchange server, (2) encrypting all communication between applications in a three-tier web server, application server, and database server, and (3) encrypting all traffic originating from the Administrators Active Directory group.
[0219] As described above, the management server in some embodiments interacts with the discovery engine 120 running on the host 200 in the data center to obtain and refresh an inventory of all processes and services running on the VMs on the host. The management plane in some embodiments then provides a rule creation interface to allow an administrator to create context-based encryption rules and / or policies for the encryption engine 124 (as well as service rules for other service engines 130). The rule creation interface allows an administrator to define high-level encryption policies (as well as other service policies) based on data collected by the discovery engine 120 and applications inventoried via context attributes collected by the context engine 110 and by the management plane's interfaces with other management server clusters.
[0220] Once high-level encryption policies (and other service policies) are defined in the management plane, the management plane provides some or all of these policies directly to management proxies (not shown) on the host 200 and / or provides some or all of these policies indirectly to these proxies via a controller configuration (e.g., a network controller). In some embodiments, the management proxies publish the received policies as rules to the context-based service rule store 140. In some embodiments, the proxies translate these policies before publishing them to the context-based service rule store 140. For example, in some embodiments, the policies are published with an AppliedTo tuple that identifies the associated service node and / or logical network. In some of these embodiments, the management proxies on the host remove the AppliedTo tuple from each service policy before pushing the policy to the service rule store 140 as a service rule. Also, as described above, the context engine 110 on the host 200 in some embodiments resolves policies based on the collected context attributes to generate rules for the service engine.
[0221] The process control (PC) engine 122 is a context-based PC engine that performs its PC operations based on PC rules that can be specified in terms of context attributes. Figure 19 illustrates some examples of such PC rules. This figure illustrates the PC rule data store 140 of some embodiments. As shown, each PC rule includes a rule identifier 1905 and a PC action 1910. In some embodiments, the PC action 1910 can be (1) allow, (2) stop and disallow, or (3) stop and abort.
[0222] Each rule identifier 1905 specifies one or more data tuples that can be used to identify a rule that matches a data message flow. As shown, the rule identifier can include context attributes such as AppID, application name, application version, user ID, group ID, threat level, resource consumption, etc. In some embodiments, the PC Engine searches the PC data store by comparing one or more message attributes (e.g., context attributes) to the rule identifier 1905 to identify the highest priority rule with a matching rule identifier. In some embodiments, the rule identifier 1905 can also include L2-L4 parameters (e.g., 5 tuple identifiers) associated with the data message flow, and the PC Engine performs its PC action on a per-flow basis. In other embodiments, the PC Engine 122 only performs its PC action to handle the event and leaves it to the firewall engine 128 to perform PC action on a per-flow basis. Thus, in some embodiments, the rule identifier 1905 of the PC Engine's PC rule does not include L2-L4 parameters.
[0223] In some embodiments, different PC Engines 122 on different hosts implement the same set of PC rules. For example, in some embodiments, different PC Engines 122 process the same PC rules on different hosts for VMs of a logical network to provide a level of security to processes running on those VMs. For this logical network, these PC Engines 122 collectively form a distributed PC Engine (i.e., a single conceptual logical PC Engine) that is spread across multiple hosts.
[0224] 19 shows three detailed examples of context-based PC rules of some embodiments. A first rule 1920 specifies that Skype version 1024 should be stopped and disallowed. In some embodiments, each time the PC Engine 122 identifies a new process event, it identifies the context attributes of the event by interacting with the context engine or by examining the records in its attribute mapping storage 223 to identify a record that specifies the context attributes of the process identifier.
[0225] A second rule 1925 specifies that all processes having a high threat level should be stopped and disallowed. As described above, the context engine 110 or the service engine 130 may interact with the threat detector 132 to evaluate the threat level associated with a process. In some embodiments, the threat detector generates a threat score that the context engine, PC engine, or other service engine quantifies into one of several categories. For example, in some embodiments, the threat detector generates a threat score from 0 to 100, and one of the engines 110 or 130 designates a score from 0 to 33 as a low threat level, a score from 34 to 66 as a medium threat level, and a score from 67 to 100 as a high threat level.
[0226] A third rule 1930 specifies that all processes that generate YouTube traffic should be stopped and suspended. In some embodiments, this rule is enforced by the PC Engine, while in other embodiments, a similar rule is enforced by the firewall engine. When the firewall engine enforces such a rule, it enforces this rule on a per-flow basis, and its action is to drop packets associated with this flow. The PC Engine can enforce this rule when it checks process events or when it is called by the SFE port 260 to perform a PC check on a particular flow.
[0227] 20 illustrates a process 2000 executed by the PC Engine 122 in some embodiments. As shown, the process 2000 begins when the PC Engine receives (2005) a process identifier from the context engine 110. The context engine relays this process ID when it receives a process notification from the GI agent 250 on the VM 205.
[0228] The process 2000 determines (2010) whether the connection state cache 225 stores a record identifying a PC action for the received process ID. Each time the PC Engine uses a PC rule to process a new process identifier, in some embodiments, the PC Engine creates a record in the connection state cache 225 to store the performed PC action so that the PC Engine can later rely on this cache for faster processing of the same process identifier. In some embodiments, each cached record in the connection state cache 225 has a record identifier defined in terms of the process identifier. In these embodiments, the process compares the received identifier with the record identifiers of the cached records to identify any records having a record identifier that matches the received process identifier.
[0229] Once the process 2000 identifies (at 2010) a record for the received process event in the connection state cache 225, the process performs (at 2015) the PC action specified in this record. If the action is to disallow or suspend, the PC engine instructs the context engine 110 to disallow or suspend the process. To do this, the context engine 110 instructs the GI agent that reported the event to disallow or suspend the process. The GI agent then instructs the OS's process subsystem to disallow or suspend the processing. After 2015, the process 2000 terminates.
[0230] If process 2000 determines (at 2010) that the connection state cache 225 does not store a record for the received process identifier, then process 2000 identifies (at 2020) one or more context attributes for this process identifier. As discussed above, the service engines of different embodiments perform this operation differently. In some embodiments, the PC Engine instructs the context engine to collect additional process attributes of the received process event, and the context engine collects this information by interacting with the GI Agent.
[0231] Once the process 2000 obtains a set of context attributes for the received data message, it uses the set of attributes to identify (2025) a PC rule in the PC rule data store 140. In some embodiments, a PC rule has a rule identifier 1505 that is defined in terms of one or more context attributes, such as an application name, an application version, a user ID, a group ID, an AppID, a threat level, a resource consumption level, etc. To identify a PC rule in the data store 140, the process in some embodiments compares the collected context attributes to the rule identifiers (e.g., rule identifiers 1905) of the PC rules to identify the highest priority rule that has an identifier that matches the collected set of attributes.
[0232] Once the process identifies (2025) the PC rule, it performs the PC action of this rule (e.g., allow, stop and disallow, stop and abort, etc.) on the received process event. If the action is disallow or abort, the PC engine instructs the context engine 110 to disallow or abort the process. To do this, the context engine 110 instructs the GI agent that reported the event to disallow or abort the process. The GI agent then instructs the process subsystem of the OS to disallow or abort the processing. After performing the PC action at 2030, the process creates (2035) a record in the connection state cache storage 225. This record identifies the PC action of the received process event. After 2035, the process ends.
[0233] As described above, the management server in some embodiments interacts with the discovery engine 120 running on the host 200 in the data center to obtain and refresh an inventory of all processes and services running on the VMs on the host. The management plane in some embodiments then provides a rule creation interface to allow an administrator to create context-based PC rules and / or policies for the PC engine 122 (as well as service rules for other service engines 130). The rule creation interface allows an administrator to define high-level PC policies (as well as other service policies) based on data collected by the discovery engine 120 and applications inventoried via context attributes collected by the context engine 110 and by the management plane's interfaces with other management server clusters.
[0234] Once the high-level PC policies (and other service policies) are defined in the management plane, the management plane provides some or all of these policies directly to management proxies (not shown) on the host 200 and / or indirectly to these proxies via a controller configuration (e.g., a network controller). In some embodiments, the management proxies publish the received policies as rules to the context-based service rule store 140. In some embodiments, the proxies translate these policies before publishing them to the context-based service rule store 140. Also, as described above, the context engine 110 on the host 200 in some embodiments resolves policies based on collected context attributes to generate rules for the service engine.
[0235] FIG. 21 illustrates an example of how service engines 130 are managed in some embodiments. The diagram illustrates multiple hosts 200 in a data center. As illustrated, each host includes several service engines 130, a context engine 110, a threat detector 132, a DPI module 135, several VMs 205, and an SFE 210. Also illustrated is a set of controllers 2110 for managing the service engines 130, the VMs 205, and the SFEs 210. As described above, the context engines 110 in some embodiments collect context attributes that are passed to a management server in the controller set over a network 2150 (e.g., over a local area network, a wide area network, a network of networks (e.g., the Internet), etc.). The controller set provides a user interface for an administrator to define context-based service rules in terms of these collected context attributes, and communicates with the hosts over the network 2150 to provide these policies. The hosts also communicatively connect to each other over this network 2150.
[0236] Many of the features and applications described above may be implemented as software processes specified as a set of instructions recorded on a computer-readable recording medium (also referred to as a computer-readable medium). When these program instructions are executed by one or more processing units (e.g., one or more processors, processor cores, or other processing units), these program instructions cause the processing unit(s) to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, RAM chips, hard drives, EPROMs, etc. Computer-readable media does not include carrier waves and electronic signals passing over wireless or wired connections.
[0237] As used herein, the term "software" includes firmware in a read-only memory or applications stored on a magnetic record that can be loaded into memory for processing by a processor. Also, in some embodiments, multiple software inventions may be implemented as sub-parts of a larger program while remaining separate software inventions. In some embodiments, multiple software inventions may also be implemented as separate programs. Finally, any combination of separate programs that together implement a software invention as described herein is within the scope of the invention. In some embodiments, a software program, when installed to operate one or more electronic systems, defines one or more specific mechanical implementations that perform the operations of the software program.
[0238] 22 conceptually illustrates a computer system 2200 upon which some embodiments of the present invention may be implemented. The computer system 2200 may be used to implement any of the hosts, controllers, and managers described above. As such, it may be used to execute any of the processes described above. The computer system includes various types of non-transitory machine-readable media and interfaces to various other types of machine-readable media. The computer system 2200 includes a bus 2205, a processing unit 2210, a system memory 2225, a read-only memory 2230, a persistent storage device 2235, input devices 2240, and output devices 2245.
[0239] The bus 2205 collectively represents all system, peripheral, and chipset buses that communicatively connect the many internal devices of the computer system 2200. For example, the bus 2205 communicatively connects the processing unit 2210 to the read-only memory 2230, the system memory 2225, and the persistent storage device 2235.
[0240] The processing unit 2210 retrieves instructions to execute and data to process from these various memory units to perform the processes of the present invention. The processing unit may be a single processor or a multi-core processor in different embodiments. The read-only memory (ROM) 2230 records static data and instructions required by the processing unit 2210 and other modules of the computer system. The persistent storage device 2235, on the other hand, is a read-write memory device. This device is a non-volatile memory unit that records instructions and data even when the computer system 2200 is off. Some embodiments of the present invention use a mass storage device (such as a magnetic or optical disk and corresponding disk drive) as the persistent storage device 2235.
[0241] Other embodiments use a removable storage device (such as a floppy disk, flash drive, etc.) as the persistent storage device. Like the persistent storage device 2235, the system memory 2225 is a read-write memory device. However, unlike the storage device 2235, the system memory is a volatile read-write memory, such as a random access memory. The system memory may store some or all of the instructions and data that the processor needs at runtime. In some embodiments, the processes of the invention are stored in the system memory 2225, the persistent storage device 2235, and / or the read only memory 2230. The processing unit 2210 retrieves instructions to execute and data to process from these various memory units to perform the processes of some embodiments.
[0242] The bus 2205 also connects to input devices 2240 and output devices 2245. The input devices enable a user to communicate information and select commands to the computer system. The input devices 2240 include alphanumeric keyboards and pointing devices (also referred to as "cursor control devices"). The output devices 2245 display images generated by the computer system. Output devices include printers, display devices such as cathode ray tubes (CRTs) or liquid crystal displays (LCDs). Some embodiments include devices such as a touch screen that function as both an input and output device.
[0243] 22, the bus 2205 also connects the computer system 2200 to a network 2265 via a network adapter (not shown). In this manner, the computer can be part of a network of computers, such as a local area network ("LAN"), a wide area network ("WAN"), an intranet, a network of networks such as the Internet, etc. Any or all of the components of the computer system 2200 can be used in connection with the present invention.
[0244] Some embodiments include electronic components such as a microprocessor, storage devices and memories that store instructions of a computer program on a machine-readable or computer-readable medium (alternatively referred to as a computer-readable storage medium, a machine-readable medium, or a machine-readable storage medium). Some examples of computer-readable media include RAM, ROM, read-only compact disks (CD-ROMs), recordable compact disks (CD-Rs), re-writable compact disks (CD-RWs), read-only digital versatile disks (e.g., DVD-ROMs, dual-layer DVD-ROMs), various recordable / re-writable DVDs (DVD-RAMs, DVD-RWs, DVD+RWs, etc.), flash memory (e.g., SD cards, mini SD cards, micro SD cards, etc.), magnetic and / or solid-state hard drives, read-only and recordable Blu-Ray® disks, ultra-high density optical disks, any other optical or magnetic media, and floppy disks. The computer-readable medium records a computer program that is executable by at least one processing unit and includes a set of instructions for performing various operations. Examples of computer programs or computer code include machine code, such as produced by a compiler, and files containing higher level code that are executed by a computer, electronic component, or microprocessor using an interpreter.
[0245] While the above description has primarily referred to microprocessors or multi-core processors executing software, some embodiments are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions stored on the circuit itself.
[0246] As used herein, the terms "computer," "server," "processor," and "memory" all refer to electronic or other technological devices. These terms do not include people or groups of people. For purposes of this specification, the terms display and displaying mean displaying on an electronic device. As used herein, the terms "computer-readable medium," "computer-readable mediums," and "machine-readable medium" are generally restricted to tangible, physical objects that store information in a form readable by a computer. These terms do not include any wireless signals, wired downloaded signals, or any other ephemeral or transitory signals.
[0247] Although the present invention has been described with reference to numerous specific details, those skilled in the art will recognize that the present invention may be embodied in other specific forms without departing from the spirit of the invention. For example, some of the figures conceptually depict processes. The specific operations of these processes may not be performed in the exact order shown and described. The specific operations may not be performed in one continuous operation stream, and different specific operations may be performed in different embodiments. Furthermore, the processes may be implemented using several sub-processes or as part of a larger macro process. Thus, those skilled in the art will appreciate that the invention is not limited to the details described above, but rather is defined by the scope of the appended claims.
Claims
1. 1. A computer-implemented method for supporting context-based services on a host computer operating on a plurality of machines and a set of one or more service engines, comprising: a context collector executed by a processor on the host computer, collecting and storing context attributes about events occurring on the set of machines; receiving an identifier associated with the event from the set of service engines; using the received identifier to identify a set of context attributes to be provided to the set of service engines for use by the service engines in processing a data message flow associated with the event; A computer-implemented method comprising:
2. 2. The computer-implemented method of claim 1, further comprising: generating a specific service rule for a specific service engine for one of the events, the specific service engine using the specific service rule to perform a service action on at least one data message flow.
3. 3. The computer-implemented method of claim 2, wherein generating the specific service rule comprises instructing the specific service engine to generate the service rule.
4. 3. The computer-implemented method of claim 2, wherein generating the specific service rule comprises generating the specific service rule in the context collector and providing the generated service rule to the specific service engine.
5. 2. The computer-implemented method of claim 1, wherein the received identifier includes a data message flow identifier, and using the data message flow identifier includes matching the data message flow identifier to match attributes associated with the collected context attributes stored by the context collector.
6. 2. The computer-implemented method of claim 1, wherein each service engine uses a provided set of context attributes to identify service rules that specify service actions that the service engine should perform on data message flows processed by the service engine.
7. 10. The computer-implemented method of claim 1, wherein the received identifiers are tokens, each token identifying a different subset of collected contextual data.
8. 8. The computer-implemented method of claim 7, further comprising generating the token and providing the token to the machine for transmission along with a data message forwarded by the machine, wherein the service engine provides the token to the context collector to receive any context attributes associated with the data message processed by the service engine.
9. 2. The computer-implemented method of claim 1, wherein the event comprises a new network connection event.
10. 2. The computer-implemented method of claim 1, wherein the event comprises a new processing event.
11. 2. The computer-implemented method of claim 1, wherein collecting the context attributes comprises receiving at least a subset of the context attributes from guest introspection agents operating on two or more machines.
12. 2. The computer-implemented method of claim 1, wherein collecting the context attributes comprises receiving a context attribute having an identifier associated with a flow.
13. 2. The computer-implemented method of claim 1, wherein each of the multiple sets of context attributes includes one or more attributes other than Layer 2 (L2), Layer 3 (L3), and Layer 4 (L4) data message header values.
14. A machine-readable medium storing a program which, when executed on at least one processing unit, performs a computer-implemented method according to any one of claims 1 to 13.
15. 1. An electronic device comprising: A set of processing units; A machine-readable medium storing a program which, when executed by at least one processing unit, performs the computer-implemented method of any one of claims 1 to 13.
2. An electronic device comprising:
Citation Information
Patent Citations
Access control system, access control method, and access control program
JP2010282242A
Method and Apparatus for Differently Encrypting Data Messages for Different Logical Networks
US20150381578A1
Event management systems
US20160164893A1
Network service control method
WO2006057048A1