Topological correlation
The topological correlation solution addresses the complexity of malware detection and connectivity issues by using dynamic hop-by-hop metrics and event analysis to automate root cause identification and remediation, improving operational efficiency in network and security environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- PALO ALTO NETWORKS INC
- Filing Date
- 2024-04-12
- Publication Date
- 2026-05-19
AI Technical Summary
Existing IT operations face challenges in efficiently identifying and mitigating malware due to the increasing complexity of network architectures and the labor-intensive process of correlating events across multiple monitoring tools, leading to higher detection and resolution times for connectivity issues.
A topological correlation solution that utilizes dynamic hop-by-hop metrics and event analysis to automate the identification of root causes in network and security environments, providing rapid remediation and reducing the time required to resolve application connectivity issues.
The solution enables efficient suppression of noise in event correlation, facilitates rapid root cause detection, and reduces the time to resolve connectivity issues by automating remediation processes, thereby enhancing operational efficiency.
Smart Images

Figure 2026515746000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 459,494, titled "APPLICATION ACCESS ANALYZER", filed on April 14, 2023; U.S. Provisional Patent Application No. 63 / 459,492, titled "SECURITY POLICY ANALYSIS - DEVOPS APPROACH", filed on April 14, 2023; and U.S. Provisional Patent Application No. 63 / 459,500, titled "TOPOLOGICAL CO - RELATION", filed on April 14, 2023, all of which are hereby incorporated by reference herein for all purposes.
Background Art
[0002] Malware is a common term generally used to refer to malicious software (e.g., including various hostile, invasive, and / or otherwise unwanted software). Malware can be in the form of code, scripts, active content, and / or other software. Examples of the use of malware include interrupting the operation of a computer and / or computer network, stealing confidential information (e.g., confidential information such as identity, financial, and / or intellectual property - related information), and / or obtaining access to a private / dedicated computer system and / or computer network. Unfortunately, as techniques are developed to detect and help mitigate malware, malicious creators find ways to avoid such efforts. Thus, there is a continuing need for improvements to techniques for identifying and mitigating malware.
Brief Description of the Drawings
[0003] Various embodiments of the present invention are disclosed in the following detailed description and accompanying drawings. [Figure 1] Figure 1 shows an exemplary environment in which a malicious application ("malware") is detected and prevented from causing harm. [Figure 2A] Figure 2A shows an embodiment of the data appliance. [Figure 2B] Figure 2B is a functional diagram of the logical components according to an embodiment of the data appliance. [Figure 3] Figure 3 shows an exemplary logical component that may be included in a system for analyzing a sample. [Figure 4A] Figure 4A provides an overview of AIOP for Secure Access Service Edge (SASE) solutions according to several embodiments. [Figure 4B] Figure 4B shows an example use case for a mobile user who is unable to connect to the corporate network due to authentication issues. [Figure 4C] Figure 4C illustrates one exemplary use case for using AIOP for Secure Access Service Edge (SASE) according to several embodiments to reduce the Mean Time to Detect (MTTD) / Mean Time to Resolve (MTTR) for access to SaaS / private applications. [Figure 5A] Figure 5A shows AIOP for SASE architectures to facilitate topological co-relation according to several embodiments. [Figure 5B] Figure 5B shows an alert engine according to several embodiments. [Figure 5C]Figure 5C shows one exemplary relationship between various entities that is automatically generated and used for correlation and PRC according to several embodiments. [Figure 5D] Figure 5D shows a Correlation and Prisma Controller Engine (CPE) according to several embodiments. [Figure 5E] Figure 5E shows one exemplary data model for an AIOP platform according to several embodiments. [Figure 6] Figure 6 shows an ML pipeline architecture for AIOP for a SASE solution to facilitate topological correlation, according to several embodiments. [Figure 7A] Figure 7A shows one exemplary use case of a significant decline in authenticated user detection according to several embodiments. [Figure 7B] Figure 7B shows one exemplary use case for anomaly detection in service-connected cloud nodes, according to several embodiments. [Figure 8A-1] Figure 8A shows the process for performing correlation on an AIOP SASE solution to facilitate a topological correlation technique, according to several embodiments. [Figure 8A-2] Figure 8A shows the process for performing correlation on an AIOP SASE solution to facilitate a topological correlation technique, according to several embodiments. [Figure 8A-3] Figure 8A shows the process for performing correlation on an AIOP SASE solution to facilitate a topological correlation technique, according to several embodiments. [Figure 8B] Figure 8B shows the process for performing correlation on an AIOP SASE solution to facilitate topological correlation techniques, according to several embodiments. [Figure 8C] Figure 8C shows the process for performing correlation on the AIOP SASE solution to facilitate topological correlation techniques, according to several embodiments. [Figure 9] Figure 9 is a flowchart illustrating a process for providing topological correlations according to several embodiments. [Figure 10] Figure 10 is another flowchart relating to a process for providing topological correlations according to several embodiments. [Modes for carrying out the invention]
[0004] The present invention can be implemented in numerous ways, including processes, apparatus, systems, compositions, computer program products embodied on computer-readable storage media, and / or instructions stored in memory, and / or processors configured to execute instructions stored and / or provided by memory coupled to the processor. In this specification, these implementations, or any other forms the present invention may take, may be referred to as techniques. Generally, the order of the steps of the disclosed process may be modified within the scope of the invention. Unless otherwise specified, components such as processors or memory described as configured to perform a task may be implemented as general-purpose components temporarily configured to perform a task at a given time, or as specific components manufactured to perform a task. As used herein, the term “processor” refers to one or more devices, circuits, and / or processing cores configured to process data, such as computer program instructions.
[0005] A detailed description of one or more embodiments of the present invention, along with accompanying drawings illustrating the principles of the present invention, is provided below. While the present invention is described in relation to such embodiments, it is not limited to any embodiment. The scope of the present invention is limited solely by the claims, and the present invention encompasses numerous alternatives, repairs, and equivalents. Numerous specific details are provided below to provide a complete understanding of the present invention. These details are provided for illustrative purposes only, and the present invention may be carried out in accordance with the claims without some or all of these specific details. For clarity, technical materials known in the art related to the present invention are not described in detail so as not to unnecessarily obscure the present invention.
[0006] A firewall generally allows authorized communications to pass through while protecting the network from unauthorized access. Typically, a firewall is a device, a set of devices, or software running on a device that provides firewall functionality for network access. For example, a firewall can be integrated into the operating system of a device (e.g., a computer, smartphone, or other type of network-enabled device). Firewalls can also be integrated or run as software applications on various types of devices or security devices, such as computer servers, gateways, network / routing devices (e.g., network routers), or data appliances (e.g., security equipment or other types of special-purpose devices), and in some implementations, specific operations can be implemented on special-purpose hardware such as ASICs or FPGAs. ru.
[0007] Firewalls typically deny or allow network transmissions based on a set of rules. These sets of rules are often referred to as policies (e.g., network policies or network security policies). For example, a firewall can filter inbound traffic by applying a set of rules or policies to prevent unwanted external traffic from reaching the protected device. Firewalls can also filter outbound traffic by applying a set of rules or policies (e.g., allow, block, monitor, notify, log, and / or other actions that may be specified in firewall rules or firewall policies, which can be triggered based on various criteria, as described here). Firewalls can also filter local network (e.g., intranet) traffic by similarly applying a set of rules or policies.
[0008] Security devices (e.g., security equipment, security gateways, security services, and / or other security devices) can perform a variety of security operations (e.g., firewalls, anti-malware, intrusion prevention / detection, proxies, and / or other security functions), network functions (e.g., routing, quality of service (QoS), workload balancing of network-related resources, and / or other network functions), and / or other security and / or network-related functions. For example, routing can be performed based on source information (e.g., IP address and port), destination information (e.g., IP address and port), and protocol information.
[0009] A basic packet filtering firewall filters network communication traffic by inspecting individual packets transmitted over the network (e.g., a packet filtering firewall or first-generation firewall, which is a stateless packet filtering firewall). A stateless packet filtering firewall typically inspects each individual packet itself and applies rules based on the inspected packet (e.g., using the combination of source and destination address information, protocol information, and port numbers of the packet).
[0010] An application firewall can also perform application layer filtering (e.g., using an application layer filtering firewall or a second-generation firewall that functions at the application level of the TCP / IP stack). An application layer filtering firewall or application firewall can generally identify a given application and protocol (e.g., web browsing using the Hypertext Transfer Protocol (HTTP), Domain Name System (DNS) requests, file transfer using the File Transfer Protocol (FTP), and various other types of applications and other protocols such as Telnet, DHCP, TCP, UDP, and TFTP (GSS)). For example, an application firewall can block unauthorized protocols attempting to communicate on standard ports (e.g., unauthorized / rogue protocols that attempt to sneak through by using non-standard ports for that protocol can generally be identified using an application firewall).
[0011] A stateful firewall can also perform stateful-based packet inspection, where each packet is examined within the context of its network transmission packet flow and the set of packets associated with it. This firewall technique is commonly referred to as stateful packet inspection because it maintains a record of all connections passing through the firewall and can determine whether a packet is the start of a new connection, part of an existing connection, or an invalid packet. For example, the state of a connection can itself be one of the criteria that trigger rules in a policy.
[0012] Advanced or next-generation firewalls, as described above, can perform stateless and stateful packet filtering and application layer filtering. Next-generation firewalls can also perform additional firewall technologies. For example, a given new firewall, often referred to as an advanced or next-generation firewall, can also identify users and content. In particular, a given next-generation firewall extends the list of applications that these firewalls can automatically identify to thousands of applications. Examples of such next-generation firewalls are commercially available from Palo Alto Networks (e.g., Palo Alto Networks' PA series firewalls). For example, Palo Alto Networks' next-generation firewalls use various identification technologies to enable enterprises and service providers to identify and control applications, users, and content—not just ports, IP addresses, and packets. Various identification technologies include Application ID (App-ID) for precise application identification, User ID (User-ID) for user identification (e.g., a user or user group), and Content ID (Content-ID) for real-time content scanning (e.g., controlling web surfing and restricting data and file transfers). These identification technologies allow businesses to securely enable application usage using business-relevant concepts, instead of following the traditional approach provided by conventional port-blocking firewalls.Also, specific-purpose hardware for next-generation firewalls (e.g., implemented as a dedicated device) generally provides a higher performance level for application inspection than software executed on general-purpose hardware (e.g., security devices provided by Palo Alto Networks, which utilize dedicated, function-specific processing tightly integrated with a single-pass software engine to minimize latency and maximize network throughput for Palo Alto Networks' PA series of next-generation firewalls).
[0013] Advanced or next-generation firewalls can also be implemented using virtualized firewalls. Examples of such next-generation firewalls are commercially available from Palo Alto Networks (Palo Alto Networks' firewalls support VMware(R) ESXi TM , TM ,
[0014] , , and NSX TM , Citrix(R) Netscaler SDX TM , KVM / OpenStack (Centos / RHEL, Ubuntu(R)), and Amazon Web Services (AWS), and support various commercial virtualization environments, including the CN series of container next-generation firewalls. For example, virtualized firewalls can support the same or exactly the same next-generation firewall and advanced threat prevention capabilities available on physical form factor devices, enabling enterprises to securely enable the inflow of applications into private, public, and hybrid cloud computing environments. Automation features such as VM monitoring, dynamic address groups, and REST-based APIs allow enterprises to dynamically monitor changes to VMs and reflect their context in security policies, thereby eliminating policy lag that can occur when VMs change.
[0014] Overview of techniques for topological correlation
[0015] Generally, existing information technology (IT) operations must traverse thousands to millions of logs and numerous devices within the enterprise infrastructure to identify application connectivity issues for a user or group of users. Troubleshooting and debugging connectivity problems typically requires domain knowledge expertise, including an understanding of network architecture, routing / switching, server configuration, complex network security policies, and vendor-specific operating systems (OS) and command-line interfaces (CLI). Thus, this significantly increases the human time and average time required to detect and resolve application connectivity issues.
[0016] For example, multiple network and security elements can generate anomalous events / alerts from their vantage points, creating an alert storm for information technology (IT). IT typically lacks adequate visibility into the actual impact of downtime issues when they occur within the enterprise (e.g., viewing the network topology as a static distance). Thus, IT personnel are challenged to isolate issues from multiple symptoms indicated by potentially numerous events / alerts (e.g., resulting in a higher number of false positives). IT personnel must then manually correlate events across different monitoring tools in an attempt to identify the root cause of the problem, which is labor-intensive, time-consuming, and results in limited correlation. This technical problem is common for heterogeneous network and security stacks (e.g., applicable to security stacks, but also generally applicable to heterogeneous network environments, thus relating to IT help desks, NetOps, network and security management, as well as SASE implementation and technical support teams).
[0017] Accordingly, novel and improved solutions for promoting topological correlations are disclosed with respect to various embodiments.
[0018] In several embodiments, novel and improved topological correlation solutions are disclosed (for example, for SASE service providers). Generally, event correlation of application and network access (e.g., access, availability, and / or performance for SaaS / private apps and / or SASE services) is based on relationships as temporal and dependency. For example, event correlation can be used for troubleshooting tunnels through interfaces, such as determining the potential cause when an interface goes down. However, existing event correlation solutions typically view dependencies and topological relationships as static distances. In contrast, the disclosed topological correlation solutions analyze such network topologies as dynamic states derived through hop-by-hop metrics and events that can occur at any of the hops. Specifically, the disclosed topological correlation solution can utilize dynamic states, such as hop-by-hop metrics and events, which may occur at any of the hops. More specifically, the disclosed topological correlation solution can use both static information and dynamic states derived based on node embeddings to correlate events. As a result, such events can then be correlated to isolate symptoms from problems and enable rapid root cause identification, as will be further described below with respect to various embodiments.
[0019] Specifically, multiple network and security elements within a network generate events / alerts that can be correlated and aggregated by the disclosed topological correlation solution to effectively and efficiently suppress and reduce noise, thereby facilitating root cause detection. For example, the disclosed topological correlation solution can provide causal information to isolate problems from symptoms, enabling automated remediation that reduces the time and man-IT time required to identify and resolve such technical issues in complex network and security computing environments.
[0020] For example, the disclosed solution may include an interface (e.g., a natural language (NL) query interface) for operators (e.g., an IT help desk, or IT / administrators for other technical support personnel / users) to detect application reachability, connectivity, and access / authorization issues, which utilizes the disclosed technology for topological correlations, as will be further described below with respect to various embodiments. The disclosed solution facilitates auto remediation. As one example, the solution provides comprehensive details of analyses and checks performed in different categories (separate domains, including user / endpoint analysis, networking analysis, and security policy analysis, as will be further described below) to an actionable verdict for queries submitted by operators. Specifically, the solution automatically discovers the network topology used by a given user (e.g., a user specified in a query) to access a given application (e.g., a SaaS / private app specified in a query), analyzes the operational status of the underlying network infrastructure, performs user authentication analysis, checks the health and reachability of the Domain Name System (DNS) and authentication (Auth) servers that the user reaches before accessing the application, and infers user- or user group-specific security policy reasonsing for any access / permission issues.
[0021] Actionable determination, root cause analysis, and pinpointing the problem significantly reduce the average time to resolve application connectivity issues. Actionable determination, root cause analysis, and problem identification also save operators the hassle and time they would otherwise have to perform by following runbooks / playbooks and debugging multiple devices, which typically require domain knowledge and expertise.
[0022] As one example, the disclosed solution may be used to check connectivity issues between (1) users, multiple users, and / or groups of users to SaaS applications from a mobile user gateway, (2) users, multiple users, and / or groups of users to private applications hosted in an on-premises data center or in a remote branch office, and (3) one or more users, multiple users, and / or groups of users to remote site connectivity to a remote branch or data center.
[0023] In some embodiments, a system / process / computer program product for topological correlation includes deriving a network topology for a cloud-based security service associated with network access to an application for a user or group of users, deriving hop-by-hop metrics for multiple events, the network topology being dynamically monitored, correlating the derived hop-by-hop metrics and events in multiple dimensions to facilitate the automatic determination of the root cause of a problem associated with network access to an application for a user or group of users based on multiple events, and performing a remediation response based on the root cause of the problem associated with network access to an application for a user or group of users based on multiple events.
[0024] In one embodiment, the disclosed topological correlation technique may be used to determine comprehensive / multidimensional correlations, as further described below.
[0025] In one embodiment, the disclosed topological correlation technique can be used to provide efficient and dynamic hop-by-hop analysis, as further described below.
[0026] In one embodiment, the disclosed topological correlation technique may be used to facilitate greater noise suppression of overlapping alerts, as will be further described below.
[0027] In one embodiment, the disclosed topological correlation technique may be used to provide a continuous live impact (e.g., blast radius) analysis, as further described below.
[0028] In one embodiment, the disclosed topological correlation technique may be used to determine the root cause of application access problems by correlating multiple data sources across multiple domains (e.g., to provide enhanced visibility into health, connectivity, and usage for networks, authentication, DNS, SaaS / private application health, security policy configurations, etc.) using AI and ML, as further described below.
[0029] In one embodiment, the disclosed topological correlation technique may be used to automatically detect anomalies and / or performance degradations in network connectivity (e.g., anomalies and / or performance degradations in network connectivity, based on the user's location / access point, based on configurable thresholds for determining reachability and / or performance degradations to a given application for the user), as further described below.
[0030] In one embodiment, the disclosed topological correlation technique can be used to generate human-consumable / understandable and actionable decision analyses that significantly reduce the average time to detect and remediate application connectivity problems, as will be further described below.
[0031] In one embodiment, the disclosed topological correlation technique can be used to perform a comprehensive analysis of various troubleshooting domains within a short period of time (e.g., minutes), which would otherwise typically require a lot of time to troubleshoot each domain, as will be further described below.
[0032] In one embodiment, the disclosed topological correlation technique may be used to perform analyses including identifying problems in network infrastructure, customer network services, client connectivity issues, SaaS / private application (app) health, and reachability issues, as further described below. For example, the disclosed topological correlation technique A can provide actionable summaries for each troubleshooting domain, and operators do not need to possess domain knowledge expertise to detect and remediate problems.
[0033] In one embodiment, the disclosed topological correlation technique can be used to automatically discover (auto-discover) the network topology used by a user to access an application, and to perform an analysis of potential application access problems. It is possible.
[0034] In one embodiment, the disclosed topological correlation technique may be used to provide a security attitude assessment by constructing a unified logical model of computations for firewall security policies.
[0035] In one embodiment, the disclosed topological correlation technique can be used to manage and maintain network topology problems, networking configuration problems, network services, and security policy tracks, which are often complex and potentially error-prone, as will be further described below. For example, the disclosed topological correlation technique can provide a comprehensive analysis of each of these domains using a convenient natural language (NL) query interface.
[0036] In one embodiment, the disclosed topological correlation technique may be used to significantly reduce the operational and support costs for enterprises and their users to access SaaS / private applications.
[0037] In one exemplary implementation, the disclosed topological correlation technique is implemented as the Prisma AI Operations (AIOP) platform, designed for use by network operations center (NOC) personnel supporting SASE customers, providing proactive service level management across customers globally, as further described below. Specifically, Prisma AIOP provides proactive monitoring, alerting, problem isolation, and playbook-driven remediation to deliver SLAs (MTTK / I, MTTR) as desired / required by customers.
[0038] In one embodiment, the disclosed Prisma AIOP can provide AI / ML-powered predictive analytics for network anomaly detection and capacity utilization to avoid outages, as further described below.
[0039] In one embodiment, the disclosed Prisma AIOP can provide a natural language (NL) interface for context troubleshooting and what-if analysis, as further described below.
[0040] In one embodiment, the disclosed Prisma AIOP can provide an automated analysis of existing security policies for the purpose of recommending best practices and / or avoiding sprawl and / or shadow security policies, as will be further described below.
[0041] In one embodiment, the disclosed Prisma AIOPs can provide a sandboxing environment for pre-analysis of security policy changes to reduce the risk of change and enhance the company's security attitude, as will be further described below.
[0042] In one embodiment, the disclosed Prisma AIOP can provide baseline and correlation of real-time and historical data / events from discrete sources for anomaly detection, root cause analysis (RCA), playbook-driven repair, and repair verification, as further described below.
[0043] In some embodiments, a system / process / computer program product for topological correlation includes monitoring access to an application over a network, generating a baseline for monitored access to the application over the network for each of several dimensions, automatically predicting network metrics for one or more of the dimensions, and generating alerts or reports with recommended configurations and / or remediation based on the predicted network metrics for one or more of the dimensions.
[0044] In addition, network traffic patterns at different office / branch locations (e.g., within a corporate network) can vary significantly in terms of the number of users, user traffic access patterns, and usage. Thus, the network for an office / branch location (e.g., a corporate site) is generally configured with sufficient bandwidth to securely protect traffic from the office / branch location through the security provider's firewall (e.g., a cloud-based firewall solution related to Secure Access Service Edge (SASE) security solutions, such as the SASE solution commercially available from Palo Alto Networks) for each application service (e.g., Software as a Service (SaaS) and / or Private Applications (Apps)). Generally, there are two fronts to ensure that each office / branch location has sufficient bandwidth to serve users within the branch: (1) overlay capacity (e.g., how much encrypted traffic is served by the firewall) and (2) underlay capacity (e.g., how much physical circuit capacity for physical internet connectivity exists from each office / branch location). Thus, to better provision the user experience, customers typically utilize long-term forecasts of site-specific bandwidth usage to facilitate the effective provisioning of additional firewall capacity and physical internet connectivity capacity.
[0045] In addition, SASE service providers (also known as cloud SASE service providers, for example) also typically plan and provide tunnel capacity, connecting multiple hops in the SASE fabric to provide sufficient connectivity across all tenants.
[0046] In several embodiments, machine learning (ML) predictions for SASE service provider capacity have been disclosed. For example, predicting bandwidth usage with a predetermined horizon (e.g., a one-month horizon) can utilize cost-optimized ML solutions. Traditionally, ML models can be built using existing algorithms that have personalized site-by-site or mesh-connection-by-mesh models related to the SASE fabric, learning site- or mesh-tunnel-specific behavior and using it to predict future usage. However, the number of such models can increase significantly over time, significantly increasing the cost and maintenance of training and inference for ML models.
[0047] Accordingly, the disclosed ML prediction for SASE service provider capacity is a novel and improved ML solution that enables cost-optimized predictions for all sites / circuits using a single ML model. Specifically, key features or sets of features are used to partition the data in the ML model so that the ML model learns different statistical distributions corresponding to different values of these features, without mixing them. In addition, the disclosed ML solution can learn and utilize correlations between different feature values, such as different branches, without reducing the accuracy of individual predictions, and in some cases improves accuracy by considering correlations, as will be further described below with respect to various embodiments.
[0048] Accordingly, novel and improved security solutions utilizing topological correlations are disclosed according to several embodiments.
[0049] These and other embodiments and examples of topological correlations are described further below.
[0050] Exemplary System Environments for Topological Correlation
[0051] Accordingly, in some embodiments, the disclosed technology includes providing a security platform configured to provide DPI capabilities (including, for example, stateful inspection), as further described below (for example, the security function / platform may be implemented using another (virtual) device / component that can implement security policies using the disclosed technology, such as a firewall (FW) / next-generation firewall (NGFW), a network sensor acting in place of a firewall, or PANOS running on a commercially available virtual / physical NGFW solution from Palo Alto Networks, or another security platform / NFGW, including, for example, Palo Alto Networks' PA Series next-generation firewalls, Palo Alto Networks' VM Series virtualized next-generation firewalls, and CN Series container next-generation firewalls, and / or other commercially available virtual-based or container-based firewalls, which may also be implemented and configured to perform the disclosed technology), for example, it may be provided in part or as a whole as a SASE security solution, where the cloud-based security solution (e.g., SASE) may be monitored using the disclosed technology for an application access analyzer.
[0052] Figure 1 shows an exemplary environment in which a malicious application (malware) is detected and prevented from causing harm. As will be described in more detail below, malware classification (for example, as performed by security platform 122) can be shared and / or refined in various ways among the various entities included in the environment shown in Figure 1. The techniques described herein can then be used to protect devices, such as endpoint client devices 104-110, from such malware.
[0053] As used herein, “malware” refers to an application that engages in behavior that a user would not authorize if fully informed, whether or not it is secret (and illegal). Exemplary examples of malware include ransomware, Trojans, viruses, rootkits, spyware, hacking tools, etc. One example of malware is a desktop / mobile application (e.g., ransomware) that encrypts data stored by a user. Another example of malware is C2 malware, similarly described above. Other forms of malware (e.g., keyloggers) can also be detected / blocked using the disclosed techniques for sample traffic-based self-learning malware detection, as further described herein.
[0054] The techniques described herein can be used with various platforms (e.g., servers, computing appliances, virtual / container environments, desktops, mobile devices, game platforms, embedded systems, etc.) and / or for the automated detection of various forms of malware (e.g., novel and / or variants of malware, such as C2 malware). In the exemplary environment shown in Figure 1, client devices 104-108 are a laptop computer, a desktop computer, and a tablet (each) located within the corporate network 140. Client device 110 is a laptop computer located outside the corporate network 140.
[0055] The data appliance 102 is configured to enforce policies regarding communication between client devices, such as client devices 104 and 106, and nodes outside the corporate network 140 (e.g., reachable via the external network 118). Exemplary such policies include those managing traffic shaping, quality of service, and traffic routing. Other exemplary policies include security policies requiring scanning for threats in incoming (and / or outgoing) email attachments, website content, files exchanged via instant messaging programs, and / or other file transfers. In some embodiments, the data appliance 102 is also configured to enforce policies regarding traffic remaining within the corporate network 140.
[0056] One embodiment of the data appliance is shown in Figure 2A. The example shown is a representation of the physical components included in the data appliance 102 in various embodiments. Specifically, the data appliance 102 includes a high-performance multi-core central processing unit (CPU) 202 and random access memory (RAM) 204. The data appliance 102 also includes storage 210 (such as one or more hard disks or solid-state units). In various embodiments, the data appliance 102 stores information (in either RAM 204, storage 210, and / or other appropriate location) used to monitor the enterprise network 140 and implement the disclosed technology. Examples of such information include application identifiers, content identifiers, user identifiers, requested URLs, IP address mappings, policy and other configuration information, signatures, hostname / URL classification information, malware profiles, and machine learning (ML) models (e.g., for sample traffic-based self-learning malware detection). The data appliance 102 may also include one or more optional hardware accelerators. For example, the data appliance 102 may include a cryptographic engine 206 configured to perform encryption and decryption operations, and one or more field-programmable gate arrays 208 configured to perform matching, act as a network processor, and / or perform other tasks.
[0057] The functionality described herein as being performed by the data appliance 102 can be provided / implemented in a variety of ways. For example, the data appliance 102 may be a dedicated device or set of devices. The functionality provided by the data appliance 102 may also be integrated or run as software on a general-purpose computer, computer server, gateway, and / or network / routing device. In some embodiments, at least some of the services described herein as being provided by the data appliance 102 are instead (or in addition to) provided to a client device (e.g., client device 104 or client device 110) by software running on the client device.
[0058] Whenever the data appliance 102 is described as performing a task, a single component, a subset of components, or all components of the data appliance 102 may work together to perform the task. Similarly, whenever a component of the data appliance 102 is described as performing a task, a subcomponent may perform the task, and / or a component may perform the task together with other components. In various embodiments, parts of the data appliance 102 are provided by one or more third parties. Depending on factors such as the amount of computing resources available to the data appliance 102, various logical components and / or features of the data appliance 102 may be omitted, and the techniques described herein will be adapted accordingly. Similarly, additional logical components / features may be included in embodiments of the data appliance 102 as applicable. One example of a component included in the data appliance 102 in various embodiments is an application identification engine configured to identify applications (for example, using various application signatures to identify an application based on packet flow analysis). For example, the application identification engine can determine the type of traffic a session is involved in, such as web browsing-social networking, web browsing-news, SSH, etc.
[0059] Figure 2B is a functional diagram of the logical components of one embodiment of the data appliance. The examples shown represent the logical components that may be included in the data appliance 102 in various embodiments. Unless otherwise specified, the various logical components of the data appliance 102 can generally be implemented in various ways, including a set of one or more scripts (e.g., written in Java, Python, etc., where applicable).
[0060] As shown in the diagram, the data appliance 102 includes a firewall and contains a management plane 232 and a data plane 234. The management plane is responsible for managing user interaction by providing a user interface for setting policies and displaying log data. The data plane is responsible for data management by performing packet processing and session processing.
[0061] The network processor 236 is configured to receive packets from client devices, such as client device 108, and provide them to the data plane 234 for processing. Whenever the flow module 238 identifies a packet as part of a new session, it generates a new session flow. Subsequent packets are identified as belonging to the session based on the flow lookup. SSL decryption is applied by the SSL decryption engine 240 where applicable; otherwise, processing by the SSL decryption engine 240 is omitted. The decryption engine 240 helps the data appliance 102 inspect and control SSL / TLS and SSH encrypted traffic, and therefore helps stop threats that might otherwise remain hidden within encrypted traffic. The decryption engine 240 can also help prevent sensitive content from leaving the corporate network 140. Decryption can be selectively controlled (e.g., enabled or disabled) based on parameters such as URL category, traffic source, traffic destination, user, user group, and port. In addition to the decryption policy (for example, specifying the session to decrypt), a decryption profile can be assigned to control various options for the session controlled by the policy. For example, the use of a specific cipher suite and encryption protocol version may be required.
[0062] The Application Identification (APP-ID) engine 242 is configured to determine the type of traffic a session is involved in. For example, the application identification engine 242 can recognize a GET request in incoming data and conclude that the session requires an HTTP decoder. In some cases, such as a web browsing session, the identified application can change, and such changes are noted by the data appliance 102. For example, a user might first browse a company wiki (classified as "Web Browsing - Productivity" based on the visited URL) and then browse a social networking site (classified as "Web Browsing - Social Networking" based on the visited URL). Different types of protocols have corresponding decoders.
[0063] Based on the decision made by the application identification engine 242, the packet is sent to the appropriate decoder, which is configured by the threat engine 244 to assemble the packet (which may be received out of order) into the correct order, perform tokenization, and extract information. The threat engine 244 also performs signature matching to determine what should happen to the packet. If necessary, the SSL cryptography engine 246 can re-encrypt the decrypted data. The packet is then forwarded using the forwarding module 248 for forwarding (e.g., to the destination).
[0064] Furthermore, as shown in Figure 2B, policy 252 is also received and stored in the management plane 232. A policy may include one or more rules, which can be specified using a domain name and / or host / server name, and the rules may apply one or more signatures or other matching criteria or heuristic methods, such as for security policy enforcement on subscriber / IP flows, based on various extracted parameters / information from the monitored session traffic flow. An exemplary policy may include a C2 malware detection policy that uses disclosed techniques for sample traffic-based self-learning malware detection. An interface (I / F) communicator 250 is provided for management communications (e.g., via (REST) API, messages, or network protocol communications, or other communication mechanisms).
[0065] Security platform
[0066] Returning to Figure 1, suppose a malicious individual (using System 120) has created malware 130 for a malicious web campaign (for example, the malware may be delivered to a user's endpoint device via a compromised website when the user visits / browses the compromised website, or via a phishing attack, etc.). The malicious individual wants a client device, such as client device 104, to run a copy of malware 130, unpack the malware executable / payload, compromise the client device, and, for example, become a bot in a botnet. The compromised client device may then be instructed to perform tasks (for example, cryptocurrency mining or participating in a denial of service attack) and to report information to an external entity, such as a command and control (C2 / C&C) server 150, and, where applicable, to receive instructions from the C2 server 150.
[0067] Assume that data appliance 102 intercepts an email sent (for example, by system 120) to a user "Alice" operating client device 104. In this example, Alice receives the email and clicks a link to a phishing / compromised site that may result in an attempt by Alice's client device 104 to download malware 130. However, in this example, data appliance 102 can perform the disclosed techniques for sample traffic-based self-learning malware detection and block access to the packed malware content from Alice's client device 104, thereby preempting and preventing any such download of malware 130 to Alice's client device 104. As further described below, data appliance 102 performs the disclosed techniques for sample traffic-based self-learning malware detection, as further described below, to detect such malware 130 and block it from harming Alice's client device 104.
[0068] In various embodiments, the data appliance 102 is configured to work in cooperation with the security platform 122. As one example, the security platform 122 may provide the data appliance 102 with a set of signatures for known malicious files (e.g., as part of a subscription). If the signature for malware 130 is included in the set of signatures (e.g., the MD5 hash of malware 130), the data appliance 102 may accordingly prevent the transmission of malware 130 to the client device 104 (e.g., by detecting that the MD5 hash of an email attachment sent to the client device 104 matches the MD5 hash of malware 130). The security platform 122 may also provide the data appliance 102 with a list of known malicious domains and / or IP addresses, enabling the data appliance 102 to block traffic between the corporate network 140 and the C2 server 150 (e.g., if the C2 server 150 is known to be malicious). A list of malicious domains (and / or IP addresses) can also help the data appliance 102 determine when one of its nodes was compromised. For example, if a client device 104 attempts to contact a C&C server 150, such an attempt is a strong indicator that client 104 is compromised by malware (and accordingly, remedial action should be taken, such as preventing client device 104 from communicating with other nodes in the corporate network 140).
[0069] As will be described in more detail below, the security platform 122 may also receive a copy of the malware 130 from the data appliance 102 to perform cloud-based security analysis to perform sample traffic-based self-learning malware detection, and the malware determination may be sent back to the data appliance 102 to enforce security policies, thereby safeguarding Alice's client device 104 from the execution of the malware 130 (for example, blocking access to the malware 130 on the client device 104).
[0070] In various embodiments, if the signature of an attachment cannot be found, the data appliance 102 may take various actions. As a first example, the data appliance 102 can fail-safe by blocking the transmission of any attachments that are not allowed-listed as benign (e.g., those that do not match the signatures of known good files). The drawback of this approach is that many legitimate attachments, which are actually benign, are unnecessarily blocked as potential malware. As a second example, the data appliance 102 can fail-danger by allowing the transmission of any attachments that are not blocked-listed as malicious (e.g., those that do not match the signatures of known bad files). The drawback of this approach is that newly created malware (that was not previously detected by platform 122) is not prevented from causing harm. As a third example, the data appliance 102 may be configured to provide a file (e.g., malware 130) to the security platform 122 for static / dynamic analysis to determine whether it is malicious and / or, if not, to classify it.
[0071] The security platform 122 stores a copy of the received sample in storage 142, and analysis is initiated (or scheduled, if applicable). One example of storage 142 is an Apache Hadoop Cluster (HDFS). The results of the analysis (and additional information about the application) are stored in database 146. If the application is determined to be malicious, the data appliance may be configured to automatically block file downloads based on the analysis results. Furthermore, a signature for the malware is generated and distributed (to data appliances such as data appliances 102, 136, and 148, for example) to automatically block future file transfer requests for downloading files determined to be malicious.
[0072] In various embodiments, the security platform 122 includes one or more dedicated commercial hardware servers running a typical server-class operating system (e.g., Linux®) (e.g., having a multi-core processor, 32G+ RAM, a Gigabit network interface adapter, and a hard drive). The security platform 122 may be implemented across a scalable infrastructure including multiple such servers, solid-state drives, and / or other applicable high-performance hardware. The security platform 122 may comprise several distributed components, including components provided by one or more third parties. For example, some or all of the security platform 122 may be implemented using Amazon Elastic Compute Cloud (EC2) and / or Amazon Simple Storage Service (S3). Furthermore, as with the data appliance 102, whenever the security platform 122 is referred to as performing a task such as storing or processing data, it should be understood that one or more subcomponents of the security platform 122 (individually or in cooperation with third-party components) may cooperate to perform that task. As one example, the security platform 122 can optionally collaborate with one or more virtual machine (VM) servers, such as VM server 124, to perform static / dynamic analysis.
[0073] One exemplary virtual machine server is a physical machine containing commercially available server-class hardware (e.g., a multi-core processor, 32+ gigabytes of RAM, and one or more gigabit network interface adapters) running commercially available virtualization software. Examples include VMware ESXi, Citrix XenServer, or Microsoft Hyper-V. In some embodiments, the virtual machine server is omitted. Furthermore, the virtual machine server may be under the control of the same entity managing the security platform 122, or it may be provided by a third party. As an example, the virtual machine server may rely on EC2, and the rest of the security platform 122 is provided by dedicated hardware owned and under the control of the operator of the security platform 122. The VM server 124 is configured to provide one or more virtual machines 126-128 to emulate client devices. The virtual machines can run various operating systems and / or versions thereof. Observed behavior resulting from running applications on the virtual machines is logged and analyzed (e.g., for indications that the application is malicious). In some embodiments, log analysis is performed by a VM server (e.g., VM server 124). In other embodiments, the analysis is performed, at least partially, by other components of the security platform 122, such as the coordinator 144.
[0074] In various embodiments, the security platform 122 makes the results of the sample analysis available to the data appliance 102 as part of a subscription, via a list of signatures (and / or other identifiers). For example, the security platform 122 may periodically (e.g., daily, hourly, or at some other interval and / or based on events configured by one or more policies) send content packages that identify malware files, including network traffic-based heuristic IPS malware detection, etc. The subscription may simply cover the analysis of files intercepted by the data appliance 102 and sent to the security platform 122 by the data appliance 102, and may also cover the signatures of malware known to the security platform 122.
[0075] In various embodiments, the security platform 122 is configured to provide security services to various entities in addition to (or, where applicable, instead of) the operators of the data appliance 102. For example, other companies, each having its own corporate networks 114 and 116 and its own data appliances 136 and 148, respectively, can contract with the operator of the security platform 122. Other types of entities can also utilize the services of the security platform 122. For example, an Internet service provider (ISP) providing internet services to a client device 110 can contract with the security platform 122 to analyze applications that the client device 110 attempts to download. As another example, the owner of the client device 110 can install software on the client device 110 that communicates with the security platform 122 (e.g., receiving a content package from the security platform 122, using the received content package to check attachments according to the techniques described herein, and then sending the application to the security platform 122 for analysis).
[0076] Sample analysis using static / dynamic analysis
[0077] Figure 3 shows an exemplary logical component that may be included in a system for analyzing samples. The analysis system 300 can be implemented using a single device. For example, the functionality of the analysis system 300 may be implemented in a malware analysis module 112 integrated into a data appliance 102. The analysis system 300 can also be implemented collectively across multiple separate devices. For example, the functionality of the analysis system 300 may be provided by a security platform 122.
[0078] In various embodiments, the analysis system 300 utilizes a list, database, or other collection (collectively shown as collection 314 in Figure 3) of known secure content and / or known malicious content. Collection 314 can be obtained in various ways, including via a subscription service (e.g., provided by a third party) and / or as a result of other processing (e.g., performed by data appliance 102 and / or security platform 122). Examples of information contained in collection 314 include URLs, domain names, and / or IP addresses of known malicious servers, URLs, domain names, and / or IP addresses of known secure servers, URLs, domain names, and / or IP addresses of known command and control (C2 / C&C) domains, signatures, hashes, and / or other identifiers of known malicious applications, signatures, hashes, and / or other identifiers of known secure applications, signatures, hashes, and / or other identifiers of known malicious files (e.g., OS exploit files), signatures, hashes, and / or other identifiers of known secure libraries, and signatures, hashes, and / or other identifiers of known malicious libraries.
[0079] In various embodiments, when a new sample is received for analysis (for example, when no existing signature associated with the sample exists in the analysis system 300), it is added to the queue 302. As shown in Figure 3, application 130 is received by system 300 and added to the queue 302.
[0080] The coordinator 304 monitors the queue 302 and, when a resource (e.g., a static analysis worker) becomes available, the coordinator 304 fetches a sample from the queue 302 for processing (e.g., fetching a copy of malware 130). In particular, the coordinator first provides the sample to the static analysis engine 306 for static analysis (305). In some embodiments, one or more static analysis engines are included within the analysis system 300, where the analysis system 300 is a single device. In other embodiments, static analysis is performed by a separate static analysis server containing multiple workers (i.e., multiple instances of the static analysis engine 306).
[0081] The static analysis engine obtains general information about the sample and includes it (along with heuristic information and other information, where applicable) in the static analysis report 308. The report may be generated by the static analysis engine or by a coordinator 304 which may be configured to receive information from the static analysis engine 306 (or by another appropriate component). As one example, static analysis of malware may include performing signature-based analysis. In some embodiments, the collected information is stored in a database record about the sample (e.g., in database 316) instead of, or in addition to, a separate static analysis report 308 being generated (i.e., a portion of the database record forms report 308). In some embodiments, the static analysis engine also forms a decision about the application (e.g., “safe”, “suspicious”, or “malicious”). For example, if an application contains even one "malicious" static feature (e.g., the application contains a hard link to a known malicious domain), the determination may be "malicious." For example, points may be assigned to each feature (e.g., based on severity if found, based on how reliable the feature is in predicting malice, etc.). Then, based on the number of points associated with the static analysis results, the static analysis engine 306 (or, if applicable, the coordinator 304) may assign a determination.
[0082] Once the static analysis is complete, the coordinator 304 finds an available dynamic analysis engine 310 to perform dynamic analysis on the application. Similar to the static analysis engine 306, the analysis system 300 may directly include one or more dynamic analysis engines. In other embodiments, dynamic analysis is performed by a separate dynamic analysis server, which includes multiple workers (i.e., multiple instances of the dynamic analysis engine 310).
[0083] Each dynamic analysis worker manages virtual machine instances (e.g., emulation / sandbox analysis of samples for malware detection, such as the C2 malware detection described above based on monitored network traffic activity). In some embodiments, the results of static analysis (e.g., performed by static analysis engine 306) are provided as input to dynamic analysis engine 310, whether in report format (308) and / or stored in database 316 or otherwise. For example, static report information may be used to help select / customize virtual machine instances used by dynamic analysis engine 310 (e.g., Microsoft Windows 7 SP2 vs. Microsoft Windows 10 Enterprise, or iOS 11.0 vs. iOS 12.0). If multiple virtual machine instances are running concurrently, a single dynamic analysis engine may manage all instances, or multiple dynamic analysis engines may be used where applicable (e.g., each managing its own virtual machine instances). During the dynamic part of the analysis, as will be described in more detail below, actions performed by the application (including network activity) are analyzed.
[0084] In various embodiments, static analysis of a sample is omitted or performed by a separate entity, where applicable. As one example, conventional static and / or dynamic analysis may be performed on a file by a first entity. Once a given file is determined to be malicious (e.g., by the first entity), the file may be provided to a second entity (e.g., an operator of security platform 122) for additional analysis of the use of malware in network activity (e.g., by the dynamic analysis engine 310).
[0085] The environment used by the analysis system 300 is instrumented / hooked so that any behavior observed while the application is running is logged as it occurs (e.g., using a customized kernel that supports hooking and logcat). Network traffic associated with the emulator is also captured (e.g., using pcap). Log / network data may be stored as temporary files on the analysis system 300. It may also be stored more permanently (e.g., using HDFS, or another suitable storage technology, or a combination of technologies such as MongoDB). A dynamic analysis engine (or another suitable component) can compare connections made by the sample to a list (314) such as domains, IP addresses, etc., and determine whether the sample communicated with (or attempted to communicate with) a malicious entity.
[0086] Similar to the static analysis engine, the dynamic analysis engine stores the results of its analysis in a record associated with the application being tested in the database 316 (and / or, if applicable, includes the results in report 312). In some embodiments, the dynamic analysis engine also forms a decision about the application (e.g., “safe,” “suspicious,” or “malicious”). For example, the determination may be “malicious” even if the application has performed only one “malicious” action (e.g., an attempt to contact a known malicious domain was made, or an attempt to extract sensitive information was observed). For another example, points may be assigned to an action performed (e.g., based on severity if found, based on how reliable the action is in predicting malice, etc.). The determination may then be assigned by the dynamic analysis engine 310 (or, if applicable, the coordinator 304) based on the number of points associated with the dynamic analysis results. In some embodiments, the final determination associated with the sample is made (e.g., by the coordinator 304) based on a combination of reports 308 and 312.
[0087] AIOP for SASE Architecture to Facilitate Application Access Visibility
[0088] Figure 4A is an overview of AIOP for Secure Access Service Edge (SASE) solutions according to several embodiments. As shown in Figure 4A, the disclosed AIOP for a SASE architecture can facilitate application access visibility, including using a single plane of glass (e.g., providing efficient visibility across network and security solutions), root cause analysis (e.g., actionable insights), and noise reduction (e.g., isolating issues from symptoms, such as accessibility to SaaS / private applications), as shown in 402 (for example, in this example, the SASE solution is the Prisma Access SASE solution, commercially available from Palo Alto Networks, headquartered in Santa Clara, California, and / or the disclosed techniques may be similarly provided for other SASE environments and commercially available solutions).
[0089] As indicated in 404, the disclosed AIOP for the SASE architecture includes various technical components, as further described below, including baselineization, correlation, prediction, formal analysis (e.g., access permission / security policies, etc.), natural language processing (NLP) for queries via the user interface (UI), and computational playbooks.
[0090] As shown in 406, the disclosed AIOP for the SASE architecture includes various data sources, including configuration information, status information, logs (e.g., firewall logs, etc.), telemetry data, and synthetic data (from various network and security solution sources, such as the GlobalProtect (GP) VPN solution commercially available from Palo Alto Networks, headquartered in Santa Clara, California, and other network and security solutions including CDSS, ADEM, SD-WAN, FLN, CIE, and PAI, as shown in Figure 4A).
[0091] Figure 4B illustrates an exemplary use case for a mobile user unable to connect to the corporate network due to authentication issues. For example, multiple computing components / entities, and network connectivity between these different computing components / entities, generally make it technically difficult for customers (e.g., customer network operations centers (NOCs) and / or IT / help desk personnel) to determine the root cause of any application connectivity problem. Specifically, the techniques disclosed for application access analyzers provide customers / customer NOCs with automated tools to analyze and detect potential SASE / Prismatic access problems for users / groups of users to access one or more applications (e.g., SaaS / private apps), as will be further described below with respect to various embodiments.
[0092] Referring to Figure 4B, this exemplary use case illustrates the challenge of manually correlating discrete events where a mobile user was unable to connect to the corporate network due to authentication issues, as shown in 420. Specifically, the manual analysis and correlation required two weeks of troubleshooting before finally diagnosing the problem using the customer's authentication server.
[0093] Figure 4C illustrates one exemplary use case for reducing the mean time to discovery (MTTD) / mean time to resolution (MTTR) for access to SaaS / private applications using AIOP for Secure Access Service Edge (SASE), according to several embodiments. Specifically, in this exemplary use case, excessive authentication (Auth) timeout failures at all SASE (e.g., Prismatic Access (PA)) locations were resolved in minutes (instead of hours or days) using the disclosed AIOP for SASE, as further described below.
[0094] Specifically, the AIOP disclosed for SASE addresses a variety of technical issues, as will be explained below. The MTTD and MTTR of application (App) access problems are typically several hours, which can increase application downtime and negatively impact a company's productivity and revenue. Troubleshooting and debugging generally require expertise in domain knowledge. Furthermore, correlating and tracing multiple factors to perform a root cause analysis (RCA) is often cumbersome and error-prone when performed manually.
[0095] For example, enterprises and / or cloud service providers with multiple hosted network services, large network infrastructures, and complex security policy configurations may encounter significant challenges in reducing the MTTR of application access issues.
[0096] As another example, identifying RCA within a corporate organization may generally require a comprehensive check across various domains, such as network connectivity, infrastructure reachability, and security policy reasoning.
[0097] Referring to Figure 4C, excessive authentication failures are detected, as shown in 430. In 432, the dynamic baseline automatically detects an abnormal drop in the mobile user count, and automated monitoring of the authentication server using the cloud probe VM determines that the authentication server is available and that the network path to the authentication server is available (e.g., up and running without performance degradation or other reachability issues) using the disclosed AIOP for SASE. In 434, ML correlation is performed using the disclosed AIOP for SASE to generate a single resolvable incident, isolating the cause as an unresponsive authentication service impacting 1200 mobile users in the US West region who are unable to connect, thereby directing the resolution to the authentication server. In 436, once the authentication timeout failure and the mobile user count at all SASE PA locations are verified to be within the normal range, the incident is resolved using the disclosed AIOP for SASE. Thus, the disclosed technology, which utilizes topological correlations, facilitates improved MTTD / MTTR for access to SaaS / private applications using AIOP for SASE, according to several embodiments.
[0098] Figure 5A illustrates AIOPs for a SASE architecture to facilitate topological correlation according to several embodiments. As shown in 502, the Prisma Access (PA) AIOps platform is designed to provide proactive service level management globally across customers and for use by NOC personnel supporting Prisma Access (PA) / SASE customers. Specifically, the diagram shown in Figure 5A illustrates a high-level architecture of the AIOPs services and their interactions, including the data processing pipelines and applications (GKE) deployed as part of the Prisma Access Insights (PAI) platform, as will be further described below.
[0099] Prisma AIOP provides proactive monitoring, alerting, problem isolation, and playbook-driven improvements to deliver SLAs (MTTK / I, MTTR) as required by customers. Prisma AIOP is built on an insight platform (e.g., called the Prisma Access Insight (PAI) platform) that provides cross-domain NOC capabilities (e.g., including Prisma SD-WAN (CGX), etc.). The PAI platform includes a publish / subscribe communication mechanism 504, as shown in Figure 5A, which facilitates communication between various components, including: The system includes an Extract Transform Load (ETL) 506 for normalizing data, an alert engine 508, a resource graph engine 510, a correlation and PRC engine (CPE) 512, a notification engine 514, and an Incident Execution Engine (IEE) 516 (for example, the IEE component can provide the following functions: (1) policy management for incidents, such as suppression policies and role management; (2) role-based access control (RBAC) for incident APIs; and (3) APIs for dashboards, single customers, and aggregates across customers; and the IEE component can also provide impact analysis, such as from CDL traffic logs to understand the number of users affected or dropped; and when an incident is created, based on customer-impacted objects (alerts) within an incident, the IEE can fetch corresponding impacts based on keys from the impact structure).The data sources include events 518, telemetry data 520 (for example, telemetry data may be collected from various sources such as firewall, AWS, GCP health data, application data, and operational data including cloud infrastructure monitoring data such as Cortex Data Lake (CDL) data, including PANOS system logs for system log data; cross-domain telemetry data may also be made available for cross-correlation), and alerts and metrics 522 for other domains.
[0100] Specifically, Figure 5A shows the high-level architecture of the AIOs services and their interactions. The engines / services described above are part of the data processing pipeline and applications (GKE) deployed as part of the PAI platform.
[0101] Figure 5B shows an alert engine according to several embodiments. As shown in 508, the alert engine includes subcomponents for alert baselineization 530, alert evaluation 532, ML detection 534, and alert information enhancement 536. ML detection is trained using an AutoTSML training component 538 that communicates with a data store shown as Google BigQuery (GBQ) / Google Cloud Store (GCS) 540.
[0102] In exemplary implementations, alerts can be reactive and predictive. An example of a reactive alert includes CPU thresholds that exceed certain anomalies. A predictive alert is a metric alert that provides a future time or event in which such behavior may occur. An example of a predictive alert includes a capacity planning alert in link utilization that provides information about when a link may become saturated. Alerts can also be threshold-based (e.g., configurable / programmable thresholds), simple baselines (e.g., averages), and can be compared to SD or ML time (AutoTSML) series-based alerts. As also shown in Figure 5B, the alert engine includes ML-based detection, which can be trained using its own infrastructure for training, model validation, deployment, etc.
[0103] For example, the disclosed alert engine service can generate alerts based on any of the telemetry data available within the system via the telemetry data source. The alert engine also includes an API, as shown in 542, which can be used to create predefined models (e.g., templates such as SD or ML). One such example is using the Influx TICK language, ProMQL, or Google MQL to pick fields and evaluation expressions. The following three templates are exemplary minimum requirements: (1) general threshold-based, (2) SMA / EMA and SD-based, and (3) MAD. All of these exemplary templates are stateful (e.g., they track alert states and state transitions). Furthermore, such alerts can persist and propagate to other services via publish / subscribe for consumption (PubSub), as shown in 504 of Figure 5B.
[0104] The following is an example implementation of an alerting output schema. [Table 1]
[0105] Figure 5C shows one exemplary relationship between various entities that are automatically generated and used for correlation and PRC according to several embodiments. Specifically, the exemplary relationship is a graph like the one shown in 550, which can be generated using a resource graph engine like the one shown in 510 of Figure 5A.
[0106] In an exemplary implementation, the Resource Graph Engine 550 is implemented as a Resource Builder Graph (RBG) that provides a relational model of a complete end-to-end (E2E) computing environment from any given user to each application (e.g., SaaS / private app). The resource graph is a hierarchical (e.g., nodes in the graph may be subgraphs with relationships) general knowledge map generated for use in correlation / causal analysis, such as to perform topological correlation techniques as described herein. At the highest level of the resource graph, a cross-domain network model is used to connect users to applications through various domains, such as SASE / Prisma access, SASE / Prisma SD-WAN, etc. Within each of these domains, there are network models with various nodes based on connectivity, such as VPN / GP, remote network (RN), proxy, SD-WAN (e.g., SD-WAN fabric), etc. In this exemplary implementation, the cross-domain network model includes control, management, and data-plane relationships between entities. Within each of these domains, each node or link contains a subgraph describing relationships across components, processes, etc. One example of a node subgraph is the relationships between IKE and IPSec-related processes in a virtual machine (VM) (e.g., the pan_task processes for IKE and IPSec in a Palo Alto Networks (PAN) VM implementation). Also, when connecting multiple domains, key identifiers (e.g., app, tenant, user ID, etc.) may be normalized. For example, a network / system administrator (admin) with domain knowledge can specify relationships via a meta-language, and RGB can then construct the graph as described above using tagged alerts / events.
[0107] Figure 5D shows a correlation and primacy control engine (CPE) according to several embodiments. As shown in 512, the CPE includes subcomponents for a data store 552 (e.g., Neo4J, or another commercial or open-source graph data store may be used similarly for this subcomponent), a primacy controller 554 (e.g., for looking up site, node, link state, and dependency information), a correlation and causality subcomponent 556, an incident data store / database (DB) subcomponent 558 (e.g., for storing incident, alert, and impact information), an alert / event communication pipeline 560 (e.g., for receiving incident ID information from CPE 556 and sending alerts, parsed SYSLOG information to CPE 556), and Redis 562 (e.g., an open-source, in-memory data structure store used as a database, cache, message broker, and streaming engine, which in this exemplary implementation may be used to provide current state and configuration information to CPE 556, as shown in Figure 5D).
[0108] In this exemplary implementation, correlation and the PRC engine (CPE) (shown, for example, as CPE512 in Figure 5A and as CPE556 in Figure 5D) facilitate correlation and deduplication. Correlation and deduplication can be used in a variety of situations. As one example, consider a scenario in which a VM (e.g., a remote network associated with a remote service provider network (RN-SPN)) goes down, and then a burst of events indicating that the tunnel is down is received from all other nodes connected to this VM. Whether such events result from planned downtime or an unplanned power outage, it is generally desirable to aggregate these related events / alerts into a single incident to facilitate more effective and efficient identification of the root cause and (if necessary) remediation / remediation. In the exemplary scenario above, the possible root cause is the VM down event. Thus, as used herein, correlation generally refers to the process of aggregating and / or clustering events / alerts based on temporal or relational boundaries to facilitate the topological correlation techniques disclosed. As used herein, deduplication generally refers to the process of identifying duplicate events and then combining those duplicate events into alerts (e.g., single / integrated alerts). Deduplication thereby reduces the number of alerts in the system (e.g., reducing noise for network / system / IT administrators / users).
[0109] Generally, there are two high-level approaches to performing the correlation and deduplication processes described above: (1) rule-based and (2) AI / ML-based. The rule-based approach is one in which system / domain experts define root causes, typically in code or using a metalanguage, in the form of an IF This Then That (IFTTT) statement. Rule-based engines can also use relationship graphs to walk incoming alerts against parts of a DAG analysis and correlate them. While this rule-based approach may work initially or for a subset of cases, over time, such rule-based approaches can become obsolete as new relationships or services are added to the infrastructure.
[0110] In contrast, ML / AI-based techniques use the relationships constructed using the aforementioned RGB as input, and the time range in which the events / alerts occurred. Clustering techniques for proximity-based, time-based, or centrality-based algorithms yield efficient and effective results. The exemplary output of the aforementioned CPE (e.g., shown as CPE512 in Figure 5A) is an incident with a collection of alerts / events and PRC instructions, which are made available to other services (e.g., other services of interest that can subscribe to such alerts / events).
[0111] Typically, correlation views time and relationships as dependencies. For example, a tunnel crosses an interface, and the interface goes down. Most systems also view dependencies and topological relationships as static distances. In contrast, in the disclosed topological correlation technique, the disclosed model views topology as a dynamic state derived through hop-by-hop metrics and events that can occur at any given hop. Such events are then correlated to isolate symptoms from problems and enable rapid root cause analysis, as will be further described herein with respect to various embodiments.
[0112] Figure 5E shows one exemplary data model for an AIOP platform according to several embodiments. Specifically, an exemplary data model for an AIOP platform is shown in 564 of Figure 5E.
[0113] The following is an example implementation of a correlation graph schema. [Table 2] TIFF2026515746000004.tif157170
[0114] The following is an example implementation of a correlated data structure. [Table 3] TIFF2026515746000006.tif153170 TIFF2026515746000007.tif156170 TIFF2026515746000008.tif158170 TIFF2026515746000009.tif158170 TIFF2026515746000010.tif169170 TIFF2026515746000011.tif122167
[0115] Figure 6 shows an ML pipeline architecture for an AIOP for a SASE solution to facilitate topological correlation, according to several embodiments. In one exemplary implementation, the disclosed AIOP for a SASE architecture performing topological correlation utilizes AI / ML techniques to perform a variety of operations, including: (1) anomaly detection (e.g., detecting tunnel degradation behavior such as round-trip time (RTT) and packet drop errors); (2) prediction (e.g., predicting memory leaks, file descriptor usage, and disk usage); and (3) causal modeling (e.g., automated and enhanced RCA for App access issues and / or other networking / performance-related issues in a SASE computing environment, as described herein with respect to various embodiments).
[0116] Referring to Figure 6, one exemplary ML pipeline architecture for AIOP for a SASE solution to facilitate topological correlation is shown in 602. In this exemplary implementation, the ML pipeline architecture can utilize open-source and / or managed service models for running the ML pipeline (e.g., Google's managed AI / ML pipeline infrastructure running KFlow for both training and service delivery purposes). Various AI / ML algorithm libraries can be used for deploying the model and inference server. The exemplary AI / ML libraries include a configurable version relating to an automated TSML library for production deployment anomaly detection for Prisma AIOP to provide near real-time inference.
[0117] Generally, as similarly described above, network administrators typically either over-provision their networks due to a lack of a data-driven approach, or only initiate capacity planning exercises after reaching or exceeding the current configured capacity limits, which results in an inadequate application experience for end users. Therefore, the disclosed AIOP solutions include providing predictive information according to several embodiments.
[0118] Exemplary use cases for prediction
[0119] Various use cases for prediction will be explained below.
[0120] As a first exemplary use case, historical bandwidth usage data can be used to train machine learning models so that capacity usage can be predicted / forecasted.
[0121] As a second exemplary use case, by forecasting information / data, network administrators can perform data-driven capacity planning more effectively and efficiently (e.g., forecasting expected bandwidth needs for a SASE customer over N months, such as the amount of bandwidth that will be required for the initial deployment of a SASE customer, based on user input including the number of users and the number of applications). For example, the forecast could be applied to predict that a given service connection (SC) will exceed its capacity in N days (e.g., 30 days), and as a result, if the capacity is not sufficiently increased in N days, users will potentially experience degradation in their private / SaaS application experience.
[0122] As a third exemplary use case, if the predicted bandwidth exceeds a user-defined threshold, the prediction can be used to automatically alert the network administrator.
[0123] As a fourth exemplary use case, forecasts could be used by a SASE provider to proactively reach customers using forecast data and proactively increase bandwidth based on forecasts of network capacity for those customers (for example, what the bandwidth usage or requirements will be when adding more than N users to a branch site, or what the expected additional bandwidth will be when enabling one or more applications such as Microsoft Teams or Zoom).
[0124] As a fifth exemplary use case, forecasts may be used to provide remediation playbooks to assist with capacity planning (for example, for SASE customers).
[0125] As a sixth exemplary use case, prediction can be applied to predict incidents (e.g., incident prediction in a given SASE deployment). For example, predictive / forecasting analysis may be applied to predict excessive authentication timeout failures caused by problems with the customer's authentication server.
[0126] As a seventh exemplary use case, predictive / forecast assessment may be used to provide what-if analysis to determine whether a user's access to a private / SaaS app is a result of a recent configuration change (for example, whether a user's access to Microsoft Outlook is impacted by a recent Active Directory (AD) group change). As another example, a customer IT administrator may use the AIOP platform's natural language (NL) interface to submit NL queries to perform multi-domain analysis for a new private / SaaS app rollout across all user locations before a given user complains about access issues.
[0127] As an eighth exemplary use case, predictive / forecast assessment can be used for enhanced security analysis (e.g., using formal modeling of a company's security policies). For example, a security administrator could verify whether a new policy would allow an engineering user group to access a private / SaaS application (e.g., Salesforce), and whether a security policy analyzer would indicate duplicate / shadow rules for the user group's access to that application. As another example, a network security (NetSec) administrator could perform a pre-change sandbox analysis to determine whether adding a new policy (e.g., to allow an HVAC system to access a patch server) would result in a compliance violation.
[0128] Use case scenarios for topological correlations
[0129] Various use cases for providing topological correlations, as well as related novel and improved techniques disclosed, will be described below.
[0130] As a first exemplary use case scenario, the disclosed topological correlation technique may be applied to facilitate effective and efficient root cause and / or other analysis of user and / or group user access to SaaS applications from a mobile user gateway.
[0131] As a second exemplary use case scenario, the disclosed topological correlation technique may be applied to effectively and efficiently facilitate root cause and / or other analysis of user and / or group user access to private applications hosted in on-premises data centers or remote branch offices.
[0132] As a third exemplary use case scenario, the disclosed topological correlation technique may be applied to facilitate effective and efficient root cause and / or other analysis of user and / or group user access to remote site connectivity to remote branches or data centers.
[0133] As a fourth exemplary use case scenario, the disclosed topological correlation techniques may be applied to facilitate the effective and efficient execution of multi-domain contextual troubleshooting analysis with actionable outcomes. For example, multi-domain contextual troubleshooting analysis may include various other domains that can be automatically analyzed using the disclosed techniques to facilitate actionable decisions (e.g., yes / no assessments, correlation and root cause analysis to identify problems and remediation actions, and / or deep domain-specific observability to facilitate faster remediation).
[0134] As a fifth exemplary use case scenario, the disclosed topological correlation technique can be applied to detect various anomalies. These include brownouts resulting from DDoS events / attacks, spikes in the amount of traffic on the tunnel (BW increases) resulting from a specific app or a specific set of endpoints (e.g., abnormally behaving laptops), and performance impacts on other users. Currently, customers lack unified visibility into security attacks and / or their impact on network behavior. One such example is a DDoS attack against a DNS server, which manifests as network capacity saturation, and network operators may plan to increase capacity as a solution. Using the incident framework in AIOP, we first build a baseline (based on seasonal time-series trends) of metrics such as DNS request / response rates and bandwidth usage regarding secure connectivity between Prisma Access and the customer's private data center to determine anomalous behavior. Then, our correlation logic correlates anomalous DNS traffic patterns with increases in bandwidth consumption. We generate incidents using the above context, enabling NetSec administrators to fundamentally cause network resource saturation resulting from DNS DDoS attacks. Furthermore, we can also provide additional context regarding the source addresses from which these DNS requests originate, allowing security administrators to block traffic, enable faster remediation, and add policies to return to normal operation.
[0135] As a sixth exemplary use case scenario, the disclosed topological correlation technique can be applied to detect a variety of anomalies, including authentication timeout failures, DNS failures, SaaS / private application unreachable, SC BGP down, SC tunnel down, etc. For example, alerts may be generated to alert customer IT about health and performance issues with individual objects such as tunnels, authentication servers, DNS, etc. However, this can result in burning multiple resources, which can be manually determined to be related events where one is the problem and the others are merely symptoms of the problem / root cause. For example, a data center tunnel may have performance or health issues that can result in BGP flapping, packet drops, poor private application experience, DNS failures, and / or authentication failures. Thus, the disclosed topological correlation techniques can be applied to correlate these discrete events across discrete data sources and deliver a single, addressable incident that is the problem / root cause (e.g., a data center connectivity degradation causing symptoms such as BGP flapping, packet drops, poor application experience, DNS failures, authentication failures, etc.). This reduces manual correlation and improves the customer's IT efficiency in troubleshooting and proactively remediating problems / root causes.
[0136] As a seventh exemplary use case scenario, the disclosed topological correlation techniques can be applied to detect a variety of anomalies, including CIE agent disconnections, user group count deviations from a baseline, and rule hit deviations. For example, a potentially malicious masquerading user may be detected when the same user / user ID logs into the corporate network from two geographically separated locations within an improbable time interval (e.g., a time travel violation when a user logs into two separate locations within a period in which such physical movement by that user would not be feasible). Customers can specify various security and authentication policies based on user groups. The cloud security service can then periodically learn about these user groups from the LDAP / AAA / AD server. If a disconnection exists between the security service and the LDAP / AAA / AD server, this results in expired user group information. This can lead to denial of access. As part of this proactive anomaly detection, the disclosed techniques may include monitoring and generating baselines for user groups, users per group, security rule hits for access to applications, etc. Such baseline deviations may then correlate with either a loss of connectivity to Active Directory (AD) or a decrease in group counts. Furthermore, such correlated anomalies can be proactively alerted to customer IT administrators to remediate the root cause before multiple users complain about such potential application access issues.
[0137] As an eighth exemplary use case scenario, the disclosed topological correlation technique can be applied to detect various anomalies, including correlating incidents to impact based on users and applications. For example, in existing IT support systems, IT administrators can typically attempt to determine the potential impact on their organization by examining the total number of user-related complaints / tickets generated around the same time period. Existing IT support systems are largely reactive when assessing the actual impact on the organization. The disclosed technique may include an incident framework with an AIOP solution that provides context about an incident by expanding on users, bandwidth, or application access, or a combination thereof, as the incident is created and the scope of impact is updated in near real-time. Furthermore, such topological correlation information, associated with network access / metrics, etc., can be used by network operators and IT administrators to facilitate the efficient and effective prioritization of the highest-impact incidents, and also enable help desks to resolve and close user trouble tickets more quickly.
[0138] Figure 7A shows one exemplary use case for a significant drop in authenticated user detection according to several embodiments. Figure 7A shows an exemplary use case process as shown in 702. In 704, a significant drop in the number of authenticated users is detected using the disclosed ML-driven early anomaly detection technique. In 706, problem isolation and deep correlation are performed using the disclosed topological correlation technique. Specifically, whether it is associated with a cloud infrastructure problem is determined by verifying that the MU portal and gateway nodes are healthy, and whether it is associated with an endpoint problem is determined by verifying that an endpoint VPN-related process (e.g., the GP VPN endpoint agent process) is down, which is determined to be the result of using an incompatible GP client software version, as shown in 706. In 708, automated remediation is performed using a computed playbook that facilitates an automated rollout to the endpoint with a compatible GP client software version. In 710, ML-driven verification is performed. Specifically, it is determined whether the user can be authenticated and whether the number of users is within the expected / normal range (for example, based on baseline analysis of the number of users).
[0139] Figure 7B shows one exemplary use case for anomaly detection in a service-connected cloud node, according to several embodiments. Figure 7B shows an exemplary use case process as shown in 720. In 722, the disclosed ML-driven early anomaly detection automatically detects a service-connected cloud node anomaly. Specifically, the anomaly detection is associated with the service-connected (SC) cloud node DP throughput being exceeded and a tunnel packet loss spike in DNS transactions (e.g., against a previously calculated baseline). In 724, problem isolation and deep correlation are performed using the disclosed topological correlation technique. Specifically, it is determined that the SC throughput and the DNS anomaly are correlated, and the ADEM report shows a low score for the private application, as shown in 724, indicating a potential endpoint-related problem. In 726, automated remediation is performed using a computed playbook that facilitates automated remediation, isolating the host using a Host Information Profile (HIP) to allow the customer to perform further investigation on the isolated user endpoint. In step 728, ML-powered validation is performed to verify that DNS spikes have returned to the baseline / normal range, and that the associated SC throughput has also returned to the baseline / normal range.
[0140] Various process embodiments for topological correlation techniques will be described further below.
[0141] Exemplary process for performing topological correlation techniques
[0142] Figures 8A–8C illustrate the process for performing correlation on the AIOP SASE solution to facilitate topological correlation techniques, according to several embodiments. In some embodiments, processes 802, 860, and 880, as shown in Figures 8A–8C, are performed using CPE512 and techniques similar to those described above, including the embodiments described with respect to Figures 4A–7B.
[0143] Referring to Figure 8A, process 802 is shown for aggregating alerts and / or critical events as incidents. In 804, the alert is processed. In 806, a lookup of the alert status is performed. In 808, the alert status is created or updated by raising or clearing. In 810, it is determined whether the alert is used in the correlation. If not, the process ends as shown in 812. Otherwise, the process proceeds to 814, and it is determined whether the alert was raised or cleared. If the alert was raised, the process proceeds to 816, and it is determined whether the parents have an open incident. If so, the process proceeds to 824, and a link from the alert to the incident is generated. Otherwise, the process proceeds to 818, and it is determined whether the dependents have an open incident. Otherwise, the process proceeds to 820, and the following is performed. In other words, an incident is created, its status is set to holding, an alert is linked, and the impact is updated. At 822, the hold timer is started. If it is determined that the family has an open incident, the process proceeds to 826, where it is determined whether there is a change in the RC. If not, the process proceeds to 828, where the following is performed: the RC is updated, an alert is linked, and the impact is updated. If so, the process proceeds to 830, where it is determined whether a status set to holding has been created. If so, the process returns to 828, where the following is performed: the RC is updated, an alert is linked, and the impact is updated. If not, the process proceeds to 832, where the following is performed: a new incident is created and linked. At 834, the hold timer is started. At 836, it is determined whether a cleared alert is linked to the incident. If not, the following is performed: an incident is created and its status is set to holding.Otherwise, the process proceeds to 840, where the following is performed: the alert and impact states are updated to clear. In 842, it is determined whether internal alerts have been cleared. If so, the process proceeds to 844, where a check for customer issues is performed. If it is related to a customer issue, the process terminates as being related to a customer incident, as shown in 846. Otherwise, the process proceeds to 848, where it is determined whether all mandatory alerts have been cleared. If so, the incident state is updated to clear, as shown in 850, and the process terminates.
[0144] Referring to Figure 8B, process 860 is shown for aggregating alerts and / or critical events as incidents. In 862, the hold timer expires. In 864, it is checked whether an incident exists for the raised alert or critical event. In 866, if not, the status is set to clear. Otherwise, a scan is performed to check for cleared / closed incidents with the same root cause (RC), as shown in 868. If yes, the raised alert or critical event is associated with its parent via linked incidents, as shown in 870, and the data store (incident data store, such as the incident DB shown in 558 of Figure 5D) is updated, as shown in 872. Otherwise, the process proceeds to 874, and the status is set to clear, as shown in 874. In 876, a customer issue check is performed. In 878, the data store (incident data store, such as the incident DB shown in 558 of Figure 5D) is updated.
[0145] Referring to Figure 8C, process 880 is shown for processing alerts to associate them with an incident. Process 882 performs processing of critical, parsed Syslog information. Process 884 creates an alert node with an event type. Process 886 determines whether an incident exists on the instance. If not, a pending incident is created, as shown in 888. Otherwise, an incident exists on the instance, and the alert is associated with the incident, as shown in 890.
[0146] Figure 9 is a flowchart relating to a process for performing a topological correlation technique according to several embodiments. In some embodiments, the process 900 shown in Figure 9 is performed by the same disclosed techniques as described above, including the embodiments described with respect to Figures 4A-7B.
[0147] In 902, the task is to derive hop-by-hop metrics for the network topology and multiple events associated with cloud-based security services and network access to applications for a user or group of users. For example, the network topology can be dynamically monitored as similarly described above with respect to Figures 4A–7B (for example, if the network topology is considered to be in a dynamic state, dynamic hop-by-hop can be performed to facilitate the deriving of hop-by-hop metrics and events associated with and / or occurring at any of the hops, thereby enabling effective and efficient correlation of such events to isolate symptoms from problems and facilitate rapid root cause analysis).
[0148] In 904, as described above with respect to Figures 4A-7B, the hop-by-hop metrics derived in multiple dimensions are correlated with events, facilitating the automatic determination of the root cause of problems associated with network access to applications for a user or group of users based on multiple events.
[0149] In 906, as described above with respect to Figures 4A-7B, a remediation response is performed based on multiple events and the root cause of the problem associated with network access to the application for a user or group of users.
[0150] Figure 10 is another flowchart relating to a process for performing a topological correlation technique according to several embodiments. In some embodiments, the process 1000 shown in Figure 10 is performed by the same disclosed techniques as described above, including the embodiments described with respect to Figures 4A-7B.
[0151] In 1002, access to applications over the network (for example, for a specified user or group of users) is monitored, similar to what was described above with respect to Figures 4A-7B.
[0152] In 1004, a baseline of monitored access to the application over the network is generated for each of the multiple dimensions, similar to what was described above with respect to Figures 4A-7B.
[0153] In 1006, as described above with respect to Figures 4A-8C, network metrics are automatically predicted for one or more of multiple dimensions. In one embodiment, ML capacity prediction / forecasting may be performed using the AIOP platform / solution described above to predict / forecast various network capacity and / or other access-related issues. As one example, the AIOP platform may be used to recommend adding network capacity for a given network for access to a private / SaaS application based on a prediction that a customer will exceed their current capacity in 30 days. As another example, the AIOP platform may be used to identify impersonating users based on the detection of unacceptable time travel (for example, a user logging into SASE / Prisma access from two different source locations / ISPs that are unlikely to occur within a given period, based on the geographical distance associated with two different source locations / ISPs, may thus be detected as anomalous user activity, and an appropriate mediation response may be recommended or automatically performed, such as isolating the user to prevent access to sensitive information / resources about the customer / company).
[0154] In 1008, as described above with respect to Figures 4A-7B, alerts or reports are generated using recommended configurations and / or remediation based on predictions of network metrics for one or more of the multiple dimensions.
[0155] While the embodiments described above have been explained in some detail for the purpose of clarifying understanding, the present invention is not limited to the details provided. Many alternative methods exist for carrying out the present invention. The disclosed embodiments are illustrative and not limiting.
Claims
1. A system including a processor and memory, The aforementioned processor, Derive hop-by-hop metrics for network topology and multiple events for cloud-based security services associated with network access to applications for a user or group of users, wherein the network topology is dynamically monitored. The derived hop-by-hop metric and the multiple events are correlated in multiple dimensions to facilitate the automatic determination of the root cause of the problem associated with the user's or group of users' access to the application over the network, and Based on the aforementioned multiple events, a remedial response is performed based on the root cause of the problem associated with the user's or group of users' access to the application via the network. It is configured in such a way, The aforementioned memory is Coupled with the aforementioned processor, and Give instructions to the aforementioned processor, It is structured in such a way. system.
2. The root cause of the problem associated with the user's or group of users' access to the application via the network is determined by using artificial intelligence and / or machine learning to use topological correlation to correlate multiple data sources across multiple domains. The system according to claim 1.
3. The root cause of the problem associated with the user's or group of users' access to the application via the network is determined by using artificial intelligence and / or machine learning to correlate multiple data sources across multiple domains, and The aforementioned multiple domains include network, authentication, DNS, SaaS / private application health, and security policy configuration. The system according to claim 1.
4. The root cause of the problem associated with the user's or group of users' access to the application over the network is determined to be related to an anomaly in network connectivity and / or performance degradation associated with the user's or group of users' access to the application over the network. The system according to claim 1.
5. The remediation response includes generating a human-consumable and actionable decision analysis that reduces the average time to detect and remediate application connectivity problems. The system according to claim 1.
6. Automatically determining the root cause of the problem associated with the user's access to the application over the network includes identifying network infrastructure problems, customer network service problems, client connectivity problems, SaaS / private application health problems, and / or other connectivity / reachability problems or performance degradation problems. The system according to claim 1.
7. The aforementioned processor further, To automatically discover the network topology used by the user or group of users to access the application, and to facilitate topological correlation for root cause analysis, It is structured in such a way. The system according to claim 1.
8. The aforementioned processor further, By constructing a unified logical model for calculations regarding security policies associated with corporate networks, we generate security attitude assessments. It is structured in such a way. The system according to claim 1.
9. The aforementioned processor further, Using the natural language query interface associated with the user's or group of users' access to the application via the network, the system processes user queries. It is structured in such a way. The system according to claim 1.
10. It is a method, The steps include deriving hop-by-hop metrics for a network topology and multiple events for a cloud-based security service associated with network access to an application for a user or group of users, wherein the network topology is dynamically monitored, and A step of correlating the derived hop-by-hop metric with the multiple events in multiple dimensions, thereby facilitating the automatic determination of the root cause of the problem associated with the user's or group of users' access to the application over the network, based on the multiple events. The steps include: performing a remedial response based on the multiple events, based on the root cause of the problem associated with the user's or group of users' access to the application over the network; Methods that include...
11. The root cause of the problem associated with the user's or group of users' access to the application via the network is determined by using artificial intelligence and / or machine learning to use topological correlation to correlate multiple data sources across multiple domains. The method according to claim 10.
12. The root cause of the problem associated with the user's or group of users' access to the application via the network is determined by using artificial intelligence and / or machine learning to correlate multiple data sources across multiple domains, and The aforementioned multiple domains include network, authentication, DNS, SaaS / private application health, and security policy configuration. The method according to claim 10.
13. The remediation response includes generating a human-consumable and actionable decision analysis that reduces the average time to detect and remediate application connectivity problems. The method according to claim 10.
14. Automatically determining the root cause of the problem associated with the user's access to the application over the network includes identifying network infrastructure problems, customer network service problems, client connectivity problems, SaaS / private application health problems, and / or other connectivity / reachability problems or performance degradation problems. The method according to claim 10.
15. The root cause of the problem associated with the user's or group of users' access to the application over the network is determined to be related to an anomaly in network connectivity and / or performance degradation associated with the user's or group of users' access to the application over the network. The method according to claim 10.
16. The above method further, A step of automatically discovering the network topology used by the user or group of users to access the application, and facilitating topological correlation for root cause analysis. The method according to claim 10, including the method described in claim 10.
17. The above method further, The step of generating security attitude assessments by constructing a unified logical model for calculations regarding security policies associated with corporate networks, The method according to claim 10, including the method described in claim 10.
18. The above method further, A step of processing a user query using a natural language query interface associated with the user's or group of users' access to the application via the network, The method according to claim 10, including the method described in claim 10.
19. A computer program stored on a non-temporary computer-readable storage medium, The aforementioned computer program includes a plurality of instructions, and when the instructions are executed, The steps include deriving hop-by-hop metrics for a network topology and multiple events for a cloud-based security service associated with network access to an application for a user or group of users, wherein the network topology is dynamically monitored, and A step of correlating the derived hop-by-hop metric with the multiple events in multiple dimensions, thereby facilitating the automatic determination of the root cause of the problem associated with the user's or group of users' access to the application over the network, based on the multiple events. The steps include: performing a remedial response based on the multiple events, based on the root cause of the problem associated with the user's or group of users' access to the application over the network; To have the computer perform this task. Computer program.
20. A system including a processor and memory, The aforementioned processor, Monitor access to applications over the network, For each of the multiple dimensions, a baseline of the monitored access to the application via the network is generated. The network metric is automatically predicted for one or more of the aforementioned multiple dimensions, and Based on the prediction of the network metric for one or more of the aforementioned dimensions, an alert or report is generated using the recommended configuration and / or remediation. It is configured in such a way, The aforementioned memory is Coupled with the aforementioned processor, and Give instructions to the aforementioned processor, It is structured in such a way. system.