Topological association

By using topology association technology and leveraging AI and ML to automatically detect and repair application connectivity issues, this technology solves the problems of high manual labor time and complex troubleshooting in existing technologies, and achieves fast and accurate fault isolation and repair.

CN120937314APending Publication Date: 2025-11-11PALO ALTO NETWORKS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480024844.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-31
Filing Date
2024-04-12
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing network and security systems face high manpower hours, complex troubleshooting, and high false positive rates when detecting and resolving application connectivity issues, making it difficult to quickly isolate the root cause.

Method used

Employing topology correlation technology, it automatically detects and repairs application connectivity issues by dynamically monitoring network topology and hop-by-hop metrics. It utilizes AI and ML for event correlation and root cause analysis, and provides a natural language query interface and automated repair.

Benefits of technology

It significantly reduces the time and manpower required to detect and repair application connectivity issues, improves the efficiency and accuracy of troubleshooting, and reduces noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937314A_ABST
    Figure CN120937314A_ABST
Patent Text Reader

Abstract

Techniques for providing topological association are disclosed. In some embodiments, a system / process / computer program product for providing topology association includes deriving hop-by-hop metrics of a network topology and a plurality of events of a cloud-based security service associated with access of a user or group of users to an application over a network, where the network topology is dynamically monitored; associating the resulting hop-by-hop metrics and the plurality of events in a plurality of dimensions to facilitate automatically determining a root cause of a question associated with access of the user or group of users to the application over the network based on the plurality of events; and performing a repair response based on the plurality of events based on a root cause of a problem associated with access of the user or group of users to the application over the network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to other applications This application claims priority to U.S. Provisional Patent Application No. 63 / 459,494, entitled “APPLICATION ACCESS ANALYZER”, filed April 14, 2023; U.S. Provisional Patent Application No. 63 / 459,492, entitled “SECURITY POLICY ANALYSIS- DEVOPS APPROACH”, filed April 14, 2023; and U.S. Provisional Patent Application No. 63 / 459,500, entitled “TOPOLOGICAL CO-RELATION”, filed April 14, 2023, all of which are incorporated herein by reference for all purposes. Background Technology

[0002] Malware is a general term used to refer to malicious software (e.g., including a wide variety of hostile, intrusive, and / or otherwise unwanted software). Malware can take the form of code, scripts, active content, and / or other software. Example uses of malware include disrupting computer and / or network operations, stealing proprietary information (e.g., confidential information such as identity, financial, and / or intellectual property-related information), and / or gaining access to private / proprietary computer systems and / or computer networks. Unfortunately, as technologies that help detect and mitigate malware develop, malicious actors have found ways to circumvent these efforts. Therefore, there is an ongoing need to identify and improve the technologies used to identify and mitigate malware. Attached Figure Description

[0003] Various embodiments of the invention are disclosed in the following detailed description and accompanying drawings.

[0004] Figure 1 The illustration shows an example of an environment in which malicious applications (“malicious software”) are detected and prevented from causing harm.

[0005] Figure 2A An embodiment of a data device is illustrated.

[0006] Figure 2B This is a functional diagram of the logic components of an embodiment of a data device.

[0007] Figure 3 The illustration shows an example of a logical component that can be included in a system used to analyze samples.

[0008] Figure 4A This is an overview of AIOPs for Secure Access Service Edge (SASE) solutions according to some embodiments.

[0009] Figure 4B The illustration shows an example use case where a mobile user is unable to connect to the company network due to authentication issues.

[0010] Figure 4C The illustration shows an example use case of using AIOPs for Secure Access Service Edge (SASE) to reduce the mean detection time (MTTD) / mean resolution time (MTTR) of access to SaaS / private apps, according to some embodiments.

[0011] Figure 5A The illustration shows an AIOPs architecture for SASE according to some embodiments to facilitate topology association.

[0012] Figure 5B The illustration depicts an alarm engine according to some embodiments.

[0013] Figure 5C The illustration shows example relationships between various entities that are automatically generated and used for association and PRC according to some embodiments.

[0014] Figure 5D The diagram illustrates the association and Prisam Controller Engine (CPE) according to some embodiments.

[0015] Figure 5E The illustration shows an example data model of the AIOPs platform according to some embodiments.

[0016] Figure 6 The illustration shows an ML pipeline architecture for AIOPs in a SASE solution, according to some embodiments, to facilitate topology association.

[0017] Figure 7A The illustration shows an example use case where the detection of authenticated users is significantly reduced according to some embodiments.

[0018] Figure 7B The illustration shows an example use case for service connection cloud node anomaly detection according to some embodiments.

[0019] Figures 8A-8C The illustration depicts a process, according to some embodiments, for performing associations for an AIOPs SASE solution to facilitate topology association techniques.

[0020] Figure 9 This is a flowchart of a process for providing topological associations according to some embodiments.

[0021] Figure 10 This is another flowchart of a process for providing topological associations according to some embodiments. Detailed Implementation

[0022] This invention can be implemented in numerous ways, including as a process; an apparatus; a system; a composition of matter; a computer program product embodied on a computer-readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on and / or provided by memory coupled to the processor. In this specification, these embodiments or any other form of the invention may be referred to as technology. Generally, the order of steps of the disclosed process can be varied within the scope of the invention. Unless otherwise stated, components described as configured to perform a task (such as processors or memory) can be implemented as general components temporarily configured to perform a task at a given time or manufactured as specific components to perform a task. As used herein, the term 'processor' refers to one or more devices, circuits, and / or processing cores configured to process data (such as computer program instructions).

[0023] The following is an appendix illustrating the principles of the invention. Figure 1 The present invention provides a detailed description of one or more embodiments thereof. The invention has been described in conjunction with such embodiments, but is not limited to any particular embodiment. The scope of the invention is limited only by the claims, and the invention includes many alternatives, modifications, and equivalents. To provide a thorough understanding of the invention, numerous specific details are set forth in the following description. These details are provided for illustrative purposes, and the invention may be practiced without some or all of these specific details. For the purpose of clarity, technical materials known in the art related to the invention have not been described in detail, so as not to unnecessarily obscure the invention.

[0024] Firewalls typically protect networks from unauthorized access while allowing authorized communication to pass through. A firewall is usually a device, a collection of devices, or software running on a device that provides firewall functionality for network access. For example, a firewall can be integrated into the operating system of a device (e.g., a computer, smartphone, or other type of network-enabled device). Firewalls can also be integrated into various types of devices (such as computer servers, gateways, network / routing devices (e.g., network routers), and data devices (e.g., security devices or other types of dedicated devices)) or executed as one or more software applications on various types of devices, and in various implementations, some operations can be implemented in dedicated hardware (such as ASICs or FPGAs).

[0025] Firewalls typically deny or allow network traffic based on a set of rules. These sets of rules are often referred to as policies (e.g., network policies or network security policies). For example, a firewall can filter inbound traffic by applying a set of rules or policies to prevent unwanted external traffic from reaching a protected device. Firewalls can also filter outbound traffic by applying a set of rules or policies (e.g., allow, block, monitor, notify, or log, and / or specify other actions in firewall rules or policies that can be triggered based on various criteria such as those described herein). Firewalls can also filter local network (e.g., intranet) traffic by similarly applying a set of rules or policies.

[0026] Security devices (e.g., security apparatuses, security gateways, security services, and / or other security devices) may include a variety of security functions (e.g., firewalls, anti-malware, intrusion prevention / detection, data loss prevention (DLP), and / or other security functions), networking functions (e.g., routing, quality of service (QoS), workload balancing of network-related resources, and / or other networking functions), and / or other functions. For example, routing functions may be based on source information (e.g., IP address and port), destination information (e.g., IP address and port), and protocol information.

[0027] Basic packet filtering firewalls (such as packet-filtering firewalls or first-generation firewalls, which are stateless packet-filtering firewalls) filter network traffic by examining individual packets transmitted over the network. Stateless packet-filtering firewalls typically examine each packet itself and apply rules based on the examined packet (e.g., using a combination of packet source and destination address information, protocol information, and port number).

[0028] Application firewalls can also perform application-layer filtering (e.g., application-layer filtering firewalls or second-generation firewalls, which operate at the application level of the TCP / IP stack). Application-layer filtering firewalls or application firewalls can typically identify certain applications and protocols (e.g., web browsing using Hypertext Transfer Protocol (HTTP), Domain Name System (DNS) requests, file transfer using File Transfer Protocol (FTP), and various other types of applications and protocols such as Telnet, DHCP, TCP, UDP, and TFTP (GSS)). For example, application firewalls can block unauthorized protocols attempting to communicate on standard ports (e.g., application firewalls can generally be used to identify unauthorized / out-of-policy protocols attempting to sneak through using non-standard ports used for that protocol).

[0029] Stateful firewalls can also perform stateful packet inspection, where each packet is examined within the context of a series of packets associated with a stream of packets traveling through the network. This firewall technique is generally referred to as stateful packet inspection because it maintains a record of all connections traversing the firewall and is able to determine whether a packet is the start of a new connection, part of an existing connection, or an invalid packet. For example, the state of a connection itself can be one of the criteria for triggering rules within a policy.

[0030] Advanced or next-generation firewalls can perform stateless and stateful packet filtering and application-layer filtering as discussed above. Next-generation firewalls can also perform additional firewall technologies. For example, some newer firewalls, sometimes referred to as advanced or next-generation firewalls, can also identify users and content. In particular, some next-generation firewalls are expanding the list of applications they can automatically identify to thousands of applications. Examples of such next-generation firewalls are commercially available from Palo Alto Networks, Inc. (e.g., Palo Alto Networks' PA series firewalls). For example, Palo Alto Networks' next-generation firewalls enable enterprises to use a variety of identification technologies to identify and control applications, users, and content, not just ports, IP addresses, and packets. These identification technologies include: APP-ID for accurate application identification, User-ID for user identification (e.g., by user or user group), and Content-ID for real-time content scanning (e.g., controlling web browsing and restricting data and file transfers). These identification technologies allow enterprises to reliably achieve application purposes using business-relevant concepts, rather than following the traditional methods provided by conventional port-blocking firewalls. Furthermore, compared to software running on general-purpose hardware, dedicated hardware for next-generation firewalls (e.g., implemented as dedicated devices) generally provides a higher level of performance for application review. Dedicated hardware for next-generation firewalls, such as security devices provided by Palo Alto Networks, Inc., uses dedicated, function-specific processing tightly integrated with a single-channel software engine to maximize network throughput while minimizing latency.

[0031] Advanced or next-generation firewalls can also be implemented using virtualized firewalls. Examples of such next-generation firewalls are commercially available from Palo Alto Networks, Inc. (e.g., Palo Alto Networks' VM family of firewalls, which supports a wide range of commercial virtualization environments, including VMware® ESXi™ and NSX™, Citrix® Netscaler SDX™, KVM / OpenStack (Centos / RHEL, Ubuntu®), and Amazon Web Services (AWS)). For example, virtualized firewalls can support similar or identical next-generation firewall and advanced threat prevention features available in physical form factor devices, allowing enterprises to securely enable applications to flow into and across their private, public, and hybrid cloud environments. Automated features such as VM monitoring, dynamic address groups, and REST-based APIs allow enterprises to proactively monitor VM changes and dynamically feed that context into security policies, thereby eliminating policy lag that can occur when VMs change.

[0032] Overview of topological association techniques.

[0033] Generally, existing information technology (IT) operations must identify application connectivity issues for users or user groups through thousands of logs and numerous devices across an enterprise infrastructure. Troubleshooting and debugging connectivity issues typically requires domain expertise, such as network architecture, routing / switching, server configuration, understanding of complex network security policies, and vendor-specific operating system (OS) and command-line interface (CLI) knowledge. Therefore, this significantly increases the human time and average time required to detect and resolve application connectivity issues.

[0034] For example, multiple network and security elements may frequently generate events / alerts that appear anomalous from their perspective, creating an alarm storm for information technology (IT). When downtime issues occur within an enterprise, IT often lacks a full understanding of the actual impact (e.g., treating the network topology as statically distant). Therefore, the challenge for IT personnel is to isolate the problem from the multiple symptoms indicated by the numerous potential events / alerts (e.g., leading to higher false positives). IT personnel must then manually correlate events across different monitoring tools to attempt to identify the root cause of the problem, which is both time-consuming and labor-intensive, yielding limited correlations. This technical issue is common to heterogeneous networks and security stacks (e.g., and while it applies to security stacks, it generally applies to heterogeneous network environments as well, and is therefore relevant to IT help desks, NetOps, network and security administrators, and SASE implementation and technical support teams).

[0035] Therefore, new and improved solutions for promoting topological associations are disclosed with respect to various embodiments.

[0036] In some embodiments, novel and improved topology correlation solutions are disclosed (e.g., for SASE service providers). Generally, event correlation of applications and network access (e.g., access, availability, and / or performance, such as for SaaS / private apps and / or SASE services) is based on time and relationships as dependencies. For example, event correlation can be used to troubleshoot tunnels on interfaces, such as determining potential causes when an interface is down. However, existing event correlation solutions typically treat dependencies and topology relationships as static distances. In contrast, the disclosed topology correlation solution analyzes this network topology as a dynamic state, derived via hop-by-hop metrics and events that can occur on any hop. Specifically, the disclosed topology correlation solution can utilize dynamic states, such as hop-by-hop metrics and events that can occur on any hop. More specifically, the disclosed topology correlation solution can correlate events using both static information and dynamically derived states based on node embeddings. Thus, such events can then be correlated to isolate symptoms from problems and enable rapid root cause identification, as will be further described below with respect to various embodiments.

[0037] Specifically, multiple network and security elements within a network generate events / alerts. The disclosed topology correlation solution can effectively and efficiently correlate these events / alerts to suppress and aggregate them, thereby reducing noise and facilitating root cause detection. For example, the disclosed topology correlation solution can provide causal information to separate problems from symptoms, enabling automated remediation. This reduces the time and human / IT man-hours required to identify and resolve such technical issues in complex network and secure computing environments.

[0038] For example, the disclosed solution may include an interface (e.g., a Natural Language (NL) query interface) to an operator (e.g., an IT / administrator, such as for an IT help desk or other technical support personnel / users) to detect application reachability, connectivity, and access / permission issues, and the solution utilizes the disclosed techniques for topology correlation, such as those further described below with respect to various embodiments. The disclosed solution facilitates automated remediation. As an example, the solution leverages comprehensive details of analyses and inspections performed across different categories (e.g., different domains, including user / endpoint analysis, network analysis, and security policy analysis, such as those further described below) to provide actionable adjudications for queries submitted by operators. Specifically, the solution automatically discovers the network topology used by a given user (e.g., one or more users specified in the query) to access a given application (e.g., a SaaS / private app specified in the query), analyzes the operational status of the underlying network infrastructure, performs user authentication analysis, checks the health and reachability of the Domain Name System (DNS) and Auth servers reached by the user before accessing the application, and performs user- or user group-specific security policy reasoning for any access / permission issues.

[0039] Actionable adjudication, root cause analysis, and correct problem determination significantly reduce the average time spent resolving application connectivity issues. Actionable adjudication, root cause analysis, and correct problem determination also save operators time and effort that would otherwise require them to follow operation manuals / action manuals and debug multiple devices, typically demanding domain knowledge and expertise.

[0040] As an example, the disclosed solution can be used to examine connectivity issues between one or more of the following: (1) a user, multiple users, and / or a group of users from a mobile user gateway to a SaaS application; (2) a user, multiple users, and / or a group of users access to a private application hosted on a premise data center or remote branch office; and (3) a user, multiple users, and / or a group of users to a remote site at a remote branch or data center.

[0041] In some embodiments, the system / process / computer program product for topology correlation includes deriving hop-by-hop metrics of the network topology and multiple events of cloud-based security services associated with a user or user group accessing an application over the network, wherein the network topology is dynamically monitored; correlating the derived hop-by-hop metrics and events across multiple dimensions to facilitate the automatic determination of the root cause of the problem associated with a user or user group accessing the application over the network based on multiple events; and performing a remediation response based on the root cause of the problem associated with a user or user group accessing the application over the network based on multiple events.

[0042] In one embodiment, the disclosed topological association technique can be used to determine comprehensive / multidimensional associations, such as those described further below.

[0043] In one embodiment, the disclosed topological association technique can be used to provide efficient and dynamic hop-by-hop analysis, as will be further described below.

[0044] In one embodiment, the disclosed topology association technique can be used to facilitate greater noise suppression of duplicate / overlapping alarms, as will be further described below.

[0045] In one embodiment, the disclosed topological correlation technique can be used to provide continuous field impact (e.g., blast radius) analysis, such as that described further below.

[0046] In one embodiment, the disclosed topology association technology can be used to associate multiple data sources across multiple domains using AI and ML to determine the root cause of application access problems (e.g., providing enhanced visibility into the health, connectivity, and usage of networks, authentication, DNS, SaaS / private app health, security policy configurations, etc.), as will be further described below.

[0047] In one embodiment, the disclosed topology association technique can be used to automatically detect network connectivity anomalies and / or performance degradation (e.g., network connectivity anomalies and / or performance degradation, such as determining the reachability and / or performance degradation of a given application for one or more users based on the location / access point of one or more users, using configurable thresholds), as will be further described below.

[0048] In one embodiment, the disclosed topology association technique can be used to generate human-consumable / understandable and actionable adjudication analysis, which greatly reduces the average time required to detect and fix application connectivity problems, as will be further described below.

[0049] In one embodiment, the disclosed topology association technique can be used to perform a detailed analysis of various troubleshooting domains in a short period of time (e.g., minutes), which would otherwise typically require many hours to troubleshoot each domain, as will be further described below.

[0050] In one embodiment, the disclosed topology association technique can be used to perform analysis that includes identifying problems in network infrastructure, customer network services, client connectivity issues, SaaS / private application (application) health, and reachability issues, as will be further described below. For example, the disclosed topology association can provide an actionable overview of each troubleshooting domain, and operators do not need domain knowledge expertise to detect and remediate one or more problems.

[0051] In one embodiment, the disclosed topology association technology can be used to automatically discover (automatically discover) the network topology that a user will use to access an application and perform analysis on potential application access problems.

[0052] In one embodiment, the disclosed technology can be used to provide security posture assessment by building a unified computational logic model for firewall security policies.

[0053] In one embodiment, the disclosed topology associations can be used to manage and maintain the tracking of network topology issues, network configuration issues, network services, and security policies, which can often be cumbersome and error-prone, as will be further described below. For example, the disclosed topology associations can provide a comprehensive analysis of each of these domains using a convenient natural language (NL) query interface.

[0054] In one embodiment, the disclosed topology can be used to significantly reduce the operational and support costs for enterprises and their users accessing their SaaS / private apps.

[0055] In one example implementation, the disclosed topology is implemented as the Prisma AI Operations (AIOPs) platform, which provides proactive service level management across global customers and is designed for use by Network Operations Center (NOC) personnel supporting SASE customers, as will be further described below. Specifically, the Prisma AIOPs platform provides proactive monitoring, alerts, issue isolation, and action manual-driven remediation to deliver SLAs (MTTK / I, MTTR) as expected / required by customers.

[0056] In one embodiment, the disclosed Prisma AIOPs can provide AI / ML-driven predictive analytics for network anomaly detection and capacity utilization to avoid outages, as further described below.

[0057] In one embodiment, the disclosed Prisma AIOPs can provide a natural language (NL) interface for contextual troubleshooting and hypothesis analysis, as further described below.

[0058] In one embodiment, the disclosed Prisma AIOPs can provide automated analysis of existing security strategies for best practice recommendations and / or to avoid propagation and / or shadow security strategies, as further described below.

[0059] In one embodiment, the disclosed Prisma AIOPs can provide a sandboxed environment for pre-change analysis of security policies to reduce the risk of change and strengthen the enterprise's security posture, as further described below.

[0060] In one embodiment, the disclosed Prisma AIOPs can provide baselines and correlations of real-time and historical data / events from discrete sources for anomaly detection, root cause analysis (RCA), action manual-driven remediation and repair verification, as further described below.

[0061] In some embodiments, the system / process / computer program product for topology association includes: monitoring access to applications via a network; generating a baseline for monitored access to applications via a network for each of multiple dimensions; automatically forecasting network metrics for one or more of the multiple dimensions; and generating alerts or reports with suggested configurations and / or remediations based on the forecasts of network metrics for one or more of the multiple dimensions.

[0062] Furthermore, network traffic patterns at different office / branch locations (e.g., corporate networks) can vary significantly in terms of user numbers, user traffic access patterns, and usage. Thus, the networks of office / branch locations (e.g., corporate sites) are generally configured with sufficient bandwidth to securely protect traffic from the office / branch location through the firewalls of security providers (e.g., cloud-based firewall solutions such as Secure Access Service Edge (SASE) security solutions, commercially available from Palo Alto Networks, headquartered in Santa Clara, California) to their respective application services (e.g., Software as a Service (SaaS) and / or private applications (Apps)). Generally, there are two aspects to ensure that each office / branch location has sufficient bandwidth to serve users within the branch: (1) coverage capacity (e.g., how much encrypted traffic the firewall serves); and (2) underlying capacity (e.g., how much physical circuitry capacity exists from each office / branch location for physical internet connectivity). Therefore, to better provide a user experience, customers generally leverage long-term forecasts of bandwidth utilization for each site to facilitate the effective provision of additional firewall capacity and physical internet connectivity capacity.

[0063] In addition, SASE service providers (also known as cloud SASE service providers) typically plan and provide capacity for multi-hop tunnels connecting the SASE architecture to provide sufficient connectivity across all tenants.

[0064] In some embodiments, machine learning (ML) predictions of capacity for SASE service providers are disclosed. For example, bandwidth usage forecasts with a predetermined range (horizon) (e.g., a one-month range) can leverage cost-optimized ML solutions. Traditionally, existing algorithms can be used to build ML models with personalized models for each site or each mesh connection in the SASE architecture. These ML models learn the specific behaviors of a site or mesh tunnel and use that behavior to predict future usage. However, the number of such models increases significantly over time, making the cost and maintenance of ML model training and inference very high.

[0065] Therefore, the disclosed ML forecasting of SASE service provider capacity is a novel and improved ML solution capable of cost-optimized forecasting for all sites / circuits using a single ML model. Specifically, key features or feature sets are used to partition the data within the ML model, allowing the ML model to learn different statistical distributions corresponding to different values ​​of these features without mixing them. Furthermore, the disclosed ML solution can learn and utilize correlations between different feature values ​​(such as different branches) without reducing the accuracy of individual forecasts, and in some cases, enhance accuracy by taking correlations into account, as will be further described below with respect to various embodiments.

[0066] Therefore, according to some embodiments, new and improved security solutions utilizing topological associations are disclosed.

[0067] These and other embodiments and examples of topological associations will be further described below.

[0068] Example system environment for topological associations.

[0069] Therefore, in some embodiments, the disclosed techniques include providing a security platform (e.g., one or more security functions / platforms implemented using a firewall (FW) / next-generation firewall (NGFW), a network sensor acting on behalf of the firewall, or another (virtual) device / component that can implement security policies using the disclosed techniques, such as PANOS that can be executed on a commercially available virtual / physical NGFW solution from Palo Alto Networks, or another security platform / NFGW, including, for example, Palo Alto Networks' PA series next-generation firewalls, Palo Alto Networks' VM series virtualized next-generation firewalls and CN series container next-generation firewalls, and / or other commercially available virtual-based or container-based firewalls that can be similarly implemented and configured to execute the disclosed techniques), which is configured to provide DPI capabilities (e.g., including stateful auditing), for example, this may be provided in part or in whole as a SASE security solution, wherein the disclosed techniques for application access analyzers can be used to monitor cloud-based security solutions (e.g., SASE), as further described below.

[0070] Figure 1 The illustration shows an example of an environment in which malicious applications (“malicious software”) are detected and prevented from causing harm. As will be described in more detail below, malware classification (e.g., performed by security platform 122) can... Figure 1 The various entities included in the environment shown share and / or improve upon each other in various ways. Furthermore, using the techniques described herein, devices (such as endpoint client devices 104 to 110) can be protected from such malware (e.g., including previously unknown malware / new variants of malware, such as C2 malware).

[0071] As used herein, “malware” refers to an application that engages in actions that users do not approve or would not approve if fully informed, regardless of whether those actions are clandestine (and illegal). Examples of malware include ransomware, Trojans, viruses, rootkits, spyware, hacking tools, etc. One example of malware is a desktop / mobile application that encrypts user-stored data (e.g., ransomware). Another example of malware is C2 malware, such as those described above. The disclosed techniques can also be used for self-learning malware detection based on sample traffic to detect and / or block other forms of malware (e.g., keyloggers), as further described herein.

[0072] The techniques described in this article can be used in conjunction with a wide variety of platforms (e.g., servers, computing devices, virtual / container environments, desktops, mobile devices, gaming platforms, embedded systems, etc.) and / or for the automated detection of a wide variety of forms of malware (e.g., new malware and / or variants of malware, such as C2 malware). Figure 1 In the example environment shown, client devices 104 to 108 (respectively) are laptops, desktop computers, and tablets located within corporate network 140. Client device 110 is a laptop located outside corporate network 140.

[0073] Data device 102 is configured to enforce policies regarding communication between client devices (such as client devices 104 and 106) and nodes outside the corporate network 140 (e.g., reachable via external network 118). Examples of such policies include those controlling traffic shaping, quality of service, and traffic routing. Other examples of policies include security policies, such as those requiring the scanning of incoming (and / or outgoing) email attachments, website content, files exchanged through instant messaging programs, and / or other file transfers for potential threats. In some embodiments, data device 102 is also configured to enforce policies regarding traffic residing within the corporate network 140.

[0074] Figure 2A Embodiments of the data device are illustrated. In various embodiments, the examples shown represent the physical components included in the data device 102. Specifically, the data device 102 includes a high-performance multi-core central processing unit (CPU) 202 and random access memory (RAM) 204. The data device 102 also includes a storage device 210 (such as one or more hard disks or solid-state storage units). In various embodiments, the data device 102 (whether in RAM 204, storage device 210, and / or other suitable locations) stores information used to monitor the enterprise network 140 and implement the disclosed techniques. Examples of such information include application identifiers, content identifiers, user identifiers, requested URLs, IP address mappings, policies and other configuration information, signatures, hostname / URL classification information, malware profiles, and machine learning (ML) models (e.g., such as self-learning malware detection based on sample traffic). The data device 102 may also include one or more optional hardware accelerators. For example, data device 102 may include an encryption engine 206 configured to perform encryption and decryption operations and one or more field-programmable gate arrays (FPGAs) 208 configured to perform matching, act as a network processor, and / or perform other tasks.

[0075] The functions described herein as being performed by data device 102 can be provided / implemented in a variety of ways. For example, data device 102 may be a dedicated device or a collection of devices. The functions provided by data device 102 may also be integrated into software on general-purpose computers, computer servers, gateways, and / or network / routing devices, or implemented as software on the aforementioned items. In some embodiments, at least some of the services described as being provided by data device 102 are provided to client devices (e.g., client device 104 or client device 110) in place of (or additionally) software executed on client devices.

[0076] Whenever data device 102 is described as performing a task, a single component, a subset of components, or all components of data device 102 may cooperate to perform the task. Similarly, whenever a component of data device 102 is described as performing a task, a sub-component may perform the task and / or a component may combine with other components to perform the task. In various embodiments, portions of data device 102 are provided by one or more third parties. Depending on factors such as the amount of computing resources available to data device 102, various logical components and / or features of data device 102 may be omitted, and the techniques described herein are adapted accordingly. Similarly, additional logical components / features may be included in embodiments of data device 102, if applicable. In various embodiments, an example of a component included in data device 102 is an application identification engine configured to identify applications (e.g., using various application signatures for identifying applications based on packet flow analysis). For example, the application identification engine may determine what type of traffic a session involves, such as web browsing-social networking; web browsing-news; SSH, etc.

[0077] Figure 2B This is a functional diagram of the logical components of an embodiment of the data device. In various embodiments, the examples shown are representations of logical components that may be included in the data device 102. Unless otherwise specified, the various logical components of the data device 102 can generally be implemented in a wide variety of ways, including as a collection of one or more scripts (e.g., written in Java, Python, etc., if applicable).

[0078] As shown, data device 102 includes a firewall and includes a management plane 232 and a data plane 234. The management plane is responsible for managing user interactions, such as by providing a user interface for configuring policies and viewing log data. The data plane is responsible for managing data, such as by performing packet processing and session processing.

[0079] Network processor 236 is configured to receive packets from client devices (such as client device 108) and provide them to data plane 234 for processing. Whenever stream module 238 identifies a packet as part of a new session, it creates a new session stream. Subsequent packets are identified as belonging to the session based on stream lookup. SSL decryption is applied by SSL decryption engine 240, if applicable. Otherwise, processing by SSL decryption engine 240 is omitted. Decryption engine 240 can help data device 102 review and control SSL / TLS and SSH encrypted traffic, and thus help stop threats that might otherwise remain hidden in encrypted traffic. Decryption engine 240 can also help prevent sensitive content from leaving corporate network 140. Decryption can be selectively controlled based on parameters such as URL classification, traffic source, traffic destination, user, user group, and port. In addition to decryption policies (e.g., decryption policies specifying which sessions to decrypt), decryption profiles can be assigned to control various options for sessions controlled by policies. For example, specific cipher suites and encryption protocol versions may be required.

[0080] The application identification (APP-ID) engine 242 is configured to determine what type of traffic a session involves. As an example, the application identification engine 242 can identify GET requests in received data and infer that the session requires an HTTP decoder. In some cases, such as web browsing sessions, the identified application can change, and such changes will be recorded by the data device 102. For example, a user might initially browse a company wiki (categorized as "Web Browsing - Productivity" based on the visited URL) and then subsequently browse a social networking site (categorized as "Web Browsing - Social Networking" based on the visited URL). Different types of protocols have corresponding decoders.

[0081] Based on the determination made by the application identification engine 242, the threat engine 244 sends the packets to the appropriate decoder, which is configured to assemble the packets (which may be received out of order) into the correct order, perform tokenization, and extract the information. The threat engine 244 also performs signature matching to determine what should happen to the packets. If necessary, the SSL encryption engine 246 can re-encrypt the decrypted data. The forwarding module 248 is used to forward the packets for transmission (e.g., to the destination).

[0082] Also Figure 2BAs shown, policy 252 is received and stored in management plane 232. A policy may include one or more rules, which may be specified using domain and / or host / server names, and may apply one or more signatures or other matching criteria or heuristics, such as those used to enforce security policies on subscriber / IP flows based on various extracted parameters / information from monitored session traffic flows. Example policies may include a C2 malware detection policy that uses disclosed techniques for self-learning malware detection based on sample traffic. An interface (I / F) communicator 250 is provided for managing communications (e.g., via a (REST) ​​API, messaging or network protocol communication, or other communication mechanisms).

[0083] Security platform.

[0084] Back Figure 1 Suppose a malicious individual (using system 120) creates malware 130, such as malware for malicious web page activities (e.g., malware delivered to a user's endpoint device via a compromised website when the user visits / browses the compromised website, or via a phishing attack, etc.). The malicious individual wants a client device (such as client device 104) to execute a copy of malware 130 to unpack the malware executable / payload, thereby compromising the client device and, for example, turning the client device into a bot in a botnet. Then, if applicable, the compromised client device can be instructed to perform tasks (e.g., cryptocurrency mining or participating in a denial-of-service attack) and report information to external entities (such as command and control (C2 / C&C) server 150) and receive instructions from C2 server 150.

[0085] Suppose that data device 102 intercepts (e.g., via system 120) an email sent to user "Alice" operating client device 104. In this example, Alice receives the email and clicks on a link to a phishing / compromised website, which may cause Alice's client device 104 to attempt to download malware 130. However, in this example, data device 102 can perform the disclosed techniques for self-learning malware detection based on sample traffic and block access from Alice's client device 104 to the packaged malware content, thereby preempting and preventing any such malware 130 from being downloaded to Alice's client device 104. As will be further described below, data device 102 performs the disclosed techniques for self-learning malware detection based on sample traffic, such as those further described below, to detect and block such malware 130 from harming Alice's client device 104.

[0086] In various embodiments, data device 102 is configured to work in cooperation with security platform 122. As an example, security platform 122 may provide data device 102 with a set of signatures of known malicious files (e.g., as part of a subscription). If the signature of malware 130 is included in the set (e.g., the MD5 hash of malware 130), data device 102 may accordingly prevent the transmission of malware 130 to client device 104 (e.g., by detecting a match between the MD5 hash of an email attachment sent to client device 104 and the MD5 hash of malware 130). Security platform 122 may also provide data device 102 with a list of known malicious domains and / or IP addresses, thereby allowing data device 102 to block traffic between enterprise network 140 and C2 server 150 (e.g., if C&C server 150 is known to be malicious). The list of malicious domains (and / or IP addresses) may also help data device 102 determine when one of its nodes has been compromised. For example, if client device 104 attempts to contact C2 server 150, such an attempt is a strong indicator that client 104 has been compromised by malware (and appropriate remedial actions should be taken, such as isolating client device 104 so that it cannot communicate with other nodes within corporate network 140).

[0087] As will be described in more detail below, security platform 122 may also receive a copy of malware 130 from data device 102 to perform cloud-based security analysis for performing self-learning malware detection based on sample traffic, and malware rulings may be sent back to data device 102 to enforce security policies thereby protecting Alice's client device 104 from the execution of malware 130 (e.g., blocking malware 130 from accessing client device 104).

[0088] In various embodiments, if no signature for an attachment is found, data device 102 can take a variety of actions. As a first example, data device 102 can achieve failsafe by blocking the transmission of any benign attachments not whitelisted (e.g., those whose signatures do not match known good files). The drawback of this approach is that, in cases where many legitimate attachments are actually benign, many of these legitimate attachments may be unnecessarily blocked as potential malware. As a second example, data device 102 can disable the danger by allowing the transmission of any malicious attachments not blacklisted (e.g., those whose signatures do not match known bad files). The drawback of this approach is that it will not be able to prevent recently created malware (previously unseen by platform 122) from causing harm. As a third example, data device 102 can be configured to provide a file (e.g., malware 130) to security platform 122 for static / dynamic analysis to determine whether it is malicious and / or otherwise classify it.

[0089] Security platform 122 stores a copy of the received sample in storage device 142 and begins (or schedules, if applicable) analysis. An example of storage device 142 is an Apache Hadoop cluster (HDFS). The analysis results (along with additional application-related information) are stored in database 146. If the application is determined to be malicious, the data device can be configured to automatically block file downloads based on the analysis results. Furthermore, a signature can be generated for the malware and distributed (e.g., to data devices such as data devices 102, 136, and 148) to automatically block future file transfer requests that download files determined to be malicious.

[0090] In various embodiments, security platform 122 includes one or more dedicated, commercially available hardware servers (e.g., having one or more multi-core processors, 32GB+ of RAM, one or more gigabit network interface adapters, and one or more hard drives) running a typical server-grade operating system (e.g., Linux). Security platform 122 can be implemented across a scalable infrastructure comprising multiple such servers, solid-state drives, and / or other suitable high-performance hardware. Security platform 122 may include several distributed components, including components provided by one or more third parties. For example, some or all of security platform 122 may be implemented using Amazon Elastic Compute Cloud (EC2) and / or Amazon Simple Storage Service (S3). Further, as with data device 102, whenever security platform 122 is referred to as performing a task (such as storing or processing data), it is to be understood that sub-components or multiple sub-components of security platform 122 (whether individually or in collaboration with third-party components) may cooperate to perform that task. As an example, security platform 122 may optionally cooperate with one or more virtual machine (VM) servers (such as VM server 124) to perform static / dynamic analysis.

[0091] An example of a virtual machine server is a physical machine that includes commercially available server-grade hardware (e.g., a multi-core processor, 32+ gigabytes of RAM, and one or more gigabit network interface adapters) running commercially available virtualization software such as VMware ESXi, Citrix XenServer, or Microsoft Hyper-V. In some embodiments, the virtual machine server is omitted. Further, the virtual machine server may be controlled by the same entity that manages security platform 122, but may also be provided by a third party. As an example, the virtual machine server may rely on EC2, where the remainder of security platform 122 is provided by dedicated hardware owned and controlled by the operator of security platform 122. VM server 124 is configured to provide one or more virtual machines 126 to 128 for emulating client devices. The virtual machines can run a wide variety of operating systems and / or versions thereof. Observed behavior resulting from applications running in the virtual machines is logged and analyzed (e.g., to indicate that the application is malicious). In some embodiments, log analysis is performed by the VM server (e.g., VM server 124). In other embodiments, the analysis is performed at least in part by other components of the security platform 122, such as the coordinator 144.

[0092] In various embodiments, as part of a subscription, security platform 122 makes its analysis results on samples available to data device 102 via a list of signatures (and / or other identifiers). For example, security platform 122 may periodically send packets of content identifying malware files, including those for heuristic IPS malware detection based on network traffic (e.g., daily, hourly, or at some other interval, and / or based on events configured by one or more policies). The subscription may cover analysis only on those files intercepted by data device 102 and sent by data device 102 to security platform 122, and may also cover signatures of malware known to security platform 122.

[0093] In various embodiments, security platform 122 is configured to provide security services to a wide variety of entities attached to (or, if applicable, instead of) the operator of data device 102. For example, other enterprises with their own respective enterprise networks 114 and 116 and their own respective data devices 136 and 148 may contract with the operator of security platform 122. Other types of entities may also use the services of security platform 122. For example, an Internet service provider (ISP) providing Internet services to client device 110 may contract with security platform 122 to analyze applications that client device 110 attempts to download. As another example, the owner of client device 110 may install software on client device 110 that communicates with security platform 122 (e.g., to receive content packets from security platform 122, use the received content packets to examine attachments according to the techniques described herein, and transmit applications to security platform 122 for analysis).

[0094] Use static / dynamic analysis to analyze the sample.

[0095] Figure 3 The illustration shows an example of a logical component that can be included in a system used for analyzing samples. Analysis system 300 can be implemented using a single device. For example, the functionality of analysis system 300 can be implemented in a malware analysis module 112 incorporated into data device 102. Analysis system 300 can also be implemented jointly across multiple different devices. For example, the functionality of analysis system 300 can be provided by security platform 122.

[0096] In various embodiments, the analysis system 300 may utilize lists, databases, or other collections of known safe and / or known harmful content (e.g., in...). Figure 3All of these are collectively shown as set 314. Set 314 can be obtained in a variety of ways, including via a subscription service (e.g., provided by a third party) and / or as the result of other processing (e.g., performed by data device 102 and / or security platform 122). Examples of information included in set 314 are: the URL, domain name, and / or IP address of a known malicious server; the URL, domain name, and / or IP address of a known secure server; the URL, domain name, and / or IP address of a known command and control (C2 / C&C) domain; the signature, hash, and / or other identifier of a known malicious application; the signature, hash, and / or other identifier of a known secure application; the signature, hash, and / or other identifier of a known malicious file (e.g., OS development documentation); the signature, hash, and / or other identifier of a known security library; and the signature, hash, and / or other identifier of a known malicious library.

[0097] In various embodiments, when a new sample is received for analysis (e.g., an existing signature associated with the sample does not exist in analysis system 300), the new sample is added to queue 302. Figure 3 As shown, application 130 is received by system 300 and added to queue 302.

[0098] Coordinator 304 monitors queue 302, and when resources (e.g., static analysis workers) become available, coordinator 304 retrieves samples from queue 302 for processing (e.g., retrieves a copy of malware 130). Specifically, coordinator 304 first provides the samples to static analysis engine 306 for static analysis. In some embodiments, one or more static analysis engines are included within analysis system 300, where analysis system 300 is a single device. In other embodiments, static analysis is performed by a separate static analysis server that includes multiple workers (i.e., multiple instances of static analysis engine 306).

[0099] The static analysis engine acquires general information about the sample and includes it (along with heuristics and other information, if applicable) in a static analysis report 308. The report may be created by the static analysis engine or by a coordinator 304 (or another suitable component), which may be configured to receive information from the static analysis engine 306. As an example, static analysis of malware may include performing signature-based analysis. In some embodiments, the collected information is stored in a database record for the sample (e.g., in database 316), in lieu of or appended to a separate static analysis report 308 (i.e., a portion of the database record forms report 308). In some embodiments, the static analysis engine also forms a ruling about the application (e.g., “safe,” “suspicious,” or “malicious”). As an example, the ruling may be “malicious” even if a “malicious” static feature exists in the application (e.g., the application includes hard links to known malicious domains). As another example, points can be assigned to each feature in the features (e.g., based on severity if a feature is found; based on how well the feature is reliable for predicting malicious behavior; etc.), and decisions can be assigned by the static analysis engine 306 (or coordinator 304, if applicable) based on the number of points associated with the static analysis results.

[0100] Once the static analysis is complete, the coordinator 304 locates an available dynamic analysis engine 310 to perform dynamic analysis on the application. Similar to the static analysis engine 306, the analysis system 300 may directly include one or more dynamic analysis engines. In other embodiments, dynamic analysis is performed by a separate dynamic analysis server that includes multiple workers (i.e., multiple instances of the dynamic analysis engine 310).

[0101] Each dynamic analysis worker manages a virtual machine instance (e.g., for simulation / sandbox analysis of samples for malware detection, such as C2 malware detection based on monitored network traffic activity as described above). In some embodiments, the results of static analysis (e.g., performed by static analysis engine 306), whether in report form (308) and / or stored in database 316 or otherwise, are provided as input to dynamic analysis engine 310. For example, static report information can be used to help select / customize the virtual machine instance used by dynamic analysis engine 310 (e.g., Microsoft Windows 7 SP2 versus Microsoft Windows 10 Enterprise, or iOS 11.0 versus iOS 12.0). In the case of multiple virtual machine instances running simultaneously, a single dynamic analysis engine can manage all instances, or multiple dynamic analysis engines can be used (e.g., each dynamic analysis engine manages its own virtual machine instance), if applicable. As will be explained in more detail below, during the dynamic portion of the analysis, the actions taken by the application, including network activity, are analyzed.

[0102] In various embodiments, static analysis of the sample is omitted or performed by a separate entity, if applicable. As an example, conventional static and / or dynamic analysis may be performed on a file by a first entity. Once (e.g., by the first entity) a given file is determined to be malicious, the file may be provided to a second entity (e.g., the operator of security platform 122) specifically for additional analysis of network activity relative to the malware (e.g., by dynamic analysis engine 310).

[0103] The environment used by the analysis system 300 is instrumented / hooked so that behaviors observed during application execution are logged as they occur (e.g., using a custom kernel that supports hooking and log debugging (logcat)). Network traffic associated with the emulator is also captured (e.g., using pcap). Log / network data can be stored on the analysis system 300 as temporary files or more permanently (e.g., using HDFS or another suitable storage technology or combination thereof, such as MongoDB). The dynamic analysis engine (or another suitable component) can compare the connections established by the sample with a list (314) of domains, IP addresses, etc., and determine whether the sample has communicated (or attempted to communicate with) a malicious entity.

[0104] Similar to the static analysis engine, the dynamic analysis engine stores the results of its analysis in records in database 316 associated with the application being tested (and / or includes the results in report 312, if applicable). In some embodiments, the dynamic analysis engine also forms a ruling about the application (e.g., “safe,” “suspicious,” or “malicious”). As an example, the ruling can be “malicious” even if the application takes a “malicious” action (e.g., an attempt to contact a known malicious domain, or an attempt to disclose sensitive information). As another example, points can be assigned to the actions taken (e.g., based on severity if an action is found; based on the reliability of the action for predicting malice; etc.), and the ruling can be assigned by the dynamic analysis engine 310 (or coordinator 304, if applicable) based on the number of points associated with the dynamic analysis results. In some embodiments, the final ruling associated with the sample is made based on a combination of reports 308 and 312 (e.g., via coordinator 304).

[0105] AIOPs architecture for SASE to facilitate application access visibility Figure 4A This is an overview of AIOPs for Secure Access Service Edge (SASE) solutions according to some embodiments. Figure 4A As shown, the disclosed AIOPs for SASE architecture can facilitate application access visibility, including root cause analysis (e.g., actionable insights) and noise reduction (e.g., isolating problems from symptoms, such as accessibility for SaaS / private apps), using a unified management view (a plane of glass) as shown in 402 (e.g., providing efficient visibility across network and security solutions). (For example, in this example, the SASE solution is the commercially available Prisma Access SASE solution from Palo Alto Networks, headquartered in Santa Clara, California, and / or the disclosed technology can be similarly provided for other SASE environments and commercially available solutions.)

[0106] As shown at 404, the disclosed AIOPs for the SASE architecture include a variety of technical components, including baselines, correlations, forecasts, formal analytics (e.g., access permissions / security policies, etc.), natural language processing (NLP) for queries via user interface (UI), and computational action manuals, such as those described further below.

[0107] As shown at point 406, the disclosed AIOPs for the SASE architecture include various data sources, including configuration information, status information, logs (e.g., firewall logs), telemetry data, and synthetic data (e.g., from various network and security solution sources, such as the GlobalProtect (GP) VPN solution commercially available from Palo Alto Networks, headquartered in Santa Clara, California, and other network and security solutions including CDSS, ADEM, SD-WAN, FLN, CIE, and PAI, etc.). Figure 4A (As shown below), and so on, as will be further described below.

[0108] Figure 4B The illustration depicts an example use case where a mobile user is unable to connect to a corporate network due to authentication issues. For instance, the presence of multiple computing components / entities and the network connectivity between these different computing components / entities typically makes it technically challenging for customers (e.g., customer network operations centers (NOCs) and / or IT / help desk personnel) to determine the root cause of any application connectivity problems. Specifically, the technology of the disclosed application access analyzer provides customers / customer NOCs with automated tools to leverage SASE / Prisma Access Analysis and detect potential problems with one or more users / user groups accessing one or more applications (e.g., SaaS / private apps), as will be further described below with respect to various embodiments.

[0109] refer to Figure 4B This example use case illustrates the challenges of manually correlating discrete events, as shown at 420, where a mobile user is unable to connect to the corporate network due to an authentication issue. Specifically, manual analysis and correlation would require two weeks of troubleshooting time to ultimately diagnose this problem with the customer's authentication server.

[0110] Figure 4C The illustration depicts an example use case of reducing the mean time to detection (MTTD) / mean time to resolution (MTTR) of access to SaaS / private apps using AIOPs for Secure Access Service Edge (SASE) according to some embodiments. Specifically, in this example use case, using the disclosed AIOPs for SASE, over-authentication (Auth) timeout failures in all SASE (e.g., Prisma Access (PA)) locations are resolved within minutes (e.g., instead of hours or days), as will be further described below.

[0111] Specifically, the disclosed AIOPs for SASE address various technical issues, as will be described here. MTTD & MTTR for application (App) access problems are typically measured in hours, increasing application downtime and negatively impacting enterprise productivity and revenue. Troubleshooting and debugging generally require domain expertise. Furthermore, manually performing correlation and tracing of multiple factors to conduct root cause analysis (RCA) is often cumbersome and error-prone.

[0112] For example, enterprises and / or cloud service providers with multiple managed network services, large network infrastructure, and complex security policy configurations may face significant challenges in reducing the MTTR of application access issues.

[0113] As another example, identifying RCA in an organization typically requires a comprehensive examination of various domains, such as network connectivity, infrastructure accessibility, and security policy reasoning.

[0114] refer to Figure 4C As shown at 430, an over-authentication failure was detected. At 432, the dynamic baseline automatically detected an abnormal drop in the mobile user count, and the automatic monitoring of the authentication server using cloud probe VMs, employing the disclosed AIOPs for SASE, determined the availability of the authentication server and the availability of network paths to it (e.g., up and running without performance degradation or other reachability issues). At 434, the disclosed AIOPs for SASE were used to perform ML correlation to generate a single actionable incident that isolated the cause to the unresponsive authentication service affecting 1200 mobile users in the western United States, thus focusing the solution on the authentication server. At 436, the disclosed AIOPs for SASE were used to resolve the incident, as it was verified that authentication timeout failures in all SASE PA locations and the mobile user count were now within normal ranges. Thus, according to some embodiments, the disclosed technique of utilizing topology correlation facilitates improved MTTD / MTTR for accessing SaaS / private apps using AIOPs for SASE.

[0115] Figure 5A The illustration depicts AIOPs for a SASE architecture, according to some embodiments, to facilitate topology association. As shown at 502, the Prisma Access (PA) AIOps platform provides proactive service-level management across global customers and is designed for use by Network Operations Center (NOC) personnel supporting Prisma Access (PA) SASE customers. Specifically, Figure 5AThe diagram shown illustrates the high-level architecture of AIOPs services and their interactions, including the data processing pipeline and applications (GKE) deployed as part of the Prisma Access Insights (PAI) platform, as will be further described below.

[0116] Prisma AIOPs provide proactive monitoring, alerts, issue isolation, and action manual-driven remediation to deliver SLAs (MTTK / I, MTTR) as requested by one or more customers. Prisma AIOPs are built on an Insights platform (e.g., the Prisma Access Insights (PAI) platform) that provides cross-domain NOC capabilities, including Prisma SD-WAN (CGX), etc. The PAI platform includes, for example, […]. Figure 5A The published / subscribed communication mechanism 504 shown facilitates communication between various components, including an Extract Transform Load (ETL) engine 506 for standardized data, an alarm engine 508, a resource graph engine 510, a correlation and PRC engine (CPE) 512, a notification engine 514, and an incident execution engine 516 (e.g., the IEE component can provide the following functions: (1) incident policy management, such as suppression policies, role management, etc.; (2) role-based access control (RBAC) for (one or more) event APIs; and (3) for dashboards, single customers, and cross-customer aggregations of (one or more) APIs; and the IEE component can also provide impact analysis, such as understanding the decline or number of affected users based on CDL traffic logs, and when an incident is created based on the affected object (alarm) in the incident, the IEE can then obtain the corresponding impact based on the keywords from the impact structure), and data sources, including events 518, telemetry data 520 (e.g., telemetry data can be collected from various sources, such as PANOS for system log data). Syslogs firewall data, including operational data such as cloud infrastructure monitoring data for AWS, GCP health data, application data, and Cortex data lake (CDL) data; cross-domain telemetry data can also be used for cross-correlation), and alerts and metrics from other domains 522.

[0117] Specifically, Figure 5A The diagram illustrates the high-level architecture of AIOPs services and their interactions. The aforementioned engine / service is part of the data processing pipeline and application (GKE) deployed as part of the PAI platform.

[0118] Figure 5BThe illustration depicts an alarm engine according to some embodiments. As shown at 508, the alarm engine includes sub-components for alarm baselineization 530, alarm evaluation 532, ML detection 534, and alarm information enrichment 536. ML detection is trained using an AutoTSML training component 538 that communicates with a data store (shown as Google Big Query (GBQ) / Google Cloud Store (GCS) 540).

[0119] In example implementations, alerts can be reactive and predictive. Examples of reactive alerts include CPU thresholds exceeding certain anomalies. Predictive alerts are alerts about metrics that provide future times or events when such behavior might occur. Examples of predictive alerts include capacity planning alerts about link utilization, which provide information about when a link might become saturated. Furthermore, alerts can be threshold-based (e.g., configurable / programmable thresholds), simple baselines (e.g., averages), and based on comparisons to SD or ML time (AutoTSML) sequences. (The last sentence appears to be incomplete and possibly refers to another context.) Figure 5B As shown, the alert engine includes ML-based detection, which can be trained using its own infrastructure for training, model validation, deployment, etc.

[0120] For example, the disclosed alarm engine service can generate alarms based on any telemetry data available in the system via telemetry data sources. The alarm engine also includes an API, as shown at 542, which can be used to create alarms via predefined models (e.g., templates like SD or ML). One such example is using the Influx TICK language or ProMQL or Google MQL for picking fields and evaluating expressions. The following three templates are the minimum requirements for examples: (1) based on a general threshold; (2) based on SMA / EMA and SD; and (3) MAD. All of these example templates are stateful (e.g., they track alarm status and status transitions). In addition, as Figure 5B As shown at position 504, such alerts can be persistently stored and propagated to other services via PubSub for consumption.

[0121] Below is an example implementation of the alarm output mode.

[0122] Figure 5C The illustration shows example relationships between various entities automatically generated and used for association and PRC according to some embodiments. Specifically, the example relationships are a graph as shown at 550, which can be used as follows: Figure 5A The resource map shown at position 510 is generated by the engine.

[0123] In the example implementation, the resource graph engine 550 is implemented as a resource builder graph (RBG), which provides a relational model of the complete end-to-end (E2E) computing environment from any given user to each application (e.g., SaaS / private app). The resource graph is hierarchical (e.g., nodes in the graph can be subgraphs with relationships), and the generated general knowledge graph is used for association / causal analysis, such as for performing topology association techniques as described herein. At the highest level of the resource graph, a cross-domain network model is used, which connects users to one or more applications (such as SASE / Prisma Access, SASE / Prisma SD-WAN, etc.) via various domains. Within each of these domains, there exists a network model with various nodes based on connectivity, such as VPN / GP, remote network (RN), proxy, SD-WAN (e.g., SD-WAN architecture), etc. In this example implementation, the cross-domain network model includes control, management, and data plane relationships between entities. In each of these domains, each node or link includes a subgraph that describes relationships across components, processes, etc. An example of a node subgraph is the relationship between IKE and IPSec related processes within a virtual machine (VM) (e.g., the IKE and IPSec pan_task process in a Palo Alto Networks (PAN) VM implementation). Furthermore, when connecting multiple domains, key identifiers (e.g., App, tenant, user ID, etc.) can be standardized. For example, a network / system administrator (admin) with domain knowledge can specify relationships via a meta-language, and RGB uses labeled alarms / events to construct the graph as described above.

[0124] Figure 5D The illustration depicts the association and Prisam Controller Engine (CPE) according to some embodiments. As shown at 512, the CPE includes a subcomponent for data storage (e.g., Neo4J or another commercially available or open-source graph data store can be similarly used for this subcomponent) 552, a Prisma controller 554 (e.g., for finding sites, nodes, link status, dependency information), an association and causation subcomponent 556, an incident data storage / database (DB) subcomponent 558 (e.g., for storing incident, alarm, and impact information), an alarm and event communication pipeline 560 (e.g., for receiving incident ID information from CPE 556 and sending alarms and parsed SYSLOG information to CPE 556), and Redis 562 (e.g., an open-source, in-memory data structure store used as a database, cache, message broker, and streaming engine; in this example embodiment, it can be used to provide CPE 556 with current status, configuration information, etc.). Figure 5D As shown in the figure.

[0125] In this example implementation, the association and PRC engine (CPE) (e.g., in Figure 5A It is shown as CPE 512 in the middle, and Figure 5D (Illustrated as CPE 556) Facilitates correlation and deduplication. Correlation and deduplication can be used in a variety of situations. As an example, consider a scenario where a VM (e.g., a remote network associated with a remote network-service provider network (RN-SPN)) goes down, and then an explosion of tunneled down events is received from all other nodes connected to that VM. Whether this is the result of planned downtime or unplanned outage, it is generally desirable to aggregate these related events / alerts into a single incident to facilitate more efficient and effective identification of root causes and remediation / response (as needed). In the example scenario above, the likely root cause is the VM downtime event. Thus, correlation as used herein generally refers to the process of aggregating and / or clustering events / alerts based on time or relational boundaries to facilitate the disclosed topology correlation techniques. Deduplication as used herein generally refers to the process of identifying duplicate events and then combining these duplicate events into alerts (e.g., single / merged alerts). Therefore, deduplication reduces the number of alerts in the system (e.g., reducing noise for the network / system / IT administrator / user).

[0126] Generally, there are two advanced approaches to perform the association and deduplication process described above: (1) rule-based; and (2) AI / ML-based. The rule-based approach involves system / domain experts typically defining root causes using either code or a meta-language in the form of if-then-that (IFTTT) statements. Rule-based engines can also use relationship graphs to walk incoming alerts and associate them with a portion of the DAG analysis. This rule-based approach may initially be effective, or work in certain situations, but it can become obsolete over time as new relationships or services are added to the infrastructure.

[0127] Conversely, ML / AI-based techniques utilize the aforementioned RGB data as input and relationships constructed based on the timeframes of event / alarm occurrences. Clustering techniques based on proximity, time-based, or centrality-based algorithms produce efficient and effective results. The aforementioned CPEs (e.g., in...) Figure 5A The example output (shown as CPE 512) is an incident with a set of alarms / events and PRC indications, which can be used by other services (e.g., other services of interest that can subscribe to such alarms / events).

[0128] Typically, association treats time and relationships as dependencies. For example, tunneling down through interfaces. Most systems also treat dependencies and topological relationships as static distances. In contrast, in the disclosed topology association techniques, the disclosed model treats the topology as a dynamic state, derived via hop-by-hop metrics and events that can occur at any hop. These events are then correlated to isolate symptoms from problems and enable rapid root cause analysis, as further described herein with respect to various embodiments.

[0129] Figure 5E The illustration shows an example data model of an AIOPs platform according to some embodiments. Specifically, in Figure 5E A sample data model for the AIOPs platform is shown at position 564.

[0130] The following is an example implementation of the association graph pattern.

[0131] The following is an example implementation of an associative data structure.

[0132] Figure 6 The illustrations depict an ML pipeline architecture for AIOPs in a SASE solution according to some embodiments to facilitate topology association. In example implementations, the disclosed AIOPs for a SASE architecture that perform topology association leverage AI / ML technologies to perform various operations, including: (1) anomaly detection (e.g., detecting tunnel degradation behaviors such as round-trip time (RTT); and detecting packet loss errors); (2) prediction (e.g., predicting memory leaks; file descriptor usage; and disk usage); and (3) causal modeling (e.g., automated and enhanced RCA for App access problems and / or other networking / performance-related problems in the SASE computing environment, such as those described herein with respect to various embodiments).

[0133] refer to Figure 6At 602, an example ML pipeline architecture for AIOPs in a SASE solution is shown for topological association. In this example implementation, the ML pipeline architecture can leverage open-source and / or managed service models to execute the ML pipeline (e.g., Google's managed AI / ML pipeline infrastructure that executes KFlow at the lower layer for both training and service purposes). Various AI / ML algorithm libraries can be used for model and inference server deployment. Example AI / ML libraries include a composable version of the AutoTSML library for anomaly detection in Prisma AIOps production deployments to provide near real-time inference.

[0134] Generally, similar to what was described above, network administrators often over-provision the network due to a lack of data-driven approaches, or only begin capacity planning after reaching or exceeding currently configured capacity limits, resulting in a poor user experience. Therefore, according to some embodiments, the disclosed AIOPs solutions include providing forecast information.

[0135] Example of use cases for forecasting The various use cases for forecasting will now be described.

[0136] As a first example use case, historical bandwidth utilization data can be used to train machine learning models so that they can forecast / predict capacity usage.

[0137] As a second example use case, forecast information / data enables network administrators to perform data-driven capacity planning more effectively and efficiently (e.g., forecasting SASE customer capacity based on user input including multiple users and multiple applications). N This includes anticipated bandwidth requirements within the next month, and the bandwidth requirements of an SASE customer's initial deployment. For example, forecasts can be applied to predict the bandwidth usage of a given service connection (SC). N If the capacity is exceeded within 30 days (e.g., 30 days), and therefore, if the capacity is within N If there is insufficient increase within a day, users may experience a degradation in the private / SaaS application experience.

[0138] As a third example use case, the forecast can be used to automatically alert network administrators when the predicted bandwidth is about to exceed a user-defined threshold.

[0139] As a fourth example use case, forecasting can be used by SASE providers to proactively provide forecast data to their customers in order to proactively increase bandwidth based on forecasts of those customers' network capacity (e.g., if increasing bandwidth to branch sites). NFor each user, the bandwidth usage or requirement will be estimated; if one or more applications are enabled, such as Microsoft Teams or Zoom, the additional bandwidth requirement will be estimated.

[0140] As a fifth example use case, forecasts can be used to provide a remediation action manual to aid capacity planning (e.g., for SASE customers).

[0141] As a sixth example use case, forecasting can be applied to predict incidents (e.g., incident prediction in a given SASE deployment). For example, forecasting / predictive analytics can be applied to predict excessive authentication timeout failures due to problems with the customer authentication server.

[0142] As a seventh example use case, predictive / forecasting assessment can be used to provide hypothetical analysis to determine whether a user's access to a private / SaaS app is the result of a recent configuration change (e.g., a user's access to Microsoft Outlook is affected by a recent Active Directory (AD) group change). As another example, customer IT administrators can use the AIOPs platform's Natural Language (NL) interface to submit NL queries to perform multi-domain analysis for a new private / SaaS app rollout across all user locations before a given user complains of an access problem.

[0143] As an eighth example use case, forecasting / predictive assessment can be used for enhanced security analysis (e.g., formal modeling of enterprise security policies). For instance, a security administrator can verify whether a new policy allows engineering user groups access to a private / SaaS app (e.g., Salesforce) and whether a security policy analyzer reveals overlapping / shadow rules for that user group's access to the application. As another example, a network security (NetSec) administrator can perform a pre-change sandbox analysis to determine whether adding a new policy (e.g., to allow HVAC systems access to a patch server) would result in a compliance violation.

[0144] Use case scenarios for topological association The various use cases for providing topological associations, as well as related new and improved techniques, will now be described.

[0145] As a first example use case scenario, the disclosed topology association technology can be applied to facilitate the effective and efficient execution of root cause and / or other analyses so that (one or more) users and / or user groups can access SaaS applications from a mobile user gateway.

[0146] As a second example use case scenario, the disclosed topology association technology can be applied to effectively and efficiently facilitate the execution of root cause and / or other analyses so that (one or more) users and / or user groups can access private applications hosted in local data centers or remote branch offices.

[0147] As a third example use case scenario, the disclosed topology association technology can be applied to facilitate the effective and efficient execution of root cause and / or other analyses for (one or more) users and / or user groups from remote site connectivity to remote branches or data centers.

[0148] As a fourth example use case scenario, the disclosed topology association techniques can be applied to facilitate the effective and efficient execution of multi-domain context troubleshooting analysis with actionable results. For example, multi-domain context troubleshooting analysis can include endpoints, SASE / Prisma Access, authentication, DNS, Layer 3 forwarding, security policies, threat prevention, private / SaaS application access, and / or various other domains that can be automatically analyzed using the disclosed techniques to facilitate actionable decisions (e.g., yes / no decisions; correlation and root cause analysis to identify issues and(one or more) remediation actions; and / or deep observability for each domain to facilitate faster remediation).

[0149] As a fifth example use case scenario, the disclosed topology correlation technology can be applied to detect various anomalies, including power throttling (brownout) due to DDoS events / attacks; a surge in traffic volume on the tunnel (BW increase) due to a specific set of apps or endpoints (e.g., a laptop behaving aberrantly); and performance impacts on other users—currently, customers lack a unified understanding of security attacks and / or their impact on network behavior. One such example is a DDoS attack on a DNS server. It manifests as network capacity saturation, and network operators may plan to increase capacity as a solution. Using the incident framework in AIOps, a baseline of metrics (based on seasonal time-series trends), such as DNS request / response rates and bandwidth usage for secure connectivity between Prisma access and the customer's private data center, is first established to identify anomalous behavior. Correlation logic then associates anomalous DNS traffic patterns with increased bandwidth consumption. Using this context to generate incidents allows NetSec administrators to find the root cause of network resource saturation due to a DNS DDoS attack. In addition, it can provide additional context for the source addresses that generated these DNS requests, allowing security administrators to add policies to block traffic, thereby enabling faster remediation and restoration to normal operating conditions.

[0150] As a sixth example use case scenario, the disclosed topology association technology can be applied to detect various anomalies, including authentication timeout failures, DNS failures, SaaS / private app unreachability, SC BGP outages, SC tunnel outages, etc. For example, alerts can typically be generated to notify a customer's IT about health and performance issues with individual objects such as tunnels, authentication servers, and DNS. However, this can lead to consuming multiple resources to manually identify these related events, one of which is the problem, while others are merely symptoms of the problem / root cause. For example, data center tunnels may have performance or health issues, which could lead to BGP flickering, packet loss, poor private app experience, DNS failures, and / or authentication failures. Thus, the disclosed topology association technology can be applied to associate these discrete events across discrete data sources and deliver a single actionable incident as a problem / root cause (e.g., data center connectivity degradation leading to symptoms such as BGP flickering, packet loss, poor app experience, DNS failure, authentication failure, etc.), thereby reducing manual association and improving the efficiency of customer IT in troubleshooting and proactively resolving problems / root causes.

[0151] As a seventh example use case scenario, the disclosed topology association technology can be applied to detect various anomalies, including CIE agent disconnections, user group count deviations from baselines, and rule hit deviations. For example, a potentially malicious spoofed user can be detected when the same user / user ID logs into the corporate network from two geographically distant locations within an infeasible time interval (e.g., a violation of a time schedule, as the user logs in to two different locations within a period of time, making such a physical schedule infeasible). Customers can specify various security and authentication policies based on user groups. The cloud security service can then periodically learn about these user groups from the LDAP / AAA / AD server. If there is a disconnection between the security service and the LDAP / AAA / AD server, this will result in outdated user group information. This may lead to denial of access. As part of this proactive anomaly detection, the disclosed technology can include monitoring and generating baselines for user groups, users within each group, and security rule hits for accessing applications. This baseline deviation may then be related to a loss of Active Directory (AD) connectivity or a decrease in group counts. Furthermore, such related anomalies can proactively alert customer IT administrators to correct the root cause before multiple users complain about such potential application access problems.

[0152] As the eighth example use case scenario, the disclosed topology correlation technology can be applied to detect various anomalies, including correlating incidents with impacts based on users and applications. For example, in existing IT support systems, IT administrators typically attempt to determine the potential impact on their organization by looking at the total number of user-related complaints / tickets submitted within the same time period. Existing IT support systems are largely reactive in assessing the actual impact on the organization. The disclosed technology can include an incident framework with AIOPs solutions that provides context about an incident by adding user, bandwidth, or application access, or a combination thereof, when the incident is created, and the scope of impact is updated near real-time. Furthermore, this topology correlation information, involving network access / metrics, can be used by network operators and IT administrators to facilitate efficient and effective prioritization of incidents with the greatest impact and enable the help desk to resolve and close user fault tickets more quickly.

[0153] Figure 7A The illustration shows an example use case where the detection of authenticated users is significantly reduced according to some embodiments. Figure 7A The example use case processing is illustrated at 702. At 704, the disclosed ML-driven early anomaly detection technique is used to detect a significant drop in the number of authenticated users. At 706, the disclosed topology correlation technique is used to perform issue isolation and deep correlation. Specifically, it is determined whether it is associated with a cloud infrastructure issue by verifying that the MU portal and gateway nodes are healthy, and whether it is associated with an endpoint issue by verifying that endpoint VPN-related processes (e.g., GP VPN endpoint proxy processes) are downtime, which is identified as a result of using an incompatible GP client software version, as shown at 706. At 708, automated remediation is performed using a compute action manual that facilitates the automatic rollout of a compatible GP client software version to the endpoint. At 710, ML-driven verification is performed. Specifically, it is determined whether users are able to authenticate and whether the number of users is within the expected / normal range (e.g., based on baseline analysis of the number of users).

[0154] Figure 7B The illustration shows an example use case for service connection cloud node anomaly detection according to some embodiments. Figure 7BThe example use case processing is illustrated at 720. At 722, the disclosed ML-driven early anomaly detection automatically detects anomalies in the Service Connectivity (SC) cloud node. Specifically, anomaly detection is correlated with exceeding Service Connectivity (SC) cloud node DP throughput and a surge in tunnel packet loss in DNS transactions (e.g., relative to a previously calculated baseline). At 724, the disclosed topology correlation technique is used to perform problem isolation and deep correlation. Specifically, it is determined that SC throughput and DNS anomalies are correlated, and ADEM reports indicate a low score for the private application to indicate potential endpoint-related issues, as shown at 724. At 726, automated remediation is performed using a computational action plan to facilitate automated remediation, which uses Host Information Profiles (HIPs) to isolate hosts, allowing customers to perform additional investigations on isolated user endpoints. At 728, ML-driven validation is performed to verify that the DNS surge returns to the baseline / normal range, and that the associated SC throughput also returns to the baseline / normal range.

[0155] Various procedural embodiments of the topology association technique will now be described further below.

[0156] Example process for performing topology association techniques Figures 8A-8C The illustration depicts a process for performing associations for an AIOPs SASE solution to facilitate topology association techniques, according to some embodiments. In some embodiments, such as Figures 8A-8C The processes 802, 860, and 880 shown are performed using CPE512 and similar techniques as described above, including those mentioned above. Figures 4A-7B The described embodiments.

[0157] refer to Figure 8AThe diagram illustrates process 802 for aggregating alerts and / or critical events into incidents. At 804, alerts are processed. At 806, an alert status lookup is performed. At 808, the creation or updating of alert statuses with issued or cleared statuses is performed. At 810, it is determined whether an alert is used for association. If not, the process then terminates as shown at 812. Otherwise, the process continues to 814 and determines whether an alert is issued or cleared. If an alert is issued, the process then continues to 816 and determines whether parents have open incidents. If yes, the process then continues to 824 and generates links from alerts to incidents. If not, the process then continues to 818 and determines whether dependents have open incidents. If no, the process then continues to 820 and performs the following actions: creating an incident, setting the status to persistent, linking the alert, and updating the impact. At 822, a persistent timer is started. If it is determined that the subordinate node has pending events, then processing continues to 826 and determines if there are any changes in the RC. If not, then processing continues to 828 and performs the following actions: updates the RC, links the alarm, and updates the impact. If yes, then processing continues to 830 and determines if a status has been created and set to hold. If yes, then processing returns to 828 and performs the following actions: updates the RC, links the alarm, and updates the impact. If not, then processing continues to 832 and performs the following actions: creates and links a new incident. At 834, a hold timer is started. At 836, it is determined whether the cleared alarm is linked to an incident. If not, then performs the following actions: creates an incident and sets its status to hold. Otherwise, processing continues to 840 and performs the following actions: updates the alarm status and impact status to cleared. At 842, it is determined whether the internal alarm has been cleared. If yes, then processing continues to 844 and performs a check on the customer issue. If it is associated with a customer issue, then the process infers that it is associated with a customer incident, as shown at 846. Otherwise, the process continues to 848 and determines whether all mandatory alarms have been cleared. If yes, then the incident status is updated to cleared, as shown at 850, and the process concludes.

[0158] refer to Figure 8BThe diagram illustrates process 860 for aggregating alarms and / or critical events into incidents. At 862, a timer is held until it expires. At 864, it is checked whether an incident has issued an alarm or critical event. At 866, if no, the status is then set to equal "cleared". Otherwise, as shown at 868, a scan is performed to check for cleared / closed incidents with the same root cause (RC). If yes, then as shown at 870, an alarm or critical event is issued and associated with the parent node via the linked incidents, and as shown at 872, the data store (e.g., the incident data store, such as...) is updated. Figure 5D The incident database is shown at point 558. Otherwise, processing continues to 874, and the status is set to equal to cleared, as shown at point 874. At 876, a check for the customer's problem is performed. At 878, the data is stored (e.g., the incident data store, such as...). Figure 5D The incident DB shown at location 558 was updated.

[0159] refer to Figure 8C The diagram illustrates the process 880 for handling alarms associated with an incident. At 882, critical parsing of Syslog information is processed. At 884, an alarm node with the event type is created. At 886, it is determined whether an incident exists on the instance. If not, a hold-on incident is then created as shown at 888. Otherwise, an incident exists on the instance, and an alarm is associated with that incident, as shown at 890.

[0160] Figure 9 This is a flowchart illustrating a process for performing topological association techniques according to some embodiments. In some embodiments, such as... Figure 9 The process 900 shown is performed using a similar disclosed technique as described above, including the techniques mentioned above. Figures 4A-7B The described embodiments.

[0161] At point 902, hop-by-hop metrics of the network topology are derived, along with multiple events associated with cloud-based security services accessed by users or user groups over the network. For example, the network topology can be monitored dynamically (e.g., dynamic hop-by-hop is performed, where the network topology is considered dynamic to facilitate the derivation of hop-by-hop metrics and events that can be associated with any hop and / or occur on any hop, allowing such events to be effectively and efficiently correlated to isolate symptoms from problems and enable rapid root cause analysis), such as those described above. Figures 4A-7B As described.

[0162] At point 904, hop-by-hop metrics and events derived from correlation across multiple dimensions are executed to facilitate the automatic identification of the root causes of problems associated with users or user groups accessing applications over the network, based on multiple events, such as those mentioned above. Figures 4A-7B As described.

[0163] At point 906, based on multiple events and the root cause of the problem related to a user or user group accessing the application over the network, a remediation response is executed, similar to the above reference. Figures 4A-7B As described.

[0164] Figure 10 This is another flowchart illustrating a process for performing topological association techniques according to some embodiments. In some embodiments, such as Figure 10 The process 1000 shown is performed using a similar disclosed technique as described above, including the above-mentioned... Figures 4A-7B The described embodiments.

[0165] At point 1002, monitor network access to the application (e.g., for a specified user or user group), similar to the above reference. Figures 4A-7B As described.

[0166] At point 1004, a baseline for monitored access to the application via the network is generated for each of the multiple dimensions, similar to the reference above. Figures 4A-7B As described.

[0167] At position 1006, automatic prediction is made for network metrics across one or more of multiple dimensions, such as those similar to those mentioned above. Figures 4A-8C As described. In one embodiment, the AIOPs platform / solution described above can be used to perform capacity ML prediction / forecasting to predict / forecast various network capacity and / or other access-related issues. As an example, the AIOPs platform can be used to recommend increasing network capacity for a given network accessing a private / SaaS application based on a prediction that a customer will exceed its current capacity within 30 days. As another example, the AIOPs platform can be used to identify spoofed users based on detected unauthorized time trips (e.g., a user logging into SASE / PrismaAccess from two different source locations / ISPs, which is impossible based on the geographic distance associated with these two different source locations / ISPs within a given time period, thus this can be detected as anomalous user activity, and appropriate mediation responses, such as isolating the user to prevent access to the customer / enterprise's sensitive information / resources, can be suggested or automated).

[0168] At point 1008, based on predictions of network metrics across one or more of multiple dimensions, an alert or report with suggested configurations and / or remediations is generated, similar to the reference above. Figures 4A-7B As described.

[0169] Although the foregoing embodiments have been described in detail for clarity of understanding, the invention is not limited to the details provided. Many alternative ways of implementing the invention are possible. The disclosed embodiments are illustrative and not restrictive.

Claims

1. A system comprising: The processor is configured as follows: This yields hop-by-hop metrics of the network topology and multiple events of cloud-based security services associated with user or user group access to applications over the network. The hop-by-hop metrics and the multiple events obtained by correlating them across multiple dimensions facilitate the automatic determination of the root cause of the problem associated with the user's or user group's access to the application via the network based on the multiple events; as well as Based on the aforementioned events and the root cause of the problem associated with the user's or user group's access to the application via the network, a remedial response is executed. as well as A memory that is coupled to the processor and configured to provide instructions to the processor.

2. The system according to claim 1, wherein, The root cause of the problem associated with the user's or user group's access to the application over the network is determined by using topology association, which is used to associate multiple data sources across multiple domains using artificial intelligence and / or machine learning.

3. The system according to claim 1, wherein, The root cause of the problem associated with the user's or user group's access to the application via the network is determined by using topology association, which is used to associate multiple data sources across multiple domains using artificial intelligence and / or machine learning, wherein the multiple domains include network, authentication, DNS, SaaS / private app health, and security policy configuration.

4. The system according to claim 1, wherein, The root cause of the problem associated with the user's or user group's access to the application via the network was determined to be related to anomalies and / or performance degradation in network connectivity associated with the user's or user group's access to the application via the network.

5. The system according to claim 1, wherein, The remediation response includes generating human-consumable and actionable adjudication analyses that reduce the average time required to detect and remediate application connectivity issues.

6. The system according to claim 1, wherein, The automatic determination of the root cause of the problem associated with the user's access to the application via the network includes identifying network infrastructure problems, customer network service problems, client connectivity problems, SaaS / private application (application) health problems, and / or other connectivity / accessibility problems or performance degradation problems.

7. The system according to claim 1, wherein, The processor is also configured to: Automatically discover the network topology used by the user or user group to access the application to facilitate topology correlations in root cause analysis.

8. The system according to claim 1, wherein, The processor is also configured to: Security posture assessments are generated by building a unified computational logic model for security policies associated with enterprise networks.

9. The system according to claim 1, wherein, The processor is also configured to: User queries are processed using a natural language query interface associated with the user's or user group's access to the application via the network.

10. A method comprising: This yields hop-by-hop metrics of the network topology and multiple events of cloud-based security services associated with user or user group access to applications over the network. The resulting hop-by-hop metrics and the multiple events are correlated across multiple dimensions to facilitate the automatic determination of the root cause of the problem associated with the user's or user group's access to the application via the network, based on the multiple events. as well as Based on the aforementioned events and the root cause of the problem associated with the user's or user group's access to the application via the network, a remedial response is executed.

11. The method according to claim 10, wherein, The root cause of the problem associated with the user's or user group's access to the application over the network is determined by using topology association, which is used to associate multiple data sources across multiple domains using artificial intelligence and / or machine learning.

12. The method according to claim 10, wherein, The root cause of the problem associated with the user's or user group's access to the application via the network is determined by using topology association, which is used to associate multiple data sources across multiple domains using artificial intelligence and / or machine learning, wherein the multiple domains include network, authentication, DNS, SaaS / private app health, and security policy configuration.

13. The method according to claim 10, wherein, The remediation response includes generating human-consumable and actionable adjudication analyses that reduce the average time required to detect and remediate application connectivity issues.

14. The method of claim 10, wherein, The automatic determination of the root cause of the problem associated with the user's access to the application via the network includes identifying network infrastructure problems, customer network service problems, client connectivity problems, SaaS / private application (application) health problems, and / or other connectivity / accessibility problems or performance degradation problems.

15. The method according to claim 10, wherein, The root cause of the problem associated with the user's or user group's access to the application via the network was determined to be related to network connectivity anomalies and / or performance degradation associated with the user's or user group's access to the application via the network.

16. The method of claim 10, further comprising: Automatically discover the network topology used by the user or user group to access the application to facilitate topology correlations in root cause analysis.

17. The method of claim 10, further comprising: Security posture assessments are generated by building a unified computational logic model for security policies associated with enterprise networks.

18. The method of claim 10, further comprising: User queries are processed using a natural language query interface associated with the user's or user group's access to the application via the network.

19. A computer program product embodied in a non-transitory computer-readable medium and comprising computer instructions for: This yields hop-by-hop metrics of the network topology and multiple events of cloud-based security services associated with user or user group access to applications over the network. The hop-by-hop metrics and the multiple events obtained by correlating them across multiple dimensions facilitate the automatic determination of the root cause of problems associated with the user's or user group's access to the application via the network, based on the multiple events; and Based on the aforementioned events and the root cause of the problem associated with the user's or user group's access to the application via the network, a remedial response is executed.

20. A system comprising: The processor is configured as follows: Monitor network access to applications; Generate a baseline for monitored access to the application via the network for each of the multiple dimensions; Automatically predict network metrics for one or more of the multiple dimensions; as well as Based on the prediction of the network metrics of one or more of the multiple dimensions, generate an alert or report with suggested configuration and / or remediation; as well as A memory that is coupled to the processor and configured to provide instructions to the processor.