Endpoint agent and related cybersecurity infrastructure
The cybersecurity endpoint agent application addresses the limitations of existing solutions by enabling remote execution of scripts and invoking cybersecurity functions on endpoint devices, thereby enhancing remediation and analysis capabilities.
Patent Information
- Application Number
- PCT/EP2024/087386
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
Existing cybersecurity solutions face challenges in efficiently performing remediation and investigative actions on endpoint devices remotely, particularly due to limitations in task-automation scripting languages like Bash and PowerShell, which require manual upload and execution of scripts and lack essential cybersecurity functions.
The development of a cybersecurity endpoint agent (EPA) application that embeds a script interpreter and provides cybersecurity functions, allowing for remote execution of scripts on endpoint devices through a WebSocket client, thereby eliminating the need for platform-specific terminal emulation or scripting languages.
Enables sophisticated remediation and analysis actions to be carried out remotely with increased efficiency, as scripts can be uploaded and executed automatically, and essential cybersecurity functions can be invoked directly, enhancing the ability to respond to and mitigate cybersecurity threats.
Smart Images

Figure EP2024087386_26062025_PF_FP_ABST
Abstract
Description
Endpoint Agent and Related Cybersecurity Infrastructure Technical Field
[0001] The present disclosure pertains generally to cybersecurity, and in particular to an endpoint agent for carrying remediation and / or investigative actions at an endpoint device that are instigated remotely, and to related cybersecurity infrastructure. Background
[0002] Cyber defence refers to technologies that are designed to protect computer systems from the threat of cyberattacks. In an active attack, an attacker attempts to alter or gain control of system resources. In a passive attack, an attacker only attempts to extract information from a system (generally whilst trying to evade detection). Private computer networks, such as those used for communication within businesses, are a common target for cyberattacks. An attacker who is able to breach (i.e. gain illegitimate access to) a private network may for example be able to gain access to sensitive data secured within it, and cause significant disruption if they are able to take control of resources as a consequence of the breach. A cyberattack can take various forms. A "syntactic" attack makes use of malicious software, such as viruses, worms and Trojan horses. A piece of malicious software, when executed by a device within the network, may be able to spread throughout the network, resulting in a potentially severe security breach. Other forms of "semantic" attack include, for example, denial-of-service (DOS) attacks which attempt to disrupt network services by targeting large volumes of data at a network; attacks via the unauthorized use of credentials (e.g. brute force or dictionary attacks); or backdoor attacks in which an attacker attempts to bypass network security systems altogether. With increasing emphasis on “remote” access, though remote desktop or virtual private network (VPN) connections and the like, further vulnerabilities and attack opportunities are created. Increasingly distributed and varied deployments create further challenges for cybersecurity monitoring and remediation systems.
[0003] Various remote access technologies allow tasks to be carried out on an endpoint device remotely from a remote console. When a remote access session is established with an endpoint device, this allows the remote console to assume a level of control over the endpoint device. For example, in a ‘remote desktop’ session, the endpoint device’s screen is sharedwith the remote console, and a remote user can perform actions such as moving the cursor and entering text. Control of the endpoint device is provided via a graphical user interface mirroring the experience of a local user.
[0004] For a skilled analyst, it may be more efficient to use some form of command-based interface, such as a terminal emulator. For example, having established a remote desktop session, the analyst might open a command line interface with administrator privileges. Effective use of such an interface requires platform-specific knowledge and experience and such interfaces can vary significantly between different platforms (Windows, Linux, Mac OS etc.)
[0005] A degree of automation may be provided through the use of executable scripts. For example, if some action requires searching through large numbers of files, or needs to be implemented across multiple endpoint devices, it may be more efficient to automate that action with a script. For example, an analyst could write a PowerShell or Bash script that includes one or more ‘for’ loops to allow some action to be iterated over a (potentially large) set of items. PowerShell is a command shell and scripting language typically utilized in Windows, whereas Bash provides a different command shell and scripting language more commonly used in Linux / Mac platforms. Summary
[0006] Scripting languages such as Bash and PowerShell are powerful tools for task automation. However, such task-automation scripting languages have various practical limitations. If a user wishes to run a script on an endpoint device, they will generally need to upload the script to the endpoint device (or devices) manually and then invoke the PowerShell or Bash interpreter on the script manually (for example, via a command entered in a terminal emulator interface). Moreover, Bash and PowerShell are designed as general system administration tools, and lack certain functions that may be required in a cybersecurity context. For example, a cybersecurity analyst might wish to isolate the endpoint device from a network if it is deemed to pose a threat (or potential threat). This would generally require the installation of a firewall or similar piece of software on the endpoint device.
[0007] Herein, a cybersecurity endpoint agent agent (EPA) application allows sophisticated remediation and / or analysis action to be carried out remotely, without relying on platform- specific terminal emulation or platform-specific task-automation scripting languages (such as PowerShell, Bash etc.).
[0008] The EPA application is provided with an embedded script interpreter and at least one software library that provides one or more predetermined cybersecurity functions and exposes those functions to the embedded script interpreter. The cybersecurity agent also includes a WebSocket client or other communication interface that allows a remote access session to be established with a remote appliance endpoint access service. Once a remote access session has been established, a remote access script can be uploaded to the cybersecurity endpoint device and run by the embedded script interpreter. The endpoint agent’s embedded software library by coding one or more function calls in the remote access script to any of the cybersecurity library function(s) exposed to the embedded interpreter. Script outputs generated in running the script at the endpoint device are transmitted back to the appliance endpoint access service. In the described embodiments, a remote access script is coded in a graphical Web interface and uploaded via WebSocket.
[0009] A first aspect herein provides an endpoint agent application for establishing, when executed on an endpoint device, a remote access session between the endpoint device and an endpoint access service. The endpoint agent application has integrated therein a session manager configured to establish a remote access session between the endpoint device and the endpoint access service a communication interface configured to receive from the endpoint access service in the remote access session a remote access script to be run on the endpoint device; a script interpreter configured to generate, in an output object local to the client device, one or more script outputs, by running the remote access script on the endpoint device, wherein the communication interface is configured to automatically transmit to the endpoint access service a copy of all script outputs written to the output object; and at least one software library configured to provide a cybersecurity function, wherein the script interpreter is configured to instigate the cybersecurity function of the software library responsive to a function call thereto in the remote access script.
[0010] In embodiments, the at least one software library may be configured to provide one or more of the following cybersecurity functions: a process termination function to terminate a process running on the endpoint device, a registry access function to access an item containedin an operating system configuration database or file of the endpoint device, a file deletion function to delete a file from a local filesystem of the endpoint device, a file transfer function to cause a copy of a file to be transmitted from a local filesystem of the endpoint device the endpoint agent access service, an isolate function to cause the endpoint device to be isolated from a communication network, a reverse isolate function to restore access to a communication network.
[0011] The script interpreter may be configured to interpret the remote access script in a platform-independent programming language, such as Python.
[0012] The communication interface may be configured to establish a full-duplex communication channel between the endpoint device and the remote access service. The remote access script may be received from the remote access service via the full-duplex communication channel and the copy of all script outputs is transmitted to the remote access service via the full-duplex communication channel.
[0013] The remote access script may be configured to cause the script interpreter to sequentially write a series of script outputs to the output object, and the communication interface may be configured to stream the series of script outputs from the output object to the remote access service via the full-duplex communication channel in a plurality of messages as the remote access script is running on the endpoint device.
[0014] For example, the full-duplex communication channel may be a WebSocket channel.
[0015] The communication interface may be configured to receive from the endpoint access service a script input and store the script input in an input object local to the endpoint device accessible to the remote access script.
[0016] The at least one software library may include a portion of compiled machine code embodying a cybersecurity function and an interface exposing the at least one cybersecurity function to the script interpreter.
[0017] The script interpreter may be configured to interpret the remote access script in an interpreted programming language, and cause the portion of compiled machine code to be executed responsive to a function call in the remote access script.
[0018] The portion of compiled machine code may, for example, have access to context data held in an internal in-memory data structure of the endpoint agent for carrying out the cybersecurity function (permitting a wider range of cybersecurity functionality).
[0019] The at least one software library may comprise a first software library that includes the portion of compiled machine code, and a second software library implemented in the interpreted programming language and configured to provide a second cybersecurity function exposed to the script interpreter.
[0020] The at least one cybersecurity function may comprise an isolate function callable to isolate the endpoint device from a network. The isolate function may be implemented by selectively blocking communications between the network and the endpoint device, whilst maintaining communication between the endpoint agent and the endpoint access service.
[0021] A second aspect herein provides a computer system for use by an analyst in establishing a remote access session with an endpoint device, the computer system comprising one or more processors configured to generate display data for rendering a graphical user interface, the graphical user interface configured to provide a session establishment mechanism for establishing the remote access session with the endpoint device. The graphical user interface is configured to provide a remediation view having an input region editable to code a remote access script and an output region. The one or more processors are configured, responsive to a single run input, to cause: (i) the remote access script to be transmitted to and executed by an endpoint agent on the endpoint device, wherein the endpoint agent is caused to automatically return one or more script outputs generated by running the remote access script locally at the endpoint device, and (ii) the one or more script outputs returned by the endpoint agent to be rendered in the output region of the remediation view.
[0022] The remote access script may be configured to cause the endpoint agent to sequentially write a series of script outputs an output object local to the endpoint device, and stream the series of script outputs from the output storage location to the remote access service via the full-duplex communication channel in a plurality of messages as the remote access script is running on the endpoint device, and the series of script outputs may be rendered sequentially in the output region as the remote access script is running on the endpoint device.
[0023] The remediation view may have at least one of: an upload region for selecting one or more files to be transferred to an input storage location of the endpoint device, and a download region for accessing one or more files transferred from an output storage location of the endpoint device.
[0024] The computer system may comprise: an analyst console comprising at least one display configured to render the graphical user interface to the analyst; and an endpoint access service configured to establish a remote access session with the endpoint device, wherein the analyst console is configured to cause the remote access script to be executed by the endpoint agent by transmitting the remote access script to the endpoint access service, and the endpoint access service is configured to receive the one or more script outputs from the endpoint device and transmit them to the analyst console.
[0025] The analyst console may be configured to establish a first WebSocket with the endpoint access service, and the endpoint access service may be configured to establish a second WebSocket with the endpoint device. The remote access script may be transmitted to the endpoint access service via the first WebSocket and proxied to the endpoint device via the second WebSocket. The one or more script outputs may be received via the second WebSocket, and proxied to the analyst console via the first WebSocket.
[0026] A further aspect herein provides executable program instructions configured so as, when executed on one or more computers, to implement any method, endpoint agent application, or system functionality disclosed herein. Brief Description of Figures
[0027] For a better understanding of the present subject matter, and to show how embodiments of the same may be carried into effect, reference is made by way of example only to the following figures in which:
[0028] Figure 1 shows, by way of context, a schematic function block diagram of a cyber defence platform;
[0029] Figure 2 shows a highly schematic representation of a network event;
[0030] Figure 3 shows a schematic block diagram of a network which may be subject to a cyber-security analysis;
[0031] Figure 4 shows a highly schematic representation of an endpoint event;
[0032] Figure 5 shows a block diagram of an endpoint device running an endpoint agent;
[0033] Figure 6 shows a high-level functional block diagram of an advanced endpoint agent;
[0034] Figure 7 shows a schematic block diagram of a cybersecurity appliance;
[0035] Figure 8 shows a schematic block diagram of an appliance endpoint access service;
[0036] Figure 9 shows a schematic block diagram of a remote access server implemented within an endpoint agent on an endpoint device;
[0037] Figure 10 shows a signalling diagram for an example remote access session;
[0038] Figure 11A shows a schematic remote access view within a remediation and analysis graphical user interface;
[0039] Figure 11B shows an alternative remote access view within a remediation and analysis graphical user interface; and
[0040] Figure 11C shows a digital estate overview within a remediation and analysis graphical user interface. Detailed Description
[0041] In example embodiments, an advanced form of endpoint agent application is deployed to an endpoint device to provide remote access functionality, allowing an analyst to investigate and remediate cybersecurity threats (or potential threats). A remediation and analysis graphical user interface (referred to as the ‘remediation UI’ for short) is provided at an analyst device (console), providing a convenient mechanism for the analyst to code and deploy a remote access script to a desired endpoint device. The endpoint agent includes an embedded script interpreter to allow scripts to be uploaded and run automatically. In the following description, unless otherwise indicated, a user generally refers to a user of an analyst device (typically an analyst instigating investigative and / or remediation actions at a remote endpoint device or devices). The endpoint agent additionally implements various cybersecurity functions, such as device isolation, process termination etc. that may be called in the script.
[0042] Actions are instigated at the endpoint agent remotely via the remediation UI, which allows the analyst to upload a remote access script, in the form of a Python script (or scripts) to the endpoint device. The endpoint agent uses an embedded Python interpreter to act as the interpreter for the script. This means that actions to be performed at the endpoint device may be defined by users in the form of valid Python code. Python is chosen as a suitable cross- platform programming language that can be run on a variety of platforms (e.g. Windows, Linux, MacOS etc.) using a suitable interpreter.
[0043] An appliance endpoint access service (AEAS) instigates remote access by sending a message to an endpoint agent. This creates a remote access session, part of which is an instance of the script interpreter on the endpoint device. At that point, scripts can be sent to the endpoint for execution, and their outputs can be sent back to the AEAS. A WebSocket connection is established between the AEAS and the endpoint agent for this purpose. A second WebSocket connection is established to allow scripts to be passed from remediation UI to the AEAS, and the output of the scripts (received from the endpoint device) to be rendered in the remediation UI. Script outputs are streamed from the endpoint device via the WebSocket connections in real-time or pseudo-real time, rendering those outputs accessible to the analyst as the script is running on the endpoint device.
[0044] WebSockets are a form of ‘push’ communications technology that allow bidirectional communication over a single TCP connection. WebSocket allows persisting connections to be established and kept open, and data can be sent in either direction without the other party to the connection having to request it. In the present context, the use of WebSockets allows remote access scripts and their outputs to be streamed bi-directionally with significantly reduced signalling overhead (compared with long-polling), providing a robust and responsive script interface for use by an analyst in a remote access session.
[0045] An analyst can code an action or set of actions in Python directly. Alternatively or additionally, taking advantage of the fact that Python is a fully-fledged programming language, a number of different modes may be presented to users which allow them to instigate remote actions in very different ways.
[0046] The Python programming language has particular benefits in this context. Python is not only a cross-platform programming language, but also provides a very extensive cross- platform library of functionality as well. Providing users with the ability to run Python meanssignificantly reducing the effort required to carry out the same or similar actions on multiple platforms. A single cross-platform library of useful analysis / remediation Python code can be maintained and made available to users. The effort involved in writing and maintaining such a library is significantly lower using Python compared to platform-specific alternatives (e.g. PowerShell, Bash etc.).
[0047] Nevertheless, it will be appreciated that other programming languages can be used, including other cross-platform programming languages. The following description refers to Python by way of example, but the described features can be implemented using alternative programming languages.
[0048] By way of example, the following modes are considered: Python script; Cross- platform commands; PowerShell; and Bash. All of these modes, including PowerShell and Bash, are implemented under-the-hood in Python.
[0049] When using the Python script mode, users are expected to enter valid Python code which will then be sent to the endpoint agent and executed by the embedded Python interpreter. The output of the script is sent back to the user’s console.
[0050] The Cross-platform commands mode provides a cross-platform command interface that aims to mimic typical Unix-style commands using a Python object-oriented interface. This can provide a more interactive experience. The Cross-platform commands mode is described in further detail below.
[0051] The PowerShell mode uses a Python script template which writes a PowerShell script the user supplies to disk and then invokes the PowerShell interpreter to execute it. The script then captures the output from PowerShell and writes it to the user’s console. The Bash mode functions in a very similar way to the PowerShell mode, except instead of invoking PowerShell, a second Python script template invokes Bash.
[0052] In some embodiments, the endpoint agent additionally operates as a sensor, to provide local monitoring and reporting of endpoint activity (e.g. recording process or user details etc.) and / or local monitoring and reporting of network traffic flowing to and / or from the endpoint device. Endpoint sensors providing local monitoring and reporting of network traffic (network endpoint sensors) may reduce reliance on other types of monitoring component (such as mirrors / TAPs) and / or complement functionality of other type(s) of monitoring component (e.g. in a deployment with “roaming” endpoints). In example embodiments,network data may be linked or otherwise associated with endpoint data locally at an endpoint device. In example embodiments, such linking may be performed locally prior to reporting, response and / or remediation. Such sensors may, for example, be deployed in combination with centralized cybersecurity infrastructure to provide sophisticated threat detection, analytics, response and / or remediation functions. Example system overview
[0053] Figure 1 shows an example of an integrated cybersecurity platform. The integrated cyber defence platform provides overarching protection for a network against cyberattacks, through a combination of comprehensive network and endpoint data collection and organization, and advanced analytics applied to the resulting output within a machine- reasoning framework.
[0054] A “Reflex” component provides an automated threat response that is built on top of the platform. For example, when a significant threat is detected, one or more endpoint devices may be automatically isolated from the network in response. The isolation functionality described herein is shared between Reflex and a ‘REMEDA’ component (see below). An isolate function within an endpoint agent can therefore be triggered automatically by Reflex in response to the detection of a significant threat, or via a function call in a remote access script uploaded to the endpoint agent (through REMEDA).
[0055] Example embodiments compile seemingly-disparate events, records etc. into cases based on recognized tactics, techniques and / or other threat / attack patterns. Cases may, for example, be populated with network data, endpoint data or a combination of endpoint and network data obtained via one or more monitoring techniques, methods, systems taught herein etc. Example embodiments provide targeted alerts and / or reporting, reducing false positives, overreporting etc.
[0056] The platform described with reference to Figure 1 has the ability to receive network and endpoint information from different sources and link that information together server- side.
[0057] However, by deploying “enhanced” endpoint agents within this system, the reliance on such server-side linking is reduced (or potentially eliminated). This is because the endpoint agent has the ability to link network and endpoint information locally, before it isreported to the platform. The enhanced form of endpoint agent is described below, with reference to Figure 5.
[0058] One form of analysis considers longer-term temporal correlations between events, and in particular different types of event such as network and endpoint events. Events that appear to be related are grouped into "cases" over time, as they arrive at an analysis engine. Each case has at least one assigned threat score, denoting the threat level indicated by its constituent events. For example, the threat score may denote one or both of confidence and severity (e.g. a high score may indicate a high confidence that an attack of material severity is occurring or has occurred). A threat score may, in that case, increase if the confidence increases, or the severity increases or both. In other implementations, separate confidence and severity scores may be computed and used within the system.
[0059] The described platform operates according to a “triangulation” model in which multiple forms of analysis may be used as a basis for threat detection. To provide effective triangulation, techniques such as anomaly detection, such as rules-based analytics and / or analytics based on supervised machine learning or other statistical methods more generally may be applied in any combination. By way of example, a particular form of threat detection analytics formulated around the "Mitre ATT&CK framework" (or any other structured source of attack knowledge) is described below. Whilst Mitre is considered, the description applies more generally to other forms of tactics / techniques, including tactic / techniques defined in alternative (e.g. bespoke) threat models, or learned though statistical analysis (such as supervised or unsupervised machine learning).
[0060] A feature of the platform is its ability to collect and link together different types of event, and in particular (i) network events and (ii) endpoint events. This occurs at various places within the system (e.g. at endpoints themselves and / or centrally), as described below.
[0061] The following description refers to network and endpoint events, but applies more generally to other event types received and processed within the system. Other types of event include “third party” events related to suspicious activity in third-party applications and / or services with which the platform is integrated, and / or events generated in an “Internet-of- things” (IoT) context (IoT events).
[0062] Network events are generated by collecting raw network data from components (sub- systems, devices, software components etc.) across a monitored network, and re-structuringthe raw network data into network events. The raw network data can for example be obtained through appropriate network tapping, to provide a comprehensive overview of activity across the network.
[0063] Endpoint events are generated using dedicated endpoint monitoring software in the form of endpoint agents that are installed on endpoints of the network being monitored. Each endpoint agent monitors local activity at the endpoint on which it is installed, and feeds the resulting data (endpoint data) into the platform for analysis.
[0064] This combination of endpoint data with network data is an extremely powerful basis for cyber defence.
[0065] In a data optimization stage, observations are captured in the form of structured, timestamped events. Both network events and endpoint events are collected at this stage and enhanced for subsequent analysis. Events generated across different data collectors are standardized, as needed, according to a predefined data model. As part of the data optimization, first stage enrichment and joining is performed. This can, to some extent at least, be performed in real-time or near-real time (processing time of around 1 second or less). That is, network and endpoint events are also enriched with additional relevant data where appropriate (enrichment data) and selectively joined (or otherwise linked together) based on short-term temporal correlations. Augmentation and joining are examples of what is referred to herein as event enhancement.
[0066] In an analytics stage, these enhanced network events are subject to sophisticated real- time analytics, by an analysis engine. This includes the use of statistical analysis techniques commonly known as “machine learning” (ML). The analysis is hypothesis-based, wherein the likelihood of different threat hypotheses being true is assessed given a set of current or historic observations.
[0067] The creation and subsequent population of cases is driven by the results of analysing incoming events. A case is created for at least one defined threat hypothesis in response to an event that is classed as potentially malicious, and populated with data of that event. That is, each case is created in response to a single event received at the analysis engine. It is noted however that the event that causes a case to be created can be a joined event, which was itself created by joining two or more separate events together, an enriched event, or both.
[0068] Once a case has been created, it may be populated with data of subsequently received events that are identified as related to the case in question (which again may be joined and / or augmented events) in order to provide a timeline of events that underpin the case.
[0069] A case may alternatively or additionally be populated with data of one or more earlier events (i.e. earlier than the event or events that triggered its creation). This is appropriate, for example, where the earlier event(s) is not significant enough in itself to warrant opening a case (e.g. because it is too common), but whose potential significance becomes apparent in the context of the event(s) that triggered the creation of the case.
[0070] An event itself does not automatically create a case. An event may be subject to analysis (which may take into account other data – such as other events and / or external datasets) and it is the result of this analysis which will dictate if it will culminate in the creation of a new case or update of an existing case. A case can be created in response to one event which meets a case creation condition, or multiple events which collectively meet a case creation condition.
[0071] Generally, the threat score for a newly-created case will be low, and it is expected that a large number of cases will be created whose threat scores never become significant (because the events driving those cases turn out to be innocuous). However, in response to a threat occurring within the network being monitored, the threat score for at least one of the cases is expected to increase as the threat develops.
[0072] Another key feature of the system is the fact that cases are only rendered available via a case user interface (UI) when their threat scores reach a significance threshold, or meet some other significance condition. In other words, although a large number of cases may be created in the background, cases are only selectively escalated to an analyst, via the case UI, when they become significant according to defined significance criteria.
[0073] Case escalation is the primary driver for actions taken in response to threats or potential threats.
[0074] Figure 1 shows a schematic block diagram of the cyber defence platform, which is a system that operates to monitor traffic flowing through a network as well as the activity at and the state of endpoints of that network in order to detect and report security threats. The cyber defence platform is implemented as a set of computer programs that perform the data processing stages disclosed herein. The computer programs are executed on one or moreprocessors of a data processing system, such as CPUs, GPUs etc. The system is shown to comprise a plurality of data collectors 102 which are also referred to herein as “coal-face producers”. The role of these components 102 is to collect network and endpoint data and, where necessary, process that data into a form suitable for cyber security analysis. One aspect of this is the collection of raw network data from components of the network being monitored and conversion of that raw data into structured events (network events), as described above. The raw network data is collected based on network tapping, for example.
[0075] Event standardisation components 104 are also shown, each of which receives the events outputted from a respective one of the coal-face producers 102. The standardisation components 104 standardise these structured events according to a predefined data model, to create standardized network and endpoint events.
[0076] The raw network data that is collected by the coal-face producers 102 is collected from a variety of different network components 100. The raw network data can for example include captured data packets as transmitted and received between components of the network, as well as externally incoming and outgoing packets arriving at and leaving the network respectively.
[0077] Additionally, structured endpoint events are collected using endpoint agents 316 executed on endpoints throughout the network. The endpoint agents provide structured endpoint events to the coal-face producers 102 and those events are subject to standardization, enrichment and correlation as above.
[0078] This is described in further detail below, with reference to Figure 3.
[0079] Once standardised, the network events are stored in a message queue 106 (event queue), along with the endpoint events. For a large-scale system, the message queue can for example be a distributed message queue. That is, a message queue 106 embodied as a distributed data storage system comprising a cluster of data storage nodes (not shown in Figure 1).
[0080] An event optimisation system 108 is shown having an input for receiving events from the message queue 106, which it processes in real-time or near real-time to provide enhanced events in the manner described below. In Figure 1, enhanced events are denoted w.esec.t, as distinct from the "raw" events (pre-enhancement) which are denoted w.raw.t. Raw events that are stored in the message queue 106 are shown down the left hand side of the messagequeue (these are the standardised, structured events provided by the standardisation components 104) whereas enhanced events are shown on the right hand side. However, it will be appreciated that this is purely schematic and that the events can be stored and managed within the message queue 106 in any suitable manner.
[0081] The event enhancement system 108 is shown to comprise an enrichment component 110 and a joining component 112. The enrichment component 106 operates to augment events from the message queue 106 with enrichment data, in a first stage enrichment. The enrichment data is data that is relevant to the event and has potential significance in a cybersecurity context. It could for example flag a file name or IP address contained in the event that is known to be malicious from a security dataset. The enrichment data can be obtained from a variety of enrichment data sources including earlier events and external information. The enrichment data used to enrich an event is stored within the event, which in turn is subsequently returned to the message queue 106 as described below. In this first stage enrichment, the enrichment data that is obtained is limited to data that it is practical to obtain in (near) real-time. Additional batch enrichment is performed later, without this limitation, as described below.
[0082] The joining component 112 operates to identify short-term, i.e. small time window, correlations between events. This makes use of the timestamps in the events and also other data such as information about entities (devices, processes, users etc.) to which the events relate. The joining component 112 joins together events that it identifies as correlated with each other (i.e. interrelated) on the timescale considered and the resulting joined user events are returned to the message queue 106. This can include joining together one or more network events with one or more endpoint events where appropriate.
[0083] In Figure 1, the joining component 112 is shown having an output to receive enriched events from the enrichment component 110 such that it operates to join events, as appropriate, after enrichment. This means that the joining component 112 is able to use any relevant enrichment data in the enriched events for the purposes of identifying short-term correlations. However, it will be appreciated that in some contexts at least it may be possible to perform enrichment and correlation in any order or in parallel.
[0084] An observation database manager 114 (storage component) is shown having an input connected to receive events from the message queue 106. The observation database manager114 retrieves events, and in particular enhanced (i.e. enriched and, where appropriate, joined) events from the message queue 106 and stores them in an observation delay line 116 (observation database). The observation delay line 116 may be a distributed database. The observation delay line 116 stores events on a longer time scale than events are stored in the message queue 106.
[0085] A batch enrichment engine 132 performs additional enrichment of the events in the observation delay line 116 over relatively long time windows and using large enrichment data sets. A batch enrichment framework 134 performs a batch enrichment process, in which events in the observation delay line 116 are further enriched. The timing of the batch enrichment process is driven by an enrichment scheduler 136 which determines a schedule for the batch enrichment process. Note that this batch enrichment is a second stage enrichment, separate from the first stage enrichment that is performed before events are stored in the observation delay line 116.
[0086] Figure 3 shows a schematic block diagram of an example network 300 which is subject to monitoring, and which is a private network. The private network 300 is shown to comprise network infrastructure, which can be formed of various network infrastructure components such as routers, switches, hubs etc. In this example, a router 304 is shown via which a connection to a public network 306 is provided such as the Internet, e.g. via a modem (not shown). This provides an entry and exit point into and out of the private network 300, via which network traffic can flow into the private network 300 from the public network 306 and vice versa. Two additional network infrastructure components 308, 310 are shown in this example, which are internal in that they only have connections to the public network 306 via the router 304. However, as will be appreciated, this is purely an example, and, in general, network infrastructure can be formed of any number of components having any suitable topology.
[0087] In addition, a plurality of endpoint devices 312a-312f are shown, which are endpoints of the private network 300. Five of these endpoints 312a-312e are local endpoints shown directly connected to the network infrastructure 302, whereas endpoint 312f is a remote endpoint that connects remotely to the network infrastructure 302 via the public network 306, using a VPN (virtual private network) connection or the like. It is noted in this respect that the term endpoint in relation to a private network includes both local endpoints and remote endpoints that are permitted access to the private network substantially as if they were a localendpoint. The endpoints 312a-312f are user devices operated by users (client endpoints), but in addition one or more server endpoints can also be provided. By way of example, a server 312g is shown connected to the network infrastructure 302, which can provide any desired service or services within private network 300. Although only one server is shown, any number of server endpoints can be provided in any desired configuration.
[0088] For the purposes of collecting raw network data, a plurality of network data capture components 314a-314c are provided. These can for example be network taps. A TAP is a component which provides access to traffic flowing through the network 300 transparently, i.e. without disrupting the flow of network traffic. Taps are non-obtrusive and generally non- detectable. A TAP can be provided in the form of a dedicated hardware TAP, for example, which is coupled to one or more network infrastructure components to provide access to the raw network data flowing through it. In this example, the taps 314a, 314b and 314c are shown coupled to the network infrastructure component 304, 308 and 310 respectively, such that they are able to provide, in combination, copies 317 of any of the raw network data flowing through the network infrastructure 302 for the purposes of monitoring. It is this raw network data that is processed into structured network events for the purpose of analysis.
[0089] Each endpoint device 312a,…,312f has an associated precise endpoint identifier (id), which is a universally unique identifier (UID) uniquely associated with the endpoint device.
[0090] Figure 2 shows a schematic illustration of certain high level structure of a network event 200.
[0091] The network event 200 is shown to comprise a timestamp 204, an entity ID 206 and network event description data (network event details) 208. The timestamp 204 and entity ID 206 constitute metadata 207 for the network event details 208.
[0092] The network event description data 208 provides a network event description. That is, details of the activity recorded by the network event that has occurred within the network being monitored. This activity could for example be the movement of a network packet or sequence of network packets through infrastructure of the network, at a particular location or at multiple locations within the network.
[0093] The network event data 208 can for example comprise one or more network event type indicators identifying the type of activity that has occurred. The entity ID 206 is an identifier of an entity involved in the activity, such as a device, user, process etc. Wheremultiple entities are involved, the network event can comprise multiple network event IDs. Two important forms of entity ID are device ID (e.g. MAC address) and network address (e.g. IP address, transport address (IP address plus port) etc.), both of which may be included in a network event.
[0094] As well as being used as part of the analysis (in conjunction with the timestamps 204), entity IDs 206 and network event description data 208 can be used as a basis for querying enrichment data sources for enrichment data.
[0095] The timestamp 204 denotes a timing of the activity by the network event 200. Such timestamps are used as a basis for associating different but related network events, together with other information in the network event 200 such as the entity ID 206 or IDs it contains.
[0096] The network event 200 can have structured fields in which this information is contained, such as a timestamp field, one or more entity ID fields and one more network event description fields.
[0097] The network event 200 is shown to comprise a network event identifier (ID) 202 which uniquely identifies the network event 200.
[0098] Returning to Figure 3, for the purpose of collecting endpoint data, endpoint monitoring software (code) is provided which is executed on the endpoints of the network 300 to monitor local activity at those endpoints. This is shown in the form of endpoint agents 316a-316g (corresponding to endpoint agents 316 in Figure 1) that are executed on the endpoints 312a-312g respectively. This is representative of the fact that endpoint monitoring software can be executed on any type of endpoint, including local, remote and / or server endpoints as appropriate. This monitoring by the endpoint agents is the underlying mechanism by which endpoint events are collected within the network 300.
[0099] Figure 4 shows a schematic illustration of a certain high level structure of an endpoint event 400.
[0100] The endpoint event 400 is shown to comprise at least one endpoint identifier, such as a device identifier (e.g. MAC address) 402 and network (e.g. IP) address 404 of the endpoint to which it relates, and endpoint event description data 406 that provides details of the local activity at the endpoint in question that triggered the creation of the endpoint event 400.
[0101] One example of endpoint activity that may be valuable from a cyber defence perspective is the opening of a connection at an endpoint. For example, a TCP / IP connection is uniquely defined by a five-tuple of parameters: source IP address (IP address of the endpoint being monitored), source port, destination IP address (IP address of an e.g. external endpoint to which the connection is being opened), destination port, and protocol. A useful endpoint event may be generated and provided to the platform for analysis when an endpoint opens a connection, in which the five-tuple defining the connection is recorded, and well as, for example, an indication of a process (application, task, etc.) executed on the endpoint that opened the connection.
[0102] As noted, one of the key features of the present cyber defence platform is its ability to link together interrelated network and endpoint events. Following the above example, by linking an endpoint event recording the opening of a connection and details of the process that opened it to network events recording the flow of traffic along that connection, it becomes possible to link specific flows of network traffic to that specific process on that endpoint.
[0103] Additional examples of endpoint information that can be captured in endpoint events include information about processes running on the endpoint (a process is, broadly, a running program), the content of files on the endpoint, user accounts on the endpoint and applications installed on the endpoint. Again, such information can be linked with any corresponding activity in the network itself, to provide a rich source of information for analysis.
[0104] Such linking can occur within the platform both as part of the real-time joining performed by the joining component 112.
[0105] However, network and endpoint events can also be linked together as part of the analysis performed by the analysis engine that is inherently able to consider links between events over longer time-scales, as will now be described.
[0106] Returning to Figure 1, the analysis engine, labelled 118, is shown having inputs connected to the event queue 106 and the observation delay line 116 for receiving events for analysis. The events received at the analysis engine 118 from the event queue 106 directly are used, in conjunction with the events stored in the observation delay line 116, as a basis for a sophisticated cyber security analysis that is applied by the analysis engine 118. Queued events as received from the message queue 106 permit real-time analysis, whilst theobservation database 116 provides a record of historical events to allow threats to be assessed over longer time scales as they develop.
[0107] The analysis applied by analysis engine 118 is an event-driven, case-based analysis as will now be described.
[0108] As indicated above, the analysis is structured around cases herein. Cases are embodied as case records that are created in an experience database 124 (which may also be a distributed database).
[0109] Case creation is driven by events that are received at the analysis engine from the message queue 106, in real-time or near-real time.
[0110] Case creation can also be driven by events that are stored in the observation delay line 116. For example, it may be that an event is only identified as potentially threat-related when that event has been enriched in the second stage enrichment.
[0111] Once created, cases are developed by matching subsequent events received from the message queue 106 to existing cases in the experience database 124.
[0112] Events stored in the observation delay line 116 may also be matched to existing cases. For example, it may be that the relevance of a historic event only becomes apparent when a later event is received.
[0113] Thus, over time, a significant case will be populated with a time sequence of interrelated events, i.e. events that are potentially related to a common security threat, and as such exhibit a potential threat pattern.
[0114] Incoming events can be matched to existing cases using defined event association criteria, as applied to the content of the events – in particular the timestamps, but also other information such as entity identifiers (device identifier, IP address etc.). These can be events in the event queue 106, the observation delay line 116, or spread across both. Three key pieces of metadata that are used as a basis for linking events in this way are: ● timestamps, ● endpoint devices, and / or specific endpoint information such as: ● endpoint host name● endpoint open sockets ● IP address.
[0115] These can be multiple pieces of metadata of each type, for example source and destination IP addresses. Such metadata of cases is derived from the event or events on which the case is based. Note the above list is not exhaustive, and the types of data can be used as a basis for event linking.
[0116] For example, events may be associated with each other based on IP address where a source IP address in one event matches a destination IP address in another, and those events are within a given time window. IP addresses provide one mechanism by which endpoint events can be matched with related network events.
[0117] As another example, open sockets on an endpoint are a valuable piece of information in this context, as they are visible to the endpoint agent on the endpoint and associate specific processes running on that endpoint with specific network connections ("conversations"). That is, a socket associated with a process running on an endpoint (generally the process that opened the socket) can be associated with a specific five-tuple at a particular moment in time. This in turn can be matched to network activity within that conversation, for example by matching the five-tuple to the header data of packets tapped from the network. This in turn allows that network activity to be matched to a specific socket and the process associated with it. The endpoint itself can be identified by host name, and the combination of host name, five tuple and time is unique (and in many cases the five tuple and time will be unique depending on the network configuration and where the communication is going). This may also make use of the time-stamps in the network and endpoint events, as the association between sockets and network connections is time limited, and terminates when a socket is closed.
[0118] As noted already, in networking, a five-tuple is a tuple of (source IP, destination IP, source port, destination port, transport protocol). This uniquely identifies a network connection within relatively small time windows. In order to match events based on network connection, a hash of the five tuple can be computed from all network data and from endpoint process connection data (data relating to the network conversations individual processes on the endpoint are engaged in). By ensuring that all endpoint data also contains the host name(derived from the endpoint software), this allows any network event to be correlated with any endpoint event (network 5 tuple hash -> endpoint 5 tuple hash -> host name) and vice versa. This provides an efficient mechanism for linking specific network connections to specific programs (processes). Such techniques can also be used to link network activity to other event description data, e.g. a specific user account on an endpoint.
[0119] As noted, each case is assigned at least one threat score, which denotes the likelihood of the threat hypothesis (or threat hypotheses) to which the case relates. Significance in this context is assessed in terms of threat scores. When the threat score for a case reaches a significance threshold or meets some other significance condition, this causes the case to be rendered accessible via a case user interface (UI) 126.
[0120] Access to the cases via the case UI 126 is controlled based on the threat scores in the case records in the experience database 124. A user interface controller (not shown) has access to the cases in the experience database 124 and their threat scores, and is configured to render a case accessible via the case UI 126 in response to its threat score reaching an applicable significance threshold.
[0121] Such cases can be accessed via the case UI 126 by a human cyber defence analyst. In this example, cases are retrieved from the experience database 124 by submitting query requests via a case API (application programming interface) 128. The case (UI) 126 can for example be a web interface that is accessed remotely via an analyst device 130.
[0122] Thus within the analysis engine there are effectively two levels of escalation:-
[0123] Case creation, driven by individual events that are identified as potentially threat- related.
[0124] Escalation of cases to the case UI 126, for use by a human analyst, only when their threat scores become significant, which may only happen when a time sequence of interrelated events has been built up over time
[0125] As an additional safeguarding measure, the user interface controller may also escalate a series of low-scoring cases related to a particular entity to the case UI 126. This is because a series of low-scoring cases may represent suspicious activity in themselves (e.g. a threat that is evading detection). Accordingly, the platform allows patterns of low-scoring cases that are related by some common entity (e.g. user) to be detected, and escalated to the case UI126. That is, information about a set of multiple cases is rendered available via the case UI 126, in response to those cases meeting a collective significance condition (indicating that set of cases as a whole is significant).
[0126] The event-driven nature of the analysis inherently accommodates different types of threats that develop on different time scales, which can be anything from seconds to months. The ability to handle threats developing on different timescales is further enhanced by the combination of real-time and non-real time processing within the system. The real-time enrichment, joining and providing of queued events from the message queue 106 allows fast- developing threats to be detected sufficiently quickly, whilst the long-term storage of events in the observation delay line 116, together with batch enrichment, provide a basis for non-real time analysis to support this.
[0127] The above mechanisms can be used both to match incoming events from the message queue 106 and events stored in the observation delay line 116 (e.g. earlier events, whose relevance only becomes apparent after later event(s) have been received) to cases. Appropriate timers may be used to determine when to look for related observations in the observation delay line 116 based on the type of observation, after an observation is made. Depending on the attacker techniques to which a particular observation relates, there will be a limited set of possible related observations in the observation delay line 116. These related observations may only occur within a particular time window after the original observation (threat time window). The platform can use timers based on the original observation type to determine when to look for related observations. The length of the timer can be determined based on the threat hypothesis associated with the case.
[0128] The analysis engine is shown to comprise a machine reasoning framework 120 and a human reasoning framework 122. The machine reasoning framework 120 applies computer- implemented data analysis algorithms to the events in the observation delay line 116, such as ML techniques.
[0129] Individual observations may be related to other observations in various ways but only a subset of these relationships will be meaningful for the purpose of detecting threats. The analysis engine 118 uses structured knowledge about attacker techniques to infer the relationships it should attempt to find for particular observation types.
[0130] This can involve matching a received event or sets of events to known tactics that are associated with known types of attack (attack techniques). Within the analysis engine 118, a plurality of analysis modules ("analytics") are provided, each of which queries the events (and possibly other data) to detect suspicious activity. Each analytic is associated with a tactic and technique that describes respective activity it can find. A hypothesis defines a case creation condition as a "triggering event”, which in turn is defined as a specific analytic result or set of analytic results that triggers the creation of a case (the case being an instance of that hypothesis). A hypothesis also defines a set of possible subsequent or prior tactics or techniques that may occur proximate in time to the triggering events (and related to the same, or some of the same, infrastructure) and be relevant to proving the hypothesis. Because each hypothesis is expressed as tactics or techniques, there may be many different analytics that can contribute information to a case. Multiple hypotheses can be defined, and cases are created as instances of those hypotheses in dependence on the analysis of the events. Tactics are high level attacker objectives like "Credential Access", whereas techniques are specific technical methods to achieve a tactic. In practice it is likely that many techniques will be associated with each tactic.
[0131] For example, it might be that after observing a browser crashing and identifying it as a possible symptom of a "Drive-by Compromise" technique (and creating a case in response), another observation proximate in time indicating the download of an executable file may be recognized as additional evidence symptomatic of "Drive-by Compromise" (and used to build up the case). Drive-by Compromise is one of a number of techniques associated with an initial access tactic.
[0132] As another example, an endpoint event may indicate that an external storage device (e.g. USB drive) has been connected to an endpoint and this may be matched to a potential “Hardware Additions” technique associated with the initial access tactic. The analysis engine 118 then monitors for related activity such as network activity that might confirm whether or not this is actually an attack targeting the relevant infrastructure.
[0133] This is performed as part of the analysis of events that is performed to create new cases and match events to existing cases. As indicated, this can be formulated around the "MITRE ATT&CK framework". The MITRE ATT&CK framework is a set of public documentation and models for cyber adversary behaviour. It is designed as a tool for cyber security experts. In the present context, the MITRE framework can be used as a basis forcreating and managing cases. In the context of managing existing cases, the MITRE framework can be used to identify patterns of suspect (potentially threat-related behaviour), which in turn can be used as a basis for matching events received at the analysis engine 118 to existing cases. In the context of case creation, it can be used as a basis for identifying suspect events, which in turn drives case creation. This analysis is also used as a basis for assigning threat scores to cases and updating the assigned threat scores as the cases are populated with additional data. However it will be appreciated that these principles can be extended to the use of any structured source of knowledge about attacker techniques. The above examples are based on tactics and associated techniques defined by the Mitre framework. The described techniques are not limited to Mitre, and can be applied with other forms of tactics / techniques, e.g. in alternative (including bespoke) threat models, or tactics / techniques that are learned via supervised or unsupervised machine learning processing (or other pattern recognition or statistical analysis methods). ‘Learned’ tactics or techniques characterize potential attacks in machine-understandable terms, which may or may not be interpretable to a human. Tactics / techniques may for example be learned by training one or more models on existing or synthetic attack data, and / or from data learned in recording human analyst behaviour.
[0134] Each case record is populated with data of the event or events which are identified as relevant to the case. Preferably, the events are captured within the case records such that a timeline of the relevant events can be rendered via the case UI 126. A case provides a timeline of events that have occurred and a description of why it is meaningful, i.e. a description of a potential threat indicated by those events.
[0135] In addition to the event timeline, a case record contains attributes that are determined based on its constituent events. Four key attributes are: ● people (users) ● processes ● devices ● network connections
[0136] A case record covering a timeline of multiple events may relate to multiple people, multiple devices and multiple users. Attribute fields of the case record are populated with these attributed based on its constituent events.
[0137] A database case schema dictates how cases are created and updated, how they are related to each other, and how they are presented at the case UI 126.
[0138] Micro services 138 are provided, from which enrichment data can be obtained, both by the batch enrichment framework 134 (second stage enrichment) and the enrichment component 110 (first stage enrichment). These can for example be cloud services which can be queried based on the events to obtain relevant enrichment data. The enrichment data can be obtained by submitting queries to the micro services based on the content of the events. For example, enrichment data could be obtained by querying based on IP address (e.g. to obtain data about IP addresses known to be malicious), file name (e.g. to obtain data about malicious file names) etc.
[0139] In addition to the case UI 126, a "hunting" UI 140 is provided via which the analyst can access recent events from the message queue 106. These can be events which have not yet made it to the observation delay line 116, but which have been subject to first stage enrichment and correlation at the event enhancement system 108. Copies of the events from the message queue 106 are stored in a hunting ground 142, which may be a distributed database and which can be queried via the hunting UI 140. This can for example be used by an analyst who has been alerted to a potential threat through the creation of a case that is made available via the case UI 126, in order to look for additional events that might be relevant to the potential threat.
[0140] In addition, copies of the raw network data itself, as obtained through tapping etc., are also selectively stored in a packet store 150. This is subject to filtering by a packet filter 152, according to suitable packet filtering criteria, where it can be accessed via the analyst device 130. An index 150a is provided to allow a lookup of packet data 150b, according to IP address and timestamps. This allows the analyst to trace back from events in the hunting ground to raw packets that relate to those events, for example.
[0141] Figure 5 shows a schematic block diagram of an endpoint device 312 on which an enhanced form of endpoint agent is executed. The enhanced endpoint agent is denoted byreference numeral 616 and may also be referred to herein as an endpoint network sensor (EPNS).
[0142] Whilst the endpoint agent 316 of Figure 1 is deployed for the purpose of endpoint activity monitoring, the EPNS 616 is additionally responsible for monitoring local network traffic. That is, in addition to collecting endpoint data, the EPNS 616 additionally monitors local network traffic to and from the endpoint device 312 in order to collect network data locally at the endpoint device 312. Local network traffic monitoring by the EPNs 616 reduces the reliance on network TAPs and other centralized network monitoring components. The description below may refer to the EPNS 616 as the endpoint agent 616 or the network sensor 616 for conciseness.
[0143] A set of processes 602 is shown to be executed on the endpoint device 312 and the EPNS collects at least some of the endpoint data by monitoring local activity by the processes 602.
[0144] One option is for the EPNS 616 to collect and send copies of all incoming / outgoing network packets for server-side processing, in the manner of a TAP or mirror (but sending a full copy of only its ‘raw’ local network traffic). In this case, the local network traffic copy would be sent to the coal face producers 102 and / or standardizers 104 of Figure 1 for pre- processing into structured events in the same way as raw network traffic received from dedicated monitoring components). However, to reduce transmission overhead, some or all of the functions of the coals face producers 102 / standardizers 104 may be performed locally by the EPNS 616 instead. In such cases, the EPNS 616 instead transmits a more concise summary of its local traffic, in the form of structured network traffic metadata. The term ‘network data’ is used broadly, unless otherwise indicated, and does not necessarily imply ‘raw’ network data (in the context of EPNS reporting, network data can take the form of more-concise network metadata summarizing local network traffic).
[0145] In the following examples, the network data collected by the EPNS 616 takes the form of network traffic metadata summarizing its local network traffic. The EPNS 616 processes the incoming and outgoing local traffic in order to extract such metadata therefrom. The extracted metadata summarizes incoming and outgoing packets of the local traffic. The incoming and outgoing packets carry process data intended for and generated by the processes 602 respectively. The EPNS 616 transmits the extracted metadata to an endpointserver 620 for further processing. The endpoint server 620 forms part of the cybersecurity platform that provides a cybersecurity service implemented remotely from the endpoint device 312.
[0146] An appliance endpoint access service (AEAS) 504 is also shown, and this also from part of the cybersecurity platform. As described in detail below, the AEAS 504 facilitates remote access sessions with endpoint device to allow remediation and / or analysis functions to instigated at the endpoint device remotely.
[0147] Herein, the terms “metadata” and “telemetry” are used interchangeably in relation to network traffic. In the present example, such telemetry includes header data of observed packets, but additionally summarizes their payload data to the extent possible. The telemetry does not include the full “raw” payload data but does summarize the data contained in the payloads of incoming and outgoing packets.
[0148] A key piece of network information is a connection or other “flow” identifier. As noted, a connection is defined by a tuple of (source IP address, source port, destination IP address, destination port, transport protocol). A “flow” generalizes the concept of a connection to connectionless protocols (see below). The opening of a connection or establishment of a flow is a key piece of information that can be used for threat detection and analytics. In the described examples, at a minimum, the EPNS 616 reports every flow that is established at the endpoint device 312 (see below for further details), preferably in combination with additional network associated with the flow. Examples of additional types of network metadata are described below.
[0149] Figure 6 shows a high-level functional block diagram of the endpoint agent 616 in one example implementation. The endpoint agent 616 is shown to comprise a network data processing component 1706, an endpoint data processing component 1708, a local threat detection component 1710, and a local threat remediation component 1712. The network data processing component 1706 receives a copy of the endpoint’s raw network traffic, and processes the raw network traffic to detect the start of new network flows and to extract structured network metadata pertaining to new or existing network flows. Respective endpoint data is associated with each flow locally by the EPNS 616, in order to provide one or more structured telemetry records 1709 containing both structured network metadata and associated structured endpoint metadata pertaining to an identified flow. Every new flow thatis identified by the EPNS 616 is reported to the endpoint server 620, along with a structured metadata summary of the network data carried in that flow and the associated endpoint data local to the EPNS 616 that has been linked to that flow locally by the EPNS 616.
[0150] In order to summarize the “raw” network data in an efficient way that is optimized for subsequent threat analysis, a data model 1701 may be rendered accessible to the EPNS 616. The data model 1701 may be stored locally at the endpoint device 312 accessible to the EPNS 616, or accessed by the EPNS 616 from a remote storage location. The data model 1701 comprises one or more network schemas 1702, which are formal data schemas applied to the raw network traffic and used to structure the network metadata in a queryable fashion. For example, different network schemes may be provided for different protocols (such as HTTP, TLS, SSH), meaning that the nature and extent of the network metadata that is generated may be different for different protocols. This may involve some form of “deep” packet analysis. The depth of the analysis may depend on factors such as the network protocol or protocols with which a given packet(s) is associated. Certain packets or parts may be disregarded if they are of no or limited analytical value. For example, encrypted packet contents may be disregarded.
[0151] For example, for packets carrying HTTP data, structured metadata elements that are extracted from the packets could include one or more of: an HTTP request method contained in a request, a response code returned by an HTTP server, a size (e.g. number of bytes) in the body of an HTTP request or response, the value of a user agent header, and HTTP authentication method used etc. HTTP is non-encrypted; hence these elements can be extracted from the plaintext application data contained in the packets. This involved full analysis of application-level data contained in the packets. For packets carried via TLS, the structured network metadata could include one or more of: TCP sequence number, a TLS server hostname, one or more TLS fingerprints, a TLS version etc.
[0152] For SSH, the extracted metadata elements could include one or more of a client SSH protocol version, a server SSH protocol version, an SSH client name, an SSH server name, comment text from an SSH client or server etc.
[0153] Other examples of extracted metadata elements include VLAN or MPLS label(s).
[0154] The selective extraction and structuring of network data performed by the EPNS 616 mirrors, at least to some extent, functions of the coal face producers 102 and standardizationcomponents 104 shown in Figure 1. Therefore, the EPNS 616 can remove the need for the coal face producers 102 and / or standardization component 104, or reduce the extent to which those components 102, 104 are relied upon (as structured network events, in the form of network traffic records, are generated locally at the endpoint device 312).
[0155] The process of extracting metadata from the local traffic at the endpoint device 312 itself has several benefits. One aim of the local processing is to reduce the amount of network data that needs to be communicated to the endpoint server 620. The metadata does not duplicate the full contents of the incoming and outgoing packets at the endpoint device 312 but provides a sufficiently comprehensive summary to nonetheless be useful in a cybersecurity threat analysis. The use of such metadata in cybersecurity is known per se however existing systems require dedicated components such as network TAPs and appliances that are generally only suitable for deployment in certain networks. Their usefulness is therefore limited to monitoring only network traffic passing through such a network to / from endpoints connected to it directly or remotely using some tunnelling mechanisms such as a VPN connection. In the modern world, with an increasing emphasis on flexible remote access, the limitations of such existing systems are becoming increasingly significant.
[0156] Another significant benefit of implementing the EPNS 616 on the endpoint 312 itself is the ability to associate mutually related network events and endpoint events locally at the endpoint 312 itself. Whilst the platform of Figure 1 that it described above is equipped to perform such linking server-side, such server-side linking is potentially less efficient and more error prone. For example, in the case that endpoint events are collected by endpoint agents and forwarded to the system whilst network traffic is monitored using tapping, the system will receive network and endpoint events from different sources. A large number of such events may be received and the infrastructure required to perform the necessary server- side processing is significant. There is also more scope for a percentage of events being lost or delayed. Significant resources are also required to match large numbers of endpoint events and network events server-side.
[0157] The EPNS 616 has the benefit of being able to link local network traffic to endpoint activity based on local timing. The local timing of local network packets received at or sent from the endpoint 312 and the local timing of incident of endpoint activity can typically be determined highly accurately at the endpoint device 312 itself.
[0158] However, this is platform dependent. In some cases, the association is done by the operating system itself, and the endpoint agent simply needs to locate that information; in others it comes down to factors such as timing and connection tuple.
[0159] The EPNS 616 provides the extracted network traffic metadata to the endpoint server 620 in a series of records transmitted to it. Records containing such network traffic metadata may be referred to herein as network traffic records. Such records, when generated locally at the endpoint device 312 by the EPNS 616, may also be referred to as endpoint records (as described later, the EPNS 616 can be deployed in a system where network traffic records are collected using a combination of local and remote network monitoring). The endpoint records generated by the EPNS 616 are associated with endpoint data collected by the EPNS 616 based on local activity monitoring at the endpoint device 312. For example, it may be the case that the endpoint records are augmented or enriched with such endpoint data locally at the endpoint device 312 or it may be that the endpoint data is contained in separate records generated locally and linked to the network traffic records by the EPNS 616. In general, any combination of augmentation, enrichment, linking and / or any other mechanism that has the effect of associating network traffic records with related endpoint data may be implemented locally at the endpoint device 312 by the EPNS 616. The associated endpoint data is likewise communicated to the endpoint server 620.
[0160] Returning to Figure 5, the endpoint 312 is shown to execute an operating system (OS) 604, on which the processes 602 and the EPNS 616 run. The processes 602 typically include instances of one or more applications 606 stored in computer storage 608 of the endpoint device 312. One function of the OS 604 is to manage the processes 602 and allocate resources to them. The OS 604 also regulates the flow of network traffic between a network interface 610 of the endpoint device 312 and the processes 602 and the EPNS 616.
[0161] In addition, the OS 604 provides a local traffic access function 612 (also shown in Figure 6) and a local activity monitoring function 614. These may, for example, be provided as part of one or more application programming interfaces (APIs) of the OS 604. The EPNS 616 uses the local traffic access function 612 in order to obtain duplicate copies of all incoming and outgoing network packets received at the network interface 610. This includes network packets sent to and from the processes 602, carrying inbound and outbound process data respectively. The EPNS 616 may also receive duplicate copies of its own incoming andoutgoing network packets. The EPNS 616 processes the duplicate packet in order to extract the network traffic metadata.
[0162] In addition, the EPNS 616 uses the local activity monitoring function 614 to monitor endpoint activity by the processes 602. Examples of the type of endpoint activity that may be monitored may include, but are not limited to, the opening of ports, the accessing of files etc. Such monitoring is used to determine endpoint data. The monitoring may be ongoing, even if the endpoint data is static. For example, processes may be monitored in order to link some activity by a process to one or more network packets. In that case, the endpoint data may take the form of a process identifier (ID) that is associated with a telemetry record(s) summarizing those packet(s) (and, conceivably other identifier(s) such as a file identifier if such information is available). Endpoint data can also include user information, such as details of a user account associated with an incident of network activity, or host information about the endpoint device 312. In that case, a telemetry record of the network activity may be associated with a user identifier. For example, the endpoint data collected by the EPNS 616 and associated with the network traffic records can comprise any combination of process, host and / or user data, for example. For example, with current operating system APIs, it is generally possible to obtain some or all of the following endpoint data for particular network packets: username, process details, parent process details, process path, process command line string. In the following examples, actions by the processes are monitored, in order to link such identifiers to network packets.
[0163] The endpoint device 312 is an endpoint of a packet-based network 630 to which it is connected to the network interface 610, and through which the incoming and outgoing network traffic flows. The packet-based network 630 could be a “closed” network such as an enterprise or corporate network (e.g. the private network 300 of Figure 3). However, it could alternatively be an “open” network (such as the internet 306). Referring to Figure 3, the network 630 could be the Internet 306 when the endpoint device 312 is “roaming”, or the private network 300 when the endpoint device is “non-roaming”.
[0164] Even when the endpoint device 312 is not currently connected to the private network 300 (whether directly or via a VPN connection), that does not necessarily mean that it poses no threat to the private network. For example, the endpoint device 312 could still contain sensitive data or, should the endpoint device 312 become infected with some form of malware, that could propagate into the private network 300 when the endpoint device 312does subsequently connect to it. There are therefore significant benefits to being able to detect cybersecurity threats even when the endpoint device 312 is roaming.
[0165] The endpoint device 312 could, for example, take the form of a user device such as a laptop or desktop computer, tablet or smart phone etc. The primary function of a user device is to provide useful functions to a user of the device. Such functions are implemented by the processes 602. The EPNS 616 is deployed on such a user device to provide secondary network and endpoint monitoring functions and submit cybersecurity data (network metadata and associated endpoint data in the present example) to the cybersecurity platform for analysis.
[0166] Referring to Figure 6, the EPNS 616 may also structure endpoint data according to the data model 1701. In this case, the data model 1701 includes one or more endpoint schemas 1704 that are used to structure “raw” endpoint data obtained via the OS. Raw endpoint data can be obtained from various sources / interfaces provided by the OS and, as noted, the nature and extent of endpoint data that is available will depend on the OS. The EPNS 616 extracts individual pieces of endpoint data from the raw endpoint data in a structured, queryable fashion, linking or otherwise associating those pieces of endpoint data with the network metadata elements to which they relate. Sources of raw endpoint data include, for example, OS event logs, alerts, performance data (e.g. CPU / memory usage by the processes 602), exceptions, browser extensions, process or thread queries etc.
[0167] Process details and / or other forms of endpoint data may be associated with an entire flow (e.g. TCP connection or UDP session) so that all the packets within the flow are linked to the processes. This is important because analysing collections of packets within a flow rather than just individual packets allows the system to reconstruct data that is spread across multiple packets. In the described examples, packets are analysed as part of a flow and process details are associated with the flow on the endpoint 312, by the EPNS 616. The server-side only ever sees the results of analysing a flow ( individual packets are not processed by the server in the present examples, rather the server only receives summary data that has already been associated client-side with a specific flow). Here, ‘analysing’ refers to the processing of packets by the EPNS 616 to detect and report all new flows to which the endpoint device 312 is party, as they are established, and to extract and transmit structured network telemetry summarizing the network data carried in flows visible at the endpoint 312, in accordance with the data model 1701 / schema(s), for use in server-side threat detection(analysing does not, in this context, refer to local threat detection; flows are not, for example, only selectively reported when the EPNS 616 considers them indicative of a threat, nor is any form of local threat detection required; if local threat detection is performed, all flows are nevertheless reported independently of such local threat detection, e.g. to facilitate server-side triangulation-based detection of threats that may not be evident from a single-point threat analysis at the client device 312).
[0168] Centralized threat detection based on data reported from multiple endpoint sensors (and any other monitoring components) is beneficial, as it allows a greater range of threats to be identified (many threats are not immediately evident when only viewed locally at a single endpoint). For example, when two endpoints separately report a common connection or other “flow” between those endpoints, each endpoint will report a set of endpoint data local to that endpoint. De-duplication processing performed server side resulting in a single flow record associated with two sets of endpoint data from the different endpoints (see below for details). This richer information source might, in turn, allow a threat to be identified that is not immediately evident at either one of the endpoints. More generally, a centralized perspective allows threats to be detected or ‘triangulated’ from multiple sources.
[0169] Nevertheless, the EPNS 616 may be configured to additionally perform some level of localised threat detection (complementing the centralized processing), denoted by the local threat detection component 1710. Local threat detection can be based on the same linked network / endpoint data that is reported to the remote cybersecurity service or a more limited set of data provided to a local threat detection component of the EPNS 616. In response to detecting a local threat, the threat detection component can take various actions, such as generating an alert (e.g. visual alert) at the endpoint device 312, or generating a threat remediation command.
[0170] Locally-detected threats may be communicated by the local threat detection component 1710 of EPNS 616 to the remote cybersecurity service (the endpoint server 620 in this case). This is separate from, and in addition to the function of reporting of local network traffic, which more closely mirrors a network tap (in combination with a data extraction system) and is ‘agnostic’ to the threat level associated with the network traffic in that network traffic is reported independently of any local threat detection (that is, all network traffic is reported ‘agnostically’ EPNS 616 at the level of detail defined in the data model / schema(s)).
[0171] Local threat remediation (1712) might, for example, comprise the endpoint agent 616 isolating the endpoint device 312 from the network 630. Another example would be terminating a process, e.g. by name or PID. “Local” in this context refers to the fact that action is taken by the endpoint agent 616 at the endpoint device 312; such local action may be instigated locally, based on local threat detection, or remotely in a remote access session. Remote Remediation and Analysis
[0172] Remote remediation and analysis (REMEDA) capability refers herein to a set of system features providing a “response service”.
[0173] The response service allows analysts to contain and resolve threats relating to (among other things) malware, data loss and abuse of management connections. In addition, REMEDA supports advanced manual detection and post-incident work through forensic endpoint analysis features.
[0174] To this end, the endpoint agent is shown to comprise a remote access server (RAS) 502 coupled to the local threat remediation component 1712. The remote access server 502 is an integrated component of the endpoint agent 616 that runs on the endpoint device; the ‘server’ terminology refers to the role of this component 502 in the context of a remote access session, in which the endpoint agent 616 ‘serves’ the appliance AEAS 504, by providing access to a range of cybersecurity functions available at the endpoint device.
[0175] Remote access is facilitated by the endpoint agent 616, which takes the form of a single application that can be installed on any number of endpoint devices 316a, 316b, 316c etc., and which conveniently integrates logic for establishing a remote access session with one or more additional cybersecurity function(s) in the same software application. In the described examples, the endpoint agent 616 integrates remote access logic (the ability to establish a remote access session with the endpoint device) with a script interpreter and a bespoke software library of cybersecurity function(s) (including threat remediation function(s) such as device isolation) together with network and endpoint monitoring and reporting logic. The endpoint agent 616 is supported by a set of appliance infrastructure that will now be described.
[0176] Figure 7 shows a highly schematic function block diagram of a cybersecurity appliance 500. The cybersecurity application 500 is a set of backend hardware and software components forming part of the cybersecurity platform of Figure 1. Whilst Figure 7 focusseson a subset of components designed to facilitate endpoint remote access, the platform supports a number of other cybersecurity features as described above.
[0177] The cybersecurity appliance 500 is shown to comprise the AEAS 504 of Figure 5, a platform application programming interface (API) 508 and an API database 510.
[0178] An analyst device 505 is provided to an analyst, who can access the cybersecurity platform via a remediation client 506 running on the analyst device. The remediation client 506 operates as an interface to the cybersecurity platform, and may be referred to as the analyst user interface (UI) or frontend. The analyst UI 506 may, for example, take the form of a Web browser, with the described functionality delivered to the analyst device 505 as a Web application. Access is provided primarily via a remediation graphical user interface (GUI) rendered at the analyst device 505. Among other things, the analyst UI 506 provides a script interface via which the analyst can write script and cause the script to be automatically uploaded to and run on a specified endpoint device. Outputs of the script are streamed from the endpoint device back to the analyst UI 506 for rendering in real-time or pseudo-real-time.
[0179] Access to the platform is provided primarily via the platform API 508. The platform API 508 provides an authentication function for authenticating the analyst and granting the analyst device 505 access to the platform. Authentication for remote access is based on an analyst account held in an authentication system 512, which may form part of the platform, or which may be managed by a third-party. Preferably, multi-factor (MFA) authentication is used. The platform API 508 also provides access to the API database 510, allowing the analyst to read audit logs pertaining to a remote access session. Audit logs are written to the API database 510 by the AEAS 504 (see below).
[0180] The platform API additionally allows an authorized analyst to retrieve and filter a list of endpoint devices within a “digital estate” accessible to them. The digital estate includes all endpoint devices on which the endpoint agent 616 is installed and to which the analyst has been granted remote access rights.
[0181] Once a remote access session has been established with an endpoint device (e.g. 316a, 316b or 316c), real-time or pseudo-real-time communication between the analyst device 505 and the endpoint device is effected by way of a pair of WebSockets: a first WebSocket 507a between the endpoint device and the AEAS 504 (EPA WebSocket), and a second WebSocket connection 507b between the AEAS 504 and the analyst UI 506 (UI WebSocket).
[0182] When the analyst instigates a request to the AEAS 504 for remote access to an identified endpoint, the AEAS 504 authenticates the request via the platform API 508.
[0183] Via the analyst UI 506, the analyst may also download any files retrieved from an endpoint device by the remote AEAS 504 using a unique, pre-authenticated URL, as described in further detail below.
[0184] Figure 8 shows further details of the AEAS 504. The AEAS 504 is shown to comprise a first WebSocket server 514 (EPA WebSocket server), which operates to establish the EPA WebSocket 507a with a corresponding WebSocket client on the endpoint device (Figure 9, 524), and a second WebSocket server 518 (UI WebSocket server), which operates to establish the UI WebSocket with the analyst UI 506.
[0185] The AEAS 504 maintains session state 516 for active remote access sessions. Once established, state data of the EPA WebSocket 507a is stored as part of the session state 516. Each active remote access session is stored in association with a remote access session id that is generated by the EPAS 504 to facilitate multiple sessions with the same endpoint. Hence, the UI WebSocket server 518 can perform a lookup to locate a specific EPA WebSocket based on a (precise endpoint id, session id) tuple.
[0186] The AEAS 504 acts as an appliance-based server for REMEDA functionality. It communicates with the analyst UI directly 506 over WebSocket, and proxies the endpoint agent WebSocket 507a to and from the UI WebSocket 507b. In so doing, data can be streamed between the analyst UI 506 and the endpoint agent 616 whilst a remote access session is ongoing.
[0187] Figure 9 shows a schematic function block diagram of the endpoint remote access server 502. As noted, the endpoint remote access server 502 is a component of the endpoint agent application 616 installed on the endpoint device, and is shown to comprise a communications interface in the form of a first WebSocket client 524 (EPA WebSocket client), a remote access session manager 525 and a script interpreter 528 supported by one or more software libraries. The EPA WebSocket client 524, remote access session manager 525, script interpreter 528 and supporting software libraries are integrated components of the same endpoint agent application 616, forming (part of) the software code of the application. In the present example, the script interpreter 528 has the form of a Python interpreter supported by the Python standard library 530 and, additionally, at least one bespoke Pythonlibrary 532 of function(s) that are implemented in Python and support actions that are commonly used in a remediation or analysis context. The following description refers to a single bespoke library 532 but the described library functions could equally be provided in multiple software libraries. The Python standard library 530 in combination with the cybersecurity library 532 provides various cybersecurity functions for use by an analyst that may be called in a remote access script uploaded by the analyst and run by the embedded Python interpreter 528.
[0188] The endpoint agent 616 may also provide one or more function(s) that are accessible via REMEDA but not necessarily implemented in Python. To enable such functions to be used within the REMEDA framework, an interface is exposed to executed Python scripts, which allows such scripts to call a portion (or portions) of code forming part of the endpoint agent 616 (outside of the RAS 502), which is implemented in another programming language such as C++ (e.g. the endpoint agent’s ‘native’ programming language). This is useful because there are some actions (such as isolating the endpoint device) which require access to internal context stored in memory of the endpoint agent 616, which would not be accessible within the Python execution environment provided by the Python interpreter 528. The local threat remediation component 1712 provides one or more additional cybersecurity function(s) that it is not feasible or convenient to implement in Python alone. Note, there may be other (non-remediation functions) that are implemented outside of Python, and / or other reasons to implement a given function outside of Python (e.g. for performance reasons; for example, it might be possible to implement a given function in Python, but more efficient to implement it in C++ with a Python interface). Whilst the description below focuses on remediation functions, the description applies equally to any application function of the endpoint agent 616 that is implemented by a non-Python portion (or portions) of code within the endpoint agent 616, and exposed to a Python script via a Python interface. The description also applies equally to programming languages other than Python, in which case a distinction is drawn between a first programming language in which a remote access script is coded, and a second programming language (e.g. native language of the endpoint agent 616) in which certain application function(s) may be implemented, and called via an interface in the first programming language.
[0189] In combination with the Python standard library 530, the bespoke Python library 532 and the local threat remediation component 1712 cooperate to provide a comprehensive set ofcybersecurity functions that are exposed to the Python interpreter 528, and may thus be called by an analysts via one or more function call coded in a remote access Python script. The range of cybersecurity functions may vary between different investigations, but will typically include functions in at least one of the following categories: ● Initial analysis, e.g.: o Copying and downloading file(s) for analysis o Exporting device firewall rules o Exporting defender exclusions details o Exporting device information ● Host isolation and cleanup, e.g.: o Killing running processes, e.g. by name or PID o Removing scheduled tasks o Stopping certain running services o Deleting key(s) in a Windows registry (or, more generally, an entry or item in a configuration database or file of any operating system) o Uninstalling certain applications o Isolating or unisolating an endpoint device from the network (the communication with the endpoint remote access service 502 continues to work even when the endpoint device is isolated). ● Host network and recent files review, e.g.: o Exporting active network connections list o Exporting list of recently created / modified / deleted files o Exporting browser histories for investigation timeline ● Identity cleanup, e.g.: o Listing active user sessions o Checking Active Directory lockout policyo Locking on-prem accounts using AD (active directory) lockout policy
[0190] It may be possible to implement certain functions, such as copying / downloading files (at least to some degree), exclusively in Python. The bespoke Python library 532 can be utilized to provide a simpler interface to ‘pure’ Python functionality (see below for a specific example). Any function of the endpoint agent 616 that is implemented exclusively in Python is inherently platform-independent.
[0191] The local threat remediation component 1712 is also implemented as a software library, in the sense that one or more functions thereof are exposed as Python functions to the Python interpreter 528. However, the local threat remediation component 1712 provides functionality that cannot be coded in Python only. Such functions typically include (for example) isolating or unisolating a device from the network. In the following examples, the local threat remediation component 1712 is exposed as a Python module (snson_epa) that is implemented within the endpoint agent 616 in C++ (see https: / / docs.python.org / 3 / extending / embedding.html#extending-embedded-python), which opens up wider access to the endpoint agent application 616 as a whole via the Python script interface. In this sense, the snson_epa module 1712 is said to provide an interface through to the endpoint agent 616. This module 1712 is used to implement functions such as isolating / unisolating the endpoint device, killing a process etc. Such functions are implemented under the hood in some other programming language, such as C++, but are exposed to the Python interpreter 528 as functions that may be called in Python. As indicated, there are broadly two motivations for implementing a support function outside of Python: the support function may require access to context held in internal agent data structures in memory, to which a Python library would not have access, and / or there may be significant performance benefits to implementing it in the agent itself, as the code within the agent is not interpreted. Regarding the latter, Python is typically implemented as an interpreted language that is not compiled into machine code, but rather is read and executed ‘on the fly’ by another computer program (the Python interpreter 528 in this case). By contrast, the endpoint agent’s native code would typically be pre-compiled into machine code that is executed directly on the underlying machine (the endpoint device in this case). In such cases, a support function that has been coded in C++ (or some other compiled language) would, in fact, be implemented in the endpoint agent 616 as a portion (or portion) of machine code, with an interface that allows a function call in Python (or some other interpretedlanguage) to trigger execution of that portion(s) of machine code. The local threat remediation component 1712 may be referred to more generally as a support module.
[0192] The local threat remediation component 1712 would typically need to be implemented differently for different platforms but those differences are hidden from the Python interpreter 528 (the library functions exposed to the python interpreter 528 called in the same way across platforms, even if those functions are implemented differently ‘under the hood’ between different platforms).
[0193] In the described examples, cybersecurity functionality, such as device isolation / unisolation (sometimes characterized as ‘firewall’ functions), is integrated within the endpoint agent itself 616, i.e. the endpoint agent itself operates as a form of firewall, rather than relying on an external firewall. Hence, the endpoint agent 616 can be characterized as a form of firewall, with an embedded Python interpreter 528 and WebSocket client 524. Communication can be safely maintained because it is selectively blocked so that outbound connections to the cybersecurity platform are allowed.
[0194] The remote access session manager 525 is responsible for creating (opening) and deleting (closing) remote access sessions as requested by the AEAS 504. State data 526 of each active (open) remote access session is stored and maintained by the remote access session manager locally at the endpoint device.
[0195] Among other things, the local remote access session state 516 can comprise one or more remote access scripts uploaded to the endpoint device by an analyst during an active remote access session. Such scripts are received via the EPA WebSocket 507a and written to the local session state 526 by the EPA WebSocket client 524.
[0196] In contrast to more conventional remote access agents, the remote access server 502 does not provide terminal emulation. Rather, it provides a way to run Python code pushed from the appliance, and returns all data written to stdout and stderr. A script written to the local session state 526 is accessible to the script interpreter 528, which runs the script and writes a sequence of script outputs to the session state 526. The script outputs are written to stdout (stdout is a form of output object, specifically a memory buffer in the script's context) which is then streamed over the network to the endpoint access service. Script outputs are streamed to the AEAS 504 via the EPA WebSocket 507a as the script is running (rather than waiting for the script to terminate and then transmitting a copy of the final stdout object),implying that a first script output may be sent to the AEAS 504 before a second script output is written to stdout. For example, each script output may be streamed to the AEAS 504 as soon as it is written to stdout (real-time) or script outputs may be streamed in relatively small batches at reasonably regular intervals (pseudo-real-time). For example, outputs may be streamed line-by-line, or individual lines may be streamed in multiple portions (e.g. each containing a single character or a few characters). Those script outputs are, in turn, proxied for rendering to the analyst UI 506 via the UI WebSocket 507b, in real-time or pseudo-real- time, allowing the analyst to see the script outputs as they are generated in real-time or pseudo-real-time. In Python, a script output is generally generated in response to a print() function call in the script, or an explicit write operation on stdout. Hence, the analyst coding the script can specify a point or points in the script at which a script output is generated, which will then be streamed back to them as the script is run by the endpoint agent 616. Exceptions are written to a separate stderr object, which is streamed back to the analyst console in the same manner.
[0197] The script runs on an input object stdin (input memory buffer in the script’s context), which is closed, precluding interactive programs.
[0198] There are some common analysis / remediation tasks which are possible using only Python, but where the built-in Python modules can be further simplified. For example, the Python code below calculates the MD5 of a file and writes it to stdout using only built-in modules: import hashlib, os def generate_file_md5(filename, blocksize=2**20): m = hashlib.md5() with open( os.path.join(filename) , "rb" ) as f: while True: buf = f.read(blocksize) if not buf: break m.update( buf ) return m.hexdigest()print(generate_file_md5(' / tmp / example.txt'))
[0199] As this is likely to be a common activity, it can be greatly simplified using the bespoke Python library 532. Using this support module, the above code might instead look something like the following: import snson_analysis print(snson_analysis.hash_file('md5', ' / tmp / example.txt')) Here, “hash_file()” is a function coded in the bespoke Python library 532 that is exposed to the Python interpreter 528.
[0200] A key set of functions provided in the snson_epa module 1712 are those of as isolating / unisolating a device from the network. As noted, these functions are exposed in Python, but implemented in the endpoint agent 616 in C++. Python code using this module 1712 to isolate the device might look something like the following: import snson_epa snson_epa.isolate_device()
[0201] Other host isolation / cleanup functions, such as those mentioned above, may be implemented in the snson_epa module 1712 in a similar manner.
[0202] In addition, an analyst has the option of uploading files to the endpoint device to a location accessible to a script, such as a temporary directory associated with the remote access session. Uploading a file involves streaming file data selected by a user in the analyst UI 506 over the UI WebSocket 507b to the AEAS 504. This data is then proxied to the corresponding EPA WebSocket507a, received by the endpoint agent, and written to disk in a pre-determined location (such as a temporary directory associated with the remote access session) that the analyst can then access via the script interface. File upload functionality may, for example, be provided using the HTML5 File API (https: / / developer.mozilla.org / en-allow file data to be processed in client-side JavaScript (within the analyst UI 506, which assumes the client role in this set-up) and transmitted directly over the UI WebSocket 507b to the AEAS 504, and then proxied to the EPA WebSocket client 524 over the EPA WebSocket 507a.
[0203] The analyst can also transfer a file from the endpoint device to their console (referred to as a ‘download’ for conciseness). File downloads are coded in the Python script. When a file download is instigated in the script, data is read from the local file in chunks, and thesechunks are transmitted over EPA WebSocket 507a to the access service 504, which is then responsible for putting them back together. To implement this functionality, a download function may be coded into the bespoke python library 532 (or in the snson_epa module 1712 if it cannot be coded purely in Python, or is more efficient to code outside of Python).
[0204] Returning briefly to Figure 8, downloaded files received by the AEAS 504 are stored in a file holding pen 521 within a file storage object 520. Each downloaded file is assigned a unique, pre-authenticated URL that is communicated to the analyst UI 506, allowing the analyst to retrieve a downloaded file from the file holding pen 521. A level of processing may be applied to files downloaded from the endpoint device before they are stored to the file holding pen 521, in order to ‘neuter’ any malicious code that may be present. Standard neutering techniques may be used. REMEDA messaging
[0205] REMEDA messaging between the endpoint agent 616, the AEAS 504 and the analyst UI 506 may be implemented using a predefined message schema. For example, a protobuf schema may be used, with each message delivered within an envelope message. The clients and server will expect to be able to decode each message as an envelope and then retrieve the contents from within, which will be one of multiple different message types. The inner message type indicates what the action is.
[0206] The envelope uses a shared session_id, which is used to identify which session the envelope is for. To allow the AEAS 504 to manage multiple sessions with a single endpoint, each message back and forth is tagged with the applicable session id. To start a session on the endpoint, the AEAS 504 sends a start message (Figure 10, T2) with a unique session_id picked by AEAS 504. This session_id will from now on be required for every message related to this session. The envelope also has an authentication token that is required in messages from the UI 506 to the AEAS 504.
[0207] Figure 10 shows a signalling flow for an example remote access session. Initially, and endpoint agent 616 registers (T1) its unique endpoint id with the AEAS 504. At this point, an EPA WebSocket 507a between the endpoint agent 616 and the AEAS 504 has been established, and the endpoint agent 616 is available for a remote access session. A remote access session is instigated by the analyst frontend 506 sending (T2) a session initiation request comprising the endpoint id to the AEAS 504. The session initiation request is sentvia a UI WebSocket 507b, which has been established between the AEAS 504 and the analyst frontend 506 by this point. In response, the AEAS 504 generates a session id and sends (T3) a session initiation request with the session id to the endpoint agent 616. The endpoint agent responds (T4) with an acknowledgement message that incudes the session id and the address of the temporary directory associated with the now-active remote access session. The session id and temporary directory address are passed (T5) from the AEAS 504 to the analyst frontend 506 in a response to the initial session initiation request.
[0208] The endpoint can provide a temporary directory for the session, which is expected to be cleaned up after the session ends (‘ / tmp / xyz’ in the example of Figure 10). The AEAS 504 knows that files and scripts that are uploaded to the endpoint agent 616 will be saved to this directory.
[0209] Now that a remote access session has been opened, the analyst frontend 506 can upload (T6) a Python remote access script to the AEAS 504 via the Ui WebSocket 507b, which is proxied (T7) to the endpoint agent 616 via the EPA WebSocket 507a. The script is uploaded in a message containing the session id.
[0210] Reference sign T8 denotes multiple streaming steps. This example considered a script whose main function call is print(‘Hello’) This trivial script is considered purely for the sake of illustration. This script, when run, causes the character string ‘Hello’ to be written to as a single line in stdout. In this particular example, that single line is divided into portions of up to two characters that are streamed back to the AEAS 504 via the EPA WebSocket 507a in three sequential messages. Those messages are, in turn, proxied back to the analyst frontend 506 for rendering in multiple steps denoted by reference sign T9. The analyst thus sees the output being rendered as ‘he’ then ‘ll’ and finally ‘o’ in essentially real-time. Each message may additionally include a flag to indicate whether the script is still running.
[0211] At step T10, a file is uploaded to the AEAS 504 from the analyst frontend 506 using the HTML file API (see above), which is passed to the endpoint agent 616 at step T11.
[0212] At step T12, the endpoint agent 616 begins transmitting a file to the AEAS 504. This is referred to at various places herein as a “download” for convenience, but it is important to note that the transmission of the file is instigated by the endpoint agent 616 in response to acommand (a call to the appropriate library function) that the analyst has included in the remote access script (it is not something that has to be requested by the AEAS 504). This exploits the bi-directional communication functionality of the EPA WebSocket 507a. The ‘downloaded’ file is streamed to the AEAS 504 via the EPA WebSocket 507a in multiple chunks (T12, T14). Those chunks are re-assembled by the AEAS 504 and the re-assembled file is stored to an addressable storage location in the file holding pen 512 (post-neutering – see above). At step T16, the address of the file is provided to the analyst frontend 506 in the form of a pre-authenticated URL, which can in turn be used to download the file from the file holding pen 521 directly.
[0213] Finally, at T20, the analyst frontend 506 sends a session termination request with the session id to the AEAS 504, which in turn sends a corresponding termination request to the endpoint agent 616, causing the remote access session to be terminated.
[0214] The AEAS 504 server relies on an authentication and authorisation mechanism implemented elsewhere within the platform: authentication and authorization tasks are delegated to the platform API 508 by passing cookie authorisation tokens (or other authentication tokens) received over WebSocket from the analyst UI 506 to a dedicated platform API endpoint for an authentication decision.
[0215] Authentication between the analyst UI 506 and the AEAS 504 is managed as follows. The analyst UI 506 requests an authentication (auth) token from the platform API 508, which has an expiry time. This token is then included in each request to AEAS 504 and can be validated without communicating to the platform API 508 (through public key cryptography). When the auth token expires (or, if possible, shortly before it is due to expire), the analyst UI 506 simply requests a new token from the platform API 508. Any suitable authentication token technology can be used, one example being JSON Web Token (JWT).
[0216] All REMEDA actions performed by users are stored by the UI WebSocket server 518 in an audit log within the API database 510. The audit logs are stored in the API database 510, and can be viewed from within the platform, via the platform API 508. All audit entries contain a timestamp, an associated analyst username and action-specific context. Audited actions include script interface commands (script contents stored as context), file upload attempts (success status, filename and hash stored as context) and file download attempts (success status, filename and hash stored as context).
[0217] As noted above, the script interface can operate in different modes. Examples modes considered below are: Python script; Cross-platform commands; PowerShell; Bash. The different modes all use Python under the hood: cross-platform commands are exposed as a Python object-oriented (OO) interface; PowerShell and Bash commands are embedded within a Python script which executes them using the platform-specific command interpreter. Python script mode
[0218] When using this mode, users are expected to enter valid Python code which will then be sent to the endpoint agent and executed by the embedded Python interpreter 528. The output of the script is sent back to the user’s console in the manner described above.
[0219] Figure 11A provides a schematic illustration of a possible graphical Web interface rendered by the analyst UI 506 at the analyst device 505, when the Python script mode is selected.
[0220] A scripting region 702 is rendered, which is an editable, multi-line field, in which the analyst can enter and edit a remote access script in Python.
[0221] A selectable “run” option 705 causes the script to be uploaded to the endpoint device as per the above description.
[0222] An output region 704 is rendered, in which the script outputs streamed back to the analyst’s console are rendered in real-time or pseudo-real-time.
[0223] A selectable ‘run’ 703 option is displayed. In response to the user selecting the run option 703, the script in the scripting region 702 is uploaded to the endpoint device and run by the endpoint agent 616 automatically on receipt. In this manner, a combined ‘upload-and- run action’ is provided: in response to a single user input, the script in the scripting region 702 is packaged and uploaded to the endpoint device and automatically run by the Python interpreter 528 embedded in the endpoint agent 616, and the output(s) of the script are automatically returned and displayed in the output region 704 (by streaming the stdout and stderr objects back to the analyst’s console via WebSocket). This stands in contrast to previous solutions, where a user would typically be required to first upload the script (using a generic file transfer function) and manually cause the uploaded script to be run. The single user input could be the selection of the run option 703, or some other single run input such as a keyboard shortcut, gesture input etc. In any event, both the transmission and execution ofthe script are triggered by the same single run input. In response to the single run input, the script output(s) are streamed back to the analyst console and displayed automatically, without the analyst having to manually retrieve them from the endpoint device.
[0224] Figure 11B shows a second example graphical user interface. In addition to the visual elements described above (rendered in a slightly different visual arrangement), a file upload region 706 and file download region 708 are rendered. The file upload region 706 allows a user to select a file or files to be uploaded to the endpoint device. The file download region identifies files received from the endpoint device and currently available in the file holding pen 521.
[0225] When a user uploads a file to the endpoint device, it is stored in the temporary directory associated with the remote access session. The user does not need to keep track of the location of uploaded files, as this is handled automatically. The user can simply reference their uploaded file(s) in the script, and the transfer and management of those files is handled automatically by the analyst UI 506 in cooperation with the endpoint agent 616. Cross-platform commands
[0226] The “Cross-platform commands” mode is a cross-platform command interface that aims to mimic typical Unix-style commands but using a Python object-oriented interface. The intention is that this is for use when a more interactive experience is desired, when a user doesn’t know exactly all the things they want to do in advance. For example, a user would use this feature if they wanted to manually explore the filesystem and then perform different actions depending on what they find. This feature is entirely implemented in Python, but the way the object-oriented interface is exposed makes it easier and more concise to use than writing a Python program, and gives the user a very similar experience to using Bash / Linux standard tools. The main advantage of this feature is that it is entirely cross-platform. This means both that it requires us only to implement it once, and also that it allows users to issue exactly the same commands (interpreted as Python) on Linux, Windows and macOS to achieve identical behaviour on all platforms. The power of this is best illustrated with an example.
[0227] Consider a user who wants to look at files in a directory, filter for ones with names containing zero or more characters followed by the word “service” and then count how many of these files there are. On Windows, using PowerShell, the user might do it this way:Write-Host ( Get-ChildItem -Path 'C:\Windows\temp' -Filter "*service"| Measure-Object ).Count
[0228] On GNU Linux and macOS, using Bash and standard tools, they might do it this way: cd / tmp ls | grep -E '.*service' | wc -l
[0229] What the user needs to write to achieve the same thing is very different on Windows and Linux / macOS. Windows is not the only problem, however. Although the commands look the same on macOS and Linux and both use Bash, the standard tools are not identical. Subtle differences exist which can yield different results for the same command. For example, the following command is valid on both GNU Linux and macOS: “thisis” | grep -Eo '.+?is'
[0230] However, the result on the two platforms is different. On GNU Linux the output is “thisis” while on macOS the output is “this”. This is because on macOS, grep performs lazy searches by default but on GNU Linux it doesn’t without the additional option, -P. As a result, users working across these platforms have to significantly change how they interact with the systems to perform identical tasks. This is time-consuming and error-prone. “Cross- platform commands” is a solution to this. By providing a suitable set of Python library functions, a command resembling the following would be all that is needed to perform the task described above on Windows, Linux and macOS: cd(tempdir()) ls() | grep(pattern='.*service') | wc()
[0231] The cross-platform mode overcomes this problem by providing a set of platform- independent, command-like Python library functions that are called in the same way irrespective of the underlying platform. Fundamentally, this is the same as the Python script mode; the only difference is that the library functions supporting the cross-platform mode are designed to provide a more interactive experience (rather than coding a multi-line script and then uploading it, the user would typically upload Python code one line at a time; the cross- platform library functions are designed to support this more interactive modality). For example, the bespoke Python library 532 might provide cross-platform “cd()”, “ls()”, “grep()” and “wc()” functions that may be called in the following manner: cd(tempdir()) ls() | grep(pattern='.*service') | wc()PowerShell
[0232] The PowerShell mode uses a Python script template which writes a PowerShell script the user supplies to disk and then invokes the PowerShell interpreter to execute it. The script then captures the output from PowerShell and writes it to the user’s console. Bash
[0233] The Bash mode functions in a very similar way to the PowerShell mode, except instead of invoking PowerShell it invokes Bash. Digital Estate Overview
[0234] Figure 11C shows an example of a digital estate view within the analyst GUI.
[0235] In order for actions to be performed on devices, such as running commands on them or isolating them, users need to have a way of seeing the available devices and identifying those they want to perform actions on. The feature that will provide this functionality is called the “Digital Estate Overview”. It allows users to view and filter a list of endpoint agents the platform knows about, and select individual devices on which actions are to be performed. An entry 722 corresponding to an endpoint device is selectable to commence a remote access session, or to return to an existing remote access session within the GUI.
[0236] The above examples consider a graphical scripting interface that allows a skilled analyst to manually write and upload scripts to an endpoint device. However, the above architecture can also be used to implement automated remediation, in which Python scripts are generated automatically. In this case, some automated software agent assumes the role of analyst, with the ability to automatically generate Python scripts, and upload and run those scripts on endpoint devices.
[0237] Among other things, the described REMEDA architecture supports the use cases listed in Table 1.TABLE 1: REMEDA use cases.
[0238] It will be appreciated that the examples described above are illustrative rather than exhaustive. In general, the functional components described above can be implemented in one or more computing devices at one or more locations within a localized or distributed computer system. A computer system comprises computing hardware which may be configured to execute any of the steps or functions taught herein. The term computing hardware encompasses any form / combination of hardware configured to execute steps or functions taught herein. Such computing hardware may comprise one or more processors, which may be programmable or non-programmable, or a combination of programmable and non-programmable hardware may be used. Examples of suitable programmable processors include general purpose processors based on an instruction set architecture, such as CPUs,GPUs / accelerator processors etc. Such general-purpose processors typically execute computer readable instructions held in memory coupled to the processor and carry out the relevant steps in accordance with those instructions. Other forms of programmable processors include field programmable gate arrays (FPGAs) having a circuit configuration programmable through circuit description code. Examples of non-programmable processors include application specific integrated circuits (ASICs). Code, instructions etc. may be stored as appropriate on transitory or non-transitory media (examples of the latter including solid state, magnetic and optical storage device(s) and the like).
Claims
Claims 1. An endpoint agent application for establishing, when executed on an endpoint device, a remote access session between the endpoint device and an endpoint access service; wherein the endpoint agent application has integrated therein: a session manager configured to establish a remote access session between the endpoint device and the endpoint access service; a communication interface configured to receive from the endpoint access service in the remote access session a remote access script to be run on the endpoint device; a script interpreter configured to generate, in an output object local to the client device, one or more script outputs, by running the remote access script on the endpoint device, wherein the communication interface is configured to automatically transmit to the endpoint access service a copy of all script outputs written to the output object; and at least one software library configured to provide a cybersecurity function, wherein the script interpreter is configured to instigate the cybersecurity function of the software library responsive to a function call thereto in the remote access script.
2. The endpoint agent of claim 1, wherein the at least one software library is configured to provide one or more of the following cybersecurity functions: a process termination function to terminate a process running on the endpoint device, a registry access function to access an item contained in an operating system configuration database or file of the endpoint device, a file deletion function to delete a file from a local filesystem of the endpoint device, a file transfer function to cause a copy of a file to be transmitted from a local filesystem of the endpoint device to the endpoint agent access service, an isolate function to cause the endpoint device to be isolated from a communication network, a reverse isolate function to restore access to a communication network.
3. The endpoint agent of any preceding claim, wherein the script interpreter is configured to interpret the remote access script in a platform-independent programming language.
4. The endpoint agent of claim 3, wherein the platform-independent programming language is Python.
5. The endpoint agent of any preceding claim, wherein the communication interface is configured to establish a full-duplex communication channel between the endpoint device and the remote access service, wherein the remote access script is received from the remote access service via the full-duplex communication channel and said copy of all script outputs is transmitted to the remote access service via the full-duplex communication channel.
6. The endpoint agent of claim 5, wherein the remote access script is configured to cause the script interpreter to sequentially write a series of script outputs to the output object, and the communication interface is configured to stream the series of script outputs from the output object to the remote access service via the full-duplex communication channel in a plurality of messages as the remote access script is running on the endpoint device.
7. The endpoint agent of claim 5 or 6, wherein the full-duplex communication channel is a WebSocket channel.
8. The endpoint agent of any preceding claim, wherein the communication interface is configured to receive from the endpoint access service a script input and store the script input in an input object local to the endpoint device accessible to the remote access script.
9. The endpoint agent of any preceding claim, wherein the at least one software library includes: a portion of compiled machine code embodying a cybersecurity function, and an interface exposing the at least one cybersecurity function to the script interpreter; wherein the script interpreter is configured to interpret the remote access script in an interpreted programming language, and cause said portion of compiled machine code to be executed responsive to a function call in the remote access script.
10. The endpoint agent of claim 9, wherein the portion of compiled machine code has access to context data held in an internal in-memory data structure of the endpoint agent for carrying out the cybersecurity function.
11. The endpoint agent of claim 9 or 10, wherein the at least one software library comprises a first software library that includes the portion of compiled machine code, and a second software library implemented in the interpreted programming language and configured to provide a second cybersecurity function exposed to the script interpreter.
12. The endpoint agent of any preceding claim, wherein the at least one cybersecurity function comprises an isolate function callable to isolate the endpoint device from a network, wherein the isolate function is implemented by selectively blocking communications between the network and the endpoint device, whilst maintaining communication between the endpoint agent and the endpoint access service.
13. A computer system for use by an analyst in establishing a remote access session with an endpoint device, the computer system comprising: one or more processors configured to generate display data for rendering a graphical user interface, the graphical user interface configured to provide a session establishment mechanism for establishing the remote access session with the endpoint device; wherein the graphical user interface is configured to provide a remediation view having: an input region editable to code a remote access script, and an output region; wherein the one or more processors are configured, responsive to a single run input, to cause: (i) the remote access script to be transmitted to and executed by an endpoint agent on the endpoint device, wherein the endpoint agent is caused to automatically return one or more script outputs generated by running the remote access script locally at the endpoint device, and (ii) the one or more script outputs returned by the endpoint agent to be rendered in the output region of the remediation view.
14. The computer system of claim 13, wherein the remote access script is configured to cause the endpoint agent to sequentially write a series of script outputs to an output object local to the endpoint device, and stream the series of script outputs from the output storage location to the remote access service via the full-duplex communication channel in a pluralityof messages as the remote access script is running on the endpoint device, wherein the series of script outputs are rendered sequentially in the output region as the remote access script is running on the endpoint device.
15. The computer system of claim 13 or 14, wherein the remediation view has at least one of: an upload region for selecting one or more files to be transferred to an input storage location of the endpoint device, and a download region for accessing one or more files transferred from an output storage location of the endpoint device.
16. The computer system of claim 15, comprising: an analyst console comprising at least one display configured to render the graphical user interface to the analyst; and an endpoint access service configured to establish a remote access session with the endpoint device, wherein the analyst console is configured to cause the remote access script to be executed by the endpoint agent by transmitting the remote access script to the endpoint access service, and the endpoint access service is configured to receive the one or more script outputs from the endpoint device and transmit them to the analyst console.
17. The computer system of claim 16, wherein the analyst console is configured to establish a first WebSocket with the endpoint access service, and the endpoint access service is configured to establish a second WebSocket with the endpoint device, wherein the remote access script is transmitted to the endpoint access service via the first WebSocket and proxied to the endpoint device via the second WebSocket, and wherein the one or more script outputs are received via the second WebSocket, and proxied to the analyst console via the first WebSocket.
18. Transitory or non-transitory media embodying executable program instructions configured so as, when executed on one or more computers, to implement the endpoint agent of any of claims 1 to 12 or the system functionality of any of claims 13 to 17.
Citation Information
Patent Citations
Compliance Management in a Local Network
US20190018965A1
Cited By
Systems and methods for real-time detection and mitigation of malicious electronic communications
US12452296B1
Systems and methods for real-time detection and mitigation of malicious electronic communications
US12641115B1