Improved cyber threat detection
The ANN-based method transforms unstructured CTI reports into machine-readable formats, addressing inefficiencies in existing systems by enhancing accuracy and reducing costs through automated threat intelligence processing.
Patent Information
- Application Number
- PCT/GB2025/051290
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-12
- Filing Date
- 2025-06-12
- Publication Date
- 2025-12-18
AI Technical Summary
Existing cyber threat detection systems face inefficiencies in processing unstructured cyber threat intelligence (CTI) reports, leading to inaccurate threat determination, high processing costs, and unreliable outputs due to limitations in current tools like Named Entity Recognition (NER) algorithms and Large Language Model (LLM) chatbots, which struggle with cybersecurity-specific terms and require significant human interaction.
A computer-implemented method using an Artificial Neural Network (ANN) to transform unstructured CTI reports into machine-readable threat intelligence (MRTI) reports, compliant with STIX, TAXII, and MITRE ATT&CK protocols, by extracting entities, determining relationships, and generating graphical or natural language representations, with customizable machine learning models for enhanced accuracy.
The method provides efficient and accurate cyber threat intelligence management by automating the processing of CTI data, reducing human intervention, and improving the reliability of threat detection, while maintaining compliance with industry standards.
Smart Images

Figure GB2025051290_18122025_PF_FP_ABST
Abstract
Description
[0001] IMPROVED CYBER THREAT DETECTION
[0002] The invention generally relates to relates to method and / or system for processing data e.g. cyber threat intelligence (CTI) data e.g. reports. More specifically, the invention relates to the analysis of CTI data for determining the security relevance and / or threat level, for generating a report that determines at least one of a level of risk and / or proposes an action in response to a determined threat. The invention can also reside in at least one of preparing training data, enhancing training data, a machine -learning model and training an artificial neural network (ANN) for implementing the method and / or system. The invention can be adapted for receiving CTI data and producing a machine readable threat intelligence (MRTI) report that comply with the OASIS Cyber Threat Intelligence (CTI) Technical Committee (TC) protocols e.g. reports that comply with at least one of: STIX (Structured Threat Information Expression) e.g. STIX2.1; TAXII (Trusted Automated Exchange of Indicator Information); and CybOX (Cyber Observable Expression).
[0003] BACKGROUND
[0004] Detecting cyber threats is crucial for online security. A key part of this detection involves threat intelligence information, which identify entities linked to suspicious activities, and can be derived, for example, from cyber threat intelligence (CTI) reports e.g. offline reports stored in formats like Word (RTM) documents, PDFs, and blogs. Determining the relevance of CTI reports is time-consuming, expensive, and impractical for all but the most advanced operations. The information and reports can be unstructured.
[0005] Known tools and services that identify malicious websites, computers, domains, etc often use multiple independent services to detect security breaches. These services can identify threats by IP address, domain name, URL, email address, file hashes, etc. The list of known orpotential threats can be significant and require impractical levels of human interaction or expensive processing cost. Network owners and security administrators use tools to detect when their users access malicious sites or computers . However, these tools need to process logs from various computers, which can be in many different formats, they rely on specific log format connectors struggle to keep up with changes from updates or new systems, making it difficult to maintain accuracy. Moreover, reports on threat intelligence can include incorrect information, causing harmless events to be wrongly flagged as threats, or vice-versa. This leads to unnecessary investigations for the system operators due to these false alarms - or worse, a missed threat.
[0006] Known solutions, like Named Entity Recognition (NER) algorithms and Large Language Model (LLM) chatbots, are not effective enough. NER algorithms struggle with cybersecurity-specific terms. Standard algorithms fail to identify specific cybersecurity threats like malware or tools. LLM chatbots, like those from OpenAI (RTM), also have limitations e.g. confidentiality, performance consistency, and accuracy. Moreover, their output can be inconsistent and unreliable. It is against this background that the present invention has been made. This invention results from efforts to overcome the problems of known methods and systems for evaluating cyber security, which suffer from inaccurate and sub-optimal threat determination. Other aims of the invention will be apparent from the following description. SUMMARY
[0007] The invention generally relates to a computer-implemented method of processing source data for determining a cyber threat, using an ANN, and generating a machine readable output dataset for producing in a natural language and / or graphical representation of the determined cyber threat. The method can be automated to transform unstructured CTI reports into at least a machine readable threat intelligence (MRTI) report that complies with at least one of (i) the OASIS Cyber Threat Intelligence (CTI) Technical Committee (TC) protocols, (ii) STIX (Structured Threat Information Expression) e g. STIX2.1, (iii) TAXII, and (iv) the MITRE ATT&CK (RTM) e.g. ATT&CK v!5.1, for enabling more efficient and accurate cyber threat intelligence management. The ANN used from training can use BERT transformers.
[0008] In one example, there resides a computer-implemented method of processing source data for determining a cyber threat, the method comprising: using an input dataset, said input dataset comprising the source data to be analysed for determining a cyber threat; processing the input dataset for (i) extracting entities and / or (ii) determining relationships between entities; and generating a machine readable output dataset for producing in a natural language and / or graphical representation of the determined cyber threat.
[0009] The method can further include: transforming the input dataset to produce a formatted dataset; and processing the formatted dataset for labelling indicators with a set of labels representing data objects to produce an object dataset, wherein entities and / or relationships are determined therefrom.
[0010] The input dataset can be at least one of received, retrieved and imported, and (i) converted from at least one of HTML, PDF and TXT format, and / or (ii) configured in a JSON file containing UTF-8 encoded text, before annotating and / or refining.
[0011] The method can use a set of predictive algorithms, which are combined to automatically process and structure cyber threat intelligence data e.g. according to the STIX standard. The method can be operated via a user interface. The user interface can enable an operator to adjust the performance metncs of ANN e.g. the machine learning models, allowing he operator to optimise the machine performance through re-labelling or re-training. For training purposes, labelled seed training data set for the cybersecurity domain was obtained. Optionally, the labelled seed training data was scaled up.
[0012] The teaching herein provides functionality that can process data e.g. CTI data report and / or a training dataset aligned to the STIX standard, using an ANN to at least one of (i) enlarge a training dataset for model training, and (ii) run different models and / or algorithms of different classes e.g. predictive transformer neural networks, pattern recognition algorithms, word look up databases. The examples herein overcome problems associated with known tools that identify malicious activity, which rely on inefficient amounts of data searching that make it difficult to maintain accuracy in their reports. The structured examples herein can effectively parse data and determine threats of relevance to the end-user. Moreover, the examples herein overcome the challenges of processing inefficient amounts of data to automatically and efficiently MRTI reports.
[0013] Producing the object dataset can includes annotating the data with: at least one of the following labels: artifact; autonomous system; directory; domain name; email address; email message; file; IPv4 Address; IPv6 Address; MAC Address; mutex; network traffic; process; software; URL; user account; Windows Registry Key; and X.509 Certificate; and / or known associated threat actors.
[0014] The object dataset can be processed to produce supplemented dataset, the method further comprising: expanding the formatted dataset and / or object dataset using an algorithm; and / or parsing overlapping annotated and / or labelled entities by keeping the longest entity.
[0015] Processing the object dataset and / or the supplemented dataset can further comprise at least one of: parsing the extracted entities; attack pattern matching, comprising: determining a spatial relationship between data objects and / or entities; and / or parsing the data objects and / or entities to identify sub-units thereof, and determining a spatial relationship between data objects and / or entities for determining relationships between data objects and / or entities extracted from the input dataset.
[0016] At least one of the transforming the input dataset, processing the formatted dataset and processing the object dataset can comprise using an artificial neural-network (ANN) using predictive models. Generating a machine readable output dataset can further comprise generating a machine readable file for graphical and / or document visualisation of the cyber threat.
[0017] In another example, there resides a machine -learning model for processing source data comprising cyber threat information, or information derived therefrom, and for determining a cyber threat level, for the use in the method claimed herein, the model comprising: a first stage comprising: determining pattern matching and / or comparisons of source data to produce an input dataset with data objects and / or entities to be evaluated; and / or supplementing the input dataset for training and evaluation for producing the formatted dataset and / or object dataset; and additionally or alternatively a second stage comprising: processing the source data and determine a spatial location for data objects and / or entities identified within the source data; and determining, at least one of a spatial relationship, patterns, and derivative patterns between derivatives of the data objects and / or entities.
[0018] The first stage can receive cyber threat intelligence (CTI) data any preceding claim and outputs the formatted dataset and / or object dataset; and / or the second stage further comprises receiving the formatted dataset and / or the object dataset and generating structured threat information report for processing as machine readable threat intelligence.
[0019] In another example, there resides a computer-implemented method of training a machine -learning model for processing cyber threat information, or information derived therefrom, and for determining a cyber threat level, in particular the machine -learning model as claimed herein, comprising: a first stage comprising: receiving the source data, and generating supplementary data therefrom using a string -searching algorithm; and additionally or alternatively a second stage comprising: receiving the formatted dataset and / or object dataset, generating embeddings, and training and comparing a plurality of ANNs and retaining the ANN that achieved the greatest performance. The longest of overlapping entities generated during embeddings can be retained. In another example, there resides a computer-implemented method of generating a training dataset for a machine -learning model, in particular the machine -learning model as claimed herein, comprising: receiving the source data, and generating supplementary data therefrom using a string -searching algorithm.
[0020] In another example, there resides a data processing system comprising means for carrying out the methods claimed herein.
[0021] In another example, there resides a computer system for processing cyber threat information, or information derived therefrom, and for determining a cyber threat level, the system comprising: a processor; and memory including executable instructions that, as a result of execution by the processor, causes the system to perform the method as claimed herein.
[0022] In another example, there resides a non-transitoiy computer-readable storage medium having stored thereon executable instructions that, as a result of being executed by a processor of a computer system, cause the computer system to perform the computer-implemented method as claimed herein.
[0023] In light of the teaching of the present invention, the skilled person would appreciate that aspects of the invention were interchangeable and transferrable between the aspects described herein, and can be combined to provide improved aspects of the invention. Further aspects of the invention will be appreciated from the following description.
[0024] DESCRIPTION OF THE FIGURES
[0025] In order that the invention can be more readily understood reference is made, by way of example, to the remaining drawings, in which:
[0026] Figure 1 is a simple flowchart of a known example of source data being processed to generate a report;
[0027] Figure 2 is a schematic of a system including a graphical user interface (GUI) and the system including a machine for implementing the processes taught herein; and
[0028] Figure 3 is an alternative arrangement of the processes of the machine of Figure 2.
[0029] Like reference numerals refer to like features.
[0030] DETAILED DESCRIPTION
[0031] Known art
[0032] Figure 1 illustrates an example of a known process s 10 in which data s 12 is passed through an LLM to produce a look-up result sl6. The process sl4 typically involves applying Named Entity Recognition (NER) algorithms e.g. hitps: / 7spacy.io / api / entityrecognizer to cyber threat intelligence (CTI), and then using zero-shot Large Language Model chatbots e.g. htips: / 7chat.openai.com / to process the data therein. Known systems use off-the-shelf Artificial Intelligence (Al) tools to apply pattern matching techniques and look-ups to identify key-words that indicate a level of threat. NERs and LLMs can be configured for repeatable and / or generic tasks e.g. benchmark tests, generic document generation, etc - but are not adaptable for special applications or customisation for specific fields of use i.e. cybersecurity. The inventors have evaluated known technologies and considered the performance to be deficient because, for example, the patterns learned by the neural networks were unable to accommodate the complicated combination of features and relationships required for determining a cybersecurity risk. To be clear, a standard implementation of un-modified algorithms e.g. prep- packaged solutions with known machine learning libraries e.g. spaCy, using the Python language to recognise certain key-words e.g. person, organisation etc was inadequate at determining the relevance of the likes of countries, companies, people etc, which were insufficiently recognised and their association with malware, tools, infrastructure etc. could not be determined.
[0033] Known LLM chatbots (e g. https: / / chat.openai.com / ) fail to recognise and / or draw associations because of, for example, a limited access to data as a consequence of confidentiality agreements, which results in unreliable performance variation. Many of the best LLM chatbots are privately hosted by organisations such as OpenAI, which creates immovable barriers to adoption for many organisations which process confidential cyber threat intelligence, or who require offline software solutions. Further, while LLM based software solutions can respond well to human inputs and provide a dynamic range of responses the consequence is that analysing CTI data produces inconsistent results - which is problematic when applied to sensitive applications such as defensive cyber operations or critical infrastructure i.e. mission-critical environments.
[0034] Moreover, the output from known LLM systems requires additional processing - by machine and / or human filtering - before a reliable or tractable report can be generated. This requires significant processing cost and is inefficient. Customising known LLMs is beyond the ability of many and access to the significant proprietary engineering of known LLMs e.g. Chat GPT4 is limited if not impossible. Overall, therefore, known systems require expensive and resource heavy extract of data and / or analysis of data to determine a cyber risk.
[0035] Improved system & method
[0036] Structure - general
[0037] The examples taught herein are described with reference to Figure 2, and Figure 3 in part, wherein a system 100 includes a graphical user interface (GUI) 102 and a machine 104. The GUI provides a portal input 106 for selecting data to be analysed, and a report output 108 for use e.g. extracting a report that provides analysis on the data. The report 108 can be presented on the GUI as a graphical 110a or document-based 110b display. The report output 108 can generate machine readable data representing the analysis of the analysed. The report 108 can be disseminated to other machines for determining and / or further analysing risk. A customisation panel 112 can be provided on the GUI to enable an operator of the system to adjust operating parameters of the machine 104 e.g. customise the analysis.
[0038] The machine 104 has a data input 120 for receiving the ‘source data’ dataset D200 from the input portal 106 e.g. the data selected for analysis by the portal input 102. Data selected for analysis as the data input 120 is used e.g. at least one of received, retrieved, imported, transformed and acquired - so it can be analysed 122 and optionally balanced using a balancer 124 that produces data that can be fed to the report output 108. The source data can be processed and / or analysed by an artificial neural network (ANN) 126, which can include at least two components, which can be separate and distinct i.e. a processing ANN 126a and an analysing ANN 126.
[0039] The machine 104 is configured to execute processes that transform the source data 120 extracted via the portal input 106 before using an analyser 122 to generate data for an output report 108. The machine is configured to process source data 120 e.g. a CTI report e.g. STIX 2.1 report for determining a cyber threat. The machine is configured e.g. via the GUI 102 to receive the source data as an input dataset D200 and at least one of (i) extract entities e g. a STIX domain object and / or (ii) determine relationships e.g. STIX relationship objects, between entities Post-processing, the machine can generate a machine readable output dataset e g. a machine readable file for producing a report in graph or written document detailing a determined cyber threat.
[0040] The balancer 124 is optional. Via the GUI 102, an operator can customise 112 a balancing 124 mechanism of the machine 104 to at least one of (i) set bespoke search and / or analysis criteria, (ii) select the processes to be applied to the source data 120, and (iii) set the order in which the processes are applied to the source data e.g. series or parallel modelling.
[0041] The processes can comprise: pre-processing sl50, in which the source data 120 received via the portal 106 is converted into a usable format i.e. a formatted dataset D210; labelling sl52, in which the formatted dataset D210is labelled with indicators using a set of labels representing data objects to produce an object dataset D220; and supplementing sl54 the object dataset D220 to produce supplemented data D230, and at least one of (i) extracting entities from the object dataset and / or (ii) determining relationships between entities.
[0042] The aforementioned processes and / or the supplemented data D230 can be subjected to machine analysis 122, e.g. using a plurality of layers on an ANN, said layers including: the labelling sl52; the supplementing sl54; category selection sl56, wherein selectable entities are used to determine a threat; determining spatial relationships si 58, wherein the distance between two points of elements of the data in a multi-dimensional space are determined; pattern recognition si 60, wherein derivations of the elements of the data are analysed for determining whether an attack pattern is present; and relationship determination s!62, wherein the analysed data from all other processes is further processed to determine relationships with the source data 120, thus producing a relationship dataset D240.
[0043] In the example of Figure 2 the processing of the source data 120 is performed by a first ANN 126a, which runs processes s!50, 152, sl54, while the analysis is performed by a second ANN 126b, which uses layers s!56, s 158, s!60, sl62. In light of the teaching herein, it can be appreciated that a first and second ANN can operate as a single ANN 126. Further, the order of the processes and / or the layers can be rearranged and / or include feedback loops. Figure 3 illustrates an alternative arrangement to Figure 2, wherein all processing is performed by the ANN 126 and the supplementation S154 of the formatted dataset following pre-processing is optional e.g. the supplementation can be provided only for training purposes. Moreover, the layers of the ANN 126 that perform category selection S156, itself an optional process, spatial determination S158 and pattern recognition e.g. attack pattern recognition SI 60 can be arranged in parallel, with the or each layer feeing into the process or relationship determination SI 60. Optional data flow paths are indicated with hashed-lines. Via the customise 112 section of the GUI, or during training of the ANN, the parameters of the ANN can be customised for optimising the performance of the machine 104 for using source data 120 to be processed and / or analysed to generate the report 108. Customisation can, for example, enable an operator of the GUI to save words and phrases e.g. create ‘custom rules’ in the GUI. A custom rule can designate a key-word or parameter of the source data as an indication of a threat actor. Thereafter, when the custom rule is uploaded, every subsequent set of data e.g. document that is uploaded via the portal 102 has every aspect scanned to determine whether there is a match. For example, the custom rule can be associated with a STIX category label. Once the ANN is trained, the balancing mechanism 124 can be by-passed Structure - detailed
[0044] The example of Figure 2 will now be described in more detail, wherein the machine 104 uses source data 120 e.g. cyber threat intelligence data that is provided or otherwise determined by a portal input 106 of the GUI 102. The GUI can be a web portal. The details of Figure 2 apply, similarly, to the machine of Figure 3.
[0045] The portal 106 can be used, for example, to upload documents, scan emails, or extract data e.g. cyber threat intelligence (CTI) reports from offline records. Large companies and government organisations have both continuous monitoring requirements and a need to process stored data e.g. to determine patterns and historical threats. The data can reside in reports in various formats e.g. MS word (RTM) documents, PDF reports, blogs, newly published reports etc. Using the machine 104, the GUI can process the reports and determine their relevance e.g. cyber threat level for the operator of the GUI. While this can be performed manually it is impractical, and the example taught herein processes machine readable data, via the GUI and machine 104, to generate a machine readable report. An operator can include, for example, a security operations centre (SOC) within a large bank, insurance company or governmental organisation.
[0046] The portal input 106 typically receives via an upload or retrieves a file in a set format e.g. txt, html, .pdf. Thereafter, pre-processing sl50 processes transform the source data 120 from various file formats into a machine readable format e.g. a single JSON file containing UTF-8 encoded text. Transforming includes using a combination of open source libraries and parsing code. By way of example, transforming PDF files uses an open source optical character recognition (OCR) library is used to identify the characters and words from the PDF file and create UTF-8 text characters and words. TXT files require minimal transforming. Transforming HTML files requires parsing to differentiate text from the non-text parts of the file. Pre-processing can be performed within the ANN 126. The pre-processing sl50 transforms the input dataset D200, or source data 120, to produce a formatted dataset e.g. a machine readable format e.g. a single JSON file containing UTF-8 encoded text.
[0047] Models are then applied to the formatted dataset. The models, when combined, enable the identification of entities within the formatted source data 120. By way of example, the entities identified can be STIX entities for producing a STIX machine readable report, which can be used to model, structure, and define different types of cyber threat information. The formatted dataset is then processed for labelling sl52. The process of identification can include a series of pattern matching rules that are applied to text blocks within the formatted data e.g. searching through the UTF-8 encoded text blocks within the JSON files. Indicators of a cyber threat are known to follow common patterns, and with enough rules and induced randomness within those rules the machine is able to capture nearly all indicators e.g. a SHA256 hashed file name can be identified using the below pattern. sha256 _pattern = "(?: [^a-fA-Fd]\\b)([a-fA-F\d]{64})(?: [^a-fA-F\d]\\b)"
[0048] The process begins with a text string within a sentence and / or paragraph, or a text string within a table, formatted within a Python dictionary e.g.
[0049] {string; "InfoSTEALER.py is downloaded from... " table; FALSE}
[0050] The output of the labelling 152 is a text string with an associated label, thus the formatted dataset is processed for labelling indicators with a set of labels. The output can be formatted within a python dictionary e.g.
[0051] {text; "InfoSTEALER.py is downloaded from... ", label; "file", start index; 0 end index; 15 score; "90"}
[0052] The labels can represent objects to produce an object dataset D220. Using the object dataset, entities and / or relationships can be determined therefrom. Entities can include, for example, a STIX domain object and / or a STIX relationship object. Patterns can be worked out for a plurality of entity types of interest, such as: other hash types, YARA rules, raw file names etc. The process of labelling can include many pattern matching algorithms e.g. REGEX algorithms e.g.
[0053] MD5 RE = re.compile(r"(?:[^a-fA-F\d]\ \b)([a-fA-Fd]{32})(? : fa-fA-Fd]\ \b)") SHA1 RE = re.compile(r"(?:[^a-fA-F\d]\ \b)([a-fA-F\d]{40})(?: fa-fA-F\d]\\b)") SHA256 RE = re.compile(r"(?: fa-fA-Fd]\ \b)([a-fA-F\d]{64})(?:^a-fA-Fd]\ \b)") SHA512 RE = re. compile(r"(?: fa-fA-Fd]\ \b)([a-fA-I d]{128})(?;[^a-fA-Fd]\ \b) ") SHA DEEP RE = re.compile(r"A\d{l,}:[A-Za-zO-9 / +]{3,}:[A-Za-zO-9 / +]{3,}$") Further, character lookups can be employed e.g. to identify known fde extensions e.g.
[0054] KNOWN FILE EXTENSIONS = [
[0055] "264",
[0056] "2g2".
[0057] "3gp",
[0058] "arf,
[0059] "asf,
[0060] "asx",
[0061] "avi",
[0062] "bik"]
[0063] Further still, predetermined rules can apply these patterns to the raw text, and at least one of: handle errors when multiple matches are found for one word; clean the text; apply a confidence score to the match for determining whether it is true; and defang the text e.g. modifying the text and / or potentially harmful links to make them safe to share.
[0064] The labelling si 52 can further include processing a text string within a sentence and / or paragraph, or a text string within a table, formatted within a Python dictionary e.g.
[0065] {string; "Unit 11 is from..." table; FALSE}
[0066] The output of this further processing can include a text string with an associated label formatted within a python dictionary e.g.
[0067] {text; "Unit 11 is from...", label; "Threat Actor", start index; 0 end index; 8 score; "90"}
[0068] This further processing can include multiple character lookups, which can be determined by an operator of the GUI who wants to customise the analysis using the customisation function 112 e.g.
[0069] CUSTOMER X RULES = [ 'Unit 11 ”,
[0070] 'Anonymous Ransomware gang",
[0071] 'BTEC group"]
[0072] Predetermined rules can apply these patterns to the raw text, wherein the extracted text string is compared to surrounding strings within the source data 120 and determines its relevance and then: rejects it, approves it, or modifies it based on the surroundings. By way of example, the process can determine whether the extracted text string is: preceded by a space; preceded by the end of a sentence, or part of one; whether the text string is surrounded by brackets; or whether the text string is at the start or end of a bracketed section and outside of the brackets. The text string can be processed, for example: to remove the "the " from the beginning of a string; determine the relationship between ownership (i.e. detecting 's at the end of the string) and the rest of the text.
[0073] Overall, the portal input 106 can be used to determine the source data 120 to be analysed and reported upon. The source data is transformed using pre-processing to produce a formatted dataset D210. The transformation can include at least one of receiving, retrieving and importing data to (i) convert the file format e.g. converted from at least one of HTML, PDF and TXT format, and / or (ii) configure the formatted dataset into a machine readable format e.g. produce a JSON file containing UTF-8 encoded text.
[0074] The formatted dataset D210 can then be annotated and / or refined i.e. The ANN 126 can process the formatted dataset directly, but can additionally be labelled to produce an object dataset D220. The labelling attributes a set of labels representing data objects to the formatted dataset. By way of example, the formatted dataset can be labelled with components of the STIX data model.
[0075] For training of the ANN 126 the object dataset D220 can include a pre-labelled dataset e.g. a cyber entity dataset e.g. from a commercial provider. The object dataset used fortraining included labels that mapped to the STIX 2.1 taxonomy. The pre-labelled dataset can be a labelled seed training data set for the cybersecurity domain. Optionally, the labelled seed training data can be scaled up.
[0076] Processing the formatted dataset D210 can include annotating the data with at least one of the following labels: artifact; autonomous system; directory; domain name; email address; email message; file; IPv4 Address; IPv6 Address; MAC Address; mutex; network traffic; process; software; URL; user account; Windows Registry Key; and X.509 Certificate; and / or known associated threat actors. Within the machine, and particularly within the ANN 126, the source data 120 is processed to apply and / or determine therein at least one of: STIX Domain Objects (SDOs); STIX Cyber-observables (SCOs); STIX Relationship Objects (SROs); STIX Extension Definition Object; STIX Meta Objects; and Patterning Language.
[0077] To further train the model of the ANN 126 the object dataset D220 can be processed to produce supplemented data D230, wherein annotations are added to the object dataset D220. The object dataset D220 and / or the formatted dataset S210 can be processed using an algorithm e.g. the Aho-Corasick algorithm. Following the automated supplementary annotations, said annotations were selectively filtered. By way of non-limiting example, the following labels were selected and saved for incorporation into the training dataset - 'identity', 'intrusion-set', 'malware', 'threat-actor', 'tool'. Alternative set of labels, or subset thereof, associated with another standard report e.g. MITRE ATT&CK can be selected.
[0078] The process of supplementing the object dataset D220 includes taking text string within a sentence / paragraph or a text string within a table, formatted within a Python dictionary e.g.
[0079] {string; "InfoSTEALER.py is downloaded from Russian servers, and is used to download InfoSTEALER malware... " table; FALSE}
[0080] The output of this further processing can include a text string with an associated label formatted within a python dictionary e.g.
[0081] {text; "InfoSTEALER.py is downloaded from Russian servers, and is used to download InfoSTEALER malware... ”, label; "malware", start index; 120 end index; 128 score; "97"}
[0082] During the training of the dataset for the embeddings two techniques were compared, namely an industry standard approach of word embeddings versus contextual string embeddings for sequence labelling, which is an algorithm first published in 2018 by A. Akbik et al. The latter has been incorporated into an open source library called Flair, and its programming library was used to generate various embeddings for the training. By way of example, embedding parameters for one model is defined below:
[0083] (embeddings): FlairEmbeddings(
[0084] (Im): LanguageModel(
[0085] (drop): Dropout(p=0.25, inplace=False)
[0086] (encoder): Embedding(275, 100)
[0087] (mn): LSTM(100, 1024)
[0088] ) Supplementing sl54 the object dataset D220 can generated overlapping embeddings i.e. overlapping entities in the supplemented data D230. For training purposes, overlapping embeddings can be parsed and / or fdtered out to inhibit technical complications, wherein the entity that was longer in length was retained.
[0089] Model training to determine the optimum implementation on an ANN model used the supplemented data D230 model training experimentation to identify the optimum implementation of an ANN 126 model for the task of extracting the remaining STIX entities was carried out. Supplemented data D220 was passed into Flair’s pre-built training programme, and a saved model was returned. The performance of the two most promising model architectures, implemented within the FLAIR library were evaluated. A first reference used a blank flair model, which used the flair embeddings ('news-forward-fast), while a second reference used a BERT-base-uncased model fine tuned on the datasets generated herein. For all models the following label dictionaries were defined: 'malware', 'indicator', 'tool', 'intrusion-set', 'threat actor' and 'vulnerability'. Each time CTI data e.g. the source data S120 is processed using the system 100 the Flair model is used and run against the CTI data to predict predetermined entities e.g. STIX entities.
[0090] By way of example, a full model specification, as implemented in Flair, is shown below. Using these parameters, training data is passed into FLAIR’s pre-built training programme and a saved model is returned, then taken and run within the machine 104 when new CTI data e.g. the source data S120 is supplied i.e. inference is drawn from the new data. Performance can be monitored and the model can be re -trained when its performance falls beneath a threshold.
[0091] Post training, model evaluation a reference dataset was established for balanced comparison e.g. the supplemented dataset D230 was parsed to retain only the sentences that have “tag p”, and saved in a markdown format. Optionally, misaligned entities were corrected, manual filtering was performed and selected sentences e.g. repeated or short were removed - and a new “md” file created. Said files were converted back to json e g. evaluation.) son and evaluation filtered.j son respectively, for management when running the models. A further file for evaluation i.e. evaluate.json was derived from the supplemented dataset D230. Evaluation was performed based on the confidence scores of these three files.
[0092] Flair CRF and Flert CRF were models that were trained and compared with the new dataset, the below results tables allowed us to identify which was best:
[0093] Model cards:
[0094] Flair-ner:
[0095] Model: "SequenceTagger(
[0096] (embeddings): FlairEmbeddings(
[0097] (Im): LanguageModel(
[0098] (drop): Dropout(p=0.25, inplace=False)
[0099] (encoder): Embedding(275, 100)
[0100] (rm): LSTM(100, 1024)
[0101] )
[0102] )
[0103] (word dropout) : WordDropout(p =0.05)
[0104] Results:
[0105] - F-score (micro) 0.7971, F-score (macro) 0.7402, Accuracy 0.6784
[0106] By class:
[0107] Results:
[0108] - F-score (micro) 0.8526, F-score (macro) 0.7833, Accuracy 0.7498
[0109] Overall, the training a machine -learning model for the ANN for processing cyber threat information, or information derived therefrom, and for determining a cyber threat level, in particular the machine -learning model comprises: a first stage comprising: receiving the source data, and generating supplementary data therefrom using a string-searching algorithm; and additionally or alternatively a second stage comprising: receiving the formatted dataset D210 and / or object dataset D220, generating embeddings, and training and comparing a plurality of ANNs and retaining the ANN that achieved the greatest performance. The longest of overlapping entities generated during embeddings can be retained.
[0110] Using the trained ANN 126, the object dataset D220 and / or the supplemented dataset D230 are analysed through a plurality of layers of the ANN The layers include at least one of the category selection sl56; determining spatial relationships sl58; pattern recognition sl60, and relationship determination sl62.
[0111] From the entities identified from the object dataset D220 and / or the supplemented dataset D230 e g. STIX entities, at least one category of entity is selected S156 for analysis. By way of uses a non-specialised open source transformer-based named entity recognition model e.g. a transformer version from the spaCy library, which is pre-trained by spaCy on the OntoNotes dataset. Further training is optional, as too is the need to supply additional training data, or perform modifications to the neural network architecture . This is because this model has been specifically trained to recognise words such: Google (RTM), Natwest (RTM), Barack Obama etc. These are the same types of words that conform to the STIX categories of identity and tool.
[0112] The process of category selection includes taking a text string within a sentence / paragraph or a text string within a table, formatted within a Python dictionary e.g.
[0113] {string; "InfoSTEALER.py is downloaded from Russian servers... " table; FALSE}
[0114] The string will be encoded, using a transformer model from an open source library, turning the words into a series of numbers.
[0115] The output of this further processing can include a text string with an associated label with the selected category e.g. 'Identity' or 'Tool', formatted within a python dictionary e.g.
[0116] {text; "InfoSTEALER.py is downloaded from Russian servers... ", label; "Identity", start index; 80 end index; 87 score; "97"}
[0117] The process of category selection S156 comprises embedding of text from the object dataset D220 and / or the supplemented dataset D230 into numerical space, which allows the algorithm of the ANN 126 e g. machine learning algorithm to compare each word / phrase to the learned words / phrases the algorithm acquired from the labelled training dataset e.g. the object dataset D220 and / or the supplemented dataset D230. If the algorithm finds a word / phrase pair that are similar enough, above a pre-set threshold, then the category e.g. identity and / or tool is taken and applied to the new word / phrase. While the model has been trained to recognise multiple categories, the entities “tool” and “identity” were found to be particularly relevant, such that predictions from other categories are filtered out.
[0118] Further, spatial relationships s 158 are determined from the entities identified within the object dataset D220 and / or the supplemented dataset D230 e.g. STIX entities. This process is part of an attack pattern matching algorithms, wherein the a computation e g single computation of the distance between two points in multi dimensional space is determined. The first point in space is a single pre-encoding of an example attack pattern sentence, the second point is a selected sentence from at least one of: the input dataset D200, formatted dataset D210, object dataset D220 and the supplemented dataset D230 e.g. the CTI data. These encodings are achieved through the use of a neural network sentence embedding algorithm e.g. ‘all-MiniLM-L6-v2’, which is a public model released on the HuggingFace website, and the code below is used: from sentence transformers import SentenceTrans former model = SentenceTransformer('sentence-transformers / all-MiniLM-L6-v2) embeddings = model, encode (sentences)
[0119] The distance between these two embeddings can then be measured with a word-distance algorithm e.g. the widely used cosine similarity algorithm. If the distance between the sentence from the CTI report and one of the example sentences is close enough then an attack pattern prediction is made for the sentence.
[0120] Another part of the attack pattern matching algorithms can comprise pattern recognition SI 60, which includes deriving a series of sentence derivations, wherein only portions of sentences are analysed, for example only the noun phrases from the sentence and / or only the how phrases from the sentence are processed. These sub-sentences are, as before, numerically embedded and the cosine similarity is calculated along with the original sentence embedding. If the averaged similarity score for all three versions of the sentence and the example sentence crosses a predetermined threshold score, then a sentence is determined to be an attack pattern.
[0121] During the determination of the spatial relationship S 158 and pattern matching S 160, the process taking text string within a sentence / paragraph or a text string within a table, formatted within a Python dictionary e.g.
[0122] {string; "Unit 11 is from Russia and uses C2 infrastructure accessed through target organizations ' SSH infrastructure... " table; FALSE} The output of this further processing can include a text string with an associated label formatted within a python dictionary e.g.
[0123] {text; "Unit 11 is from Russia and uses C2 infrastructure accessed through target organizations ' SSH infrastructure... ”, label; "ATT&CK PATTERN: T1563.001", start index; 40 end index; 65 score; ”90”}
[0124] The labelled data D220 and / or the supplemented data D230 is processed by the ANN 126 i.e. at least one of the category selection process S156, spatial analysis S158 and pattern recognition S160 to determine a relationship SI 62 and produce a relationship dataset D240.
[0125] Within the relationship analysis S160, a series of rules are applied to the extracted entities to predict the most likely relationships between those entities. By way of example, threat actor relationships with the CTI data e.g. the input dataset D200 and / or the formatted dataset D210 are used to identify the most commonly mentioned threat actor and assume that the CTI data is centred on that one threat actor. Thereafter, further relationships can be determined with other entities e.g. ‘targets’, ‘alias’ etc. A plurality of rules can be applied to determine, heuristically, the main threat actor and its relationships with other entities in the CTI data.
[0126] Further rules determine relationships between sentences within the input dataset D200 and / or the formatted dataset. In other words, sentences can be evaluated in at least one of: isolation; relationships with other sentences; relationships with neighbouring sentences; relationships with a plurality of neighbouring sentences; and relationships within groups of sentences can be determined. Evaluation determines associations between entities e.g. malware with the nearest mentioned threat actor. Weightings can be applied to the relationships between sentences.
[0127] Relationship analysis SI 60 can capture a text string with an associated label. Relationship analysis can be applied to the output of at least one of the previous processes i.e. pre-processing S150, labelling S152, supplementing S154, category selection S 156, spatial analysis S158 and pattern recognition SI 60. Relationship analysis SI 60 can be applied to the output of all previous components in which the labels for each sentence have been determined. The input data can be in the same dictionary format as the other components output:
[0128] {text; "Unit 11 is from... ” label; "Threat Actor" sentence index or table idx; start index; 0 end index; 8 score; "90" sub type (optional); }
[0129] The output of this relationship analysis SI 60 i.e. a relationship dataset D240 can a list of predicted relationships within the document, in the following format list: [ (source ID, type, target ID), ... ]
[0130] Relationship analysis SI 60 processes the entities extracted from the text, sentences and table indicators, and builds up a list of relationships between the objects. Multiple rules determine relationships between entities for generating a reports from the relationship dataset D240 produced by the ANN 126 output. A probabilistic approach can attach a likelihood to each relationship. A score for each relationship can be maintained and updated based on a number of features from the text e.g. the most mentioned malware and threat actor if there isn't one in the same sentence, an Intrusion Set more tagged than an Actor, tagging objects that have the same sentence index, linking any unlinked malware names to the closest primary threat-actor, linking unlinked indicators to the closest primary malware, linking any unlinked attack patterns to the closest primary malware. The relationship dataset D240 can be formatted to produce an output dataset D250, which provides machine readable data for a report 108 within the GUI 102.
[0131] Overall, the machine 104 can be operated via the GUI 102 to extract from the data provided to the system 100 e.g. CTI data, such as emails, reports, etc. Entities e.g. STIX entities, such as STIX 2. 1 entities are extracted from the data and analysed in an entirely automated and accurate manner using the ANN 126 of the machine. The entities are extracted from the documents provided e.g. customer documents stored in multiple formats, and returned in an output dataset D250 e.g. a JSON file format for automated utilisation by other software solutions.
[0132] The system 100 provides an automated analysis and / or annotation tool that is secure and confidential. Through a customisation 112 function in the GUI 102, the output dataset D250 can be customised for individual organisations, teams within the organisations, and individuals within those teams.
[0133] The system 100 i.e. the GUI 102 and / or the machine 104 can be hosted on a server and accessed via a web browser. Documents can be added manually and / or the GUI 102 can be directed to a location for retrieving documents. The output dataset can be available as a download e g. an individual structured JSON files as a download option. Additionally or alternatively, an application programming interface (API) hosted on a server can be provided for users to send their files with an API call, wherein they are processed and a structured output dataset D250 e g. JSON document is returned directly to them through their API system. The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.” The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified unless clearly indicated to the contrary. Thus, as a non-limiting example, a reference to “A and / or B,” when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A without B (optionally including elements other than B); in another embodiment, to B without A (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0134] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e. “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.
[0135] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0136] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03. Use of ordinal terms such as “first,” “second,” “third,” etc., in the claims to modify a claim element does not by itself connote any priority, precedence, or order of one claim element over another or the temporal order in which acts of a method are performed, but are used merely as labels to distinguish one claim element having a certain name from another element having a same name (but for use of the ordinal term) to distinguish the claim elements.
[0137] The invention also consists in any individual features described or implicit herein or shown or implicit in the drawings or any combination of any such features or any generalisation of any such features or combination.
Claims
CLAIMS1. A computer-implemented method of processing source data for determining a cyber threat, the method comprising: using an input dataset, said input dataset comprising the source data to be analysed for determining a cyber threat; processing the input dataset for (i) extracting entities and / or (ii) determining relationships between entities; and generating a machine readable output dataset for producing in a natural language and / or graphical representation of the determined cyber threat.
2. The method of claim 1, further including: transforming the input dataset to produce a formatted dataset; and processing the formatted dataset for labelling indicators with a set of labels representing data objects to produce an object dataset, wherein entities and / or relationships are determined therefrom.
3. The method of claim 1 or 2, wherein the input dataset is at least one of received, retrieved and imported, and(i) converted from at least one of HTML, PDF and TXT format, and / or(ii) configured in a JSON file containing UTF-8 encoded text, before annotating and / or refining.
4. The method of claim any preceding claim, wherein producing the object dataset includes annotating the data with: at least one of the following labels: artifact; autonomous system; directory; domain name; email address; email message; file; IPv4 Address; IPv6 Address; MAC Address; mutex; network traffic; process; software; URL; user account; Windows Registry Key; and X.509 Certificate; and / or known associated threat actors.
5. The method of any of claims 2, 3 or 4, wherein the object dataset is processed to produce supplemented dataset, the method further comprising: expanding the formatted dataset and / or object dataset using an algorithm; and / or parsing overlapping annotated and / or labelled entities by keeping the longest entity.
6. The method of any preceding claim, wherein processing the object dataset and / orthe supplemented dataset further comprises at least one of: parsing the extracted entities;attack pattern matching, comprising: determining a spatial relationship between data objects and / or entities; and / or parsing the data objects and / or entities to identify sub-units thereof, and determining a spatial relationship between data objects and / or entities, for determining relationships between data objects and / or entities extracted from the input dataset.
7. The method of any preceding claim, wherein at least one of the transforming the input dataset, processing the formatted dataset and processing the object dataset comprises using an artificial neural -network (ANN) using predictive models.
8. The method of any preceding claim, wherein generating a machine readable output dataset further comprises generating a machine readable file for graphical and / or document visualisation of the cyber threat.
9. A machine -learning model for processing source data comprising cyber threat information, or information derived therefrom, and for determining a cyber threat level, for the use in the method of any of claims 1 to 8, the model comprising: a first stage comprising: determining pattern matching and / or comparisons of source data to produce an input dataset with data objects and / or entities to be evaluated; and / or supplementing the input dataset for training and evaluation for producing the formatted dataset and / or object dataset; and additionally or alternatively a second stage comprising: processing the source data and determine a spatial location for data objects and / or entities identified within the source data; and determining, at least one of a spatial relationship, patterns, and derivative patterns between derivatives of the data objects and / or entities.
10. The model of claim 9, wherein the first stage receives cyber threat intelligence (CTI) data any preceding claim and outputs the formatted dataset and / or object dataset; and / or the second stage further comprises receiving the formatted dataset and / or the object dataset and generating structured threat information report for processing as machine readable threat intelligence.
11. A computer-implemented method of training a machine-learning model for processing cyber threat information, or information derived therefrom, and for determining a cyber threat level, in particular the machine -learning model of claim 9 or 10, comprising: a first stage comprising: receiving the source data, and generating supplementary data therefrom using a string-searching algorithm; and additionally or alternatively a second stage comprising: receiving the formatted dataset and / or object dataset, generating embeddings, and training and comparing a plurality of ANNs and retaining the ANN that achieved the greatest performance.
12. The method of claim 11, wherein the longest of overlapping entities generated during embeddings were retained.
13. A computer-implemented method of generating a training dataset for a machine-learning model, in particular the machine -learning model of claim 8, comprising: receiving the source data, and generating supplementary data therefrom using a string-searching algorithm.
14. A data processing system comprising means for carrying out the method of any of claims 1 to 8, or 11 to 13.
15. A computer system for processing cyber threat information, or information derived therefrom, and for determining a cyber threat level, the system comprising: a processor; and memory including executable instructions that, as a result of execution by the processor, causes the system to perform the method of any of claims 1 to 8, or 11 to 13.
16. A non-transitory computer-readable storage medium having stored thereon executable instructions that, as a result of being executed by a processor of a computer system, cause the computer system to perform the computer-implemented method of any of claims 1 to 8, or 11 to 13.
Citation Information
Patent Citations
Text data-oriented threat intelligence knowledge graph construction method
CN110717049A
Cybersecurity threat modeling and analysis with text miner and data flow diagram editor
US20220038490A1
Estimation apparatus, estimation method and program
US20230008765A1