Threat intelligence extraction method and device, equipment and storage medium

By identifying the type of network threat analysis report and determining the target model in multiple intelligence extraction models, TTP intelligence extraction is performed on the network threat analysis report, which solves the problem of low extraction accuracy in the existing technology, realizes automated and accurate TTP intelligence extraction, and improves the timely deployment of security defense capabilities.

CN120179799APending Publication Date: 2025-06-20BEIJING HONGTENG INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311714723.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The accuracy of TTP intelligence extraction in the existing technology of the network threat analysis report is low, resulting in the inability to deploy security defense capabilities in a timely manner, bringing uncertainty to security defense work.

Method used

By identifying the type of network threat analysis report, the target model is determined in multiple intelligence extraction models, and the report is extracted and processed to obtain TTP intelligence information. This method trains the pre-trained language model based on the training data set, and the training data set is built on multiple attack technology and tactical groups, which include technology and tactics, expression statements and association information.

Benefits of technology

It realizes automatic extraction of network threat TTP intelligence, improves the accuracy and accuracy of extraction, and can be used for unstructured reports, ensuring the timely deployment of security defense capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179799A_ABST
    Figure CN120179799A_ABST
Patent Text Reader

Abstract

The invention discloses a threat intelligence extraction method and device, equipment and a storage medium, and relates to the field of network threat analysis, and the method comprises the steps: obtaining a network threat analysis report; identifying the type of the network threat analysis report; determining a target model in a plurality of intelligence extraction models according to the type; wherein the intelligence extraction model is obtained by training a pre-training language model based on a training data set, the training data set is constructed based on a plurality of attack technique and tactics groups, and the attack technique and tactics groups comprise technique and tactics, expression statements of the technique and tactics and associated information of the technique and tactics; and performing extraction processing on the network threat analysis report by using the target model to obtain TTP intelligence information. According to the method and the device, the problem of relatively low accuracy of TTP information extraction of the network threat analysis report is solved, more comprehensive information extraction is realized, and the accuracy of TTP information extraction is also improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network threat analysis technology, and in particular, to a threat intelligence extraction method, device, equipment and storage medium. Background Art

[0002] With the explosive growth of network threat attacks, the sharing of TTP (Tactics Techniques and Procedures) intelligence used by attackers in network threat analysis reports is crucial for network security construction. Currently, for the extraction of TTP intelligence from network threat analysis reports, it is easy to encounter the situation where the relevant security defense capabilities cannot be deployed in a timely manner due to the untimely extraction of TTP intelligence, which brings serious uncertainties to security defense work.

[0003] In the current related technologies, although network threat analysis reports can be processed relatively simply, these projects are all of a research nature, and their output accuracy and precision are not very satisfactory, and their applicability to actual applications is poor. Summary of the Invention

[0004] The main purpose of this application is to provide a threat intelligence extraction method, device, equipment and storage medium, aiming to solve the technical problem of low accuracy in the extraction of TTP intelligence from network threat analysis reports in related technologies.

[0005] To achieve the above purpose, this application adopts the following technical solutions:

[0006] In a first aspect, this application provides a threat intelligence extraction method, which includes:

[0007] Obtain a network threat analysis report;

[0008] Identify the type of the network threat analysis report;

[0009] Determine a target model from multiple intelligence extraction models according to the type; among them, the intelligence extraction model is obtained by training a pre-trained language model based on a training data set, the training data set is constructed based on multiple attack technique and tactic groups, and the attack technique and tactic group includes technique and tactic, the expression statement of the technique and tactic, and the association information of the technique and tactic;

[0010] Use the target model to perform extraction processing on the network threat analysis report to obtain TTP intelligence information.

[0011] Optionally, in the above threat intelligence extraction method, before the step of determining the target model from multiple intelligence extraction models according to the type, the method further includes:

[0012] Obtain a training data set;

[0013] Obtain multiple pre-trained language models;

[0014] Train the multiple pre-trained language models respectively according to the training data set to obtain multiple test models;

[0015] Screen the multiple test models according to the training data set to obtain multiple trained intelligence extraction models.

[0016] Optionally, in the above threat intelligence extraction method, the step of obtaining the training data set includes:

[0017] Obtain sample data under different attack scenarios, and the sample data includes attack data and attack description text data;

[0018] According to the attack description text data, use the N-Gram algorithm and anaphora resolution technology to perform TTP information tagging to obtain techniques, tactics and their expression statements;

[0019] According to the attack description text data, use natural language processing technology to establish a network attack entity library. The network attack entity library includes techniques, tactics and their associated information, and the associated information includes commands, tools and codes used by the techniques and tactics;

[0020] According to the attack data, establish multiple attack technique and tactic groups based on prior knowledge. The attack technique and tactic groups include main techniques and tactics, expression statements of the main techniques and tactics, and associated information of the main techniques and tactics;

[0021] Construct a training data set according to the multiple attack technique and tactic groups.

[0022] Optionally, in the above threat intelligence extraction method, the step of using natural language processing technology to establish a network attack entity library according to the attack description text data includes:

[0023] Perform part-of-speech tagging and semantic dependency analysis on the attack description text data to obtain the associated information used by each technique and tactic;

[0024] Establish the association relationship between each technique and tactic and the associated information to obtain the network attack entity library.

[0025] Optionally, in the above threat intelligence extraction method, the attack technique and tactic group further includes secondary techniques and tactics, expression statements of the secondary techniques and tactics, and associated information of the secondary techniques and tactics;

[0026] The step of constructing a training data set according to the multiple attack technique and tactic groups includes:

[0027] For each attack technique and tactic group, construct positive example data according to the main techniques and tactics, expression statements of the main techniques and tactics, and associated information of the main techniques and tactics;

[0028] Construct negative example data according to secondary technical and tactical details, their expression statements, and associated information of secondary technical and tactical details;

[0029] Combine positive example data and negative example data of multiple attack technical and tactical groups to obtain a training data set.

[0030] Optionally, in the above threat intelligence extraction method, the types include multiple language types, the multiple pre-trained models include multiple pre-trained language models supporting various language types, the multiple test models include multiple test models supporting various language types, and the multiple intelligence extraction models include multiple intelligence extraction models corresponding one-to-one to multiple language types;

[0031] The steps of determining a target model from multiple intelligence extraction models according to the type include:

[0032] Match the intelligence extraction model corresponding to the type in multiple intelligence extraction models to obtain the target model.

[0033] Optionally, in the above threat intelligence extraction method, the steps of screening multiple test models according to the training data set to obtain multiple trained intelligence extraction models include:

[0034] Determine the Macro-F1 value of each test model based on the Macro F1 evaluation metric;

[0035] Determine multiple optimal test models supporting various language types according to the Macro-F1 values of multiple test models;

[0036] Conduct distillation training on multiple optimal test models respectively to obtain corresponding multiple intelligence extraction models.

[0037] In a second aspect, the present application provides a threat intelligence extraction device, which includes:

[0038] A report acquisition module, configured to acquire a network threat analysis report;

[0039] A report identification module, configured to identify the type of the network threat analysis report;

[0040] A model selection module, configured to determine a target model from multiple intelligence extraction models according to the type; wherein, the intelligence extraction model is obtained by training a pre-trained language model based on a training data set, the training data set is constructed based on multiple attack technical and tactical groups, and the attack technical and tactical groups include technical and tactical details, their expression statements, and associated information of technical and tactical details;

[0041] An intelligence extraction module, configured to perform extraction processing on the network threat analysis report using the target model to obtain TTP intelligence information.

[0042] In a third aspect, the present application provides a threat intelligence extraction device, which includes a processor and a memory. A threat intelligence extraction program is stored in the memory. When the threat intelligence extraction program is executed by the processor, the above-mentioned threat intelligence extraction method is implemented.

[0043] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by one or more processors, the above-mentioned threat intelligence extraction method is implemented.

[0044] One or more of the above technical solutions provided by the present application may have the following advantages or at least achieve the following technical effects:

[0045] A threat intelligence extraction method, device, equipment and storage medium proposed by the present application identify the type of network threat analysis report, determine a target model among multiple intelligence extraction models, and use the target model to extract and process the network threat analysis report to obtain TTP intelligence information, realizing the automatic extraction of network threat TTP intelligence, and can perform intelligence extraction on unstructured reports to obtain structured result information; on this basis, a training data set constructed by using multiple attack techniques and tactics groups is used to train a pre-trained language model to obtain an intelligence extraction model. Each attack technique and tactics group includes techniques and tactics, expression statements of the techniques and tactics, and associated information of the techniques and tactics, taking into account multiple aspects of information of the techniques and tactics, avoiding the loss of important TTP intelligence due to the lack of description information of the techniques and tactics or the dispersion of the description information of the techniques and tactics in the network threat analysis report, and can achieve more comprehensive intelligence extraction. Moreover, the trained intelligence extraction model has the advantage of high accuracy, thus improving the accuracy of TTP intelligence extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these provided drawings without creative efforts.

[0047] Figure 1 It is a schematic flowchart of the first embodiment of the threat intelligence extraction method of the present application;

[0048] Figure 2 It is a schematic hardware structure diagram of the threat intelligence extraction device involved in the present application;

[0049] Figure 3 It is a schematic flowchart of S500-S540 in the second embodiment of the threat intelligence extraction method of the present application;

[0050] Figure 4 For Figure 3 a detailed process schematic diagram of S510 in

[0051] Figure 5 a schematic diagram of functional modules of the first embodiment of the threat intelligence extraction device of the present application.

[0052] The realization of the purpose of the present application, functional features and advantages will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments

[0053] To make the purpose, technical solution and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0054] It should be noted that in the present application, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or system including such element. In the present application, the suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of description of the present application, and have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. In addition, the technical solutions of each embodiment can be combined with each other, however, on the basis that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0055] Analysis of related technologies reveals that in the current TTP intelligence extraction technology for network threat analysis reports, due to the lack of standard structured language description, the extraction process is complex, time-consuming and laborious. If this problem is not solved, it is easy to occur that the relevant security defense capabilities cannot be deployed in time due to untimely TTP intelligence extraction, which will bring serious uncertainties to security defense work.

[0056] Currently, the TRAM project (ThreatReport ATT&CKMapper, a project that marks ATT&CK techniques involved in security reports written in natural language) implemented based on machine learning algorithms (ML) and the TTP extraction project implemented based on information retrieval (IR) technology can handle network threat analysis reports relatively simply. However, these projects are all research-oriented, and their output precision and accuracy are not very satisfactory, and their applicability to actual applications is poor. Moreover, these projects are only designed for English reports and cannot be applied to Chinese reports with complex and variable description methods.

[0057] In view of the technical problem of low precision in the extraction of TTP intelligence from network threat analysis reports in related technologies, the present application provides a threat intelligence extraction method, device, equipment, and storage medium.

[0058] The following combines the accompanying drawings to detail the threat intelligence extraction method, device, equipment, and storage medium provided by the present application through specific embodiments and implementation manners.

[0059] Embodiment 1

[0060] Referring to Figure 1 , a first embodiment of the threat intelligence extraction method of the present application is proposed. The threat intelligence extraction method is applied to a threat intelligence extraction device.

[0061] A threat intelligence extraction device refers to a terminal device or a network device capable of achieving network connection. The threat intelligence extraction device can be a terminal device such as a mobile phone, a computer, a tablet computer, a portable computer, an embedded industrial control computer, etc., or a network device such as a server or a cloud platform.

[0062] As Figure 2 shown, it is a schematic diagram of the hardware structure of the threat intelligence extraction device. The threat intelligence extraction device may include: a processor 1001, such as a CPU (Central Processing Unit, central processor), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005.

[0063] Specifically, the communication bus 1002 is used to implement connection and communication between these components; the user interface 1003 is used to connect to the client and conduct data communication with the client. The user interface 1003 may include an output unit and an input unit; the network interface 1004 is used to connect to the background server and conduct data communication with the background server. The network interface 1004 may include an input / output interface; the memory 1005 is used to store various types of data, which may include, for example, instructions for any application or method in the threat intelligence extraction device, as well as application-related data. The memory 1005 may be a built-in memory; optionally, the memory 1005 may also be a storage device independent of the processor 1001. Continuing to refer to Figure 2 , the memory 1005 may include an operating system, a network communication module, a user interface module, and a threat intelligence extraction program; the processor 1001 is used to call the threat intelligence extraction program stored in the memory 1005 and perform the following operations:

[0064] Obtain a network threat analysis report;

[0065] Identify the type of the network threat analysis report;

[0066] Determine a target model from multiple intelligence extraction models according to the type; wherein, the intelligence extraction model is obtained by training a pre-trained language model based on a training data set, and the training data set is constructed based on multiple attack technique and tactic groups, and the attack technique and tactic group includes techniques and tactics, expression statements of the techniques and tactics, and association information of the techniques and tactics;

[0067] Use the target model to perform extraction processing on the network threat analysis report to obtain TTP intelligence information.

[0068] Based on the above threat intelligence extraction device, the threat intelligence extraction method of this embodiment will be described in detail below in combination with Figure 1 the flow schematic diagram shown. The method may include the following steps:

[0069] S100: Obtain a network threat analysis report.

[0070] Specifically, network threats include various types of network attacks. A network attack refers to any type of attack action against a computer information system, infrastructure, computer network, or personal computer device, such as actions of destruction, disclosure, modification, rendering software or services dysfunctional, obtaining or accessing computer data without authorization, etc. The network threat analysis report may include an unstructured analysis report related to the attacker, attack event data, tools, codes, parameters, etc. used by the attacker to conduct the attack.

[0071] In the specific implementation process, the threat intelligence extraction device can monitor computer devices, computer networks, etc. through a network security monitoring program or device, collect all network attack behaviors suffered, and optionally include basic analysis of these network attack behaviors to obtain a network threat analysis report.

[0072] S200: Identify the type of the network threat analysis report.

[0073] Specifically, the types include the language type used in the report. Of course, it can also include other different types when classifying for intelligence extraction, such as the security monitoring object in the report being a device or a network, the security monitoring tool in the report being a device or a program, etc. These are specific types that can classify the network threat report, which are not limited here.

[0074] Since different intelligence extraction models are selected for network threat analysis reports described in different languages, the threat intelligence extraction device can first identify the language type of the network threat analysis report.

[0075] S300: Determine a target model from multiple intelligence extraction models according to the type; among them, the intelligence extraction model is obtained by training a pre-trained language model based on a training data set, and the training data set is constructed based on multiple attack technique and tactic groups, and the attack technique and tactic group includes techniques and tactics, expression statements of the techniques and tactics, and association information of the techniques and tactics.

[0076] Specifically, since the network threat analysis report has different types, the intelligence extraction model can have one or more models corresponding to different types. The pre-trained language model is a representation learning of natural language, which converts natural language into a data expression form that can be processed by a machine. It can be a general language representation model trained through a large amount of corpus. In this embodiment, the pre-trained language model is used to perform downstream task training, that is, intelligence extraction training, in order to obtain a trained intelligence extraction model.

[0077] The training data set is not directly obtained sample data, but is obtained after processing the sample data. Processing the sample data can obtain attack technique and tactic group data. The techniques and tactics in the attack technique and tactic group refer to the techniques and tactics used by the attacker, and the expression statements of the techniques and tactics correspondingly include the expressions of techniques and the expressions of tactics. The association information of the techniques and tactics also includes the association information of techniques and the association information of tactics.

[0078] In this embodiment, first, the tactics and techniques used by the attacker, the expression statements of the tactics and techniques, and the associated information of the tactics and techniques are marked or identified from the sample data to form an attack tactics and techniques group. Then, multiple such attack tactics and techniques groups are used to form a training data set. By processing the sample data, the specificity of the training data set is ensured, so that the intelligence extraction model trained based on this training data set can be applicable to various actual situations. For example, some network threat analysis reports lack tactics description information or the tactics description information in the network threat analysis reports is scattered, etc.

[0079] In the specific implementation process, after the threat intelligence extraction device obtains the type of the network threat analysis report, it can select a corresponding intelligence extraction model from multiple pre-trained intelligence extraction models as the target model.

[0080] S400: Use the target model to perform extraction processing on the network threat analysis report to obtain TTP intelligence information.

[0081] Specifically, the TTP intelligence information includes the attack tactics, attack techniques, and attack implementation processes used by the network threat attacker. Among them, the attack tactics refer to the short-term tactical goals of the attacker during the attack, such as initial access, execution, privilege escalation, defense evasion, credential access, exfiltration, etc.; the attack techniques refer to how the attacker achieves the tactical goals by taking actions, such as modifying the registry, modifying domain policies or group policies, disabling or modifying tools, disabling or modifying the system firewall, brute force cracking, spoofing, etc.; the attack implementation process refers to the specific methods for the attacker to implement the tactics or techniques. For example, the way to implement the defense evasion tactic: hide the content of the executable file or file on the system or during transmission through encryption, encoding, compression, or other means, so as to obfuscate the content of the executable file or file; another example is the way to implement the technique of modifying the system firewall: use a script to enable the remote desktop connection, start the Microsoft protection service, set rules in the system firewall to allow the attacker to connect, modify the value in the registry, and allow users to connect without authentication.

[0082] In the specific implementation process, the extraction processing of the network threat analysis report can be to input the network threat analysis report into the target model, and then the target model outputs the TTP intelligence information contained in the network threat analysis report. Subsequently, this structured TTP intelligence information can be used in production.

[0083] Different from Named Entity Recognition (NER) in traditional information processing technologies, the extraction of TTP intelligence is not based on identifying and extracting specific cyber-attack names. For example, some methods use NER or IR technologies for cyber-attack identification and information extraction. In this embodiment, the extraction of TTP intelligence actually involves first cognitively understanding natural language through security domain knowledge, and then converting the understood content into Tactics, Techniques, and Procedures, which can subsequently be linked to the corresponding ontology design model.

[0084] In this embodiment, by using standard structured language description technology, TTP intelligence used by attackers can be automatically identified and extracted from complex, unstructured cyber-threat analysis reports, achieving the automatic extraction of TTP intelligence information, facilitating subsequent analysis of cyber-threat situations, or providing data support for subsequent automatic cyber-threat analysis.

[0085] The threat intelligence extraction method provided in this embodiment identifies the type of cyber-threat analysis report, determines the target model among multiple intelligence extraction models, and uses the target model to extract and process the cyber-threat analysis report to obtain TTP intelligence information, achieving the automatic extraction of cyber-threat TTP intelligence. It can perform intelligence extraction on unstructured reports to obtain structured result information. On this basis, a pre-trained language model is trained using a training data set constructed by multiple attack tactics groups. Each attack tactics group includes tactics, expression statements of the tactics, and associated information of the tactics, considering multiple aspects of information of the tactics, avoiding the loss of important TTP intelligence due to the lack of tactics description information or the dispersion of tactics description information in the cyber-threat analysis report, and enabling more comprehensive intelligence extraction. Moreover, the trained intelligence extraction model has the advantage of high accuracy, thereby improving the accuracy of TTP intelligence extraction.

[0086] Embodiment Two

[0087] Based on the same technical concept, referring to Figures 3 to 4 , a second embodiment of the threat intelligence extraction method of this application is proposed. This threat intelligence extraction method is applied to a threat intelligence extraction device. The threat intelligence extraction method of this embodiment will be described in detail below.

[0088] Furthermore, as shown in the process schematic diagram in Figure 3 , before step S300 "Determine the target model among multiple intelligence extraction models according to the type", this method may further include:

[0089] S510: Obtain the training data set.

[0090] Specifically, the training data set is constructed based on multiple attack technique and tactic groups, and an attack technique and tactic group includes a technique and tactic, a statement sentence of the technique and tactic, and associated information of the technique and tactic. Among them, the technique and tactic is the abbreviation of an attack tactic and an attack technique, and the associated information may be information such as a key tool, command, code, etc. associated with the technique and tactic and usable for an information retrieval technique.

[0091] S520: Obtain multiple pre-trained language models.

[0092] Specifically, the pre-trained language model may adopt a pre-trained language model based on the Transformer architecture and the self-attention mechanism. Among them, the Transformer architecture may include an input layer, an encoding layer, a decoding layer, and an output layer; the self-attention mechanism refers to an algorithm that enables a machine to notice the correlation between different parts of the entire input. The self-attention mechanism can reduce the dependence of the model on external information and is better at capturing the internal correlation of data or features. The pre-trained language model adopts the Transformer architecture and the self-attention mechanism, and its model may further include an attention mechanism layer (Attention layer).

[0093] In this embodiment, the pre-trained language model may adopt a BERT (Bidirectional Encoder Representation from Transformer) model or an ERNIE (Enhanced Representation through kNowledge IntEgration) model. Among them, the BERT model is based on the Transformer architecture. When inputting a sentence, the encoding layer of the Transformer will output the vector representation of each word in the sentence, and the bidirectionality is due to the bidirectionality of the encoding layer of the Transformer. The ERNIE model is based on the deep learning framework PaddlePaddle and can model various different types of knowledge, such as grammar knowledge, semantic knowledge, and entity knowledge, etc., so as to improve the performance of natural language understanding tasks.

[0094] In a specific application, the threat intelligence extraction device may initially select multiple pre-trained language models suitable for the application scenario of this application from the model library according to the differences in language tasks, word segmentation, and masking techniques adopted by different pre-trained language models.

[0095] S530: Train the multiple pre-trained language models respectively according to the training data set to obtain multiple test models.

[0096] Specifically, the training of the pre-trained language model has two stages: the pre-training stage and the fine-tuning stage. In the pre-training stage, the training data can be divided from the training dataset, and then the training data is used to pre-train multiple pre-trained language models respectively. Then, the parameters of the model with the best training result are selected, that is, after hyperparameter automatic optimization, the model under the optimal parameters is used as the trained test model. Based on the aforementioned network threat analysis reports, there can be multiple types, so during the model training process, multiple pre-trained language models can also be selected corresponding to different types, so as to obtain multiple test models for different types.

[0097] In this embodiment, after initially selecting the pre-trained language models, these pre-trained language models can be actually tested and trained, and then further selected according to the test results of the trained models and the hardware conditions required in actual applications.

[0098] In order to obtain a model more suitable for use in the threat intelligence extraction scenario, during the training of the pre-trained language model, when performing hyperparameter automatic optimization, the value ranges of the following key hyperparameters can be defined through a script:

[0099] 1. Define the maximum value length of the smallest unit Token: MAX_SEQ_LENGTH, and the optional values are 128, 256, and 512;

[0100] 2. Define the length of batch training and parameter update: BATCH_SIZE, and the optional values are 8, 16, and 32;

[0101] 3. Select an optimizer, define a loss function and a learning rate to control the learning speed: Since it is transfer learning, a smaller learning rate value is required. The optional values of the learning rate are 1e-5, 2e-5, 3e-5, and 5e-5; the defined loss function is the BCEWithLogitsLoss function, and the optimizer can use the AdamW optimization algorithm of BERT;

[0102] 4. Define the number of training epochs Epoch: Use the automatic stop method, and the training of the model will automatically stop if the training results do not improve for N consecutive times.

[0103] S540: Screen multiple test models according to the training dataset to obtain multiple trained intelligence extraction models.

[0104] Specifically, the training data is divided from the training data set, and the test data can be divided from the training data set accordingly. The test data is used to perform actual application tests on multiple test models to obtain the test results of each test model. Then, the most suitable test model can be further selected based on the test results and the hardware conditions required in the actual application. As an intelligence extraction model under a category, there can be multiple types of corresponding network threat analysis reports, so there can be multiple intelligence extraction models. Afterwards, the trained multiple intelligence extraction models can be stored in the threat intelligence extraction device so that they can be called when executing step S300 to determine the target model.

[0105] In an optional implementation manner of this embodiment, if Figure 4 In the detailed flow chart shown, step S510 "obtaining a training data set" may include:

[0106] S511: Obtain sample data under different attack scenarios, where the sample data includes attack data and attack description text data.

[0107] Specifically, the attack data and attack description text data can be obtained according to the actual attack scenario to obtain sample data under different attack scenarios. The attack data may include attacker information, attack tactics and attack techniques, etc. The attack description text data may be similar to the following text:

[0108] "Disabled antivirus (AV) programs such as Windows Defender add (Disabled antivirus (AV) programs such as Windows Defender add

[0109] "HKLM\Software\Policies\Microsoft\Windows Defender" / vDisableAntiVirus (disable antivirus) / t REG_DWORD / d 1 / f add

[0110] "HKLM\Software\Policies\Microsoft\Windows Defender" / vDisableAntiSpyware (disable anti-spyware) / t REG_DWORD / d 1 / f add

[0111] "HKLM\Software\Policies\Microsoft\Windows Defender\MpEngine" / vMpEnablePus (modify the value of MpEnablePus) / t REG_DWORD / d 0 / f”

[0112] In the above description text, the implementation process of the attack event is defined as disabling the antivirus program. The attack event is achieved by disabling the antivirus, disabling the anti-spyware, and modifying the value of MpEnablePus. These specific methods belong to the attack technique of Modify Registry and correspond to the attack tactic of Defense Evasion, so as to obtain attack data and attack description text data.

[0113] S512: According to the attack description text data, use the N-Gram algorithm and anaphora resolution technology to mark TTP information, and obtain the tactics and their expression statements.

[0114] Specifically, the N-Gram algorithm is an algorithm based on a statistical language model. It performs a sliding window operation of size N on the content in the text by bytes, forming a sequence of byte segments of length N, also called the N-Gram sliding window technology. The anaphora resolution technology can identify the content that refers to the same entity in a text, usually names and pronouns. Combining the N-Gram algorithm and the anaphora resolution technology to mark the attack description text data, extract TTP information from the context of the attack description text, and obtain the tactics and their expression statements. In practical applications, the value of N can also be set to mark multiple statements, so that the subsequent training of the pre-trained language model can use these multiple statements.

[0115] In this embodiment, the TTP information is extracted from the context of the attack description text data by using the N-Gram sliding window and anaphora resolution, so as to add the complete expression statements of the tactics to the training dataset, and solve the problem that important TTP intelligence may be lost due to the dispersion of the tactic description information in the network threat analysis report.

[0116] S513: According to the attack description text data, use natural language processing technology to establish a network attack entity library. The network attack entity library includes tactics and their associated information, and the associated information includes the commands, tools, and codes used by the tactics.

[0117] Specifically, use natural language processing technology to collect associated information such as key tools, commands, and codes that can be used for information retrieval technology, and then establish the association relationship between the tactics and these associated information, so as to establish a network attack entity library. For example, for the following attack description text data:

[0118] "The fetched payload is supposed to be saved in%Profile%\update.dll.Eventually,the fetched file is spawned with the following commands:(Extracted payload should be saved in%Profile%\update.dll.Finally,the extracted file is generated using the following commands)"

[0119] Command #1: rundll32.exe%Profile%\update.dll,#1

[0120] 5pOygllrsNaAYqx8JNZSTouZNjo+j5XEFHzxqllqpQ==

[0121] Command #2: rundll32.exe%Profile%\update.dll,#1

[0122] 5oGygYVhos+laqBINdFaVJSfMiwhh4LCn4="

[0123] Based on the description text, two commands can be collected. Both of these commands are associated with the specific attack technique of Rundll32 and belong to the attack tactic of defense evasion, that is, the attack techniques and tactics and their associated information are obtained."

[0124] In this embodiment, by establishing a network attack entity library, it is possible to extract the attack techniques and tactics information used by attackers from multiple perspectives such as tools, command lines, and code snippets, solving the problem that important TTP intelligence may be lost due to the lack of attack techniques and tactics description information in network threat analysis reports."

[0125] S514: Based on the attack data, establish multiple attack technique and tactic groups based on prior knowledge. The attack technique and tactic groups include the main attack techniques, the expression statements of the main attack techniques, and the associated information of the main attack techniques."

[0126] Specifically, the main attack technique is the attack technique used by the attacker to achieve the main purpose when conducting a specific attack."

[0127] In this embodiment, after executing steps S512 and S513, the attack techniques and tactics and their expression statements, as well as the attack techniques and tactics and their associated information, are obtained. Then, based on the above content, attack technique and tactic groups can be established with prior knowledge, and the main attack techniques in the attack technique and tactic groups can be distinguished. For example, for the sample data in step S511 above, the attack technique and tactic groups shown in Table 1 below can be established:"

[0128] Table 1

[0129]

[0130] It can be seen that the attack technique and tactic group includes the main techniques and tactics, the expression statements of the main techniques and tactics, and the associated information of the main techniques and tactics. Repeat the above steps to establish multiple attack technique and tactic groups for subsequent use.

[0131] S515: Construct a training data set according to multiple attack technique and tactic groups.

[0132] Specifically, a group of training data can be constructed with one attack technique and tactic group, or a group of training data can be constructed with multiple attack technique and tactic groups. The amount of data in each group of training data is not limited. However, in order to better use the training data set for model training, it can be ensured that the number of positive example data in the training data set is not less than 50.

[0133] In this embodiment, by collecting and optimizing the training data in the training set, the final training data set is obtained, which can make the training data set applicable to the training process of the pre-trained language model in the specific scenario of threat intelligence extraction, ensuring the applicability of the model, and thus ensuring the applicability of the threat intelligence extraction method. At the same time, the attack technique and tactic group includes the technique and tactic information and its related information in multiple aspects. The training data is complete and comprehensively considered, which can make the model trained based on this training data set have a higher accuracy.

[0134] In an alternative implementation manner of this embodiment, step S513, "establish a network attack entity library according to the attack description text data by using natural language processing technology", may include:

[0135] S513.1: Perform part-of-speech tagging and semantic dependency analysis on the attack description text data to obtain the associated information used by each technique and tactic;

[0136] S513.2: Establish the association relationship between each technique and tactic and the associated information to obtain the network attack entity library.

[0137] Specifically, techniques such as Part-Of-Speech tagging (POS tagging) and Semantic Dependency Parsing (SDP) in natural language processing can be used to collect associated information such as key tools, commands, codes, etc. that can be used in information retrieval technology. Then, the association relationship between this associated information and technical and tactical information is established, and the above operations are performed on multiple technical and tactical information respectively, so that a network attack entity library can be constructed. Among them, POS tagging refers to a text data processing technique that marks the part of speech of words in a corpus according to their meanings and context. Semantic Dependency Parsing refers to a technique that analyzes the semantic associations between various language units in a sentence and presents the semantic associations in a dependency structure.

[0138] In this embodiment, by using two specific language processing techniques in combination, the associated information used by each technical and tactical information can be accurately obtained, which is convenient for establishing more and accurate association relationships subsequently, so as to construct a complete and practical network attack entity library, providing data support for establishing an attack technical and tactical group subsequently. For example, for the attack description text data exemplified in the foregoing step S513, an attack technical and tactical group as shown in Table 2 below can be established:

[0139] Table 2

[0140]

[0141] It can be seen that this attack technical and tactical group includes the main technical and tactical information, the expression statements of the main technical and tactical information, and the associated information of the main technical and tactical information.

[0142] Furthermore, the attack technical and tactical group can also include secondary technical and tactical information, the expression statements of the secondary technical and tactical information, and the associated information of the secondary technical and tactical information;

[0143] Correspondingly, step S515 "Construct a training data set according to multiple attack technical and tactical groups" may include:

[0144] S515.1: For each attack technical and tactical group, construct positive example data according to the main technical and tactical information, the expression statements of the main technical and tactical information, and the associated information of the main technical and tactical information;

[0145] S515.2: Construct negative example data according to the secondary technical and tactical information, the expression statements of the secondary technical and tactical information, and the associated information of the secondary technical and tactical information;

[0146] S515.3: Combine the positive example data and negative example data of multiple attack technical and tactical groups to obtain a training data set.

[0147] Specifically, the secondary attack techniques and tactics are some attack techniques attached to the main attack techniques and tactics, that is, the aforementioned main techniques and tactics, and are mainly some tool-based techniques. For example, PowerShell (T1059.001) in Execution (TA0002).

[0148] In the specific implementation process, the secondary techniques and tactics in each group of attack techniques and tactics can be set as optional. That is to say, the group of attack techniques and tactics can only include the main techniques and tactics, the statement sentences of the main techniques and tactics, and the associated information of the main techniques and tactics; or it can also include the main techniques and tactics, the statement sentences of the main techniques and tactics, and the associated information of the main techniques and tactics, as well as the secondary techniques and tactics, the statement sentences of the secondary techniques and tactics, and the associated information of the secondary techniques and tactics.

[0149] In this embodiment, continuing with the sample data in step S511, when establishing a group of attack techniques and tactics based on prior knowledge and distinguishing the main attack techniques and tactics from the secondary attack techniques and tactics in the group, the following group of attack techniques and tactics as shown in Table 3 can be established:

[0150] Table 3

[0151]

[0152] Continuing, for the group of attack techniques and tactics in Table 3, a positive example data and a negative example data can be constructed respectively. Similarly, for multiple groups of attack techniques and tactics, a training data set including multiple positive example data or a training data set including multiple positive example data and multiple negative example data can be established. In actual application, it is optional to ensure that the data volume of the positive example data is greater than that of the negative example data.

[0153] In this embodiment, by distinguishing the main and secondary attack techniques and tactics, the problem of distorted prediction results caused by uneven distribution of techniques and tactics in the data set can be solved.

[0154] In an optional implementation manner of this embodiment, the types of network threat analysis reports can include multiple language types, multiple pre-training models include multiple pre-training language models supporting various language types, multiple test models include multiple test models supporting various language types, and multiple intelligence extraction models include multiple intelligence extraction models corresponding one-to-one to various language types;

[0155] Correspondingly, step S300 "Determine the target model from multiple intelligence extraction models according to the type" can include:

[0156] S310: Match the intelligence extraction model corresponding to the type in multiple intelligence extraction models to obtain the target model.

[0157] In this embodiment, taking the types of network threat analysis reports including Chinese and English as an example, multiple pre-trained models include at least one pre-trained language model supporting Chinese and at least one pre-trained language model supporting English, that is, there are multiple pre-trained language models; correspondingly, multiple test models obtained by training the pre-trained language models include at least one test model supporting Chinese and at least one test model supporting English, that is, there are multiple test models; after screening the multiple test models respectively according to the training dataset, an intelligence extraction model supporting Chinese and an intelligence extraction model supporting English can be correspondingly obtained, that is, an intelligence extraction model corresponding one by one to multiple language types.

[0158] In the specific implementation process, the appropriate pre-trained language model can be selected according to the language used in the target network threat analysis report to be extracted. For example, in this embodiment, Chinese and English are supported, and the pre-trained language model supporting Chinese and the pre-trained language model supporting English can be selected. In addition, based on the possible differences in language tasks, word segmentation and masking techniques adopted by different pre-trained language models, the pre-trained language model suitable for this embodiment can be initially selected. For example:

[0159] Select the following pre-trained language models supporting Chinese:

[0160] Bert-Base-Chinese, Bert-Wwm-Chinese, Bert-Wwm-Ext-Chinese, Roberta-Wwm-Ext, Roberta-Wwm-Ext-Large, ERNIE-3.0-Base-Zh, and ERNIE-3.0-Xbase-Zh, etc.;

[0161] Select the following pre-trained language models supporting English:

[0162] Bert-Base-Cased, Bert-Base-Uncased,

[0163] Bert-Large-Uncased-Whole-Word-Masking,

[0164] Bert-Large-Cased-Whole-Word-Masking, and ERNIE-2.0-Large-En, etc.;

[0165] Then, train the above pre-trained language models respectively according to the training dataset to obtain the corresponding test models supporting Chinese and the test models supporting English;

[0166] Finally, screen multiple test models according to the training dataset, and the selection can be made by combining the actual test results and the actual hardware conditions. Accordingly, an intelligence extraction model that supports Chinese and an intelligence extraction model that supports English are obtained, that is, two intelligence extraction models are obtained, so that when the threat intelligence extraction device determines that the type of the network threat analysis report is a Chinese report, the intelligence extraction model that supports Chinese is selected as the target model, or when the type of the network threat analysis report is an English report, the intelligence extraction model that supports English is selected as the target model.

[0167] In an alternative implementation manner of this embodiment, step S540 "screen multiple test models according to the training dataset to obtain multiple trained intelligence extraction models" may include:

[0168] S541: Determine the Macro-F1 value of each test model based on the Macro F1 evaluation index;

[0169] S542: Determine multiple optimal test models that support various language types according to the Macro-F1 values of the multiple test models;

[0170] S543: Perform distillation training on the multiple optimal test models respectively to obtain corresponding multiple intelligence extraction models.

[0171] Specifically, the test result of the test model can be the Macro-F1 value determined based on the Macro F1 evaluation index. Based on this Macro-F1 value, automatic optimization of the model hyperparameters can be achieved, and the test model under the parameters corresponding to the optimal training result is selected to be determined as the optimal test model. Then, distillation training can be performed on the optimal test model to achieve transfer learning and reduce the complexity of the intelligence extraction model. Among them, when using the Macro F1 to define the evaluation index, first calculate the precision, recall and their F1 values for each category under the test model, and then obtain the Macro-F1 value of the test model by taking the average.

[0172] Distillation training can use knowledge distillation to transfer the dark knowledge in a complex model to a simple model. The complex model serves as the teacher model, which has powerful capabilities and performance, while the simple model, as the student model, is more compact. During the distillation training stage, distillation is performed on the optimal test model. Specifically, this model can be used as the teacher model, and a simple model similar to the Transformer architecture and self-attention mechanism can be established as the student model. The k-q matrix and v-v matrix of the Attention layer of the student model are processed with KL loss against the corresponding matrices of the teacher model, and the prediction layer is distilled, enabling the student model to approximate or even exceed the teacher model as much as possible. Thus, a model with lower complexity can be used to obtain a similar prediction effect, and finally, an intelligence extraction model that can be used in the production environment and is applicable to the threat intelligence extraction scenario of this embodiment can be obtained based on the student model.

[0173] Optionally, multiple intelligence extraction models are trained based on the above steps S510 - S540. After determining the target model among the multiple intelligence extraction models, and then using the target model to extract and process the network threat analysis report to obtain TTP intelligence information, it is also possible to obtain manual feedback results for the model output and perform backtracking loop training to continuously optimize and improve the accuracy of the intelligence extraction model.

[0174] It should be noted that for more implementation details in the specific implementation of the above method steps, reference can be made to the description of the specific implementation in Embodiment 1. For the sake of simplicity of the specification, it will not be repeated here.

[0175] The threat intelligence extraction method provided in this embodiment uses the training dataset after distinguishing the primary and secondary attack techniques and tactics to train the pre-trained language model based on the Transformer architecture and self-attention mechanism, and obtains the intelligence extraction model through transfer learning to automatically identify and extract TTP intelligence from unstructured Chinese and English network threat analysis reports. It can not only perform automated extraction but also support the recognition of more techniques and tactics, improving the accuracy of TTP extraction and ensuring a lower false alarm rate. Based on this, this method can help the panoramic attack and defense knowledge graph quickly and accurately accumulate a large amount of attack and defense implementation processes, thereby establishing corresponding security defense capabilities, and can also transmit TTP intelligence information to empower the intelligence cloud, thus ensuring the full coverage of the IOC pain pyramid model and providing high-quality simulation content packages for the local security brain simulator intrusion service.

[0176] Embodiment 3

[0177] Based on the same inventive concept, refer to Figure 5, a first embodiment of the threat intelligence extraction device of the present application is proposed. The threat intelligence extraction device can be a virtual device and is applied to threat intelligence extraction equipment.

[0178] The following Figure 5 With reference to the schematic diagram of the functional modules shown, the threat intelligence extraction device provided in this embodiment will be described in detail. The device may include:

[0179] A report acquisition module, configured to acquire network threat analysis reports;

[0180] A report identification module, configured to identify the type of network threat analysis reports;

[0181] A model selection module, configured to determine a target model from multiple intelligence extraction models according to the type; wherein, the intelligence extraction models are obtained by training a pre-trained language model based on a training data set, and the training data set is constructed based on multiple attack technique and tactic groups, and the attack technique and tactic groups include techniques and tactics, expression statements of the techniques and tactics, and associated information of the techniques and tactics;

[0182] An intelligence extraction module, configured to use the target model to perform extraction processing on the network threat analysis report to obtain TTP intelligence information.

[0183] Furthermore, the device may further include:

[0184] A data set acquisition module, configured to acquire a training data set;

[0185] An initial model acquisition module, configured to acquire multiple pre-trained language models;

[0186] A model training module, configured to train multiple pre-trained language models respectively according to the training data set to obtain multiple test models;

[0187] A model acquisition module, configured to screen multiple test models according to the training data set to obtain multiple trained intelligence extraction models.

[0188] Furthermore, the data set acquisition module may include:

[0189] A technique and tactic identification unit, configured to perform TTP information tagging on the attack description text data by using the N-Gram algorithm and anaphora resolution technology to obtain techniques and tactics and their expression statements;

[0190] An entity library establishment unit, configured to establish a network attack entity library by using natural language processing technology according to the attack description text data. The network attack entity library includes techniques and tactics and their associated information, and the associated information includes commands, tools, and code snippets used by the techniques and tactics;

[0191] A technical and tactical association unit for establishing multiple attack technical and tactical groups based on prior knowledge according to attack data. The attack technical and tactical groups include main technical and tactical, statement sentences of the main technical and tactical, and associated information of the main technical and tactical.

[0192] A dataset acquisition unit for constructing a training dataset according to multiple attack technical and tactical groups.

[0193] Furthermore, an entity library establishment unit is also used for performing part-of-speech tagging and semantic dependency analysis on attack description text data to obtain the associated information used by each technical and tactical; establishing the association relationship between each technical and tactical and the associated information to obtain a network attack entity library.

[0194] Furthermore, the attack technical and tactical groups also include secondary technical and tactical, statement sentences of the secondary technical and tactical, and associated information of the secondary technical and tactical; the dataset acquisition unit is also used for each attack technical and tactical group to construct positive example data according to the main technical and tactical, statement sentences of the main technical and tactical, and associated information of the main technical and tactical; construct negative example data according to the secondary technical and tactical, statement sentences of the secondary technical and tactical, and associated information of the secondary technical and tactical; combine the positive example data and negative example data of multiple attack technical and tactical groups to obtain a training dataset.

[0195] Furthermore, the types of network threat analysis reports include multiple language types, the multiple pre-trained models include multiple pre-trained language models supporting various language types, the multiple test models include multiple test models supporting various language types, and the multiple intelligence extraction models include multiple intelligence extraction models corresponding one-to-one to various language types; the model selection module may include:

[0196] A model matching unit for matching the intelligence extraction model corresponding to the type in multiple intelligence extraction models to obtain a target model.

[0197] Furthermore, the model acquisition module may include:

[0198] An index calculation unit for determining the Macro-F1 value of each test model based on the Macro F1 evaluation index;

[0199] A model optimization unit for determining multiple optimal test models supporting various language types according to the Macro-F1 values of multiple test models;

[0200] A distillation processing unit for performing distillation training on multiple optimal test models respectively to obtain corresponding multiple intelligence extraction models.

[0201] It should be noted that the functions that can be realized by each module in the threat intelligence extraction device provided in this embodiment and the corresponding technical effects achieved can refer to the descriptions of the specific implementation manners in each embodiment of the threat intelligence extraction method of this application. For the sake of brevity of the specification, it will not be elaborated here.

[0202] Embodiment 4

[0203] Based on the same inventive concept, with reference to Figure 2 the schematic diagram of the hardware structure, this embodiment provides a threat intelligence extraction device, which may include a processor and a memory. A threat intelligence extraction program is stored in the memory. When the threat intelligence extraction program is executed by the processor, all or part of the steps of each embodiment of the threat intelligence extraction method of this application are implemented.

[0204] Specifically, the threat intelligence extraction device refers to a terminal device or a network device capable of implementing a network connection, which may be a terminal device such as a mobile phone, a computer, a tablet computer, a portable computer, an embedded industrial control computer, etc., or a network device such as a server or a cloud platform.

[0205] It can be understood that the threat intelligence extraction device may further include a communication bus, a user interface, and a network interface. Among them, the communication bus is used to implement connection communication between these components; the user interface is used to connect to the client and communicate with the client. The user interface may include an output unit such as a display screen, a speaker, etc., and an input unit such as a keyboard, a microphone, etc.; the network interface is used to connect to the background server and communicate with the background server. The network interface may include an input / output interface, such as a standard wired interface, a wireless interface such as a Wi-Fi interface; the memory is used to store various types of data. These data may include, for example, instructions for any application or method in the threat intelligence extraction device, as well as application-related data. The memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, flash memory, a magnetic disk, or an optical disc, etc.; optionally, the memory may also be a storage device independent of the processor; the processor is used to call the threat intelligence extraction program stored in the memory and execute the threat intelligence extraction method as described above. The processor may be an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic controller (PLC), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components, and is used to execute all or part of the steps of each embodiment of the threat intelligence extraction method as described above.

[0206] It should be noted that Figure 2The hardware structure shown does not constitute a limitation on the threat intelligence extraction device of the present application, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0207] Embodiment 5

[0208] Based on the same inventive concept, this embodiment provides a computer-readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic memory, a magnetic disk, an optical disk, a server, etc. A computer program is stored on the storage medium, and the computer program can be executed by one or more processors. When the computer program is executed by the processor, all or part of the steps of each embodiment of the threat intelligence extraction method of the present application can be implemented.

[0209] It should be noted that the serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments. The above embodiments are only optional embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application under the inventive concept of the present application, or directly or indirectly applied to other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A threat intelligence extraction method, characterized in that, The method includes: Obtaining a network threat analysis report; Identifying the type of the network threat analysis report; Determining a target model from multiple intelligence extraction models according to the type; wherein, the intelligence extraction model is obtained by training a pre-trained language model based on a training data set, and the training data set is constructed based on multiple attack technique and tactic groups, and the attack technique and tactic group includes a technique and tactic, a statement sentence of the technique and tactic, and associated information of the technique and tactic; Performing extraction processing on the network threat analysis report by using the target model to obtain TTP intelligence information.

2. The threat intelligence extraction method according to claim 1, characterized in that, Before the step of determining a target model from multiple intelligence extraction models according to the type, the method further includes: Obtaining a training data set; Obtaining multiple pre-trained language models; Respectively training the multiple pre-trained language models according to the training data set to obtain multiple test models; Screening the multiple test models according to the training data set to obtain multiple trained intelligence extraction models.

3. The threat intelligence extraction method according to claim 2, characterized in that, The step of obtaining the training data set includes: Obtaining sample data under different attack scenarios, where the sample data includes attack data and attack description text data; Performing TTP information marking on the attack description text data by using the N-Gram algorithm and anaphora resolution technology to obtain techniques and tactics and their statement sentences; Establishing a network attack entity library by using natural language processing technology according to the attack description text data, where the network attack entity library includes techniques and tactics and their associated information, and the associated information includes commands, tools, and codes used by the techniques and tactics; Based on prior knowledge, establishing multiple attack technique and tactic groups according to the attack data, where the attack technique and tactic group includes a main technique and tactic, a statement sentence of the main technique and tactic, and associated information of the main technique and tactic; Constructing a training data set according to the multiple attack technique and tactic groups.

4. The threat intelligence extraction method according to claim 3, characterized in that, The step of establishing a network attack entity library by using natural language processing technology according to the attack description text data includes: Performing part-of-speech tagging and semantic dependency analysis on the attack description text data to obtain associated information used by each technique and tactic; Establishing an association relationship between each technique and tactic and the associated information to obtain a network attack entity library.

5. The threat intelligence extraction method according to claim 3, characterized in that, The attack technique and tactic group further includes a secondary technique and tactic, a statement sentence of the secondary technique and tactic, and associated information of the secondary technique and tactic; The step of constructing a training data set according to the multiple attack technique and tactic groups includes: For each attack technique and tactic group, constructing positive example data according to the main technique and tactic, the statement sentence of the main technique and tactic, and the associated information of the main technique and tactic; Constructing negative example data according to the secondary technique and tactic, the statement sentence of the secondary technique and tactic, and the associated information of the secondary technique and tactic; Combining the positive example data and the negative example data of the multiple attack technique and tactic groups to obtain a training data set.

6. The threat intelligence extraction method according to claim 2, characterized in that, The type includes multiple language types, the multiple pre-trained models include multiple pre-trained language models supporting various language types, the multiple test models include multiple test models supporting various language types, and the multiple intelligence extraction models include multiple intelligence extraction models corresponding one-to-one to the multiple language types; The step of determining a target model from multiple intelligence extraction models according to the type includes: Matching an intelligence extraction model corresponding to the type from the multiple intelligence extraction models to obtain a target model.

7. The threat intelligence extraction method according to claim 6, characterized in that, The step of screening the multiple test models according to the training data set to obtain multiple trained intelligence extraction models includes: Determining the Macro-F1 value of each test model based on the Macro F1 evaluation index; Determining multiple optimal test models supporting various language types according to the Macro-F1 values of the multiple test models; Respectively performing distillation training on the multiple optimal test models to obtain corresponding multiple intelligence extraction models.

8. A threat intelligence extraction device, characterized in that, The device includes: A report acquisition module for acquiring a network threat analysis report; A report identification module for identifying the type of the network threat analysis report; A model selection module for determining a target model from multiple intelligence extraction models according to the type; wherein, the intelligence extraction model is obtained by training a pre-trained language model based on a training data set, and the training data set is constructed based on multiple attack techniques and tactics groups, and the attack techniques and tactics groups include techniques and tactics, expression statements of the techniques and tactics, and associated information of the techniques and tactics; An intelligence extraction module for using the target model to perform extraction processing on the network threat analysis report to obtain TTP intelligence information.

9. A threat intelligence extraction device, characterized in that, The threat intelligence extraction device includes a processor and a memory, and a threat intelligence extraction program is stored on the memory. When the threat intelligence extraction program is executed by the processor, the threat intelligence extraction method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium. When the computer program is executed by one or more processors, the threat intelligence extraction method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Threat intelligence structured analysis method and device based on large language model

    CN121792190A