Threat intelligence text classification method and device and electronic equipment

By constructing a deep residual convolutional network based on convolutional neural networks and residual networks, the threat intelligence text is converted into vectors and classified, which solves the problem of low classification accuracy of threat intelligence text in the existing technology and achieves higher classification accuracy and feature extraction capabilities.

CN120744110APending Publication Date: 2025-10-03HILLSTONE NETWORKS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410383145.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

When using regular expressions to screen and classify threat intelligence text in existing technologies, the accuracy is low, and traditional convolutional neural networks are prone to network degradation problems, resulting in unsatisfactory classification results.

Method used

A deep residual convolutional network based on convolutional neural networks and residual networks is used to perform vector conversion and classification on threat intelligence text. By filtering out text with fewer than a preset number of characters or garbled characters, and using loss indicators for preliminary classification, combined with title and character splicing processing, feature extraction and generalization capabilities are improved.

Benefits of technology

It improves the accuracy of threat intelligence text classification, avoids information loss, and enhances the effectiveness of feature extraction and classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744110A_ABST
    Figure CN120744110A_ABST
Patent Text Reader

Abstract

The invention discloses a threat intelligence text classification method and device and electronic equipment. The method comprises the steps of obtaining N threat intelligence texts to be classified; filtering M target texts in the N threat intelligence texts to obtain S threat intelligence texts; performing vector conversion on each threat intelligence text in the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text; and inputting the target vector corresponding to each threat intelligence text into a target classification model for classification processing, and outputting category information to which each threat intelligence text belongs, the target classification model being a model constructed based on a convolutional neural network and a residual network. Through the method and the device, the problem that the accuracy of classifying the threat intelligence texts is low due to the fact that the threat intelligence texts used for recording the network security threats are screened and classified by using regular expressions in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and more specifically, to a method and device for classifying threat intelligence text, and an electronic device. Background Art

[0002] Threat intelligence is evidence-based knowledge encompassing context, mechanisms, indicators, implications, and actionable recommendations. Furthermore, threat intelligence describes existing or impending threats or compromises to assets and can be used to inform decision-makers' responses to these threats or compromises. Therefore, in the cybersecurity field, timely and accurate collection, organization, and refinement of threat intelligence facilitates security technicians' ability to predict and prevent threatening behavior, improving the efficiency of emergency response.

[0003] In related technologies, when organizing and classifying threat intelligence, regular expressions are generally used to filter the collected description texts of various types of security incidents, and then a convolutional neural network is used to classify the filtered text to obtain the threat intelligence text classification results. However, the accuracy of using regular expressions to filter text in related technologies generally depends on the writing of regular expressions. Therefore, the filtered text is likely to contain usable threat intelligence information, such as security advisories and security announcements. These articles are mainly text descriptions of threat events and certain threatening behaviors, which are easily filtered out by regular matching, resulting in the loss of important information. Moreover, the convolutional neural network method is used to classify the filtered text. Because traditional convolutional neural network designs are difficult to break through the limitations of the number of neural network layers, and as the network depth increases, network degradation problems are prone to occur, which will lead to unsatisfactory final classification results of the model, and thus lead to low accuracy in classifying threat intelligence text.

[0004] Currently, no effective solution has been proposed to the problem that regular expressions are used in related technologies to filter and classify threat intelligence texts used to record network security threats, resulting in low accuracy in classifying threat intelligence texts. Summary of the Invention

[0005] The main purpose of this application is to provide a method and device for classifying threat intelligence text, and an electronic device, so as to solve the problem in the related art of using regular expressions to screen and classify threat intelligence texts used to record network security threats, resulting in low accuracy in classifying threat intelligence texts.

[0006] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for classifying threat intelligence texts is provided. The method includes: obtaining N threat intelligence texts to be classified, wherein the threat intelligence texts are used to record information related to the threat behavior of the target network, and N is a positive integer; filtering M target texts in the N threat intelligence texts to obtain S threat intelligence texts, wherein the target texts include texts with less than a preset number of characters and / or garbled texts, and M and S are both positive integers less than N; performing vector conversion on each threat intelligence text in the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text; inputting the target vector corresponding to each threat intelligence text into a target classification model for classification processing, and outputting the category information to which each threat intelligence text belongs, wherein the target classification model is a model constructed based on a convolutional neural network and a residual network.

[0007] Furthermore, performing vector conversion on each of the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text includes: determining whether each of the S threat intelligence texts contains a compromise indicator, wherein the compromise indicator is used to determine whether the target network has been attacked; if there is at least one threat intelligence text among the S threat intelligence texts that does not contain the compromise indicator, then treating the at least one threat intelligence text as a first-category text, and treating the threat intelligence texts other than the at least one threat intelligence text among the S threat intelligence texts as a second-category text; performing vector conversion on each threat intelligence text in the first-category text and the second-category text to obtain a target vector corresponding to each threat intelligence text.

[0008] Furthermore, performing vector conversion on each threat intelligence text in the first category of text and the second category of text to obtain a target vector corresponding to each threat intelligence text includes: obtaining the title information of each threat intelligence text in the first category of text and the second category of text; truncating the characters that exceed the truncation length in each threat intelligence text in the first category of text and the second category of text to obtain T target characters in each threat intelligence text, wherein the length of the T target characters is equal to the truncation length, and T is a positive integer greater than 1; splicing the title information of each threat intelligence text and the T target characters in each threat intelligence text to obtain the spliced ​​information corresponding to each threat intelligence text; performing vector conversion on the spliced ​​information corresponding to each threat intelligence text to obtain the target vector corresponding to each threat intelligence text.

[0009] Furthermore, the characters exceeding the truncation length in each threat intelligence text in the first category of text and the second category of text are truncated to obtain T target characters in each threat intelligence text, including: judging whether the length of the characters in each threat intelligence text in the first category of text and the second category of text is greater than the truncation length; if the length of the characters in each threat intelligence text in the first category of text and the second category of text is greater than the truncation length, directly truncating the characters exceeding the truncation length in each threat intelligence text in the first category of text and the second category of text to obtain T target characters in each threat intelligence text; if there is at least one threat intelligence text in the first category of text and the second category of text whose length is not greater than the truncation length, using the target symbol to pad the characters in the at least one threat intelligence text to obtain the padded characters corresponding to the at least one threat intelligence text, wherein the length of the padded characters is equal to the truncation length; and using the padded characters corresponding to the at least one threat intelligence text as the T target characters corresponding to the at least one threat intelligence text.

[0010] Furthermore, the target classification model is obtained by: obtaining K sample data including threat intelligence texts in multiple formats, and constructing a training set for training the model based on the K sample data, wherein K is a positive integer greater than 1; adding a residual module to the convolutional neural network to obtain a residual network, and constructing an original classification model for classifying the threat intelligence text based on the residual network; using the training set to train the original classification model to obtain the target classification model.

[0011] Furthermore, based on the K sample data, constructing a training set for training a model includes: if the K sample data include threat intelligence text in a first type format, converting the threat intelligence text in the first type format in the K sample data into threat intelligence text in a third type format to obtain Y threat intelligence texts in the third type format, wherein Y is a positive integer less than or equal to K; if the K sample data include threat intelligence text in a second type format, then extracting the body part of the threat intelligence text in the second type format according to the target field in the threat intelligence text in the second type format in the K sample data, wherein the target field includes at least one of the following: a field for representing web page characters and a field for representing web page tags; storing the body part of the threat intelligence text in the second type format as a third type format to obtain W threat intelligence texts in the third type format, wherein W is a positive integer less than or equal to K; constructing the training set for training a model based on the Y threat intelligence texts in the third type format and / or the W threat intelligence texts in the third type format.

[0012] Furthermore, the original classification model is trained using the training set to obtain the target classification model, including: inputting each sample data in the training set into the original classification model for classification processing, and outputting the category information corresponding to each sample data; based on the category information corresponding to each sample data and the actual category information corresponding to each sample data, a loss function is calculated; according to the loss function, the model parameters of the original classification model are updated to obtain updated model parameters; based on the updated model parameters, the target classification model is obtained.

[0013] Furthermore, after inputting the target vector corresponding to each threat intelligence text into the target classification model for classification processing and outputting the category information to which each threat intelligence text belongs, the method also includes: determining the threat behavior information in the target network according to the category information to which each threat intelligence text belongs; generating prompt information based on the threat behavior information in the target network, and sending the prompt information to the target object, wherein the prompt information is used to prompt the target object to process the threat behavior in the target network.

[0014] In order to achieve the above-mentioned purpose, according to another aspect of the present application, a classification device for threat intelligence texts is provided. The device includes: a first acquisition module for acquiring N threat intelligence texts to be classified, wherein the threat intelligence texts are used to record information related to the threat behavior of the target network, and N is a positive integer; a first processing module for filtering M target texts among the N threat intelligence texts to obtain S threat intelligence texts, wherein the target texts include texts with a number of characters less than a preset number and / or garbled texts, and M and S are both positive integers less than N; a first conversion module for performing vector conversion on each threat intelligence text among the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text; a first output module for inputting the target vector corresponding to each threat intelligence text into a target classification model for classification processing, and outputting the category information to which each threat intelligence text belongs, wherein the target classification model is a model constructed based on a convolutional neural network and a residual network.

[0015] Furthermore, the first conversion module includes: a first judgment unit, used to judge whether each threat intelligence text in the S threat intelligence texts contains a compromise indicator, wherein the compromise indicator is used to determine whether the target network has been attacked; a first determination unit, used to treat the at least one threat intelligence text as a first category text if there is at least one threat intelligence text in the S threat intelligence texts that does not contain the compromise indicator, and to treat the threat intelligence texts in the S threat intelligence texts other than the at least one threat intelligence text as a second category text; a first conversion unit, used to perform vector conversion on each threat intelligence text in the first category text and the second category text to obtain a target vector corresponding to each threat intelligence text.

[0016] Furthermore, the first conversion unit includes: a first acquisition submodule, used to obtain the title information of each threat intelligence text in the first category of text and the second category of text; a first processing submodule, used to truncate the characters in each threat intelligence text in the first category of text and the second category of text that exceed the truncation length, and obtain T target characters in each threat intelligence text, wherein the length of the T target characters is equal to the truncation length, and T is a positive integer greater than 1; a second processing submodule, used to splice the title information of each threat intelligence text and the T target characters in each threat intelligence text, and obtain the spliced ​​information corresponding to each threat intelligence text; the first conversion submodule, used to perform vector conversion on the spliced ​​information corresponding to each threat intelligence text, and obtain the target vector corresponding to each threat intelligence text.

[0017] Furthermore, the first processing submodule includes: a first judgment subunit, used to judge whether the length of the characters in each threat intelligence text in the first category of text and the second category of text is greater than the truncation length; a first processing subunit, used to directly truncate the characters in each threat intelligence text in the first category of text and the second category of text that exceed the truncation length if the length of the characters in each threat intelligence text in the first category of text and the second category of text is greater than the truncation length, so as to obtain T target characters in each threat intelligence text; a second processing subunit, used to use the target symbol to pad the characters in the at least one threat intelligence text if the length of the characters in at least one threat intelligence text in the first category of text and the second category of text is not greater than the truncation length, so as to obtain the padded characters corresponding to the at least one threat intelligence text, wherein the length of the padded characters is equal to the truncation length; a first determination subunit, used to use the padded characters corresponding to the at least one threat intelligence text as the T target characters corresponding to the at least one threat intelligence text.

[0018] Furthermore, the target classification model is obtained through the following modules: a second processing module, used to obtain K sample data including threat intelligence texts in multiple formats, and based on the K sample data, construct a training set for training the model, wherein K is a positive integer greater than 1; a third processing module, used to add a residual module to the convolutional neural network to obtain a residual network, and construct an original classification model for classifying the threat intelligence text based on the residual network; a first training module, used to train the original classification model using the training set to obtain the target classification model.

[0019] Furthermore, the second processing module includes: a second conversion unit, which is used to convert the threat intelligence text in the first type format in the K sample data into a threat intelligence text in a third type format if the K sample data include threat intelligence text in the first type format, and obtain Y threat intelligence texts in the third type format, wherein Y is a positive integer less than or equal to K; a first extraction unit, which is used to extract the body part of the threat intelligence text in the second type format according to the target field in the threat intelligence text in the second type format in the K sample data if the K sample data include threat intelligence text in the second type format, wherein the target field includes at least one of the following: a field for representing web page characters and a field for representing web page tags; a first storage unit, which is used to store the body part of the threat intelligence text in the second type format as a third type format, and obtain W threat intelligence texts in the third type format, wherein W is a positive integer less than or equal to K; a first construction unit, which is used to construct the training set for training the model based on the Y threat intelligence texts in the third type format and / or the W threat intelligence texts in the third type format.

[0020] Furthermore, the first training module includes: a first output unit, used to input each sample data in the training set into the original classification model for classification processing, and output the category information corresponding to each sample data; a first calculation unit, used to calculate the loss function based on the category information corresponding to each sample data and the actual category information corresponding to each sample data; a first processing unit, used to update the model parameters of the original classification model according to the loss function to obtain updated model parameters; a second determination unit, used to obtain the target classification model based on the updated model parameters.

[0021] Furthermore, the device also includes: a first determination module, which is used to input the target vector corresponding to each threat intelligence text into the target classification model for classification processing, output the category information to which each threat intelligence text belongs, and then determine the threat behavior information in the target network according to the category information to which each threat intelligence text belongs; a fourth processing module, which is used to generate prompt information based on the threat behavior information in the target network, and send the prompt information to the target object, wherein the prompt information is used to prompt the target object to process the threat behavior in the target network.

[0022] In order to achieve the above-mentioned purpose, according to another aspect of the present application, an electronic device is provided, which includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the above-mentioned threat intelligence text classification methods.

[0023] Through this application, the following steps are adopted: obtaining N threat intelligence texts to be classified, wherein the threat intelligence texts are used to record information related to the threat behavior of the target network, and N is a positive integer; filtering M target texts among the N threat intelligence texts to obtain S threat intelligence texts, wherein the target text includes texts with less than a preset number of characters and / or garbled texts, and both M and S are positive integers less than N; performing vector conversion on each threat intelligence text among the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text; inputting the target vector corresponding to each threat intelligence text into a target classification model for classification processing, and outputting category information to which each threat intelligence text belongs, wherein the target classification model is a model constructed based on a convolutional neural network and a residual network, which solves the problem in the related art of using regular expressions to screen and classify threat intelligence texts used to record network security threats, resulting in low accuracy in classifying threat intelligence texts. Moreover, since regular expressions are used to filter threat intelligence texts in the prior art, some usable threat intelligence information is easily filtered out. However, the conditional screening method used in the present application first filters out the target text in the threat intelligence text. Compared with the method of using regular expressions to filter texts in the prior art, information loss during the screening and filtering of threat intelligence can be avoided. In addition, the present application uses a deep residual convolutional network constructed based on a convolutional neural network and a residual network to extract features of the threat intelligence text content and complete the classification of the threat intelligence text. Compared with traditional convolutional neural networks, it has stronger feature extraction and generalization capabilities, thereby achieving the purpose of improving the accuracy of classifying threat intelligence texts. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0025] Figure 1 This is a flowchart of a method for classifying threat intelligence text according to an embodiment of the present application;

[0026] Figure 2 This is a schematic diagram of classifying threat intelligence text according to the classification criteria in an embodiment of the present application;

[0027] Figure 3 This is a flowchart of an optional threat intelligence text classification method provided in accordance with an embodiment of the present application;

[0028] Figure 4 is a schematic diagram of a device for classifying threat intelligence text according to an embodiment of the present application;

[0029] Figure 5 is a schematic diagram of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0031] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display and analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set up between this system and the relevant user or organization. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving the consent information fed back by the aforementioned user or organization.

[0034] The embodiments or examples of the present disclosure are not exhaustive, but are merely illustrations of some embodiments or examples, and are not intended to be specific limitations on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment or example can be implemented as an independent example, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment or example can also be implemented as an independent example, and the order of the steps in a certain embodiment or example can be arbitrarily exchanged. In addition, the optional methods or optional examples in a certain embodiment or example can be arbitrarily combined; in addition, the various embodiments or examples can be arbitrarily combined. For example, some or all steps of different embodiments or examples can be arbitrarily combined, and a certain embodiment or example can be arbitrarily combined with the optional methods or optional examples of other embodiments or examples.

[0035] The present invention will be described below in conjunction with preferred implementation steps. Figure 1 Flowchart of the method for classifying threat intelligence text according to an embodiment of the present application. Figure 1 As shown, the method includes the following steps:

[0036] Step S101: Obtain N threat intelligence texts to be classified, wherein the threat intelligence texts are used to record information related to threat behaviors of a target network, and N is a positive integer.

[0037] For example, the N threat intelligence texts described above may be multiple pre-collected threat intelligence texts to be classified. Furthermore, the threat intelligence texts may be texts related to threat intelligence, and the threat intelligence texts may record some existing or future threats or hazards to the network (corresponding to the target network described above). Furthermore, the threat intelligence texts may also be pre-collected descriptions of various types of security events, and these descriptions may include text descriptions of threat events and certain threatening behaviors.

[0038] In some examples, before executing step 102, if the N threat intelligence texts include threat intelligence texts in a first type format, the threat intelligence texts in the first type format in the N threat intelligence texts are converted into threat intelligence texts in a third type format to obtain multiple threat intelligence texts in the third type format; if the N threat intelligence texts include threat intelligence texts in a second type format, the body part of the threat intelligence text in the second type format is extracted based on the target field in the threat intelligence text in the second type format in the N threat intelligence texts, wherein the target field is a target field including at least one of the following: a field for representing web page characters and a field for representing web page tags; the body part of the threat intelligence text in the second type format is stored in the third type format to obtain multiple threat intelligence texts in the third type format, and the threat intelligence text in the third type format is used as the threat intelligence text to be classified.

[0039] Step S102: Filter M target texts among N threat intelligence texts to obtain S threat intelligence texts, wherein the target texts include texts with fewer than a preset number of characters and / or garbled texts, and M and S are both positive integers less than N.

[0040] For example, the preset number can be 50 characters, and after collecting multiple threat intelligence texts to be classified (corresponding to the N threat intelligence texts), the texts with fewer characters (i.e., less than 50 characters) and garbled characters in the pre-collected multiple threat intelligence texts (corresponding to the N threat intelligence texts) can be classified as target texts. Then, these target texts (corresponding to the M target texts) can be filtered out from the collected multiple threat intelligence texts (corresponding to the N threat intelligence texts), and these target texts (corresponding to the M target texts) can be removed from the collected multiple threat intelligence texts (corresponding to the N threat intelligence texts) to obtain the S threat intelligence texts.

[0041] Step S103: Perform vector conversion on each threat intelligence text in the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text.

[0042] For example, the filtered text (corresponding to the S threat intelligence texts) may be subjected to vector conversion processing to obtain a text vector corresponding to each threat intelligence text (corresponding to the target vector).

[0043] In step S104, the target vector corresponding to each threat intelligence text is input into the target classification model for classification processing, and the category information to which each threat intelligence text belongs is output. The target classification model is a model constructed based on a convolutional neural network and a residual network.

[0044] For example, a deep residual convolutional network (corresponding to the target classification model mentioned above) can be constructed using a convolutional neural network and a residual network. Then, the text vector corresponding to each threat intelligence text (corresponding to the target vector mentioned above) can be input into the pre-constructed deep residual convolutional network (corresponding to the target classification model mentioned above) to obtain the classified threat intelligence text.

[0045] Through the above steps S101 to S104, since regular expressions are used to filter threat intelligence texts in the prior art, some usable threat intelligence information is easily filtered out. However, the conditional screening method used in the present application first filters out the target text in the threat intelligence text. Compared with the method of using regular expressions to filter texts in the prior art, information loss during the screening and filtering of threat intelligence can be avoided. In addition, the present application uses a deep residual convolutional network constructed based on a convolutional neural network and a residual network to extract features of the threat intelligence text content and complete the classification of the threat intelligence text. Compared with traditional convolutional neural networks, it has stronger feature extraction and generalization capabilities, thereby achieving the purpose of improving the accuracy of classifying threat intelligence texts.

[0046] Optionally, in the threat intelligence text classification method provided in an embodiment of the present application, the target classification model is obtained by: obtaining K sample data including threat intelligence texts in multiple formats, and constructing a training set for training the model based on the K sample data, where K is a positive integer greater than 1; adding a residual module to the convolutional neural network to obtain a residual network, and constructing an original classification model for classifying the threat intelligence text based on the residual network; using the training set to train the original classification model to obtain the target classification model.

[0047] For example, the K sample data mentioned above can be threat intelligence text samples collected in PDF or HTML format, and the original classification model mentioned above can be a model for classifying threat intelligence text constructed by adding a residual module to a convolutional neural network. The pre-built original classification model can then be trained using the collected threat intelligence text samples in PDF or HTML format to obtain a trained model (corresponding to the target classification model mentioned above).

[0048] Through the above scheme, the original classification model can be easily trained using the sample data in the training set to obtain a more accurate classification model.

[0049] Optionally, in the threat intelligence text classification method provided in an embodiment of the present application, constructing a training set for training a model based on K sample data includes: if the K sample data include threat intelligence text in a first type format, then converting the threat intelligence text in the first type format in the K sample data into threat intelligence text in a third type format to obtain Y threat intelligence texts in the third type format, where Y is a positive integer less than or equal to K; if the K sample data include threat intelligence text in a second type format, then extracting the body part of the threat intelligence text in the second type format based on the target field in the threat intelligence text in the second type format in the K sample data, where the target field includes at least one of the following: a field for representing web page characters and a field for representing web page tags; storing the body part of the threat intelligence text in the second type format as a third type format to obtain W threat intelligence texts in the third type format, where W is a positive integer less than or equal to K; constructing a training set for training a model based on Y threat intelligence texts in the third type format and / or W threat intelligence texts in the third type format.

[0050] For example, the above-mentioned first type of format can be PDF format, the above-mentioned second type of format can be HTML format, and the above-mentioned third type of format can be TXT format. For example, if there is threat intelligence text in PDF format in the collected sample data, the threat intelligence text in PDF format can be directly converted into threat intelligence text in TXT format that is easy to read and write. If there is threat intelligence text in HTML format in the collected sample data, the main text part can be extracted from the threat intelligence text in HTML format based on the web page characters, web page tags and other fields embedded in the threat intelligence text in HTML format (corresponding to the above-mentioned target fields), and stored in TXT format (corresponding to the above-mentioned third type of format). The converted threat intelligence texts in TXT format can then be aggregated to form the above-mentioned training set.

[0051] Through the above solution, the collected threat intelligence texts in different formats can be quickly and accurately converted into threat intelligence texts in TXT format that are easy to read and write.

[0052] Optionally, in the classification method of threat intelligence text provided in the embodiment of the present application, the training set is used to train the original classification model to obtain the target classification model, including: inputting each sample data in the training set into the original classification model for classification processing, and outputting the category information corresponding to each sample data; based on the category information corresponding to each sample data and the actual category information corresponding to each sample data, calculating the loss function; updating the model parameters of the original classification model according to the loss function to obtain updated model parameters; and obtaining the target classification model based on the updated model parameters.

[0053] For example, the above-mentioned loss function can be a cross-entropy loss function, and the cross-entropy loss function can also be referred to as the cross-entropy loss. Moreover, in the process of training the model (corresponding to the above-mentioned original classification model), the cross-entropy loss is generally not directly calculated using the K sample data in the training set (corresponding to the above-mentioned K sample data). Instead, the K sample data is first divided into multiple small batches, and the number of data in the small batch can be V. Then, for the V sample data, V cross entropies are calculated, and the V cross entropies are summed and divided by V to obtain the average cross-entropy loss of the small batch data. Then, based on the calculated average cross-entropy loss, the gradient of each parameter is calculated using the backpropagation method, and the model parameters are updated using the Adam method (an optimization algorithm used to train neural networks and other machine learning models). Corresponding calculations are performed on all small batches of data divided by K samples, and the model parameters are updated. This process is repeated multiple times until the model converges, i.e., the model training is completed. For example, after the deep residual network (corresponding to the original classification model mentioned above), a fully connected layer can be used to map features into outputs of different categories; then the Softmax function (a commonly used activation function) can be used to normalize the output results, and the normalized results and the true labels can be used to calculate the cross-entropy loss (corresponding to the above-mentioned loss function); the gradient of the cross-entropy loss is then calculated using the backpropagation method, and the model parameters are continuously updated according to the gradient; and the model parameters are continuously updated until convergence to obtain a trained model (corresponding to the above-mentioned target classification model).

[0054] In summary, by using activation functions and calculating cross-entropy loss, the original classification model can be trained quickly and accurately, thereby improving the accuracy of the classification model in classifying threat intelligence text.

[0055] Optionally, in the threat intelligence text classification method provided in an embodiment of the present application, each threat intelligence text in the S threat intelligence texts is vector-converted to obtain a target vector corresponding to each threat intelligence text, including: determining whether each threat intelligence text in the S threat intelligence texts contains a compromise indicator, wherein the compromise indicator is used to determine whether the target network has been attacked; if there is at least one threat intelligence text in the S threat intelligence texts that does not contain a compromise indicator, then at least one threat intelligence text is treated as a first-category text, and the threat intelligence texts other than at least one threat intelligence text in the S threat intelligence texts are treated as second-category texts; each threat intelligence text in the first-category text and the second-category text is vector-converted to obtain a target vector corresponding to each threat intelligence text.

[0056] For example, the filtered texts (corresponding to the S threat intelligence texts mentioned above) can be roughly divided according to whether the threat intelligence texts can match the compromise indicators. The texts without compromise indicators can be included in the first category of texts, and the texts containing compromise indicators can be included in the second category of texts. Then, each threat intelligence text in the first and second categories can be vectorized and the text vector corresponding to each threat intelligence text can be obtained (corresponding to the target vector mentioned above).

[0057] Through the above solution, threat intelligence text can be quickly and accurately classified according to the compromise indicators.

[0058] Optionally, in the threat intelligence text classification method provided in an embodiment of the present application, each threat intelligence text in the first category of text and the second category of text is vector-converted to obtain a target vector corresponding to each threat intelligence text, including: obtaining the title information of each threat intelligence text in the first category of text and the second category of text; truncating the characters in each threat intelligence text in the first category of text and the second category of text that exceed the truncation length to obtain T target characters in each threat intelligence text, wherein the length of the T target characters is equal to the truncation length, and T is a positive integer greater than 1; splicing the title information of each threat intelligence text and the T target characters in each threat intelligence text to obtain the spliced ​​information corresponding to each threat intelligence text; and performing vector-conversion on the spliced ​​information corresponding to each threat intelligence text to obtain the target vector corresponding to each threat intelligence text.

[0059] For example, you can first determine whether there is a text title at the beginning of the body of the threat intelligence text. If so, you can separate the title and the body with a first preset string (for example, 'SEP'); if not, you can insert the title at the beginning of the body and separate it with the first preset string (for example, 'SEP'). Then you can count the character length characteristics of each threat intelligence text, and take the median of the character length as the truncation length. For a specific threat intelligence text, you can truncate the characters in the text that exceed the truncation length to obtain multiple truncated characters (corresponding to the T target characters mentioned above); if the number of characters is less than the truncation length, you can use a second preset string (for example, 'PAD') symbol to pad to the truncation length. Then, splice the title of each threat intelligence text and the multiple truncated characters (corresponding to the T target characters mentioned above), and perform vector conversion on the spliced ​​content corresponding to each threat intelligence text to obtain the text vector corresponding to each threat intelligence text (corresponding to the target vector mentioned above).

[0060] Through the above solution, the title and character content of the threat intelligence text can be conveniently spliced, and the text vector corresponding to the threat intelligence text can be quickly and accurately obtained.

[0061] Optionally, in the threat intelligence text classification method provided in the embodiment of the present application, the characters exceeding the truncation length in each threat intelligence text in the first category text and the second category text are truncated to obtain T target characters in each threat intelligence text, including: determining whether the length of the characters in each threat intelligence text in the first category text and the second category text is greater than the truncation length; if the length of the characters in each threat intelligence text in the first category text and the second category text is greater than the truncation length, directly truncating the characters exceeding the truncation length in each threat intelligence text in the first category text and the second category text to obtain T target characters in each threat intelligence text; if there is at least one threat intelligence text in the first category text and the second category text whose length is not greater than the truncation length, using the target symbol to padded the characters in the at least one threat intelligence text to obtain the padded characters corresponding to the at least one threat intelligence text, wherein the length of the padded characters is equal to the truncation length; and using the padded characters corresponding to the at least one threat intelligence text as the T target characters corresponding to the at least one threat intelligence text.

[0062] For example, for a specific threat intelligence text, characters that exceed the truncation length in the text can be truncated, and characters that are less than the truncation length can be padded to the truncation length with a preset string (for example, the 'PAD' symbol, corresponding to the target symbol mentioned above). Specifically, for example, the median of the character length can be taken as the truncation length, and then it can be determined whether the length of the characters in the threat intelligence text (corresponding to the first and second types of texts mentioned above) is greater than the truncation length; if the length of the characters in the threat intelligence text is greater than the truncation length, the characters that exceed the truncation length in the threat intelligence text can be truncated; if the length of the characters in the threat intelligence text is not greater than the truncation length, the preset string (for example, the 'PAD' symbol, corresponding to the target symbol mentioned above) can be padded to the truncation length.

[0063] Through the above solution, multiple characters in the threat intelligence text can be quickly and accurately obtained for truncation processing.

[0064] Optionally, in the threat intelligence text classification method provided in an embodiment of the present application, after inputting the target vector corresponding to each threat intelligence text into the target classification model for classification processing and outputting the category information to which each threat intelligence text belongs, the method also includes: determining the threat behavior information in the target network based on the category information to which each threat intelligence text belongs; generating prompt information based on the threat behavior information in the target network, and sending the prompt information to the target object, wherein the prompt information is used to prompt the target object to process the threat behavior in the target network.

[0065] For example, it is possible to determine whether there is threatening behavior in the network (corresponding to the above-mentioned target network) based on the classified threat intelligence text, and if it is determined that there is threatening behavior in the network (corresponding to the above-mentioned target network), a prompt message can be sent to the network maintenance personnel (corresponding to the above-mentioned target object), and the network maintenance personnel (corresponding to the above-mentioned target object) can be reminded to deal with the threatening behavior in the network (corresponding to the above-mentioned target network); if it is determined that there is no threatening behavior in the network (corresponding to the above-mentioned target network), the network (corresponding to the above-mentioned target network) can continue to be monitored, and whether there is threatening behavior in the network (corresponding to the above-mentioned target network).

[0066] Through the above solution, threatening behaviors in the network can be identified quickly and accurately, thereby ensuring the security of the network.

[0067] For example, in actual applications, threat intelligence is mostly unstructured text data, and the data sources are complex and diverse, including various open sources, commercial and professional intelligence sources, etc. In order to extract effective threat intelligence content from massive texts for decision-making, a set of effective processing solutions is needed to track and process threat intelligence in a timely, efficient and accurate manner. Therefore, in this embodiment, a method for classifying and arranging threat intelligence texts that is more resource-friendly while ensuring classification accuracy is provided. Moreover, the method provided by this embodiment can avoid the waste of computing resources and the loss of information as much as possible when screening and filtering threat intelligence. In addition, this embodiment uses a deep residual convolutional network (i.e., the target classification model in the aforementioned embodiment) to extract features from text content and complete information classification. Compared with traditional convolutional neural networks, this method has stronger feature extraction capabilities and generalization capabilities. In addition, the Bert pre-training model (a pre-training model) used in the prior art can achieve better feature extraction effects by fine-tuning a small amount of text, but the Bert model is based on the design of the Transformer architecture (a network architecture based on the attention mechanism), which makes the model have strict restrictions on text length when performing text feature extraction. Longer texts cannot be directly input. The general pre-trained Bert model limits the input length to 512, and threat intelligence articles are usually large in length. The input length of 512 is difficult to meet the classification requirements of most threat intelligence texts. Therefore, the method provided by this embodiment can support longer text inputs compared to the pre-training model based on the Transformer architecture. This is because the complexity of the Transformer architecture for text length when performing encoding calculations is O(N^2), while the complexity of the convolution calculation for text length is O(N), which means that the deep residual convolutional neural network used in this embodiment can support longer text inputs.

[0068] In addition, the complete technical solution of this embodiment includes the following contents:

[0069] 1. Based on the content and structure of threat intelligence texts, we can combine threat intelligence professional knowledge and actual application scenarios to develop classification standards for threat intelligence texts. For example, Figure 2 This is a schematic diagram of dividing threat intelligence text according to the classification criteria in the embodiment of the present application. Figure 2 The content shown divides the threat intelligence text.

[0070] 2. Determine the source of threat intelligence text data and collect the text. Most of the collected text data is in PDF or HTML format.

[0071] 3. Convert PDF-formatted data into TXT files that are easy to read and write.

[0072] 4. The text data in HTML format contains a large number of redundant fields (i.e., the target fields in the aforementioned embodiment), such as embedded web page characters, web page tags and other data. This step can use regular matching to extract the main text part and store it in TXT format.

[0073] 6. For texts in TXT format, conditional screening is used to remove texts with fewer characters and garbled characters. Such texts are ambiguous in meaning and contain less valuable information, and are classified as unusable texts (i.e., the target texts in the aforementioned embodiment).

[0074] 5. Extract the loss indicators contained in the text from the main text through regular expression matching and grammatical rules, and roughly divide the text into categories based on whether the loss indicators match the text. Texts without loss indicators are classified as the first category of articles (i.e., the first category of text in the above embodiment), and texts containing loss indicators are classified as the second category of articles (i.e., the second category of text in the above embodiment).

[0075] 7. For the remaining first-category articles (i.e., the first-category texts in the aforementioned embodiment) and second-category articles (i.e., the second-category texts in the aforementioned embodiment), the text content can be manually reviewed according to the established category classification standards, and the text can be included in the category to which it belongs to complete the construction of the training set.

[0076] 8. In the training set, random replication, synonym replacement and other methods can be used to perform data enhancement operations on categories with a small number of samples.

[0077] 9. In the training set, determine whether there is a text title at the beginning of the text. If so, separate the title from the text with 'SEP'; if not, insert the title at the beginning of the text and separate it with 'SEP'.

[0078] 10. In the training set, the character length characteristics of each sample are counted, and the median of the character length is taken as the truncation length. For a specific text, the characters in the text that exceed the truncation length are truncated, and the characters that are less than the truncation length are padded with the 'PAD' symbol (i.e., the target symbol in the aforementioned embodiment) to the truncation length.

[0079] 11. Modify the original deep residual network and build a one-dimensional deep residual network framework suitable for using text as input (i.e., the target classification model in the aforementioned embodiment).

[0080] Furthermore, the one-dimensional convolutional deep residual network (1D-DRN) incorporates a residual structure into a one-dimensional convolutional neural network, avoiding the severe network degradation problem caused by increasing model depth. This allows for the design of larger network structures. This approach allows for greater freedom in increasing model depth, enabling stronger feature extraction capabilities and ultimately achieving better recognition results.

[0081] One-dimensional convolutional neural networks (1D-CNNs) are a variant of convolutional neural networks. They are primarily used to process one-dimensional sequence data, such as audio and text. Compared to traditional fully connected neural networks, 1D-CNNs can better process local relationships in sequence data, and therefore perform well in tasks such as speech recognition, natural language processing, and time series prediction.

[0082] 12. Use a fully connected layer after the deep residual network to map features into outputs of different categories.

[0083] 13. Use the Softmax function (a commonly used activation function) to normalize the output results, and calculate the cross entropy loss between the normalized results and the true labels.

[0084] 14. Use the backpropagation method to calculate the gradient of the cross entropy loss and continuously update the model parameters based on the gradient.

[0085] 15. Input the training data into the model, and continuously update the model parameters until convergence. During the model training phase, the Adam method (an optimization algorithm used to train neural networks and other machine learning models) was selected for optimization to accelerate the model's convergence. Simultaneously, the network structure was adjusted as the model's recognition performance improved. Specifically, based on the one-dimensional deep residual network structure, the size of the convolution kernel and the depth of the model were fine-tuned, significantly improving the model's recognition readiness rate in actual tasks.

[0086] 16. Perform consistent pre-processing operations on texts automatically collected from multiple data sources before model training, input the processed data into the trained model, perform recognition screening, and automatically organize and classify the text based on the recognition results. For example, Figure 3This is a flowchart of an optional threat intelligence text classification method provided in an embodiment of the present application, which can be used as follows: Figure 3 The threat intelligence text classification method shown is used to classify the threat intelligence text. In addition, Figure 3 GloVe is the abbreviation of Global Vectors for Word Representation, which is a technology used to represent words as vectors.

[0087] In this embodiment, a large amount of text content can be first obtained from multiple threat intelligence sources, and then conditional judgment is used to filter out useless data. Regular matching is used to perform a rough screening of whether the large amount of text content contains IoC (Indicators of Compromise). Then, combined with threat intelligence professional domain knowledge, the text type is carefully divided. According to the division criteria, the threat intelligence text is manually reviewed and annotated to generate a training set. A one-dimensional deep residual network is used to train a text recognition model. The trained model is used to automatically identify and classify the automatically collected threat intelligence data, and finally an overall process and solution for classified high-quality threat intelligence data is obtained.

[0088] Therefore, when screening and classifying threat intelligence text, this embodiment can first summarize and divide the text of the intelligence source in combination with professional domain knowledge, so that the classification results are more in line with the actual application scenario. Subsequently, a one-dimensional deep residual network is introduced. This network structure can solve the problem of network degradation due to the increase in network depth, so that a deeper network can be designed, so that when performing text information classification and screening, it has stronger feature extraction capabilities and recognition accuracy. At the same time, the network uses convolutional layers to extract text features. Compared with most pre-trained language models based on the Transformer architecture (a network architecture based on the attention mechanism), the method of this embodiment has smaller restrictions on the length of text input while ensuring recognition accuracy. It can use longer text as input, which speeds up the training of the model and reduces resource consumption.

[0089] In summary, the threat intelligence text classification method provided in the embodiment of the present application obtains N threat intelligence texts to be classified, wherein the threat intelligence texts are used to record information related to the threat behavior of the target network, and N is a positive integer; M target texts in the N threat intelligence texts are filtered to obtain S threat intelligence texts, wherein the target text includes text with less than a preset number of characters and / or garbled text, and M and S are both positive integers less than N; each threat intelligence text in the S threat intelligence texts is vector-converted to obtain a target vector corresponding to each threat intelligence text; the target vector corresponding to each threat intelligence text is input into a target classification model for classification processing, and the category information to which each threat intelligence text belongs is output, wherein the target classification model is a model constructed based on a convolutional neural network and a residual network, which solves the problem in the related art of using regular expressions to screen and classify threat intelligence texts used to record network security threats, resulting in low accuracy in classifying threat intelligence texts. Moreover, since regular expressions are used to filter threat intelligence texts in the prior art, some usable threat intelligence information is easily filtered out. However, the conditional screening method used in the present application first filters out the target text in the threat intelligence text. Compared with the method of using regular expressions to filter texts in the prior art, information loss during the screening and filtering of threat intelligence can be avoided. In addition, the present application uses a deep residual convolutional network constructed based on a convolutional neural network and a residual network to extract features of the threat intelligence text content and complete the classification of the threat intelligence text. Compared with traditional convolutional neural networks, it has stronger feature extraction and generalization capabilities, thereby achieving the purpose of improving the accuracy of classifying threat intelligence texts.

[0090] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0091] The embodiment of the present application also provides a device for classifying threat intelligence text. It should be noted that the device for classifying threat intelligence text in the embodiment of the present application can be used to execute the method for classifying threat intelligence text provided in the embodiment of the present application. The following introduces the device for classifying threat intelligence text provided in the embodiment of the present application.

[0092] Figure 4 Schematic diagram of a device for classifying threat intelligence text according to an embodiment of the present application. Figure 4 As shown, the device includes: a first acquisition module 401 , a first processing module 402 , a first conversion module 403 and a first output module 404 .

[0093] Specifically, the first acquisition module 401 is configured to acquire N threat intelligence texts to be classified, wherein the threat intelligence texts are used to record information related to threat behaviors of a target network, and N is a positive integer;

[0094] A first processing module 402 is configured to filter M target texts from the N threat intelligence texts to obtain S threat intelligence texts, wherein the target texts include texts with fewer than a preset number of characters and / or garbled texts, and M and S are both positive integers less than N.

[0095] A first conversion module 403 is configured to perform vector conversion on each of the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text;

[0096] The first output module 404 is used to input the target vector corresponding to each threat intelligence text into the target classification model for classification processing, and output the category information to which each threat intelligence text belongs, wherein the target classification model is a model constructed based on a convolutional neural network and a residual network.

[0097] In summary, the threat intelligence text classification device provided in the embodiment of the present application obtains N threat intelligence texts to be classified through the first acquisition module 401, wherein the threat intelligence text is used to record information related to the threat behavior of the target network, and N is a positive integer; the first processing module 402 filters M target texts in the N threat intelligence texts to obtain S threat intelligence texts, wherein the target text includes text with less than a preset number of characters and / or garbled text, and M and S are both positive integers less than N; the first conversion module 403 performs vector conversion on each threat intelligence text in the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text; the first output module 404 inputs the target vector corresponding to each threat intelligence text into the target classification model for classification processing, and outputs the category information to which each threat intelligence text belongs, wherein the target classification model is a model constructed based on convolutional neural networks and residual networks, which solves the problem in the related art of using regular expressions to screen and classify threat intelligence texts used to record network security threats, resulting in low accuracy in classifying threat intelligence texts. Moreover, since regular expressions are used to filter threat intelligence texts in the prior art, some usable threat intelligence information is easily filtered out. However, the conditional screening method used in the present application first filters out the target text in the threat intelligence text. Compared with the method of using regular expressions to filter texts in the prior art, information loss during the screening and filtering of threat intelligence can be avoided. In addition, the present application uses a deep residual convolutional network constructed based on a convolutional neural network and a residual network to extract features of the threat intelligence text content and complete the classification of the threat intelligence text. Compared with traditional convolutional neural networks, it has stronger feature extraction and generalization capabilities, thereby achieving the purpose of improving the accuracy of classifying threat intelligence texts.

[0098] Optionally, in the threat intelligence text classification device provided in an embodiment of the present application, the first conversion module includes: a first judgment unit, used to judge whether each threat intelligence text in the S threat intelligence texts contains a compromise indicator, wherein the compromise indicator is used to determine whether the target network has been attacked; a first determination unit, used to treat at least one threat intelligence text as a first category text if there is at least one threat intelligence text in the S threat intelligence texts that does not contain a compromise indicator, and to treat the threat intelligence texts other than at least one threat intelligence text in the S threat intelligence texts as a second category text; a first conversion unit, used to perform vector conversion on each threat intelligence text in the first category text and the second category text to obtain a target vector corresponding to each threat intelligence text.

[0099] Optionally, in the threat intelligence text classification device provided in an embodiment of the present application, the first conversion unit includes: a first acquisition submodule, used to obtain the title information of each threat intelligence text in the first category text and the second category text; a first processing submodule, used to truncate the characters in each threat intelligence text in the first category text and the second category text that exceed the truncation length to obtain T target characters in each threat intelligence text, wherein the length of the T target characters is equal to the truncation length, and T is a positive integer greater than 1; a second processing submodule, used to splice the title information of each threat intelligence text and the T target characters in each threat intelligence text to obtain the spliced ​​information corresponding to each threat intelligence text; the first conversion submodule, used to perform vector conversion on the spliced ​​information corresponding to each threat intelligence text to obtain the target vector corresponding to each threat intelligence text.

[0100] Optionally, in the threat intelligence text classification device provided in the embodiment of the present application, the first processing submodule includes: a first judgment subunit, used to judge whether the length of the characters in each threat intelligence text in the first category text and the second category text is greater than the truncation length; a first processing subunit, used to directly truncate the characters in each threat intelligence text in the first category text and the second category text that exceed the truncation length if the length of the characters in each threat intelligence text in the first category text and the second category text is greater than the truncation length, to obtain T target characters in each threat intelligence text; a second processing subunit, used to use the target symbol to fill the characters in the at least one threat intelligence text if the length of the characters in at least one threat intelligence text in the first category text and the second category text is not greater than the truncation length, to obtain the padded characters corresponding to the at least one threat intelligence text, wherein the length of the padded characters is equal to the truncation length; a first determination subunit, used to use the padded characters corresponding to the at least one threat intelligence text as the T target characters corresponding to the at least one threat intelligence text.

[0101] Optionally, in the classification device for threat intelligence text provided in an embodiment of the present application, the target classification model is obtained through the following modules: a second processing module, used to obtain K sample data including threat intelligence texts in multiple formats, and based on the K sample data, construct a training set for training the model, where K is a positive integer greater than 1; a third processing module, used to add a residual module to the convolutional neural network to obtain a residual network, and construct an original classification model for classifying the threat intelligence text based on the residual network; a first training module, used to train the original classification model using the training set to obtain a target classification model.

[0102] Optionally, in the threat intelligence text classification device provided in the embodiment of the present application, the second processing module includes: a second conversion unit, which is used to convert the threat intelligence text in the first type format in the K sample data into a threat intelligence text in a third type format if the K sample data include threat intelligence text in the first type format, and obtain Y threat intelligence texts in the third type format, wherein Y is a positive integer less than or equal to K; a first extraction unit, which is used to extract the main text part of the threat intelligence text in the second type format according to the target field in the threat intelligence text in the second type format in the K sample data if the K sample data include threat intelligence text in the second type format, wherein the target field includes at least one of the following: a field for representing web page characters and a field for representing web page tags; a first storage unit, which is used to store the main text part of the threat intelligence text in the second type format as a third type format, and obtain W threat intelligence texts in the third type format, wherein W is a positive integer less than or equal to K; and a first construction unit, which is used to construct a training set for training a model based on Y threat intelligence texts in the third type format and / or W threat intelligence texts in the third type format.

[0103] Optionally, in the classification device for threat intelligence text provided in an embodiment of the present application, the first training module includes: a first output unit, used to input each sample data in the training set into the original classification model for classification processing, and output the category information corresponding to each sample data; a first calculation unit, used to calculate the loss function based on the category information corresponding to each sample data and the actual category information corresponding to each sample data; a first processing unit, used to update the model parameters of the original classification model according to the loss function to obtain updated model parameters; a second determination unit, used to obtain the target classification model based on the updated model parameters.

[0104] Optionally, in the threat intelligence text classification device provided in the embodiment of the present application, the device also includes: a first determination module, which is used to input the target vector corresponding to each threat intelligence text into the target classification model for classification processing, output the category information to which each threat intelligence text belongs, and then determine the threat behavior information in the target network according to the category information to which each threat intelligence text belongs; a fourth processing module, which is used to generate prompt information based on the threat behavior information in the target network, and send the prompt information to the target object, wherein the prompt information is used to prompt the target object to process the threat behavior in the target network.

[0105] The execution process of each module of the threat intelligence text classification device and its beneficial effects can be referred to the relevant content of the aforementioned method embodiment, and will not be repeated here.

[0106] The classification device for threat intelligence text includes a processor and a memory. The above-mentioned first acquisition module 401, first processing module 402, first conversion module 403 and first output module 404 are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.

[0107] The processor contains a kernel, which retrieves the corresponding program unit from the memory. You can set one or more kernels, and adjust kernel parameters to improve the accuracy of threat intelligence text classification.

[0108] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0109] An embodiment of the present invention provides a computer-readable storage medium having a program stored thereon, which implements the method for classifying threat intelligence text when executed by a processor.

[0110] An embodiment of the present invention provides a processor, which is used to run a program, wherein the program executes the threat intelligence text classification method when running.

[0111] like Figure 5 As shown, an embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and runnable on the processor. When the processor executes the program, the following steps are implemented: obtaining N threat intelligence texts to be classified, wherein the threat intelligence texts are used to record information related to the threat behavior of the target network, and N is a positive integer; filtering M target texts among the N threat intelligence texts to obtain S threat intelligence texts, wherein the target texts include texts with less than a preset number of characters and / or garbled texts, and both M and S are positive integers less than N; performing vector conversion on each threat intelligence text in the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text; inputting the target vector corresponding to each threat intelligence text into a target classification model for classification processing, and outputting category information to which each threat intelligence text belongs, wherein the target classification model is a model constructed based on a convolutional neural network and a residual network.

[0112] When the processor executes the program, the following steps are also implemented: performing vector conversion on each threat intelligence text in the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text, including: judging whether each threat intelligence text in the S threat intelligence texts contains a compromise indicator, wherein the compromise indicator is used to determine whether the target network has been attacked; if there is at least one threat intelligence text in the S threat intelligence texts that does not contain the compromise indicator, then treating the at least one threat intelligence text as a first category text, and treating the threat intelligence texts in the S threat intelligence texts other than the at least one threat intelligence text as a second category text; performing vector conversion on each threat intelligence text in the first category text and the second category text to obtain a target vector corresponding to each threat intelligence text.

[0113] When the processor executes the program, the following steps are also implemented: performing vector conversion on each threat intelligence text in the first category of text and the second category of text to obtain a target vector corresponding to each threat intelligence text, including: obtaining the title information of each threat intelligence text in the first category of text and the second category of text; truncating the characters that exceed the truncation length in each threat intelligence text in the first category of text and the second category of text to obtain T target characters in each threat intelligence text, wherein the length of the T target characters is equal to the truncation length, and T is a positive integer greater than 1; splicing the title information of each threat intelligence text and the T target characters in each threat intelligence text to obtain the spliced ​​information corresponding to each threat intelligence text; performing vector conversion on the spliced ​​information corresponding to each threat intelligence text to obtain the target vector corresponding to each threat intelligence text.

[0114] When the processor executes the program, it also implements the following steps: truncating the characters in each threat intelligence text in the first category of text and the second category of text that exceed the truncation length to obtain T target characters in each threat intelligence text, including: judging whether the length of the characters in each threat intelligence text in the first category of text and the second category of text is greater than the truncation length; if the length of the characters in each threat intelligence text in the first category of text and the second category of text is greater than the truncation length, directly truncating the characters in each threat intelligence text in the first category of text and the second category of text that exceed the truncation length to obtain T target characters in each threat intelligence text; if there is at least one threat intelligence text in the first category of text and the second category of text whose length is not greater than the truncation length, using the target symbol to pad the characters in the at least one threat intelligence text to obtain the padded characters corresponding to the at least one threat intelligence text, wherein the length of the padded characters is equal to the truncation length; and using the padded characters corresponding to the at least one threat intelligence text as the T target characters corresponding to the at least one threat intelligence text.

[0115] When the processor executes the program, the following steps are also implemented: the target classification model is obtained by: obtaining K sample data including threat intelligence texts in multiple formats, and constructing a training set for training the model based on the K sample data, wherein K is a positive integer greater than 1; adding a residual module to the convolutional neural network to obtain a residual network, and constructing an original classification model for classifying the threat intelligence text based on the residual network; using the training set to train the original classification model to obtain the target classification model.

[0116] When the processor executes the program, the following steps are also implemented: if the K sample data include threat intelligence text in the first type format, the threat intelligence text in the first type format in the K sample data is converted into threat intelligence text in the third type format to obtain Y threat intelligence texts in the third type format, wherein Y is a positive integer less than or equal to K; if the K sample data include threat intelligence text in the second type format, the text part of the threat intelligence text in the second type format is extracted according to the target field in the threat intelligence text in the second type format in the K sample data, wherein the target field includes at least one of the following: a field for representing web page characters and a field for representing web page tags; the text part of the threat intelligence text in the second type format is stored in the third type format to obtain W threat intelligence texts in the third type format, wherein W is a positive integer less than or equal to K; based on the Y threat intelligence texts in the third type format and / or the W threat intelligence texts in the third type format, the training set for training the model is constructed.

[0117] When the processor executes the program, the following steps are also implemented: the original classification model is trained using the training set to obtain the target classification model, including: each sample data in the training set is input into the original classification model for classification processing, and the category information corresponding to each sample data is output; based on the category information corresponding to each sample data and the actual category information corresponding to each sample data, a loss function is calculated; the model parameters of the original classification model are updated according to the loss function to obtain updated model parameters; based on the updated model parameters, the target classification model is obtained.

[0118] When the processor executes the program, the following steps are also implemented: after inputting the target vector corresponding to each threat intelligence text into the target classification model for classification processing and outputting the category information to which each threat intelligence text belongs, the method also includes: determining the threat behavior information in the target network based on the category information to which each threat intelligence text belongs; generating prompt information based on the threat behavior information in the target network, and sending the prompt information to the target object, wherein the prompt information is used to prompt the target object to process the threat behavior in the target network.

[0119] The devices in this article can be servers, PCs, PADs, mobile phones, etc.

[0120] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of the threat intelligence text classification method provided in any of the aforementioned embodiments.

[0121] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0122] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0123] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0125] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0126] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0127] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0128] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0129] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0130] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A method for classifying threat intelligence text, characterized in that: include: Obtain N threat intelligence texts to be classified, wherein the threat intelligence texts are used to record information related to threat behaviors of a target network, and N is a positive integer; Filtering M target texts from the N threat intelligence texts to obtain S threat intelligence texts, wherein the target texts include texts with fewer than a preset number of characters and / or garbled texts, and M and S are both positive integers less than N; Performing vector conversion on each of the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text; The target vector corresponding to each threat intelligence text is input into the target classification model for classification processing, and the category information to which each threat intelligence text belongs is output, wherein the target classification model is a model constructed based on convolutional neural network and residual network.

2. The method according to claim 1, characterized in that Performing vector conversion on each of the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text includes: Determining whether each of the S threat intelligence texts contains a compromise indicator, wherein the compromise indicator is used to determine whether the target network has been attacked; If at least one threat intelligence text among the S threat intelligence texts does not contain the compromise indicator, the at least one threat intelligence text is regarded as a first-category text, and the threat intelligence texts other than the at least one threat intelligence text among the S threat intelligence texts are regarded as second-category texts; Vector conversion is performed on each threat intelligence text in the first category of text and the second category of text to obtain a target vector corresponding to each threat intelligence text.

3. The method according to claim 2, characterized in that Performing vector conversion on each threat intelligence text in the first category of text and the second category of text to obtain a target vector corresponding to each threat intelligence text includes: Obtaining title information of each threat intelligence text in the first category of text and the second category of text; Truncating characters exceeding a truncation length in each threat intelligence text in the first category of text and the second category of text to obtain T target characters in each threat intelligence text, where the length of the T target characters is equal to the truncation length, and T is a positive integer greater than 1; The title information of each threat intelligence text and T target characters in each threat intelligence text are concatenated to obtain the concatenated information corresponding to each threat intelligence text; Perform vector conversion on the concatenated information corresponding to each threat intelligence text to obtain the target vector corresponding to each threat intelligence text.

4. The method according to claim 3, characterized in that Characters exceeding the truncation length in each threat intelligence text in the first category of text and the second category of text are truncated, and T target characters in each threat intelligence text are obtained, including: Determining whether a length of characters in each threat intelligence text in the first category of text and the second category of text is greater than the truncation length; If the length of characters in each threat intelligence text in the first category of text and the second category of text is greater than the truncation length, directly truncate the characters in each threat intelligence text in the first category of text and the second category of text that exceed the truncation length to obtain T target characters in each threat intelligence text; If the length of characters in at least one threat intelligence text in the first category of text and the second category of text is not greater than the truncation length, padding the characters in the at least one threat intelligence text with the target symbol to obtain padded characters corresponding to the at least one threat intelligence text, wherein the length of the padded characters is equal to the truncation length; The padded characters corresponding to the at least one threat intelligence text are used as T target characters corresponding to the at least one threat intelligence text.

5. The method according to any one of claims 1 to 4, characterized in that The target classification model is obtained in the following way: Obtain K sample data including threat intelligence texts in multiple formats, and construct a training set for training a model based on the K sample data, where K is a positive integer greater than 1; Adding a residual module to a convolutional neural network to obtain a residual network, and constructing an original classification model for classifying the threat intelligence text based on the residual network; The original classification model is trained using the training set to obtain the target classification model.

6. The method according to claim 5, characterized in that Based on the K sample data, constructing a training set for training the model includes: If the K sample data include threat intelligence text in a first format, convert the threat intelligence text in the first format in the K sample data into threat intelligence text in a third format to obtain Y threat intelligence texts in the third format, where Y is a positive integer less than or equal to K; If the K sample data include threat intelligence text in the second type format, extracting a text portion of the threat intelligence text in the second type format according to a target field in the threat intelligence text in the second type format in the K sample data, wherein the target field includes at least one of the following: a field for representing web page characters and a field for representing web page tags; Storing the body of the threat intelligence text in the second format in a third format to obtain W threat intelligence texts in the third format, where W is a positive integer less than or equal to K; Based on the Y threat intelligence texts in the third type format and / or the W threat intelligence texts in the third type format, the training set for training the model is constructed.

7. The method according to claim 5, characterized in that The original classification model is trained using the training set to obtain the target classification model, including: Input each sample data in the training set into the original classification model for classification processing, and output the category information corresponding to each sample data; Based on the category information corresponding to each sample data and the true category information corresponding to each sample data, the loss function is calculated; Updating the model parameters of the original classification model according to the loss function to obtain updated model parameters; Based on the updated model parameters, the target classification model is obtained.

8. The method according to claim 1, characterized in that After inputting the target vector corresponding to each threat intelligence text into the target classification model for classification processing and outputting category information to which each threat intelligence text belongs, the method further includes: Determining threat behavior information in the target network based on the category information to which each threat intelligence text belongs; Prompt information is generated based on the threat behavior information in the target network, and the prompt information is sent to the target object, wherein the prompt information is used to prompt the target object to handle the threat behavior in the target network.

9. A device for classifying threat intelligence text, characterized in that: include: A first acquisition module is configured to acquire N threat intelligence texts to be classified, wherein the threat intelligence texts are used to record information related to threat behaviors of a target network, and N is a positive integer; A first processing module is configured to filter M target texts from the N threat intelligence texts to obtain S threat intelligence texts, wherein the target texts include texts with fewer than a preset number of characters and / or garbled texts, and M and S are both positive integers less than N; A first conversion module is configured to perform vector conversion on each of the S threat intelligence texts to obtain a target vector corresponding to each threat intelligence text; The first output module is used to input the target vector corresponding to each threat intelligence text into the target classification model for classification processing, and output the category information to which each threat intelligence text belongs, wherein the target classification model is a model constructed based on a convolutional neural network and a residual network.

10. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the threat intelligence text classification method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Text classification method and device

    CN108536815A

  • Threat index analysis method and analysis device

    CN114297377A

  • Threat intelligence classification method and device, electronic equipment and storage medium

    CN115495744A

  • Threat intelligence classification method and device, electronic equipment and storage medium

    CN117312943A