IOC information extraction method and device, equipment and medium

Through large language model collaboration in multi-agent collaborative architecture, IOC information is screened, extracted and formatted to convert IOC information, which solves the problem of time-consuming and labor-intensive model training and low extraction accuracy in natural language processing technology, and achieves efficient and accurate IOC information extraction.

CN120185848APending Publication Date: 2025-06-20上海市大数据中心 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510099268.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, when realizing IOC information extraction based on natural language processing, model training involves time-consuming and labor-intensive manual labeling, and the accuracy of IOC information extraction is low.

Method used

The multi-agent collaboration architecture is adopted, and multiple large language models are used to filter, extract and format the IOC information, and through functionally independent agent cooperation, complex tasks are decomposed and task processing efficiency is improved.

Benefits of technology

It realizes efficient and accurate IOC information extraction, reduces dependence on manual annotation, improves extraction accuracy, and is suitable for diverse intelligence sources and new IOC categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120185848A_ABST
    Figure CN120185848A_ABST
Patent Text Reader

Abstract

The invention discloses an IOC information extraction method and device, equipment and a medium, and relates to the technical field of information security and artificial intelligence, and the method comprises the steps: screening out a paragraph text containing IOC information from threat intelligence through a first agent in a multi-agent collaborative architecture; extracting IOC information in an initial format in the paragraph text by using a second agent in the multi-agent collaborative architecture; performing format conversion on the IOC information in the initial format by utilizing a fourth agent in the multi-agent collaborative architecture to obtain target IOC information; wherein the first agent, the second agent and the fourth agent are large language models. The method is suitable for IOC information extraction scenes such as network security information processing in the fields of finance, medical treatment and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of information security and artificial intelligence, and particularly to an IOC information extraction method, device, equipment, and medium. Background Art

[0002] In network security, threat intelligence is an important resource for defending against cyberattacks, and the key data contained in threat indicators (IOC: Indicators of Compromise) is an important basis for describing cyberattack behaviors. Traditional IOC information extraction methods usually rely on named entity recognition (NER) technology in natural language processing (NLP), using models such as BERT or LSTM+CRF. However, this IOC information extraction method has the following problems:

[0003] 1. Require a large amount of manual annotation: Model training depends on manual annotation, which is time-consuming and laborious.

[0004] 2. Text extraction errors: Threat intelligence may come from PDFs or web pages, and information may be incomplete during text extraction due to issues such as spaces or line breaks.

[0005] 3. Defang operation interference: For security reasons, information such as URLs and IPs in threat intelligence may be modified to forms such as example[.]com, and traditional models are difficult to correctly identify. Summary of the Invention

[0006] In view of this, this application provides an IOC information extraction method, device, equipment, and medium, mainly aiming to solve the technical problems that when implementing IOC information extraction based on natural language processing in the prior art, model training involves manual annotation, which is time-consuming and laborious, and the accuracy of IOC information extraction is relatively low.

[0007] According to one aspect of this application, an IOC information extraction method is provided, and this method includes:

[0008] Using the first agent in the multi-agent collaboration architecture to screen out the paragraph text containing IOC information from threat intelligence;

[0009] Using the second agent in the multi-agent collaboration architecture to extract the IOC information in the initial format from the paragraph text;

[0010] Using the fourth agent in the multi-agent collaboration architecture to obtain the target IOC information by performing format conversion on the IOC information in the initial format;

[0011] Wherein, the first agent, the second agent, and the fourth agent are large language models.

[0012] According to another aspect of the present application, an IOC information extraction device is provided, and the device includes:

[0013] A first agent module, configured to use the first agent in the multi-agent collaborative architecture to screen out the passage text containing IOC information from the threat intelligence;

[0014] A second agent module, configured to use the second agent in the multi-agent collaborative architecture to extract the IOC information in the initial format from the passage text;

[0015] A fourth agent module, configured to use the fourth agent in the multi-agent collaborative architecture to obtain the target IOC information by performing format conversion on the IOC information in the initial format;

[0016] Wherein, the first agent, the second agent, and the fourth agent are large language models.

[0017] According to another aspect of the present application, a computer storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned IOC information extraction method is implemented.

[0018] According to still another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the program, the above-mentioned IOC information extraction method is implemented.

[0019] By means of the above technical solutions, compared with the prior art in which when implementing IOC information extraction based on natural language processing, model training involves manual annotation which is time-consuming and laborious, and the accuracy of IOC information extraction is relatively low, in the present application, the first agent in the multi-agent collaborative architecture is used to screen out the passage text containing IOC information from the threat intelligence; the second agent in the multi-agent collaborative architecture is used to extract the IOC information in the initial format from the passage text; the fourth agent in the multi-agent collaborative architecture is used to obtain the target IOC information by performing format conversion on the IOC information in the initial format; wherein, the first agent, the second agent, and the fourth agent are large language models. It can be seen that through the collaborative work of multiple functionally independent intelligent agents Agent, the functions of each intelligent agent Agent are clear, the complex task is decomposed into independent subtasks, the task processing efficiency is improved, and there is no need for model training with a large amount of manual marking, so as to achieve efficient and accurate IOC information extraction.

[0020] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0022] Figure 1 The flowchart of an IOC information extraction method provided by an embodiment of the present application is shown;

[0023] Figure 2 The flowchart of another IOC information extraction method provided by an embodiment of the present application is shown;

[0024] Figure 3 The structural schematic diagram of an IOC information extraction device provided by an embodiment of the present application is shown;

[0025] Figure 4 The structural schematic diagram of another IOC information extraction device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0027] Aiming at the technical problems in the prior art that when implementing IOC information extraction based on natural language processing, model training involves manual annotation, which is time-consuming and laborious, and the accuracy of IOC information extraction is relatively low. This embodiment provides an IOC information extraction method. Through the cooperation of multiple intelligent agents with independent functions, the functions of each intelligent agent are clear, and complex tasks are decomposed into independent subtasks, improving the task processing efficiency, so as to achieve efficient and accurate IOC information extraction. As Figure 1 shown, the above method includes the following steps:

[0028] Step 101: Use the first intelligent agent in the multi-agent cooperation architecture to screen out the paragraph text containing IOC information from the threat intelligence.

[0029] In this embodiment, the large language model can cover a large amount of cross-domain corpus due to pre-training, including threat intelligence-related content, and has the ability to directly understand and extract IOC information. Therefore, in this embodiment, without fine-tuning the model, the performance of the large language model is fully utilized, without additional training or annotation data. Using the architecture of multi-agent, complex tasks are decomposed into independent sub-tasks, and each agent is responsible for completing them, improving the task processing efficiency, and further enhancing the accuracy of IOC information extraction.

[0030] Step 102: Use the second agent in the multi-agent collaborative architecture to extract the IOC information in the initial format from the paragraph text.

[0031] In this embodiment, the second agent is the IOCExtractor Agent and also a large language model. Its task is to extract IOC information such as IP addresses, Email, URLs, Domains, and file hashes from the selected paragraph text. According to the requirements of the actual application scenario, the content of the IOC information is not specifically limited here. Specifically, by setting system prompts and sample examples, the second agent can directly utilize the ability of the large language model to understand and generate natural language without additional training or annotation data, and extract the IOC information in the initial format from the paragraph text.

[0032] Step 103: Use the fourth agent in the multi-agent collaborative architecture to obtain the target IOC information by performing format conversion on the IOC information in the initial format.

[0033] In this embodiment, the first agent, the second agent, and the fourth agent are large language models. According to the requirements of the actual application scenario, the large language models here can be the same multiple large language models or different multiple large language models. The large language models here are different from the existing natural language processing models such as BERT or LSTM+CRF, which do not require a large amount of manual annotation for model training. Instead, by setting system prompts and sample examples for each large language model, complex tasks are decomposed into independent sub-tasks to achieve task distribution and improve task processing efficiency.

[0034] For this embodiment, according to the above solution, the first agent in the multi-agent collaboration architecture can be used to screen out the paragraph text containing IOC information from the threat intelligence; the second agent in the multi-agent collaboration architecture can be used to extract the IOC information in the initial format from the paragraph text; the fourth agent in the multi-agent collaboration architecture can be used to obtain the target IOC information by performing format conversion on the IOC information in the initial format. Among them, the first agent, the second agent, and the fourth agent are large language models. Compared with the existing technical solution for IOC information extraction based on natural language processing, where model training involves time-consuming and laborious manual annotation and the accuracy of IOC information extraction is relatively low, the agents in this embodiment can achieve information flow through a preset working logic and the instruction execution mechanism of the large language model, decompose complex tasks into independent subtasks, improve task processing efficiency, and do not require a large amount of manually labeled model training, thus realizing efficient and accurate IOC information extraction.

[0035] Furthermore, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the specific implementation process of this embodiment, another IOC information extraction method is provided. Since existing large language models need to prepare a large amount of threat intelligence and IOC information extraction results to form training samples, and then perform model fine-tuning on the basis of the base model to make it a target model for IOC information extraction, but model fine-tuning requires a large amount of computing power. Therefore, this embodiment can solve the technical problems in the existing technical solution for IOC information extraction based on natural language processing, where model training involves time-consuming and laborious manual annotation and the accuracy of IOC information extraction is relatively low. This embodiment can achieve efficient and accurate IOC information extraction through the collaborative work of multiple functionally independent agents. As Figure 2 shown, it is applicable to network security intelligence processing in fields such as finance and healthcare (for example, scenarios such as threat intelligence platforms, network security situation awareness systems, and enterprise security operation centers SOC), realizes cross-domain adaptation, and the IOC information extraction method based on the collaborative work of multiple agents of the large language model includes:

[0036] Step 201: Use the first agent in the multi-agent collaboration architecture to perform paragraph segmentation processing on the threat intelligence to obtain multiple paragraph texts.

[0037] Step 202: Perform IOC information recognition on the multiple paragraph texts, screen out the paragraph texts containing IOC information from the multiple initial paragraph texts, and send them to the second agent.

[0038] In implementation, the IOC information includes key data such as the attacker's IP address, Email, URL, Domain, and file hash values (such as SHA256, SHA1, MD5). The first intelligent agent is the SnippetSelector Agent, which is also a large language model. Its task is to identify and extract the paragraph text containing IOC information from threat intelligence through field screening. Specifically, by setting system prompts and sample examples, the first intelligent agent can eliminate paragraph text unrelated to IOC information through preliminary screening, enabling the second intelligent agent to reduce the processing of invalid text, quickly extract IOC key data, improve the extraction efficiency, and achieve the purpose of effectively reducing the workload of subsequent task processing.

[0039] Furthermore, the SnippetSelector Agent uses the large language model LLM to perform paragraph segmentation processing on the input threat intelligence and screen out the paragraph text that may contain IOC information. For example, eliminate the background description paragraphs and only retain the content related to the attacker's IP and URL. After determining that the paragraph text output by the SnippetSelector Agent contains IOC information, pass the paragraph text to the second intelligent agent, the IOCExtractor, to achieve the dynamic decision-making of the SnippetSelector Agent. Among them, the optimization strategies for screening can include improving the screening accuracy of the paragraph text containing IOC information by combining regular expressions, context scoring mechanisms, etc.

[0040] Furthermore, the output text of each intelligent agent in the multi-agent collaborative architecture includes status information, which is the execution result information of the current intelligent agent's task. Specifically, based on the preset working logic between different intelligent agents Agent in the information flow mechanism and the instruction execution mechanism of the large language model, task distribution is realized, enabling each intelligent agent Agent to generate corresponding output results based on the input data and the carried status information (the task execution result of the previous intelligent agent Agent), and the output result includes the task execution result of the current intelligent agent Agent, thereby realizing information flow and the transmission of status information. Among them, the status information is used to identify the task execution result of the current intelligent agent Agent. For example, "screening completed" (corresponding to the processing stage of the first intelligent agent Agent) or "verification passed / successful" (corresponding to the processing stage of the third intelligent agent).

[0041] According to the needs of the actual application scenario, set the system prompt for the first agent Agent, specifically, fill in the semantic description of the corresponding IOC type according to the IOC type extracted by Critieria, and when there are new IOC types to be extracted later, add the semantic description of the new IOC type directly to the system prompt of the first agent Agent. An example is:

[0042] You are SnippetSelector, an agent specialized in identifying and extracting relevant paragraphs from long texts based on user needs. Your task is to read through the provided text and select paragraphs that match the user's specified criteria. Do not modify the selected paragraphs; output them exactly as they appear in the original text.

[0043] (You are SnippetSelector, an agent specialized in identifying and extracting relevant paragraphs from a long text based on user needs. Your task is to read through the provided text and select the paragraphs that meet the user-specified criteria. Do not modify the selected paragraphs; output them as they are in the original text.)

[0044] Instructions:

[0045] 1. Carefully read the user's criteria.

[0046] 2.Scan the provided text and identify paragraphs that meet the criteria.

[0047] 3. Output the identified paragraphs without any modifications.

[0048] 4. The output is constructed with the selected paragraphs and does not generate any explanation.

[0049] 5. If no paragraph meets the criteria, output "---".

[0050] User's Criteria

[0051] {criteria}

[0052] Text

[0053] {content (input threat intelligence content)}

[0054] Step 203: Use the second agent in the multi-agent collaboration architecture to extract the IOC information in the initial format from the paragraph text.

[0055] To illustrate the specific implementation of step 203, as a preferred embodiment, step 203 includes: if the paragraph text includes rewritten information, perform a restoration process on the rewritten information to obtain the original paragraph text; and extract the IOC information in the initial format from the original paragraph text to obtain the IOC information in the initial format of the threat intelligence.

[0056] In implementation, the second agent IOCExtractor extracts the corresponding IOC information according to the output result from the first agent SnippetSelector according to the information category of the IOC information. The information categories of the IOC information include: IP address (supporting IPv4 and IPv6), Email, URL, Domain, file hash value (SHA256, SHA1, MD5). Further, for the Defang form (such as http[:] / / example[.]com), the large language model can automatically perform a restoration process on the rewritten information to obtain the text information in the original format.

[0057] According to the requirements of the actual application scenario, set the system prompt for the second agent Agent, and ioc_json_schema is the semantic description of the IOC type to be extracted. The example is as follows:

[0058] ##Role

[0059] You are a cybersecurity expert specializing in threat intelligence research, with extensive knowledge of IoCs (Indicators of Compromise) related to malware families.

[0060] ##Task

[0061] 1. Verify if the article contains any IoC information, such as filehash values (sha256, sha1, md5), domains, emails, ipv4, ipv6, and URLs.

[0062] 2. If no IoC information is found, return an empty JSON.

[0063] 3. If IoC information is present, extract it and organize it use the following JSON schema:

[0064] ```json

[0065] {ioc_json_schema}

[0066] ```

[0067] ##Note

[0068] 1. Only output the IoC information extracted from the original text; do not include sample values.

[0069] 2. The original text may have an incorrect format. Ignore it and extract all the values.

[0070] 3. Provide only the JSON result.

[0071] ##Original Text

[0072] {content}

[0073] ##IOC already extracted

[0074] {correct_ioc_list} (List of IOC information with correct content)

[0075] ##IOC may contain errors

[0076] {incorrect_ioc_list} (List of IOC information with incorrect content)

[0077] Step 204: Use the third agent in the multi-agent collaborative architecture to verify the IOC information in the initial format. If the verification is successful, send the IOC information in the initial format to the fourth agent; if the verification fails, generate correction information for the IOC information in the initial format and send it to the second agent.

[0078] In the implementation, the verification process includes format verification and context verification. The third agent IOCValidator includes verification rules for format verification (for example, verification code) and a large language module for context verification, which is used to verify the output result of the second agent IOCExtractor. The input data is the paragraph text screened by the first agent and the IOC information extracted by the second agent. The rationality of the IOC information is verified by comparing the context of the IOC information in the paragraph text, and the format of the IOC information is verified based on the format rules. If the output result is a verification failure, correction information is generated based on the error content and fed back to the second agent IOCExtractor, which is corrected by the second agent IOCExtractor; if the output result is a verification success, the IOC information in the initial format and the status information of successful verification are directly sent to the fourth agent. Among them, format verification includes the legality of the IP address, the length of the hash value, etc., and the context verification is to verify the context of the IOC information position in the paragraph text to check whether the extracted IOC information is reasonable.

[0079] According to the needs of the actual application scenario, set the system prompt for the third agent Agent. The example is:

[0080] ##Role

[0081] You are an expert in verifying Indicators of Compromise (IOCs) extracted from cybersecurity intelligence reports.

[0082] ##Inputs:

[0083] A raw intelligence text.(Paragraph text)

[0084] A list of extracted IOCs from a prior step.

[0085] ##Tasks:

[0086] Review the list of extracted IOCs against the raw intelligence text.

[0087] For each IOC in the list: (Traverse each IOC information in the list)

[0088] Confirm if it is correctly extracted.

[0089] Identify any errors (e.g., misidentified or incorrectly formatted IOCs).

[0090] Determine if there are any missing IOCs that were not extracted but should be present based on the raw intelligence text.

[0091] ##Outputs:

[0092] Return two separate lists:

[0093] Correct IOCs: A list of IOCs that are accurately extracted.

[0094] Incorrect IOCs: A list of IOCs that are incorrect or improperly extracted, along with an explanation of the issue for each.

[0095] If any IOCs are missing, respond "This list may be incomplete"

[0096] If the IOC extraction is fully correct, confirm that the extraction is accurate and provide the correct IOC list.

[0097] Example Output Format

[0098] Correct IOCs: [List of correct IOCs]

[0099] Incorrect IOCs: [List of incorrect IOCs with explanations]

[0100] Overall Validation: "The extraction is accurate." OR "The extraction has errors."

[0101] To illustrate the specific implementation of step 203, as a preferred embodiment, step 203 further includes: re-extracting IOC information from the paragraph text according to the paragraph text from the first agent and the correction information of the IOC information in the initial format to obtain the IOC information in the initial format of the threat intelligence.

[0102] In implementation, the third agent, as an error correction mechanism, uses the third agent IOCValidator to verify and feedback the mechanism for correcting the IOC information extraction result, which can improve the accuracy of IOC information extraction through the verification and feedback process. Therefore, when the second agent IOCExtractor re-extracts IOC information from the corresponding passage text according to the correction information from the third agent IOCValidator, it passes the output result and status information to the third agent IOCValidator again until the third agent IOCValidator verifies successfully, and then passes the output result and status information of the third agent IOCValidator to the fourth agent. It can be seen that based on the information flow mechanism of this implementation, it can ensure that each agent works collaboratively, providing high flexibility, dynamically adjusting the data processing path based on the third agent, thereby improving the task processing efficiency and accuracy.

[0103] It should be noted that when the third agent sends the correction information to the second agent, the second agent re-extracts IOC information from the passage text according to the passage text and correction information from the first agent. According to the requirements of the actual application scenario, the correction information includes the IOC information in the initial format extracted by the second agent before (which can be all IOC information or the IOC information of the error content) and the corresponding correction information. Therefore, the passage text from the first agent and the correction information from the third agent are used as the input data of the second agent, so that the second agent can quickly locate the position of the IOC information (of the error content) and re-extract IOC information from the passage text.

[0104] Step 205: Use the fourth agent in the multi-agent collaborative architecture to obtain the target IOC information by performing format conversion on the IOC information in the initial format.

[0105] In implementation, the fourth agent IOCExporter receives the IOC information in the initial format transmitted from the third agent, determines the preset template and preset format according to the downstream application format, and thus performs data formatting output on the IOC information verified successfully by the third agent to obtain the target IOC information. Among them, the preset format is JSON or CSV to adapt to the downstream application format for downstream use.

[0106] An example of the output result in JSON format is:

[0107]

[0108]

[0109] Step 206: When the threat intelligence includes a new IOC type, change the system instructions and sample examples of the agents in the multi-agent collaboration architecture according to the new IOC type.

[0110] In implementation, each agent in the multi-agent collaboration architecture contains an independent system prompt and few-shot examples. When the threat intelligence includes a new IOC type, there is no need to relabel the data and retrain the model, that is, there is no need to re-optimize and train the large language model in each agent. Instead, by directly utilizing the performance of the large language model and adjusting the system instructions and sample examples in the agent, the expansion and verification of the new IOC type can be achieved without changing the core logic. It can be seen that the multi-agent collaboration architecture in this implementation can be applicable to various threat intelligence formats, quickly expand and support new IOC categories, and thus has high expansion performance.

[0111] By applying the technical solution of this embodiment, an IOC information extraction method based on a large language model and multi-Agent collaboration, the multi-agent collaboration architecture (an IOC information extraction architecture supporting cross-domain expansion and multi-format output) includes a first agent SnippetSelector, a second agent IOCExtractor, a third agent IOCValidator, and a fourth agent IOCExporter. The first agent SnippetSelector is used to segment and screen the threat intelligence, and based on the large language model of the second agent, extract IOC information from the refined paragraph text. The mechanism of using the third agent IOCValidator to verify and feedback to correct the IOC extraction result ensures that the fourth agent IOCExporter can output the verified IOC information in the specified format. Compared with the existing technical solutions for IOC information extraction based on natural language processing, where model training involves manual annotation, is time-consuming and laborious, the accuracy of IOC information extraction is relatively low, and the expansion performance is insufficient. This embodiment can, based on the large language model in the multi-agent collaboration architecture, through context understanding, exclude irrelevant data, improve the effect and accuracy of IOC information extraction. At the same time, without model training, it can be applicable to diverse intelligence sources and the needs of new IOC categories, enhancing flexibility. Further, based on the modular design of the multi-agent collaboration architecture, each agent runs independently, facilitating subsequent maintenance and optimization. In addition, through experimental results, it is obtained that compared with traditional NER methods (such as BERT+CRF), the extraction accuracy on the threat intelligence dataset is increased by 15%.

[0112] Further, as Figure 1 a specific implementation of the method, the embodiment of the present application provides an IOC information extraction device, such as Figure 3As shown, the device includes: a first proxy module 31, a second proxy module 32, and a fourth proxy module 34.

[0113] The first proxy module 31 is used to screen out the paragraph text containing IOC information from the threat intelligence by using the first agent in the multi-agent collaborative architecture.

[0114] The second proxy module 32 is used to extract the IOC information in the initial format from the paragraph text by using the second agent in the multi-agent collaborative architecture.

[0115] The fourth proxy module 34 is used to obtain the target IOC information by performing format conversion on the IOC information in the initial format by using the fourth agent in the multi-agent collaborative architecture; wherein, the first agent, the second agent, and the fourth agent are large language models.

[0116] In a specific application scenario, such as Figure 4 As shown, the device further includes: a third proxy module 33.

[0117] The third proxy module 33 is used to perform verification processing on the IOC information in the initial format by using the third agent in the multi-agent collaborative architecture. If the verification is successful, the IOC information in the initial format is sent to the fourth agent; if the verification fails, correction information of the IOC information in the initial format is generated and sent to the second agent; wherein, the verification processing includes format verification and context verification.

[0118] In a specific application scenario, the output text of each agent in the multi-agent collaborative architecture includes status information, and the status information is the execution result information of the current agent performing the task.

[0119] In a specific application scenario, the first proxy module 31 includes: a segmentation sub-module 311 and a screening sub-module 312.

[0120] The segmentation sub-module 311 is used to perform paragraph segmentation processing on the threat intelligence by using the first agent in the multi-agent collaborative architecture to obtain multiple paragraph texts.

[0121] The screening sub-module 312 is used to identify IOC information in the multiple paragraph texts, screen out the paragraph texts containing IOC information from the multiple initial paragraph texts, and send them to the second agent.

[0122] In a specific application scenario, the second proxy module 32 includes: a restoration sub-module 321 and an extraction sub-module 322.

[0123] The reduction sub-module 321 is configured to, if the paragraph text includes rewritten information, perform reduction processing on the rewritten information to obtain the original paragraph text.

[0124] The extraction sub-module 322 is configured to extract IOC information in an initial format from the original paragraph text to obtain the IOC information in the initial format of the threat intelligence.

[0125] In a specific application scenario, the second agent module 32 further includes: a correction sub-module 323.

[0126] The correction sub-module 323 is configured to re-extract IOC information from the paragraph text according to the paragraph text from the first agent and the correction information of the IOC information in the initial format to obtain the IOC information in the initial format of the threat intelligence.

[0127] In a specific application scenario, the device is further configured to, when the threat intelligence includes a new IOC type, change the system instructions and sample examples of the agents in the multi-agent collaboration architecture according to the new IOC type.

[0128] It should be noted that for other corresponding descriptions of each functional unit involved in the IOC information extraction device provided in the embodiments of the present application, reference can be made to Figure 1 and Figure 2 the corresponding descriptions therein, which will not be elaborated herein.

[0129] Based on the above as Figure 1 and Figure 2 shown in the method, correspondingly, the embodiments of the present application further provide a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the IOC information extraction method as shown in Figure 1 and Figure 2 shown.

[0130] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, and the software product can be stored in a storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present application.

[0131] Based on the above as Figure 1 、 Figure 2 shown in the method, and Figure 3 、 Figure 4In the virtual device embodiment shown, to achieve the above object, an embodiment of the present application further provides a computer device, which may specifically be a personal computer, a server, a network device, etc. The physical device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the IOC information extraction method as shown in Figure 1 and Figure 2 the figure.

[0132] Optionally, the computer device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, etc. The user interface may include a display screen and an input unit such as a keyboard. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a Bluetooth interface, a WI-FI interface), etc.

[0133] Those skilled in the art can understand that the structure of the computer device provided in this embodiment does not limit the physical device, and it may include more or fewer components, or combine some components, or have different component arrangements.

[0134] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the computer device, and supports the operation of the information processing program and other software and / or programs. The network communication module is used to implement communication between the components inside the storage medium, and communication between other hardware and software in the physical device.

[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform, or by hardware. By applying the technical solution of the present application, compared with the existing technical solution for IOC information extraction based on natural language processing, where model training involves manual annotation, is time-consuming and laborious, the accuracy of IOC information extraction is relatively low, and the scalability is insufficient, this embodiment can be based on the large language model in the multi-agent collaborative architecture, through context understanding, exclude irrelevant data, improve the IOC information extraction effect and accuracy. At the same time, without model training, it can be applicable to the needs of diverse intelligence sources and new IOC categories, enhancing flexibility; further, based on the modular design of the multi-agent collaborative architecture, each agent runs independently, facilitating subsequent maintenance and optimization.

[0136] Those skilled in the art can understand that the attached drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the attached drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the devices in the implementation scenario can be distributed in the devices in the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed and located in one or more devices different from the present implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0137] The above serial numbers of the present application are only for description and do not represent the advantages or disadvantages of the implementation scenario. The above-disclosed are only several specific implementation scenarios of the present application. However, the present application is not limited thereto, and any changes that can be thought of by those skilled in the art should fall within the protection scope of the present application.

Claims

1. A method for extracting IOC information, characterized in that: include: Using the first agent in the multi-agent collaborative architecture, the paragraph text containing IOC information is filtered out from the threat intelligence; Utilizing a second agent in a multi-agent collaborative framework, extracting IOC information in an initial format from the paragraph text; Using a fourth agent in the multi-agent collaborative architecture, the target IOC information is obtained by converting the format of the IOC information in the initial format; Among them, the first agent, the second agent and the fourth agent are large language models.

2. The method according to claim 1, characterized in that: After the step of extracting the IOC information in the initial format from the paragraph text using the second agent in the multi-agent collaborative framework, the method further includes: Using the third agent in the multi-agent collaborative architecture, the IOC information in the initial format is verified and processed. If the verification succeeds, the IOC information in the initial format is sent to the fourth agent; if the verification fails, the correction information of the IOC information in the initial format is generated and sent to the second agent; The verification process includes format verification and context verification.

3. The method according to claim 1 or 2, characterized in that: The output text of each agent in the multi-agent collaborative architecture includes state information, and the state information is the execution result information of the task executed by the current agent.

4. The method according to claim 1, characterized in that: The step of using the first agent in the multi-agent collaborative architecture to filter out paragraph text containing IOC information from threat intelligence includes: Using the first agent in the multi-agent collaborative architecture to perform paragraph segmentation processing on the threat intelligence to obtain multiple paragraph texts; IOC information is identified on the multiple paragraph texts, paragraph texts containing IOC information are screened out from the multiple initial paragraph texts, and the paragraph texts are sent to the second agent.

5. The method according to claim 1, characterized in that The step of extracting the IOC information in the initial format from the paragraph text using the second agent in the multi-agent collaborative framework comprises: If the paragraph text includes rewriting information, the rewriting information is restored to obtain the original paragraph text; The IOC information in an initial format is extracted from the original paragraph text to obtain the IOC information in an initial format of the threat intelligence.

6. The method according to claim 1 or 2, characterized in that: The step of extracting the IOC information in the initial format from the paragraph text using the second agent in the multi-agent collaborative framework also includes: The IOC information is re-extracted from the paragraph text according to the paragraph text from the first intelligent agent and the correction information of the IOC information in the initial format to obtain the IOC information in the initial format of the threat intelligence.

7. The method according to claim 1 or 2, characterized in that: The method further comprises: When the threat intelligence includes a newly added IOC type, the system instructions and sample examples of the agents in the multi-agent collaborative architecture are changed according to the newly added IOC type.

8. An IOC information extraction device, characterized in that: include: A first agent module is used to filter out paragraph text containing IOC information from threat intelligence using a first agent in a multi-agent collaborative architecture; A second agent module is used to extract the IOC information in the initial format from the paragraph text using the second agent in the multi-agent collaborative architecture; A fourth agent module is used to obtain target IOC information by converting the format of the IOC information in the initial format using the fourth agent in the multi-agent collaborative architecture; Among them, the first agent, the second agent and the fourth agent are large language models.

9. A computer storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the IOC information extraction method according to any one of claims 1 to 7 is implemented.

10. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the program, the IOC information extraction method according to any one of claims 1 to 7 is implemented.