An APT open source intelligence intelligent analysis system and method empowered by a large model

The APT open-source intelligence intelligent analysis system, powered by a large model, solves the challenges of pre-emptive prevention and data structuring in APT defense by breaking down tasks through text segmentation and multi-turn dialogue modules, combined with anomaly handling and disambiguation techniques, thus achieving efficient and accurate intelligence analysis.

CN120996018APending Publication Date: 2025-11-21SUN YAT SEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510934759.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies rely too heavily on "known facts" in APT defense, making it difficult to achieve proactive prevention. The structured processing of open-source threat intelligence is inefficient, and the ability to analyze multi-source heterogeneous data is insufficient. Long text processing in LLMs suffers from truncation and information fragmentation issues.

Method used

The APT open-source intelligence intelligent analysis system, powered by a large model, divides the input text into blocks through a text segmentation module, breaks down the multi-turn dialogue module into 6 sub-tasks, each with a single objective, and combines an anomaly handling module to dynamically adjust parameters and countrylayer API to disambiguate, thereby achieving entity alignment and deduplication.

Benefits of technology

It improves the efficiency of structured processing of APT intelligence, reduces the possibility of information confusion and formatting errors, solves the problems of truncation and information fragmentation in long text processing of LLMs, and enhances the ability to handle cross-document naming ambiguity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996018A_ABST
    Figure CN120996018A_ABST
Patent Text Reader

Abstract

The application discloses a large model empowerment APT open source intelligence intelligent analysis system and method, and the system comprises a text division module, which is used for assembling input text into text blocks with a length not exceeding a preset length, and a first round of dialogue adopts a block length of M times the maximum token number, wherein M is a positive integer; a multi-round dialogue module, which is used for adopting a multi-round instruction dialogue guiding mechanism, decomposing an APT intelligence extraction task into six sequentially executed subtasks, and executing only a single subtask in each round of dialogue and returning structured data; and an exception processing module, which is used for monitoring and processing abnormal events in a flow, dynamically optimizing text division parameters or starting manual takeover according to an exception type, and finally recording an exception processing track. Through appropriate dialogue mechanism design and Prompt design, the application completes natural language OSINT APT related entity and relationship extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cyberspace security, specifically to a large-scale model-enabled open-source APT intelligence intelligent analysis system and method. Background Technology

[0002] Advanced Persistent Threat (APT) is a long-term, covert computer intrusion process targeting a specific target. The instigators are often state-sponsored, well-funded, and technically proficient. They exploit specific vulnerabilities to gain long-term control of the target computer, thereby stealing data or destroying the target. It is a serious cyber threat.

[0003] To effectively combat APTs, existing technical approaches primarily focus on "in-process response" and "post-incident attribution." These include constructing attribution graphs based on host system logs and developing innovative algorithms for rapid identification and blocking during attacks; using innovative identification methods based on communication traffic to identify APT communications and block data theft; and capturing APT software using honeypots and profiling APT groups using sandboxing and disassembly techniques to determine attack attribution. While these methods offer good defensive capabilities in certain scenarios, they also face challenges: an over-reliance on "known facts," meaning they cannot achieve "prevention." Attackers utilize zero-day exploits and supply chain contamination to achieve immediate strikes and long-term infiltration, bypassing most existing defense systems and posing a serious challenge to traditional detection technologies.

[0004] In recent years, open-source threat intelligence (OSINT) has gradually become an important part of APT defense systems. Global security vendors (FireEye, Kaspersky, 360, and QiAnXin, etc.) regularly release APT OSINT; various hacker forums, CVE vulnerability databases, and threat intelligence platforms also regularly provide a large amount of OSINT. Analyzing and utilizing these OSINTs allows for rapid understanding of the latest attack group dynamics, attack tactics, and IOCs (Indicators of Compromises), significantly improving the ability to detect, automatically defend against, and quickly respond to APT attacks. However, applying OSINT to APT defense still faces the following challenges:

[0005] Structured OSINT is often considered a trade secret, while open-source OSINT is mostly provided in unstructured natural language (such as threat reports, incident descriptions, etc.). Before using OSINT in these natural languages ​​to automate APT analysis, a cumbersome entity-relationship extraction process is required.

[0006] Traditional named entity recognition algorithms applied to OSINT have low accuracy, slow speed, and are prone to confusion with and omission of entities and relationships.

[0007] The lack of multi-source heterogeneous OSINT extraction and analysis systems, the prominent problem of data silos, and the weak automated collection and intelligent analysis capabilities make it difficult to support the construction of a large-scale APT defense and tracing system.

[0008] The development and advancement of Large Language Models (LLMs) have brought new hope to solving the problems of unstructured data entity recognition, semantic understanding, and manual parsing in OSINT. LLMs possess powerful language modeling and contextual understanding capabilities. With appropriate prompts, they can automatically perform tasks such as named entity recognition, attack event extraction, indicator extraction (e.g., IP address, domain name, CVE), and semantic classification and attribution inference for APT-related texts. Especially when dealing with intelligence texts from diverse sources, in various formats, and with complex security terminology, LLMs demonstrate strong robustness, high generalization ability, and flexible adaptation to context. The open-source API of LLMs can also be easily integrated into intelligent analysis system frameworks, automatically and automatically extracting relevant intelligence indicators from OSINT with almost no manual intervention.

[0009] However, LLMs can only process a limited amount of text at a time, with the input length limit often exceeding the output length limit. Excessively long inputs often result in excessively long outputs, triggering detachment and incomplete output, which affects formatted parsing. Conversely, excessively short inputs lead to fragmented information, increasing processing costs and making subsequent entity alignment and fusion more difficult.

[0010] Chinese invention application No. 202410051928.6 discloses "A method, apparatus, medium and electronic device for language model analysis of threats", which includes: acquiring threat intelligence and determining the type of threat intelligence; selecting an information extraction method for language model analysis of threat intelligence according to the type of threat intelligence; applying the selected information extraction method to extract information from the threat intelligence; and obtaining the information extraction results to obtain a threat knowledge graph of the corresponding threat intelligence. Summary of the Invention

[0011] To address the technical problems existing in the background art, this invention provides a large-model-enabled APT open-source intelligence intelligent analysis system and method. The technical solution adopted by this invention includes:

[0012] The first aspect of this invention provides a large-model-enabled open-source APT intelligence intelligent analysis system, the system comprising:

[0013] The text segmentation module is used to assemble the input text into text blocks of no more than a preset length. The first round of dialogue uses a block length that is M times the maximum number of tokens output, where M is a positive integer.

[0014] The multi-turn dialogue module is used to break down the APT intelligence extraction task into 6 sequentially executed sub-tasks using a multi-turn instruction-based dialogue guidance mechanism. Each round of dialogue executes only a single sub-task and returns structured data.

[0015] The exception handling module is used to monitor exception events in the processing flow, dynamically optimize text segmentation parameters or initiate manual takeover based on the exception type, and finally record the exception handling trajectory.

[0016] As a preferred approach, the first round of dialogue uses a block length that is three times the maximum number of tokens output.

[0017] As a preferred embodiment, the multi-turn dialogue module includes:

[0018] The first sub-task module is used to extract basic information from the document;

[0019] The second sub-task module is used to extract the list of APT organizations and their aliases;

[0020] The third sub-task module is used to iteratively extract basic information from each APT organization.

[0021] The fourth sub-task module is used to extract IOCs information from APT organizations;

[0022] The fifth sub-task module is used to extract the TTPs (Tactical Techniques) of the APT organization;

[0023] The sixth sub-task module is used to extract geopolitical background information of APT organizations;

[0024] The alignment and disambiguation module is used to perform entity alignment, disambiguation, and deduplication on the information extracted by all subtask modules through the countrylayer API.

[0025] As a preferred embodiment, the basic information of the document includes:

[0026] Document type, document title, publication date, author, threat entities involved, report language, and report publisher.

[0027] As a preferred embodiment, the basic information of the APT organization includes:

[0028] List of organization aliases, list of victim countries, list of victim industries, list of victim sectors, attribution party, attribution party confidence level, first active time, last active time, and current status.

[0029] As a preferred embodiment, the IOCs information of the APT organization includes:

[0030] IP list, domain list, URL list, malicious file list, malicious email address list, malicious file directory, CVE standard vulnerability.

[0031] As a preferred embodiment, the APT organization's TTPs tactics include:

[0032] List of technologies used, list of tactics used, list of tools used, and a summary of the use of TTPs.

[0033] As a preferred embodiment, the geopolitical background information of the APT organization includes:

[0034] Geographical background, organizational background, motivations, and other relevant background information.

[0035] As a preferred embodiment, the exception handling module includes:

[0036] An anomaly detection module is used to monitor output integrity and subject-object logic anomalies;

[0037] The first exception handling module is used to reduce chunk_length to no less than the maximum number of output tokens when responding to an incomplete output exception.

[0038] The second exception handling module is used to increase the text block overlap length (overlap_length) when responding to a subject-object separation exception.

[0039] The third exception handling module is used to initiate a manual takeover mechanism when other exceptions occur or the number of times the same exception is repeatedly triggered reaches a preset threshold.

[0040] The exception logging module is used to record exception events in log form and write them to the exception database for exception analysis.

[0041] A second aspect of this invention provides a large-model-enabled intelligent analysis method for APT open-source intelligence, the method comprising:

[0042] Input text and dynamically divide it into chunks. In the first round, the chunk_length is 3 times the maximum number of tokens output.

[0043] The six-stage, multi-round dialogue task is executed sequentially, and JSON structured output is obtained in each stage.

[0044] The countrylayer API is used to perform entity alignment, disambiguation, and deduplication on the JSON structured output obtained at each stage;

[0045] Perform anomaly detection and handling on the JSON structured output after disambiguation by the countrylayer API. When an output anomaly is detected:

[0046] If the output is truncated, reduce the chunk_length and re-divide the chunks;

[0047] If the subject and object are separated, increase the overlap_length and re-divide the blocks;

[0048] If the anomaly is triggered repeatedly, it will be handed over to manual handling.

[0049] Abnormal events are logged and written to an anomaly database for anomaly analysis.

[0050] Compared with the prior art, the beneficial effects of this invention are:

[0051] This invention solves the problems of truncation and information fragmentation in long text processing in LLMs by using dual-parameter adaptive adjustment and three-level anomaly response.

[0052] This invention breaks down APT intelligence extraction into 6 progressive atomic tasks, with each round of dialogue forcing a single task objective and JSON structured output, reducing the possibility of obfuscation and formatting errors.

[0053] This invention solves the technical problem of cross-document naming ambiguity by using the countrylayer API for disambiguation. Attached Figure Description

[0054] Figure 1 This embodiment provides a module connection diagram for an APT open-source intelligence intelligent analysis system empowered by a large model;

[0055] Figure 2 This is a flowchart of the multi-turn dialogue module provided in this embodiment;

[0056] Figure 3 This embodiment provides a flowchart of an APT open-source intelligence intelligent analysis method empowered by a large model. Detailed Implementation

[0057] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the invention.

[0058] It should be understood that the described embodiments are merely some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of the embodiments of this application.

[0059] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0060] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0061] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. The invention will be further described below with reference to the accompanying drawings and embodiments.

[0062] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0063] Example 1

[0064] Please refer to Figure 1 This embodiment provides a large-model-enabled APT open-source intelligence intelligent analysis system, the system comprising:

[0065] The text segmentation module is used to assemble the input text into text blocks of no more than a preset length. The first round of dialogue uses a block length that is M times the maximum number of tokens output, where M is a positive integer.

[0066] The multi-turn dialogue module is used to break down the APT intelligence extraction task into 6 sequentially executed sub-tasks using a multi-turn instruction-based dialogue guidance mechanism. Each round of dialogue executes only a single sub-task and returns structured data.

[0067] The exception handling module is used to monitor exception events in the processing flow, dynamically optimize text segmentation parameters or initiate manual takeover based on the exception type, and finally record the exception handling trajectory.

[0068] In one specific embodiment, the first round of dialogue uses a block length that is three times the maximum number of tokens output.

[0069] In one specific embodiment, the multi-turn dialogue module includes:

[0070] The first sub-task module is used to extract basic information from the document;

[0071] The second sub-task module is used to extract the list of APT organizations and their aliases;

[0072] The third sub-task module is used to iteratively extract basic information from each APT organization.

[0073] The fourth sub-task module is used to extract IOCs information from APT organizations;

[0074] The fifth sub-task module is used to extract the TTPs (Tactical Techniques) of the APT organization;

[0075] The sixth sub-task module is used to extract geopolitical background information of APT organizations;

[0076] The alignment and disambiguation module is used to perform entity alignment, disambiguation, and deduplication on the information extracted by all subtask modules through the countrylayer API.

[0077] In one specific embodiment, the basic document information includes:

[0078] Document type, document title, publication date, author, threat entities involved, report language, and report publisher.

[0079] In one specific embodiment, the basic information of the APT organization includes:

[0080] List of organization aliases, list of victim countries, list of victim industries, list of victim sectors, attribution party, attribution party confidence level, first active time, last active time, and current status.

[0081] In one specific embodiment, the IOCs information of the APT organization includes:

[0082] IP list, domain list, URL list, malicious file list, malicious email address list, malicious file directory, CVE standard vulnerability.

[0083] In one specific embodiment, the APT organization's TTPs tactics include:

[0084] List of technologies used, list of tactics used, list of tools used, and a summary of the use of TTPs.

[0085] In one specific embodiment, the APT organization's geopolitical background information includes:

[0086] Geographical background, organizational background, motivations, and other relevant background information.

[0087] In one specific embodiment, the exception handling module includes:

[0088] An anomaly detection module is used to monitor output integrity and subject-object logic anomalies;

[0089] The first exception handling module is used to reduce chunk_length to no less than the maximum number of output tokens when responding to an incomplete output exception.

[0090] The second exception handling module is used to increase the text block overlap length (overlap_length) when responding to a subject-object separation exception.

[0091] The third exception handling module is used to initiate a manual takeover mechanism when other exceptions occur or the number of times the same exception is repeatedly triggered reaches a preset threshold.

[0092] The exception logging module is used to record exception events in log form and write them to the exception database for exception analysis.

[0093] Example 2

[0094] Please refer to Figure 1 This embodiment provides a large-model-enabled APT open-source intelligence intelligent analysis system, the system comprising:

[0095] The text segmentation module is used to assemble the input text into text blocks of no more than a preset length. The first round of dialogue uses a block length that is M times the maximum number of tokens output, where M is a positive integer.

[0096] The multi-turn dialogue module is used to break down the APT intelligence extraction task into 6 sequentially executed sub-tasks using a multi-turn instruction-based dialogue guidance mechanism. Each round of dialogue executes only a single sub-task and returns structured data.

[0097] The exception handling module is used to monitor exception events in the processing flow, dynamically optimize text segmentation parameters or initiate manual takeover based on the exception type, and finally record the exception handling trajectory.

[0098] It should be noted that the text segmentation module can assemble blocks not exceeding a specified length (chunk_length). The first round of dialogue uses a longer input block length (3 times the maximum output token) to improve efficiency. If incomplete output is detected, an exception is triggered, and the exception handling module modifies the chunk_length, replacing it with a shorter input block length (at least not shorter than the maximum output token count). Since text segmentation can potentially affect information integrity, i.e., subject-object separation may occur (the subject is in one block, while the corresponding object (the information to be extracted) is in the next block), the Prompt design requires LLMs to trigger an exception if they find an object without a subject. The exception handling module then adjusts the overlap_length, re-segments the blocks, and makes the request again.

[0099] The multi-turn dialogue module is the core of this invention. To improve the accuracy and controllability of LLMs in entity and relation extraction tasks within the OSINT scenario, a multi-turn instruction-based dialogue guidance mechanism is employed. This mechanism breaks down the complex information extraction task into six sub-tasks, corresponding to six prompt designs. In each round of dialogue, the model is guided to complete only one explicit task, effectively reducing problems such as model output confusion, format errors, or task overlap. This design ensures that the large model can still stably complete structured output under zero-sample or few-sample conditions, improving the overall reliability and scalability of information extraction. The multi-turn dialogue architecture is as follows: Figure 2 As shown, it includes the following 6 tasks.

[0100] (1) Extract basic document information and return it in JSON format, including document type (APT report, threat intelligence, malware report, cybersecurity knowledge, vulnerability exploitation, or others), document title, publication time, author, threat entities involved, report language, and report publisher (government, individual, organization, or unknown).

[0101] (2) Extract the APT organizations mentioned in the document and return them in list format, avoiding the inclusion of different aliases of the same APT organization in the list. The returned data should be in JSON format, including the list of APT organizations and APT aliases.

[0102] Based on the results returned in (2), the iteration involves a list of APT organizations, requiring LLMs to precisely locate the following information:

[0103] (3) Extract the basic information of the APT organization and return it in JSON format.

[0104] (4) Extract the IOCs of the APT organization and return the list of involved IPs, domains, URLs, malicious files, malicious email addresses, malicious file directories and CVE standard vulnerabilities in JSON format. If there is a summary of the use of IOCs, please also return it.

[0105] (5) Extract the TTPs involved in the APT organization and return the list of technologies, tactics and tools used in JSON format. If there is a summary of the TTPs used, please also return it.

[0106] (6) Extract the geopolitical background of the APT organization and return the geopolitical background, organizational background, motivation and other relevant background information in JSON format.

[0107] Through multiple rounds of dialogue, it is ensured that LLMs complete only one task per round and return the specified information in a formatted manner. At the same time, each round of dialogue is designed with an anomaly identifier, requiring the model to determine whether there are any anomalies in the extraction. If so, it is set to True, and human intervention can be carried out in a timely manner.

[0108] In one specific embodiment, the first round of dialogue uses a block length that is three times the maximum number of tokens output.

[0109] In one specific embodiment, the multi-turn dialogue module includes:

[0110] The first sub-task module is used to extract basic information from the document;

[0111] The second sub-task module is used to extract the list of APT organizations and their aliases;

[0112] The third sub-task module is used to iteratively extract basic information from each APT organization.

[0113] The fourth sub-task module is used to extract IOCs information from APT organizations;

[0114] The fifth sub-task module is used to extract the TTPs (Tactical Techniques) of the APT organization;

[0115] The sixth sub-task module is used to extract geopolitical background information of APT organizations;

[0116] The alignment and disambiguation module is used to perform entity alignment, disambiguation, and deduplication on the information extracted by all subtask modules through the countrylayer API.

[0117] In one specific embodiment, the basic document information includes:

[0118] Document type, document title, publication date, author, threat entities involved, report language, and report publisher.

[0119] In one specific embodiment, the basic information of the APT organization includes:

[0120] List of organization aliases, list of victim countries, list of victim industries, list of victim sectors, attribution party, attribution party confidence level, first active time, last active time, and current status.

[0121] In one specific embodiment, the IOCs information of the APT organization includes:

[0122] IP list, domain list, URL list, malicious file list, malicious email address list, malicious file directory, CVE standard vulnerability.

[0123] In one specific embodiment, the APT organization's TTPs tactics include:

[0124] List of technologies used, list of tactics used, list of tools used, and a summary of the use of TTPs.

[0125] In one specific embodiment, the APT organization's geopolitical background information includes:

[0126] Geographical background, organizational background, motivations, and other relevant background information.

[0127] In one specific embodiment, the exception handling module includes:

[0128] An anomaly detection module is used to monitor output integrity and subject-object logic anomalies;

[0129] The first exception handling module is used to reduce chunk_length to no less than the maximum number of output tokens when responding to an incomplete output exception.

[0130] The second exception handling module is used to increase the text block overlap length (overlap_length) when responding to a subject-object separation exception.

[0131] The third exception handling module is used to initiate a manual takeover mechanism when other exceptions occur or the number of times the same exception is repeatedly triggered reaches a preset threshold.

[0132] The exception logging module is used to record exception events in log form and write them to the exception database for exception analysis.

[0133] Example 3

[0134] Please refer to Figure 3 This embodiment provides a large-model-enabled intelligent analysis method for APT open-source intelligence, the method comprising:

[0135] S1: Input text and dynamically divide it into chunks. In the first round, the chunk_length is 3 times the maximum number of tokens output.

[0136] S2: Sequentially execute a six-stage, multi-round dialogue task, obtaining JSON structured output for each stage;

[0137] S3: Perform entity alignment, disambiguation, and deduplication on the JSON structured output obtained at each stage using the countrylayer API;

[0138] S4: Perform anomaly detection and handling on the JSON structured output after disambiguation by the countrylayer API; when an output anomaly is detected:

[0139] If the output is truncated, reduce the chunk_length and re-divide the chunks;

[0140] If the subject and object are separated, increase the overlap_length and re-divide the blocks;

[0141] If the anomaly is triggered repeatedly, it will be handed over to manual handling.

[0142] Abnormal events are logged and written to an anomaly database for anomaly analysis.

[0143] It should be noted that although LLMs have been formatted and returned, entity alignment and deduplication are still required. For example, different documents may use different names for the same organization, leading to inconsistent LLM extraction results. For instance, the organization "APT38" may be written as "apt 38" or "apt-38" in some documents, requiring deduplication. To address this, this invention first standardizes all extracted organization names to lowercase. For organizations named with "apt" and numbers, regular expressions are used to extract number-aligned entities; for organizations named purely in English, spaces are removed and case is converted, ensuring the uniqueness of their names. Another entity requiring deduplication is the country (region) name, such as the United States, America, USA, and US, all of which refer to the United States. Therefore, the countrylayer API is introduced for deduplication. This API can automatically parse country names based on different abbreviations or full names.

[0144] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A large model empowered APT open source intelligence intelligent analysis system, characterized in that, The system comprises: a text division module for assembling input text into text blocks not exceeding a preset length, and a first round of dialogue using a block length of M times the maximum number of tokens output, wherein M is a positive integer; a multi-round dialogue module for using a multi-round imperative dialogue guiding mechanism to decompose an APT intelligence extraction task into six sequentially executed subtasks, with only a single subtask being executed in each round of dialogue and structured data being returned; an exception handling module for monitoring abnormal events in the processing flow, dynamically optimizing text division parameters or starting manual takeover according to the type of the abnormal event, and finally recording the abnormal processing track.

2. The large model empowered APT open source intelligence intelligent analysis system according to claim 1, wherein, The first round of dialogue uses a block length of 3 times the maximum number of tokens output.

3. The large model-enabled APT open source intelligence intelligent analysis system according to claim 1, characterized in that, The multi-round dialogue module comprises: a first subtask module for extracting document basic information; a second subtask module for extracting APT organization lists and aliases; a third subtask module for iteratively extracting basic information of each APT organization; a fourth subtask module for extracting IOCs information of the APT organization; a fifth subtask module for extracting TTPs tactics and techniques of the APT organization; a sixth subtask module for extracting geopolitical background information of the APT organization; an alignment and disambiguation module for disambiguating and aligning entities of information extracted by all subtask modules through the countrylayer API.

4. The large model-enabled APT open source intelligence intelligent analysis system according to claim 3, characterized in that, The document basic information comprises: document type, document title, publication time, author, threat entity involved, report language, and report publisher.

5. The large model-enabled APT open source intelligence intelligent analysis system according to claim 3, characterized in that, The APT organization basic information comprises: organization alias list, victim country list, victim industry list, victim industry list, attribution party, attribution party confidence, first active time, last active time, and current status.

6. The large model empowered APT open source intelligence intelligent analysis system according to claim 3, characterized in that, The IOCs information of the APT organization comprises: IP list, domain name list, URL list, malicious file list, malicious email address list, malicious file directory, and CVE standard vulnerability.

7. The large model empowered APT open source intelligence intelligent analysis system according to claim 3, characterized in that, The TTPs tactics and techniques of the APT organization comprise: used technology list, used tactics list, used tool list, and summary of TTPs usage.

8. The large model empowered APT open source intelligence intelligent analysis system according to claim 3, characterized in that, The geopolitical background information of the APT organization comprises: geopolitical background, organizational background, motivation, and other background information involved.

9. The large model empowered APT open source intelligence intelligent analysis system according to claim 1, wherein, The exception handling module comprises: an exception detection module for monitoring output integrity and subject-object logic exceptions; a first exception handling module for reducing chunk_length to not less than the maximum number of tokens output in response to an incomplete output exception; a second exception handling module for increasing the text block overlap length overlap_length in response to a subject-object separation exception; a third exception handling module for starting a manual takeover mechanism when other exceptions occur or the same exception is repeatedly triggered a number of times reaching a preset threshold; an exception recording module for recording abnormal events in the form of logs and writing them into an exception database for exception analysis.

10. A large model empowered APT open source intelligence intelligent analysis method, characterized in that, The method comprises: inputting text and dynamically dividing it into blocks, with a chunk_length of 3 times the maximum number of tokens output in the first round; sequentially executing a six-stage multi-round dialogue task, with JSON structured output being obtained in each stage; Entity alignment and disambiguation for JSON structured output of each stage by countrylayer API disambiguation; Abnormality detection and handling for JSON structured output after countrylayer API disambiguation; when detecting output abnormality: If the output is truncated, reduce chunk_length and rechunk; If the subject and object are separated, increase overlap_length and rechunk; If the abnormality is repeatedly triggered, hand over to manual processing; Record the abnormal event in the form of a log and write it into the exception database for exception analysis.

Citation Information

Patent Citations

  • A method, device, medium and electronic device for analyzing language model of threats

    CN117786088B