APT organization attack portrait generation reasoning method and device and electronic equipment
By annotating and fine-tuning the big model, the target big model is generated to process Chinese threat intelligence, solving the problems of low accuracy and high labor costs in the existing technology, achieving more accurate attack image generation and reasoning, and helping security personnel better deal with complex attack threats.
Patent Information
- Application Number
- CN202510115200.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing attack image generation reasoning methods have low accuracy when dealing with Chinese threat intelligence, and deep learning methods require a large amount of manual annotation of data, resulting in high labor costs. Large models lack precise knowledge support when facing domain-specific problems, which is prone to infer incorrect reasoning results.
By obtaining reference data, the big model is marked and fine-tuned, and the target model is obtained, which is used to identify and extract the analytical text in entity and relationships, build a knowledge graph and generate an APT organization attack portrait, and finally conduct attack inference.
It significantly reduces labor costs, improves the accuracy of attack portraits, and can more effectively reveal the behavior patterns and attack strategies of APT organizations, helping security personnel more accurately identify and respond to complex attack threats.
Smart Images

Figure CN120069066A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network information security technology, and in particular, to a method, device, and electronic device for generating and inferring an APT organization attack profile. Background Art
[0002] With the rapid development of Internet technology, network attack means such as network intrusion, worm infection, ransomware, and distributed denial of service attack have emerged frequently, bringing huge security risks to society and enterprises. Among these potential threats, an Advanced Persistent Threat (APT) attack is a continuous and effective network attack activity launched by complex attackers with political and national backgrounds against specific and high-value target organizations. If an APT attack is successfully launched, it often causes more serious damage than ordinary network attacks. APT attacks have seriously threatened national security, and the research on the defense against APT attacks is extremely urgent.
[0003] The main technical points for generating an attack profile include entity recognition and relationship extraction. In the field of threat intelligence, named entities and relationships are highly professional but lack a unified naming standard. Moreover, with the development of technology, the scale of threat intelligence entities has grown rapidly, and they are usually outside the existing thesaurus, bringing greater challenges to entity recognition and relationship extraction.
[0004] Currently, entity recognition and relationship extraction mainly rely on deep learning methods for processing. However, the existing deep learning methods still face some key problems: First, the accuracy of information extraction is not satisfactory, especially when dealing with Chinese threat intelligence, the accuracy is relatively low. According to existing research, the F1 value of the current mainstream deep learning methods in Chinese threat intelligence information extraction is about 70%. Second, deep learning methods usually require a large amount of manually labeled data for training, which greatly increases the labor cost of data annotation and limits their promotion and implementation in practical applications.
[0005] When using a large model for attack inference, although it has powerful language understanding and generation capabilities, it still lacks expertise in the specific field of threat intelligence. Due to the highly professional and dynamically updated characteristics of threat intelligence, when facing domain-specific problems, the large model often produces the so-called "hallucination" phenomenon due to the lack of accurate knowledge support, that is, the generated inference results seem reasonable but are actually wrong. This phenomenon not only affects the accuracy of the inference results but also may lead to incorrect threat analysis and judgment, thus bringing risks to the formulation and implementation of subsequent security strategies.
[0006] For the problem of poor accuracy existing in the existing attack profile generation and inference methods, no effective solution has been proposed yet. Summary of the Invention
[0007] The present invention provides a method, apparatus and electronic device for generating and inferring an APT organization attack portrait, so as to solve the defect of poor accuracy existing in the existing attack portrait generation and inference methods.
[0008] In a first aspect, the present invention provides a method for generating and inferring an APT organization attack portrait, including: Obtaining reference data and performing data annotation on the reference data through a large model; Fine-tuning the large model with the reference data to obtain a target large model; Obtaining a text to be analyzed and performing entity recognition and relationship extraction on the text to be analyzed through the target large model to obtain entity results and relationship results; Constructing a knowledge graph based on the extracted entity results and relationship results and generating an APT organization attack portrait; Performing attack inference on the APT organization attack portrait to obtain an inference result.
[0009] According to the method for generating and inferring an APT organization attack portrait provided by the present invention, obtaining reference data and performing data annotation on the reference data through a large model includes: Obtaining reference data; the reference data includes APT attack analysis data of APT analysis reports from different sources; Using prompt engineering to enable the large model to extract entities and relationships in the reference data to obtain reference entities and reference relationships; Performing data cleaning processing on the reference entities and reference relationships to obtain the final reference data.
[0010] According to the method for generating and inferring an APT organization attack portrait provided by the present invention, fine-tuning the large model with the reference data to obtain a target large model includes: Sorting out the reference data and adjusting the data format to obtain a standard data set; Invoking the LLaMAFactory framework and performing LoRA fine-tuning on the large model with the standard data set to obtain the target large model.
[0011] According to the method for generating and inferring an APT organization attack portrait provided by the present invention, performing LoRA fine-tuning on the large model with the standard data set to obtain the target large model includes: Invoking the standard data set and setting the path information and format information of the standard data set; Obtaining the path information of the standard data set, setting the storage path of the training result and the fine-tuning type; the fine-tuning type is LoRA; Retrieve the standard data set according to the path information, and perform training and fine-tuning on the large model according to the fine-tuning type, output the training result, and obtain the target large model.
[0012] According to an APT organization attack portrait generation and inference method provided by the present invention, entity recognition and relationship extraction are performed on the text to be analyzed through the target large model, and entity results and relationship results are obtained, including: Input the text to be analyzed into the target large model, and through prompt engineering technology, guide the target large model to perform entity recognition and relationship extraction in the text to be analyzed, and obtain the entity results and the relationship results.
[0013] According to an APT organization attack portrait generation and inference method provided by the present invention, the entity results include attackers, victims, and attack means.
[0014] According to an APT organization attack portrait generation and inference method provided by the present invention, a knowledge graph is constructed based on the extracted entity results and relationship results, including: Obtain relevant libraries, and connect to the graph database through the relevant libraries; Create nodes and relationships of the knowledge graph based on the relevant libraries and the graph database; the nodes correspond to the entity results, and the relationships correspond to the relationship results.
[0015] According to an APT organization attack portrait generation and inference method provided by the present invention, attack inference is performed on the APT organization attack portrait to obtain an inference result, including: Deploy the GraphRAG framework and load the target large model; Combine the text to be analyzed and the APT organization attack portrait, and perform inference on the text to be analyzed in the form of question-and-answer interaction to obtain an inference result.
[0016] In a second aspect, the present invention also provides an APT organization attack portrait generation and inference device, including: An acquisition module for acquiring reference data and performing data annotation on the reference data through a large model; A fine-tuning module for fine-tuning the large model through the reference data to obtain a target large model; An analysis module for acquiring the text to be analyzed and performing entity recognition and relationship extraction on the text to be analyzed through the target large model to obtain entity results and relationship results; A generation module for constructing a knowledge graph based on the extracted entity results and relationship results and generating an APT organization attack portrait; An inference module for performing attack inference on the APT organization attack portrait to obtain an inference result.
[0017] In a third aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the APT organization attack portrait generation and inference method as described in the first aspect above.
[0018] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the APT organization attack portrait generation and inference method as described in the first aspect above.
[0019] In a fifth aspect, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the APT organization attack portrait generation and inference method as described in the first aspect above.
[0020] Compared with the prior art, the present invention has the following beneficial effects: The APT organization attack portrait generation and inference method provided by the present invention, by combining large models, significantly reduces the labor cost compared with the traditional deep learning-based method. By giving full play to the advantages of large models in semantic understanding, it is possible to construct a more accurate and comprehensive APT organization attack portrait. Through in-depth analysis and inference of the APT organization attack portrait, it is possible to reveal the behavior patterns, attack strategies and potential targets of APT organizations from multiple dimensions, helping security personnel to more accurately identify and respond to complex attack threats, providing strong data support, and thus formulating more effective and personalized security protection measures, solving the problem of poor accuracy in existing attack portrait generation and inference methods.
[0021] In addition, the present invention selects a target large model obtained through fine-tuning instead of using a general model such as GPT. The reason is that the fine-tuned target large model can be optimized for a specific field and can extract key information in APT attack intelligence more accurately. The fine-tuning process enables the model to better adapt to the professional knowledge in the security field, avoiding the limitations of general models when facing professional tasks. At the same time, using a locally fine-tuned model reduces the exposure of sensitive information, reduces potential security risks, and ensures data privacy and system security. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0023] Figure 1 is a flowchart of the APT organization attack portrait generation and reasoning method provided by the present invention; Figure 2 is a schematic diagram of generating and reasoning the attack portrait in the embodiment of the present invention; Figure 3 is a structural block diagram of the APT organization attack portrait generation and reasoning device provided by the present invention; Figure 4 is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners
[0024] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0025] The present invention provides an APT organization attack portrait generation and reasoning method, Figure 1 which is a flowchart of the APT organization attack portrait generation and reasoning method provided by the present invention. As Figure 1 shown, the method includes the following steps: Step S101, obtain reference data and perform data annotation on the reference data through a large model; Step S102, fine-tune the large model with the reference data to obtain a target large model; Step S103, obtain the text to be analyzed and perform entity recognition and relationship extraction on the text to be analyzed through the target large model to obtain entity results and relationship results; Step S104, construct a knowledge graph based on the extracted entity results and relationship results and generate an APT organization attack portrait; Step S105, perform attack reasoning on the APT organization attack portrait to obtain reasoning results.
[0026] In this method, first, reference data with the function of historical experience analysis is obtained, and these reference data are annotated by a large model. Through the powerful semantic understanding ability of the large model, key entities and relationships in the reference data can be identified, and annotation results are automatically generated. Then, the large model is fine-tuned with the reference data processed by data annotation to improve the performance of the large model in specific tasks, and a target large model is obtained. The target large model obtained through fine-tuning already has sufficient knowledge in threat intelligence. Based on this, text to be analyzed related to APT organization intelligence is obtained, and entity recognition and relationship extraction are performed on the text to be analyzed through the target large model to improve the accuracy of entity recognition and relationship extraction, and entity results and relationship results are obtained. Based on the above-extracted entity results and relationship results, a knowledge graph is constructed and an APT organization attack portrait is generated. This knowledge graph can more vividly and accurately represent the association between entities and relationships in the text to be analyzed. Finally, attack reasoning is performed on the APT organization attack portrait to analyze potential attack patterns and threat chains, and reasoning results are obtained. In the above process, by combining a large model, compared with traditional deep learning-based methods, the labor cost is significantly reduced. By giving full play to the advantages of the large model in semantic understanding, a more accurate and comprehensive APT organization attack portrait can be constructed. Through in-depth analysis and reasoning of the APT organization attack portrait, the behavior patterns, attack strategies, and potential targets of APT organizations can be revealed from multiple dimensions, helping security personnel more accurately identify and respond to complex attack threats, providing strong data support, and thus formulating more effective and personalized security protection measures, solving the problem of poor accuracy existing in the existing attack portrait generation and reasoning methods.
[0027] In addition, the present invention selects the target large model obtained through fine-tuning instead of using general models such as GPT. The reason is that the fine-tuned target large model can be optimized for a specific field and can more accurately extract key information in APT attack intelligence. The fine-tuning process enables the model to better adapt to the professional knowledge in the security field, avoiding the limitations of general models when facing professional tasks. At the same time, using a locally fine-tuned model reduces the exposure of sensitive information, reduces potential security risks, and ensures data privacy and system security.
[0028] Figure 2 is a schematic diagram for generating and reasoning an attack portrait in an embodiment of the present invention, as Figure 2As shown, in some of these embodiments, in step S101, reference data is obtained, and the reference data is data-annotated by a large model, including: obtaining reference data; the reference data includes APT attack analysis data from APT analysis reports from different sources; using prompt engineering to enable the large model to extract entities and relationships in the reference data to obtain reference entities and reference relationships; performing data cleaning processing on the reference entities and reference relationships to obtain the final reference data.
[0029] Specifically, first, web crawler technology is used to crawl relevant data from multiple sources, including security reports, vulnerability information, and relevant analysis articles on APT attacks, etc. Then, the large model is applied to automatically annotate the crawled data. Through the powerful semantic understanding ability of the large model, the key entities and relationships in the data can be identified, and the annotation results can be automatically generated. Through this step, a standard data set for fine-tuning the large model can be obtained.
[0030] Exemplarily, web crawler technology is used to crawl APT analysis reports released by different security vendors, and APT attack analysis blogs appearing in security forums are collected. The requests library in Python is used to implement data crawling. The get() function in the requests library is used to send a GET request to the target website and obtain the server response data. When the returned status_code is 200, it indicates success, that is, the required page information has been successfully crawled. Then, the BeautifulSoup library is used to implement page parsing. Through BeautifulSoup, HTML documents can be easily parsed, and elements such as titles, links, paragraphs, images, tables, and lists can be extracted from them and converted into Python objects, making them easy to operate and process.
[0031] Then, using the powerful semantic understanding ability of the large model, prompt engineering is used to enable the large model to extract entities and relationships in the text. An example of a prompt is "Suppose you are a security engineer. Now, in order to generate an attack profile of an APT organization, you are required to help extract the entities and relationships in the given text to build a knowledge graph. Entities include: attacker (Attacker), tool (tools), region (region), industry (industry). Relationships include: use (use), develop (develop), targetted (targetted), etc. After receiving the information, please extract the entities and relationships according to the requirements." Check the extracted information above, exclude errors or irrelevant content, and ensure that the final data quality meets the requirements. The higher the data quality, the better the performance of the target large model after fine-tuning. Therefore, in order to generate a more accurate attack profile, this step can be performed by a small amount of manual inspection of the annotated data to exclude the incorrect data.
[0032] In some of these embodiments, in step S102, the large model is fine-tuned with reference data to obtain a target large model, including: sorting out the reference data and adjusting the data format to obtain a standard data set; calling the LLaMAFactory framework and performing LoRA fine-tuning on the large model with the standard data set to obtain the target large model.
[0033] In this embodiment, the large model is fine-tuned with the standard data set to obtain a target large model, including: calling the standard data set and setting the path information and format information of the standard data set; obtaining the path information of the standard data set, setting the storage path of the training result and the fine-tuning type; the fine-tuning type is LoRA; retrieving the standard data set according to the path information and performing training fine-tuning on the large model according to the fine-tuning type, outputting the training result, and obtaining the target large model.
[0034] Exemplarily, first, the collected reference data is sorted out into a standard data set conforming to the Alpaca format. Specifically, the entire standard data set is a list of json objects. The processed data set is in the Alpaca format, which is: {"instruction": "Recently, the Red Raindrop Team of QiAnXin Threat Intelligence Center discovered multiple malicious code attack activities suspected of targeting Android users in the Republic of Korea during daily advanced threat monitoring.", "input": null, "output": "Entity: ('Republic of Korea', 'Region')"}. The standard data set is a json file, and the file content is multiple pieces of data in the above format. Then, the llama model is downloaded through platforms such as Hugging Face or ModelScope. Finally, the LLaMA Factory framework is used to perform LoRA fine-tuning on the large model. Specifically as follows: First, place the standard data set in the specified data directory. Then, update the data / dataset_info.json file, add the path and format information of the standard data set to ensure that the standard data set is correctly registered and can be recognized by subsequent training processes. Then, run the instruction llamafactory-cli train -h in the command line to start the fine-tuning process of the large model. During operation, the location of the standard data set, the storage path of the training result need to be specified, and the fine-tuning type is set to LoRA. Through these parameter configurations, the specified standard data set can be loaded for training, and the training result can be saved in the specified folder for subsequent analysis and use. The LoRA fine-tuning technology makes the fine-tuning process more efficient and consumes less resources by adjusting the low-rank parameters of the large model, while improving the performance of the large model on specific tasks.
[0035] In some of these embodiments, in step S103, entity recognition and relationship extraction are performed on the text to be analyzed through the target large model to obtain entity results and relationship results, including: inputting the text to be analyzed into the target large model, and through prompt engineering techniques, guiding the target large model to perform entity recognition and relationship extraction in the text to be analyzed to obtain entity results and relationship results. Among them, the entity results include attackers, victims, and attack methods.
[0036] Exemplarily, input the intelligence text related to the APT organization to be analyzed into the fine-tuned target large model, and use prompt engineering techniques to guide the target large model to extract entities (such as attackers, victims, attack methods, etc.) and the relationships between entities in the input intelligence text. Taking the analysis of the APT32 organization as an example, first collect the threat intelligence about APT32. Then input the intelligence text into the target large model, and through the prompt "Assume you are a security engineer. Now, in order to generate the attack profile of the APT organization, you are required to help extract the entities and relationships in the given text to build a knowledge graph. Entities include: attacker (Attacker), tool (tools), region (region), industry (industry). Relationships include: use, develop, targetted, etc. After receiving the information, please extract the entities and relationships therein as required." to extract the entities and relationships in the intelligence.
[0037] On this basis, in step S104, a knowledge graph is constructed based on the extracted entity results and relationship results, including: obtaining a relevant library and connecting to a graph database through the relevant library; creating nodes and relationships of the knowledge graph based on the relevant library and the graph database; the nodes correspond to the entity results, and the relationships correspond to the relationship results.
[0038] Specifically, a knowledge graph is constructed based on the extracted entity results and relationship results, thereby forming an attack profile of the APT organization. The construction of the knowledge graph is implemented through Neo4j. Neo4j is a high-performance and scalable NoSQL database based on a graph database. It uses a graph model to store and process data, where the data is represented in the form of nodes, relationships, and attributes, and each node and relationship can have any number of attributes. This data model is very suitable for processing data with complex relationships, such as applications like social networks, recommendation systems, network security, and knowledge graphs.
[0039] Exemplarily, first, start the server using the command neo4j.bat console in the command line. After startup, it can be accessed in the browser through the URL http: / / localhost:7474 / . To generate a knowledge graph using Python with Neo4j, relevant libraries such as py2neo need to be installed. Then, connect to the Neo4j database through the function Graph() in the py2neo library. The Graph() function connects to the Neo4j database. After successful connection, operations can be performed on the database, providing a connection and interaction interface for the subsequent creation of the knowledge graph. Then, create nodes through the function Node() in the py2neo library and create relationships through the function Relationship(). Each node represents an entity in the knowledge graph (such as an APT organization, attacker, victim, etc.). Through the Relationship() function, relationships can be created between nodes, such as the attack behavior between the attacker and the victim, the dependency relationship between vulnerabilities and attack means, etc. By creating these nodes and relationships, a structured graph data model can be constructed to display the attack paths and behavior patterns of APT organizations.
[0040] In some of these embodiments, in step S105, perform attack reasoning on the APT organization attack profile to obtain an inference result, including: deploying the GraphRAG framework and loading the target large model; combining the text to be analyzed and the APT organization attack profile, and reasoning about the text to be analyzed in the form of a question-and-answer interaction to obtain an inference result.
[0041] Specifically, use the GraphRAG framework for attack reasoning to analyze potential attack patterns and threat chains. As an extension of RAG technology, the GraphRAG framework introduces structured domain knowledge into the reasoning process by constructing a knowledge graph, enabling more fine-grained and accurate knowledge retrieval and reasoning. In the reasoning of APT organization attack behaviors, the GraphRAG framework can more accurately identify the attack patterns and behavior chains of APT organizations by graphically modeling information such as attack paths, attack means, and victim targets.
[0042] Exemplarily, first, locally deploy the GraphRAG framework and load the fine-tuned target large model. Then, place the text file of the APT organization-related intelligence to be analyzed into the input directory of the GraphRAG framework system. After starting the GraphRAG framework, the system will automatically read and load the intelligence data in the input directory and prepare for inference analysis. Then, through the question-and-answer interaction method, infer the relevant intelligence of the APT organization. After the GraphRAG framework loads the intelligence data, the user can, through the question-and-answer interaction method, ask the system inference questions about the APT organization. For example, the user can ask "What attack methods does this APT organization use?" or "What are the attack targets of the APT organization?". Through this question-and-answer form, the GraphRAG framework will combine the large model and the knowledge graph for inference and output detailed analysis results such as the behavior patterns, attack strategies, and target information of the APT organization.
[0043] In summary, this method gives full play to the advantages of the large model and constructs a more accurate attack portrait of the APT organization. Compared with traditional attack portrait generation and inference methods, the present invention has significant advantages: on the one hand, it greatly reduces the labor cost; on the other hand, by combining the powerful inference ability of the large model, it improves the accuracy of the attack portrait and can more effectively reveal the attack patterns of the APT organization. Through this innovative method, not only the inference accuracy is improved, but also the understanding and response ability to complex attack behaviors are enhanced.
[0044] The present invention also provides an APT organization attack portrait generation and inference device. The APT organization attack portrait generation and inference device provided by the present invention will be described below. The APT organization attack portrait generation and inference device described below can be mutually corresponding and referred to the APT organization attack portrait generation and inference method described above. Figure 3 is the structural block diagram of the APT organization attack portrait generation and inference device provided by the present invention, as Figure 3 shown. The device includes: An acquisition module 301, configured to acquire reference data and perform data annotation on the reference data through a large model; A fine-tuning module 302, configured to fine-tune the large model through the reference data to obtain a target large model; An analysis module 303, configured to acquire the text to be analyzed and perform entity recognition and relationship extraction on the text to be analyzed through the target large model to obtain entity results and relationship results; A generation module 304, configured to construct a knowledge graph based on the extracted entity results and relationship results and generate an APT organization attack portrait; An inference module 305, configured to perform attack inference on the APT organization attack portrait to obtain an inference result.
[0045] When this device is in use, first, the acquisition module 301 acquires reference data with the function of historical experience analysis, and annotates these reference data through a large model. With the powerful semantic understanding ability of the large model, it can identify key entities and relationships in the reference data and automatically generate annotation results. Then, the fine-tuning module 302 fine-tunes the large model with the reference data after data annotation processing to improve the performance of the large model in specific tasks and obtain the target large model. The target large model obtained through fine-tuning already has sufficient knowledge in threat intelligence. Based on this, the analysis module 303 acquires the text to be analyzed related to APT organization intelligence, and performs entity recognition and relationship extraction on the text to be analyzed through the target large model to improve the accuracy of entity recognition and relationship extraction, and obtains entity results and relationship results. The generation module 304 constructs a knowledge graph and generates an APT organization attack portrait based on the above-mentioned extracted entity results and relationship results. This knowledge graph can more vividly and accurately represent the association between entities and relationships in the text to be analyzed. Finally, the reasoning module 305 performs attack reasoning on the APT organization attack portrait, analyzes potential attack patterns and threat chains, and obtains reasoning results. In the above process, by combining a large model, compared with traditional deep learning-based methods, the labor cost is significantly reduced. By giving full play to the advantages of the large model in semantic understanding, a more accurate and comprehensive APT organization attack portrait can be constructed. Through in-depth analysis and reasoning of the APT organization attack portrait, the behavior patterns, attack strategies and potential targets of the APT organization can be revealed from multiple dimensions, helping security personnel to more accurately identify and respond to complex attack threats, providing strong data support, and thus formulating more effective and personalized security protection measures, solving the problem of poor accuracy in existing attack portrait generation and reasoning methods.
[0046] Figure 4 An example of the entity structure diagram of an electronic device is as Figure 4 shown. The electronic device may include: a processor 401, a communications interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communications interface 402, and the memory 403 complete mutual communication through the communication bus 404. The processor 401 can call the logical instructions in the memory 403 to execute the APT organization attack portrait generation and reasoning method, and this method includes: Acquire reference data and perform data annotation on the reference data through a large model; Fine-tune the large model with the reference data to obtain the target large model; Acquire the text to be analyzed and perform entity recognition and relationship extraction on the text to be analyzed through the target large model to obtain entity results and relationship results; Construct a knowledge graph based on the extracted entity results and relationship results, and generate an APT organization attack profile; Conduct attack reasoning on the APT organization attack profile to obtain reasoning results.
[0047] In addition, when the logic instructions in the above-mentioned memory 403 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0048] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the APT organization attack profile generation and reasoning method provided by the above-mentioned various methods. The method includes: Obtain reference data, and perform data annotation on the reference data through a large model; Fine-tune the large model through the reference data to obtain a target large model; Obtain the text to be analyzed, and perform entity recognition and relationship extraction on the text to be analyzed through the target large model to obtain entity results and relationship results; Construct a knowledge graph based on the extracted entity results and relationship results, and generate an APT organization attack profile; Conduct attack reasoning on the APT organization attack profile to obtain reasoning results.
[0049] On yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the APT organization attack profile generation and reasoning method provided by the above-mentioned various methods. The method includes: Obtain reference data, and perform data annotation on the reference data through a large model; Fine-tune the large model through the reference data to obtain a target large model; Obtain the text to be analyzed, and perform entity recognition and relationship extraction on the text to be analyzed through the target large model to obtain entity results and relationship results; Construct a knowledge graph based on the extracted entity results and relationship results, and generate an APT organization attack profile; Perform attack reasoning on the APT organization attack profile to obtain reasoning results.
[0050] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0051] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0052] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating an attack profile of an APT organization, characterized in that: include: Acquire reference data, and annotate the reference data using a large model; Fine-tune the large model using the reference data to obtain a target large model; Acquire the text to be analyzed, and perform entity recognition and relationship extraction on the text to be analyzed through the target large model to obtain entity results and relationship results; Building a knowledge graph based on the extracted entity results and the relationship results, and generating an attack profile of an APT organization; Perform attack reasoning on the attack portrait of the APT organization to obtain a reasoning result.
2. The APT organization attack profile generation reasoning method according to claim 1 is characterized in that: Acquire reference data and annotate the reference data using a large model, including: Acquire reference data; the reference data includes APT attack analysis data from APT analysis reports from different sources; The large model extracts entities and relationships in the reference data through prompt word engineering to obtain reference entities and reference relationships; Data cleaning is performed on the reference entity and the reference relationship to obtain final reference data.
3. The APT organization attack profile generation reasoning method according to claim 1 is characterized in that: Fine-tuning the large model using the reference data to obtain a target large model includes: Sorting the reference data, adjusting the data format, and obtaining a standard data set; The LLaMAFactory framework is called, and the large model is fine-tuned by LoRA using the standard data set to obtain the target large model.
4. The APT organization attack profile generation reasoning method according to claim 3 is characterized in that: Fine-tuning the large model using the standard data set to obtain the target large model includes: Calling the standard data set, and setting the path information and format information of the standard data set; Obtain the path information of the standard data set, set the storage path of the training results and the fine-tuning type; the fine-tuning type is LoRA; The standard data set is retrieved according to the path information, and the large model is trained and fine-tuned according to the fine-tuning type, and the training result is output to obtain the target large model.
5. The APT organization attack profile generation reasoning method according to claim 1 is characterized in that: The target large model is used to perform entity recognition and relationship extraction on the text to be analyzed to obtain entity results and relationship results, including: The text to be analyzed is input into the target large model, and the target large model is guided to perform entity recognition and relationship extraction in the text to be analyzed through prompt word technology to obtain the entity result and the relationship result.
6. The APT organization attack profile generation reasoning method according to claim 5 is characterized in that: The entity results include attackers, victims, and attack methods.
7. The APT organization attack profile generation reasoning method according to claim 6 is characterized in that: Constructing a knowledge graph based on the extracted entity results and the relationship results, including: Acquire a related library, and connect to a graphic database through the related library; The nodes and relationships of the knowledge graph are created based on the relevant library and the graph database; the nodes correspond to the entity results, and the relationships correspond to the relationship results.
8. The APT organization attack profile generation reasoning method according to claim 1 is characterized in that: Perform attack reasoning on the attack profile of the APT organization to obtain reasoning results, including: Deploy the GraphRAG framework and load the target large model; The text to be analyzed is combined with the attack portrait of the APT organization, and the text to be analyzed is inferred through a question-and-answer interaction to obtain an inference result.
9. A device for generating and reasoning an attack profile of an APT organization, characterized in that: include: An acquisition module, used to acquire reference data and annotate the reference data using a large model; A fine-tuning module, used to fine-tune the large model using the reference data to obtain a target large model; An analysis module is used to obtain a text to be analyzed, and perform entity recognition and relationship extraction on the text to be analyzed through the target large model to obtain entity results and relationship results; A generation module, used to construct a knowledge graph based on the extracted entity results and the relationship results, and generate an attack portrait of an APT organization; The reasoning module is used to perform attack reasoning on the attack profile of the APT organization to obtain a reasoning result.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the APT organization attack profile generation reasoning method as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Cooperative enhancement method oriented to APT knowledge graph and large language model
CN118802369A
Well engineering intelligent question-answering system and method
CN119088929A
Multi-agent driven industrial software component assembling method and system
CN119166193A
Data security analysis method and intelligent calculation data security workstation
CN119249440A
Method and device for optimizing memory ability of large model agent
CN119293259A