Attack tool code knowledge extraction method and device and computer equipment
By building a special instruction set and soft prompt vector optimization large language model, the difficulty of identifying complex entities under the condition of lack of labeling data is solved, and efficient and accurate extraction of attack tool code knowledge under a small amount of resources is achieved.
Patent Information
- Application Number
- CN202510518441.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-12
AI Technical Summary
The existing open source tool code knowledge extraction methods are difficult to accurately identify complex attack entities under the condition of lack of labeling data, and large language models have problems such as information conflict and high computing resource consumption in the generation task.
Build a special instruction set and soft prompt vector, combine pre-trained large language model and prefix network layer, optimize the generative information extraction model through a small amount of supervision and fine-tuning data sets, and guide the model to accurately extract attack tool code knowledge.
With a smaller resource consumption, high accuracy and standardized extraction of complex attack tool codes is achieved, reducing dependence on large-scale annotation data and computing resources, and adapting to the knowledge extraction requirements of different types of attack tool codes.
Smart Images

Figure CN120471150A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, in particular to the field of code semantic understanding and threat intelligence extraction, and specifically to a method, device and computer equipment for extracting attack tool code knowledge. Background Art
[0002] With the expansion of the network attack surface and the escalation of offensive and defensive confrontations, cybersecurity threats are proliferating, characterized by a diverse array of attack behaviors and tools. Research shows that a significant proportion of hackers and security developers use open source tools or release their tools to open source platforms, and the number of these tools and developers is continuously growing. Open source hacking tool code contains a wealth of information related to attack behaviors. Timely discovery and analysis of this information is crucial for early detection of new threats, analysis of attack intent, and the development of defensive measures. Therefore, the ability to automatically identify and extract this attack-related entity information at the code level has significant practical application value. Knowledge extraction techniques for open source software source code primarily analyze the code structure to extract essential entities such as classes, methods, and variables. However, research on knowledge extraction for open source software in specific domains is relatively scarce. In particular, open source code, such as hacking tools, often involves complex entities such as attack payloads, target systems, and exploitation methods, making accurate identification difficult using structured representations alone.
[0003] The shortcomings of existing knowledge extraction methods are summarized as follows:
[0004] (1) Current code entity recognition technologies mainly rely on intermediate representations such as the code's abstract syntax tree and code attribute graph to identify and extract basic entities (such as classes and variable names) in the code. However, these methods are difficult to handle more complex code entity recognition scenarios, such as cross-domain entities and entities that require contextual semantics for reasoning.
[0005] (2) Traditional rule or template matching methods lack generalization capabilities and cannot generalize to unknown new samples. Methods based on statistical machine learning require a large amount of high-quality labeled data, but collecting and labeling data in practical applications is often difficult.
[0006] (3) Although large language models have great advantages in generative tasks and can easily complete knowledge extraction tasks based on a small number of samples and prompts, they cannot completely avoid problems such as information conflict and making something out of nothing in the generated results. At the same time, as the number of model parameters continues to increase, the model training process also faces severe challenges in terms of computing resources and time consumption.
[0007] In summary, existing methods may suffer from a lack of labeled data when the open source tool code has complex semantics, many nested and discontinuous entities, which can easily lead to difficulties in extracting code knowledge under the condition of scarce labeled data. Summary of the Invention
[0008] In view of this, the present invention provides a method, apparatus and computer equipment for extracting code knowledge from attack tools to solve the problem of difficulty in extracting code knowledge under the condition of scarce annotated data.
[0009] In a first aspect, the present invention provides a method for extracting attack tool code knowledge, the method comprising:
[0010] Obtain open source attack tool code and build an instruction set for the attack tool code knowledge extraction task based on the open source attack tool code;
[0011] Add a prefix network layer to the pre-trained large language model and generate a soft hint vector based on the prefix network layer;
[0012] Obtain the supervised fine-tuning dataset corresponding to the attack tool code, and build a structure-optimized generative information extraction model based on the instruction set, soft hint vectors, supervised fine-tuning dataset, pre-trained large language model, and prefix network layer;
[0013] The open source attack tool code is input into the structure-optimized generative information extraction model to identify the attack tool code knowledge.
[0014] The present invention provides a method for extracting attack tool code knowledge. This method constructs a specialized instruction set that effectively guides the model to accurately extract the required knowledge, preventing the generated content from deviating from expectations and significantly improving the accuracy and standardization of the extraction results. Adding a prefix network layer to a pre-trained large language model increases the number of trainable parameters, avoids training instability, and adapts to complex attack tool code knowledge extraction tasks. Based on a small amount of annotated supervised fine-tuning dataset, the method combines the instruction set and soft hint vectors with the pre-trained model to fully utilize the knowledge learned during the pre-training phase at a low resource cost, extracting structured knowledge from massive amounts of unstructured open source code. Compared to traditional methods, this method reduces the reliance on large-scale annotated data and computing resources. Open source attack tool code is input as files, with a simple and standardized input format. The method can be flexibly applied to different open source attack tool codes. By adjusting the instruction set and fine-tuning the model, it can adapt to the knowledge extraction needs of attack tool codes of different types and characteristics, demonstrating strong versatility and practicality. The powerful understanding and generation capabilities of the pre-trained large language model are leveraged to address the difficulty of complex code entity recognition in the absence of annotated data.
[0015] In an optional embodiment, constructing an instruction set for an attack tool code knowledge extraction task includes:
[0016] Set the code knowledge entity corresponding to the attack tool code in the extraction task;
[0017] An instruction set for attack tool code knowledge extraction task is constructed based on code knowledge entities. The instruction set includes task description, entity definition, output format, sample examples and error prevention.
[0018] This paper proposes a method for extracting knowledge from attack tool code. By defining the code knowledge entity corresponding to the attack tool code in the extraction task, the extraction scope is precisely defined, allowing the model to focus on the key knowledge within the entity. By combining the task description of the instruction set with the entity definition, the model can accurately understand the extraction requirements, avoid invalid or erroneous extraction, and significantly improve the accuracy and relevance of the extraction results.
[0019] In an optional embodiment, the pre-trained large language model includes multiple Transformer layers, a prefix network layer is added to the pre-trained large language model, and a soft hint vector is generated based on the prefix network layer, including:
[0020] Add an additional linear layer as a prefix network layer before each Transformer layer of the pre-trained large language model;
[0021] The prefix network layer parameters are initialized using reparameterization technology to generate the corresponding soft hint vector.
[0022] The present invention provides a method for extracting knowledge from attack tool code, which adds an additional linear layer as a prefix network layer before each Transformer layer of a pre-trained large language model, introducing more trainable parameters for the model during downstream task adaptation. These parameters can be optimized for the attack tool code knowledge extraction task, allowing the model to learn feature representations that are more in line with task requirements, thereby enhancing the model's learning and processing capabilities for specific tasks and improving the extraction effect. The prefix network layer parameters are initialized using reparameterization technology, effectively avoiding the training instability problem that may be caused by random initialization. A stable training process helps the model converge faster, reduces fluctuations and anomalies during training, ensures that the model can continuously and stably learn and optimize, and thus guarantees the performance of the final model in the attack tool code knowledge extraction task.
[0023] In an optional embodiment, before obtaining the supervised fine-tuning dataset corresponding to the attack tool code knowledge, the attack tool code knowledge extraction method further includes:
[0024] Adding examples of extraction task context to pre-trained large language models based on instruction sets.
[0025] The present invention provides a method for extracting attack tool code knowledge. By adding extraction task context examples to a pre-trained large language model based on an instruction set, abstract instructions are converted into concrete and intuitive cases, helping the model to more clearly grasp the goals, requirements, and output forms of the attack tool code knowledge extraction task.
[0026] In an optional implementation, obtaining a supervised fine-tuning dataset corresponding to the attack tool code includes:
[0027] The code knowledge entities corresponding to the attack tool codes are annotated based on expert experience as a supervised fine-tuning dataset.
[0028] The present invention provides a method for extracting attack tool code knowledge. This method, based on a supervised fine-tuning dataset annotated by expert experience, effectively avoids annotation errors caused by insufficient understanding of attack tool code knowledge, ensures data accuracy and professionalism, and provides high-quality annotated samples for model training. Experts are familiar with the application and characteristics of attack tool codes in real-world scenarios, and the annotated data more closely reflects the actual requirements for extracting attack tool code knowledge. Fine-tuning the model based on this dataset allows it to better learn the knowledge extraction patterns and features required for practical applications, making the trained model more targeted and practical in real-world use and effectively solving practical problems.
[0029] In an optional embodiment, a structured generative information extraction model is constructed based on an instruction set, a soft hint vector, a supervised fine-tuning dataset, a pre-trained large language model, and a prefix network layer, including:
[0030] The instruction set, soft hint vector, and supervised fine-tuning dataset are input into the pre-trained large language model through the prefix network layer to output the generated results, and the contrastive loss decoding strategy is used to optimize the generated results.
[0031] Calculate the maximum likelihood estimation loss between the optimized generation result and the extracted task context example, and update the parameters of the prefix network layer based on the maximum likelihood estimation loss to generate an optimized soft prompt vector sequence;
[0032] The instruction set, extracted task context examples, optimized soft prompt vector sequence and supervised fine-tuning dataset are concatenated and re-input into the pre-trained large language model and prefix network layer as new input to obtain a structurally optimized generative information extraction model.
[0033] The present invention provides a method for extracting attack tool code knowledge, organically combining an instruction set, soft hint vectors, a supervised fine-tuning dataset, a pre-trained large language model, and a prefix network layer, with each component playing a unique role. The instruction set specifies the task specifications, the soft hint vectors guide model generation, the supervised fine-tuning dataset provides a labeling basis, and the pre-trained large language model and prefix network layer collaborate to process the input. The collaborative work of these multiple components enables the model to fully understand the attack tool code knowledge extraction task, significantly enhancing its task processing capabilities. A contrastive loss decoding strategy is used to optimize the generated results, avoiding monotonous repetition or instability in the generated content. By calculating the maximum likelihood estimation loss between the optimized generated results and the extraction task context examples, the prefix network layer parameters are updated based on the loss value to generate an optimized soft hint vector sequence. This precise parameter update method enables the model to be continuously adjusted and optimized according to task requirements, effectively avoiding blind training and improving the model's performance in the attack tool code knowledge extraction task. The instruction set, the extraction task context examples, the optimized soft hint vector sequence, and the supervised fine-tuning dataset are concatenated and re-input into the model, forming an iterative optimization process. The model continuously learns and improves through multiple input and output, and can better adapt to changes in code knowledge extraction scenarios of different attack tools.
[0034] In an optional embodiment, the instruction set, the extracted task context example, the optimized soft prompt vector sequence, and the supervised fine-tuning dataset are concatenated and re-input as new input to the pre-trained large language model and the prefix network layer to obtain a structure-optimized generative information extraction model, including:
[0035] The instruction set, the extracted task context examples, the optimized soft prompt vector sequence, and the supervised fine-tuning dataset are concatenated and re-input into the pre-trained large language model and prefix network layer as new input to generate the output sequence probability distribution;
[0036] The loss is calculated as the difference between the output sequence probability distribution and the supervised fine-tuning dataset. The prefix network layer parameters are updated inversely based on the loss, and the prefix network layer is fine-tuned based on the updated prefix network layer parameters to obtain a generative information extraction model with optimized structure.
[0037] The present invention provides a method for extracting attack tool code knowledge. The method combines an instruction set, an extraction task context example, an optimized soft prompt vector sequence, and a supervised fine-tuning dataset as new input, integrating multi-dimensional information such as task specifications, example guidance, optimization prompts, and annotated data. The synergistic effect of multi-source data helps the pre-trained large language model and prefix network layer to more comprehensively and deeply understand the attack tool code knowledge extraction task, strengthen the model's ability to capture task characteristics and regularities, and lay the foundation for accurate extraction. The loss is calculated based on the difference between the output sequence probability distribution and the supervised fine-tuning dataset. The annotated data is directly used as a reference standard to accurately measure the deviation between the model output and actual requirements. Based on this loss, the prefix network layer parameters are reversely updated, and the model structure and parameters can be adjusted in a targeted manner to optimize the model in a direction that meets the task requirements, avoid ineffective training, and improve training efficiency and model performance. By continuously updating the prefix network layer parameters based on the loss and performing fine-tuning, an iterative optimization process is formed. The prefix network layer is repeatedly fine-tuned to make the model structure deeply adapted to the attack tool code knowledge extraction task. The optimized prefix network layer can generate soft hint vectors more stably, guiding the model to output more reliable results.
[0038] In an optional embodiment, the open source attack tool code is input into a structure-optimized generative information extraction model to identify attack tool code knowledge, including:
[0039] The open source attack tool code is organized into a file-based format, and the model input format is constructed by combining the soft prompt vectors, instruction sets, and preset start and end markers in the structure-optimized generative information extraction model.
[0040] Based on the model input format, soft prompt vectors, instruction sets, open source attack tool codes in file form, and preset start and end markers are input into the structure-optimized generative information extraction model to identify the attack tool code knowledge entities and corresponding attribute values.
[0041] The present invention provides a method for extracting attack tool code knowledge, which organizes the repository code into a file-based format, and combines soft prompt vectors, instruction sets, and preset start and end marks to construct a unified model input format, so that code data enters the model in a standardized form. This standardized input method facilitates the model to quickly identify and process data, reduces the complexity and time cost of data preprocessing, and improves the efficiency of overall knowledge extraction. The soft prompt vector, instruction set, and repository code are input into the model together, and the multiple factors work together. The three are combined with the code data to enable the model to more accurately understand the attack tool code knowledge extraction task, thereby identifying more accurate attack tool code knowledge entities and corresponding attribute values, and improving the accuracy and reliability of the extraction results.
[0042] In a second aspect, the present invention provides an attack tool code knowledge extraction device, the device comprising:
[0043] The code acquisition and instruction set construction module is used to obtain the open source attack tool code and construct the instruction set for the attack tool code knowledge extraction task based on the open source attack tool code;
[0044] A soft hint vector generation module is used to add a prefix network layer to the pre-trained large language model and generate a soft hint vector based on the prefix network layer;
[0045] An information extraction model construction module is used to obtain a supervised fine-tuning dataset corresponding to the attack tool code and build a structured generative information extraction model based on the instruction set, soft hint vectors and supervised fine-tuning dataset, a pre-trained large language model, and a prefix network layer;
[0046] The recognition module is used to input the open source attack tool code into the structure-optimized generative information extraction model to identify the attack tool code knowledge.
[0047] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to thereby execute the attack tool code knowledge extraction method of the above-mentioned first aspect or any corresponding embodiment thereof.
[0048] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to cause a computer to execute the attack tool code knowledge extraction method of the first aspect or any corresponding embodiment thereof.
[0049] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, wherein the computer instructions are used to cause a computer to execute the attack tool code knowledge extraction method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 1 is a flow chart of a method for extracting attack tool code knowledge according to an embodiment of the present invention;
[0052] Figure 2is a flow chart of another attack tool code knowledge extraction method according to an embodiment of the present invention;
[0053] Figure 3 1 is a flow chart of another attack tool code knowledge extraction method according to an embodiment of the present invention;
[0054] Figure 4 1 is a flow chart of another attack tool code knowledge extraction method according to an embodiment of the present invention;
[0055] Figure 5 2 is a schematic diagram of a structure-optimized generative information extraction model according to an embodiment of the present invention;
[0056] Figure 6 is a schematic diagram of a model input and output example according to an embodiment of the present invention;
[0057] Figure 7 is a structural block diagram of an attack tool code knowledge extraction device according to an embodiment of the present invention;
[0058] Figure 8 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0059] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0060] Existing security research primarily focuses on extracting security intelligence from threat intelligence reports or technical blogs, neglecting the code information of open source attack tools widely used by hackers and security professionals. However, due to the poor standardization of open source community data and the complex structure of code, extracting structured code knowledge from it is difficult. Currently, most knowledge extraction research on open source software focuses on a single type of data, and the extracted software entities are primarily limited to basic entities such as classes and variables in the code, making it difficult to extract complex code entities.
[0061] As the parameters of large language models have grown to hundreds of billions and they have begun to show early signs of general artificial intelligence, researchers have gradually applied large language models to various tasks. Knowledge extraction based on large language models leverages the models' powerful understanding and generation capabilities, guiding them to generate structured information based on templates or prompts. While large language models offer significant advantages in generative tasks, easily enabling knowledge extraction from a small number of samples and prompts, they cannot completely avoid issues such as information conflicts and fabrication in the generated results.
[0062] While the large number of parameters in large language models brings powerful understanding and reasoning capabilities, the model training process also faces severe computing resource and time challenges.
[0063] An embodiment of the present invention provides a method for extracting attack tool code knowledge. By constructing a structurally optimized generative information extraction model structure, a large pre-trained model adapted to downstream tasks can be obtained with a small number of samples and less resource consumption, achieving the effect of accurately extracting attack tool code knowledge under the condition of scarce labeled data.
[0064] According to an embodiment of the present invention, an embodiment of a method for extracting knowledge from attack tool code is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0065] In this embodiment, a method for extracting attack tool code knowledge is provided, which can be used in the above-mentioned computer device. Figure 1 FIG. 1 is a flow chart of a method for extracting attack tool code knowledge according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0066] Step S101: obtain open source attack tool code, and construct an instruction set for an attack tool code knowledge extraction task based on the open source attack tool code.
[0067] Specifically, we use the official API provided by GitHub (a cloud-based code hosting platform) and Python's Requests library (a Python third-party library) to collect repositories related to network attacks in the open source community, crawl the repository metadata and full code, obtain the open source attack tool code, and store the open source attack tool code in the database.
[0068] Knowledge extraction (KE) is an important task in natural language processing. Its goal is to automatically extract structured knowledge from multi-source, unstructured data based on the entity types and relationships defined in the ontology model.
[0069] In order to understand attack tool codes and tasks and achieve the goal of extracting structured attack tool code knowledge from massive unstructured open source codes, it is necessary to construct an instruction set for the attack tool code knowledge extraction task to effectively guide the pre-trained large language model based on the code and text knowledge learned in the pre-training phase.
[0070] Step S102: Add a prefix network layer to the pre-trained large language model, and generate a soft hint vector based on the prefix network layer.
[0071] Specifically, the pre-trained large language model is also called the Code Large Language Model (Code-LLM). The Code Large Model is based on the natural language large model architecture and is a large-scale language model obtained by pre-training with massive code data. It has the ability to understand and generate code.
[0072] To introduce additional parameters to guide adaptation to downstream tasks, a linear layer is added as a prefix network layer before each Transformer layer in the pre-trained large language model. This effectively increases the number of trainable parameters during downstream task adaptation. These additional adjustable parameters continuously optimize task-specific soft prompts during training. Using these optimized soft prompts to guide the model in generating task outputs effectively mitigates the instability caused by discrete prompts being significantly influenced by the prompt template.
[0073] Step S103: Obtain a supervised fine-tuning dataset corresponding to the attack tool code, and build a structure-optimized generative information extraction model based on the instruction set, soft hint vector, supervised fine-tuning dataset, pre-trained large language model, and prefix network layer.
[0074] Specifically, fine-tuning refers to further training a pre-trained model on a specific new dataset and fine-tuning the model parameters. This can fully utilize the general knowledge learned by the pre-trained model, while performing targeted optimization for new tasks, thereby improving the model's performance on new tasks.
[0075] Based on expert experience, the code knowledge entities corresponding to the attack tool code are annotated as a supervised fine-tuning dataset, i.e., the training sample set. The instruction set, soft hints, and code files are simultaneously used as model inputs. The loss is calculated based on the difference between the generated results and the annotated data to update the prefix network layer parameters. The supervised fine-tuning model's multi-layer prefix network layer yields a fine-tuned attack tool code knowledge extraction model, i.e., a structurally optimized generative information extraction model.
[0076] In step S104, the open source attack tool code is input into the structure-optimized generative information extraction model to identify the attack tool code knowledge.
[0077] Specifically, open source attack tool codes are randomly obtained from the database, and then input into a structure-optimized generative information extraction model to identify attack tool code knowledge.
[0078] The attack tool code knowledge extraction method provided in this embodiment constructs a specialized instruction set that effectively guides the model to accurately extract the required knowledge, preventing the generated content from deviating from expectations and significantly improving the accuracy and standardization of the extraction results. Adding a prefix network layer to a pre-trained large language model increases the number of trainable parameters, avoids training instability, and adapts to complex attack tool code knowledge extraction tasks. Based on a small amount of annotated supervised fine-tuning dataset, the instruction set and soft hint vectors are combined with the pre-trained model to fully utilize the knowledge learned during the pre-training phase at a low resource cost, extracting structured knowledge from massive amounts of unstructured open source code. Compared to traditional methods, this method reduces the reliance on large-scale annotated data and extensive computing resources. The open source attack tool code is input as a file, with a simple and standardized input format. It can be flexibly applied to different open source attack tool codes. By adjusting the instruction set and fine-tuning the model, it can adapt to the knowledge extraction needs of attack tool codes of different types and characteristics, demonstrating its high versatility and practicality. Leveraging the powerful understanding and generation capabilities of the pre-trained large language model, it solves the problem of complex code entity recognition difficulties in the absence of annotated data.
[0079] In this embodiment, a method for extracting attack tool code knowledge is provided, which can be used in the above-mentioned computer device. Figure 2 FIG. 1 is a flow chart of a method for extracting attack tool code knowledge according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:
[0080] Step S201: obtain open source attack tool code, and construct an instruction set for the attack tool code knowledge extraction task based on the open source attack tool code.
[0081] Specifically, the above step S201 includes:
[0082] Step S2011: setting the code knowledge entity corresponding to the attack tool code in the extraction task.
[0083] Specifically, based on expert experience and the characteristics of open source attack tools, the code knowledge entity types for knowledge extraction are limited to eight categories, including operating systems, vulnerabilities, attack payloads, sensitive files, sensitive paths, sensitive links, network protocols and encoding methods.
[0084] Step S2012: construct an instruction set for the attack tool code knowledge extraction task based on the code knowledge entity. The instruction set includes task description, entity definition, output format, sample examples, and error prevention.
[0085] Specifically, the code knowledge extraction instruction set is defined from five aspects: task description, entity definition, output format, sample example, and error prevention. In unsupervised pre-training, the large language model acquires a generalization and pattern recognition ability, which is called contextual learning ability, and can quickly adapt to or identify the tasks to be done based on examples. However, if it is based only on contextual examples, the output generated by the model may not meet expectations. By providing the model with semantically clear instructions to drive and limit the model generation process, the format and content of the model output can be effectively standardized. Therefore, this embodiment defines an instruction set for the attack tool code knowledge extraction task, and adds a sample example field to the instruction set. The specific field description is shown in Table 1 below.
[0086] Table 1 Instruction set field description
[0087]
[0088] Step S202: Add a prefix network layer to the pre-trained large language model, and generate a soft hint vector based on the prefix network layer.
[0089] Specifically, the pre-trained large language model includes multiple Transformer layers, and the above step S202 includes:
[0090] Step S2021: Add an additional linear layer as a prefix network layer before each Transformer layer of the pre-trained large language model.
[0091] like Figure 5 As shown in the figure, an additional linear layer is added as a prefix network layer before each Transformer layer in the pre-trained large language model, effectively increasing the number of trainable parameters of the pre-trained large language model during downstream task adaptation. These additional adjustable parameters continuously optimize task-specific soft prompts during training. Using the optimized soft prompts to guide the pre-trained large language model in generating task outputs effectively avoids the instability caused by the large influence of the prompt pre-trained large language template on discrete prompts.
[0092] Step S2022: Initialize the prefix network layer parameters using a reparameterization technique to generate a corresponding soft hint vector.
[0093] Specifically, reparameterization techniques are primarily used to address the problem of estimating the gradient of random variables, enabling the training of neural networks containing random layers using the backpropagation algorithm. In deep learning, particularly in variational autoencoders and diffusion models, reparameterization enables models to better handle random processes, thereby improving their training efficiency and stability.
[0094] Soft Prompt Vector is a technique in prompt learning, mainly used to guide the performance of pre-trained language models on specific tasks. Soft prompt vectors embed trainable continuous vectors into the model input. These vectors are optimized through gradient optimization methods during training to adapt to specific tasks and datasets.
[0095] This embodiment uses a reparameterization technique to initialize prefix network layer parameters and generate corresponding soft hint vectors.
[0096] Step S203: adding an extraction task context example to the pre-trained large language model based on the instruction set.
[0097] Specifically, based on the eight code knowledge entity types defined in the instruction set (operating system, vulnerability, attack payload, sensitive file, sensitive path, sensitive link, network protocol, and encoding method), typical code snippets containing these entities were selected from open source attack tool codes. For example, for the "vulnerability" entity type, a section of attack code containing a common vulnerability (such as SQL injection) was selected as an example. This ensured that the examples comprehensively covered all entities, allowing the model to learn how different entities are represented in the code.
[0098] Design input code examples and corresponding correct output results according to the task description requirements in the instruction set. Output results must strictly follow the specified output format and be presented in a structured form, such as JSON, clearly listing entity types and corresponding attribute values. For example, for a code segment containing a specific attack payload, the output result would be {"entity type":"attack payload","attribute value":"specific attack payload name and parameters"}, so that the model can intuitively understand the expected output format of the task.
[0099] Ways to add examples include:
[0100] Directly embedding the instruction set: Designed input and output examples are added directly to the instruction set as sample example fields. When the instruction set is fed into a pre-trained large language model, these examples are also received by the model and serve as a reference for the model's understanding of the task. For example, in the sample example section of the instruction set, multiple different code input examples and their corresponding correct outputs are listed in sequence, allowing the model to learn the task rules through multiple examples.
[0101] Pre-trained guidance input: Before feeding the instruction set and soft hint vectors into the pre-trained large language model, examples of the extracted task context are fed into the model as guidance information. This gives the model a preliminary understanding of the task scenario and expected results before it encounters specific task instructions and data, helping it adapt more quickly to subsequent extraction tasks.
[0102] Extracting task context examples visualizes the task descriptions and entity definitions in the instruction set. For example, the definition of the "sensitive file" entity in the instruction set is relatively abstract. By showing a code snippet containing the sensitive file path and name, along with the corresponding output, the model can more clearly understand the specific manifestation of "sensitive file" in the actual code, thereby strengthening its understanding of the instructions.
[0103] The instruction set specifies the output format, and the examples demonstrate how to output in that format through specific cases. After learning from the examples, the model can follow the instruction set's output format requirements in actual extraction tasks, generating standardized and unified results, which facilitates subsequent processing and analysis of the extraction results.
[0104] Step S204: Obtain the supervised fine-tuning dataset corresponding to the attack tool code, and build a structure-optimized generative information extraction model based on the instruction set, soft hint vector, supervised fine-tuning dataset, pre-trained large language model, and prefix network layer. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.
[0105] Step S205: Input the existing open source attack tool code into the structure-optimized generative information extraction model to identify the attack tool code knowledge. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.
[0106] The attack tool code knowledge extraction method provided in this embodiment significantly enriches threat intelligence libraries by extracting attack tool information, including attack payloads, target systems, and vulnerabilities, from open-source attack tool code, thus aiding attack detection and tracing. Combining the advantages of instruction learning, contextual learning, and soft prompting, this method guides the model to understand the semantics of input text and downstream tasks based on knowledge learned from large-scale data during the pre-training phase. This achieves high-quality knowledge extraction while minimizing resource consumption and manual annotation costs.
[0107] In this embodiment, a method for extracting attack tool code knowledge is provided, which can be used in the above-mentioned computer device. Figure 3 FIG. 1 is a flow chart of a method for extracting attack tool code knowledge according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:
[0108] Step S301: Obtain the open source attack tool code and construct an instruction set for the attack tool code knowledge extraction task based on the open source attack tool code. Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.
[0109] Step S302: Add a prefix network layer to the pre-trained large language model and generate a soft hint vector based on the prefix network layer. Figure 2 Step S202 of the illustrated embodiment will not be described in detail here.
[0110] Step S303: Obtain a supervised fine-tuning dataset corresponding to the attack tool code, and build a structure-optimized generative information extraction model based on the instruction set, soft hint vector, supervised fine-tuning dataset, pre-trained large language model, and prefix network layer.
[0111] Specifically, the above step S303 includes:
[0112] Step S3031: annotate the code knowledge entities corresponding to the attack tool code based on expert experience as a supervised fine-tuning dataset.
[0113] Specifically, experts with deep theoretical knowledge and extensive practical experience in the field of network security, especially in attack tool code analysis, will be selected. These experts must be familiar with the principles and code structures of various attack tools, as well as common attack methods and corresponding code knowledge entity types, such as operating systems, vulnerabilities, and attack payloads, to ensure they have the ability to accurately identify and annotate code knowledge entities. Before the annotation work begins, experts will be organized for centralized training to clarify the annotation standards and specifications for attack tool code knowledge entities. Combined with the definitions and examples of code knowledge entity types in the instruction set, the experts will be explained in detail the identification points, annotation requirements, and output format of each type of entity to ensure that different experts follow consistent standards during the annotation process.
[0114] The collected open source attack tool code samples are divided into multiple labeling task units according to certain rules (such as code size, attack type, etc.) and assigned to different experts or expert groups to ensure that the labeling work is carried out efficiently and orderly.
[0115] Experts conduct a detailed line-by-line analysis of assigned code samples. Based on their experience and the standards learned through training, they identify the various code knowledge entities involved in the code. For example, when analyzing an attack script, experts need to determine the operating system type used, the vulnerable code snippets present, the attack payloads carried, and other information, and annotate them according to the prescribed format. For each identified entity, they record detailed information such as its code location, entity type, and specific attribute values.
[0116] The obtained supervised fine-tuning dataset is used as training samples for pre-training large language models.
[0117] In step S3032, the instruction set, soft hint vector, and supervised fine-tuning dataset are input into the pre-trained large language model through the prefix network layer to output the generated result, and the contrast loss decoding strategy is used to optimize the generated result.
[0118] Specifically, the constructed instruction set, the generated soft hint vector, and the supervised fine-tuning dataset annotated based on expert experience are integrated. During the input process, these data are first fed into the prefix network layer. The prefix network layer is composed of an additional linear layer before each Transformer layer of the pre-trained large language model. It can perform preliminary processing and feature extraction on the input data, providing the pre-trained large language model with input that better meets the task requirements. The data processed by the prefix network layer is then fed into the pre-trained large language model, and after calculation and processing by the multiple layers of Transformer layers within the model, the preliminary generation results are finally output.
[0119] A contrastive loss decoding strategy is used to optimize the initial generated results. This strategy, based on the probability distribution of the model-generated sequences, assumes that the generated sequence should be selected from the largest candidate in the probability distribution, while maintaining the required differentiation between the currently generated token and the previously generated content. This strategy also avoids the instability or unreadable generation caused by random sampling algorithms. Common generative models use a greedy decoding strategy, directly selecting the word with the highest probability as the output. However, this decoding strategy is relatively conservative and can result in monotonous and repetitive generated text.
[0120] Specifically, the strategy is for the generated content x0,…,x t-1 , at time step t, label x t The calculation process is shown in formula (1).
[0121]
[0122] Among them, V (k) is the mark x t The candidate set is based on the model p θ (v|x <t ) predicts the k labels with the highest probability; α is the adjustment parameter, s is the similarity function, h v 、 Represents the label v and label x respectively t The embedding vector of The maximum value of the similarity between the tag v and the previously generated sequence tag. If the tag v is very similar to the previous sequence tag, its overall probability value is reduced, thereby reducing the possibility of being decoded.
[0123] The generation process is shown in formula (2):
[0124]
[0125] Among them, W0 is the initial context word sequence, T is the length of the word sequence determined during generation, and when the EOS mark (stop mark) appears at a certain time t, the word sequence generation is stopped, w t represents the tth word, w 1:T Represents the sequence from the 1st word to the Tth word, w 1:0 Indicates that there is no previous word at the beginning of the sequence, that is, the initial state of the sequence is empty, P(w 1:T |W0) is to generate the first word w1 to the Tth word w T The probability of the sequence, P(w t ∣w 1:t-1 , W0) represents the process from t=1 to t-1, each step generates word w t probability.
[0126] Step S3033: Calculate the maximum likelihood estimation loss between the optimized generation result and the extracted task context example, and update the parameters of the prefix network layer based on the maximum likelihood estimation loss to generate an optimized soft prompt vector sequence.
[0127] Specifically, the optimized generation results are compared with pre-prepared examples of the extraction task context, and the difference between the model output and the expected result is measured by calculating the maximum likelihood loss between the two. The maximum likelihood loss quantifies the degree of deviation in the model prediction by evaluating the probability of the generation result appearing under the given context example, providing a basis for subsequent parameter updates. Based on the maximum likelihood loss, the parameters of the prefix network layer are updated. The backpropagation algorithm starts from the output layer of the large language model, propagates the loss value backward along the network layer, calculates the contribution of each parameter to the loss, and then adjusts the parameter value. During the parameter update process, the sequence of soft prompt vectors generated by the prefix network layer is also optimized. These adjustable parameters are continuously adjusted during training, so that the soft prompt vector can better guide the model to generate output that meets the task requirements, effectively avoiding the instability caused by the large influence of the prompt template on discrete prompts.
[0128] Specifically, the word embedding is performed on the tag sequence of length n to obtain the embedding representation Where e is the dimension of the embedding space. After applying this embedding table, we get a soft-coded representation P e By concatenating the vectors embedded with the context examples and instruction set text, we can get the complete input [P e ;I e ;X e], I e Embedding vectors representing context examples and instruction sets. This new input will enter the multi-layer stacked decoder layer and prefix network layer of the pre-trained language model for training and inference.
[0129] In step S3034, the instruction set, the extracted task context example, the optimized soft prompt vector sequence, and the supervised fine-tuning dataset are concatenated and re-input as new input to the pre-trained large language model and the prefix network layer to obtain a generative information extraction model with optimized structure.
[0130] In some optional implementations, step S3034 includes:
[0131] In step a1, the instruction set, the extracted task context examples, the optimized soft prompt vector sequence, and the supervised fine-tuning dataset are concatenated and re-input as new input to the pre-trained large language model and prefix network layer to generate the output sequence probability distribution.
[0132] Specifically, after the concatenated new input data enters the pre-trained large language model and prefix network layer, it is first processed by the prefix network layer. The linear layer in the prefix network layer performs feature transformation on the input data, extracting features relevant to the attack tool code knowledge extraction task and generating more targeted soft prompts. These soft prompts, along with the input data, enter the pre-trained large language model. Within the model's multi-layer Transformer layers, a self-attention mechanism performs deep semantic analysis on the input data, capturing the associations and features between code knowledge entities. After multiple layers of calculation and transmission, the model's output layer generates an output sequence probability distribution, which represents the model's predicted probability for each possible output token.
[0133] For example, for the output sequence probability distribution T generated by the pre-trained model and the prefix network layer, the hidden layer output of the last layer is obtained On this basis, the fully connected layer transforms the vector to achieve more stable convergence. Based on the output vector of the output layer decoder and the previously generated code fragment, the token with the highest score calculated by the decoding algorithm is compared and searched as the next input. A new character or token is sampled from it. When the EOS token appears at a certain time t, the word sequence generation stops.
[0134] In step a2, the loss is calculated based on the difference between the output sequence probability distribution and the supervised fine-tuning dataset. The prefix network layer parameters are updated inversely based on the loss, and the prefix network layer is fine-tuned based on the updated prefix network layer parameters to obtain a generative information extraction model with optimized structure.
[0135] We used a supervised fine-tuning dataset as a baseline for comparison. This dataset contains accurate code knowledge entity information annotated based on expert experience. We compared the probability distribution of the output sequences generated by the model with the annotated results in the supervised fine-tuning dataset, and calculated the difference between the two to measure the degree of error in the model's predictions.
[0136] A suitable loss function (such as the cross-entropy loss function) is used to calculate the difference between the probability distribution of the output sequence and the supervised fine-tuning dataset. Taking the cross-entropy loss function as an example, its calculation process is as follows: for each labeled example in the supervised fine-tuning dataset, the cross-entropy value between the model's predicted probability distribution and the true annotation is calculated. The cross-entropy values of all examples are then summed or averaged to obtain the final loss value. This loss value quantifies the deviation between the model output and the true result. The smaller the loss value, the closer the model's prediction is to the true annotation, and the better the model performance.
[0137] The prefix network layer is fine-tuned by repeatedly repeating the above process of input, loss calculation, and parameter update. In each iteration, the model adjusts the prefix network layer's parameters based on the newly calculated loss value, optimizing the soft hint vector sequence. As training progresses, the prefix network layer gradually adapts to the characteristics of the attack tool code knowledge extraction task, generating more effective soft hints and guiding the model to output more accurate results.
[0138] After multiple rounds of fine-tuning, the parameters of the prefix network layer were optimized and the overall structure of the model was adjusted to form a structurally optimized generative information extraction model.
[0139] The model pseudo code of the training process of step S303 is shown in Table 2 below:
[0140] Table 2 Model pseudo code
[0141]
[0142] Step S304: input the open source attack tool code into the structure-optimized generative information extraction model to identify the attack tool code knowledge.
[0143] Specifically, the above step S304 includes:
[0144] Step S3041: organize the open source attack tool code into a file-based format, and construct a model input format by combining the soft prompt vector, instruction set, and preset start and end markers in the structure-optimized generative information extraction model.
[0145] Specifically, the preset start and end marks are represented by [BOS] marks and [EOS] marks respectively.
[0146] The open source attack tool code is input into the model as a file. The embedding vector of the task description in the instruction set is used to initialize the soft prompt vector sequence. The model input is constructed in the format of soft prompt vector + instruction set + [BOS] + open source attack tool code in file form + [EOS].
[0147] In step S3042, based on the model input format, the soft hint vector, instruction set, open source attack tool code in file format, and preset start and end markers are input into the structure-optimized generative information extraction model to identify the attack tool code knowledge entity and the corresponding attribute value.
[0148] For example, Figure 6 This is an example diagram of the input and output of a structure-optimized generative information extraction model. The input of the model is a soft prompt vector, an instruction set, a stored open source attack tool code snippet, and preset start and end markers. The output is the identified attack tool code knowledge entity and the corresponding attribute value.
[0149] The attack tool code knowledge extraction method provided in this embodiment provides a structurally optimized generative knowledge extraction model. By constructing a multi-layer prefix network layer, a pre-trained model with frozen parameters, and an output layer based on a contrastive loss decoding strategy, it fully unleashes the potential of a pre-trained large language model and optimizes model generation. By adding a trainable prefix encoding layer to an existing pre-trained language model and adjusting the sampling strategy of the output layer, the pre-trained model can be better transferred to downstream tasks. The open source attack tool code is organized into a file-based format and a unified model input format is constructed by combining soft hint vectors, an instruction set, and preset start and end markers, allowing code data to enter the model in a standardized form. This standardized input method facilitates the model's rapid data recognition and processing, reduces the complexity and time cost of data preprocessing, and improves the overall efficiency of knowledge extraction. The soft hint vectors, instruction set, and open source attack tool code are input into the model together. These multiple factors work synergistically. Combined with the code data, these three elements enable the model to more accurately understand the attack tool code knowledge extraction task, thereby identifying more precise attack tool code knowledge entities and corresponding attribute values, improving the accuracy and reliability of the extraction results.
[0150] As one or more specific application embodiments of the present invention, combined with Figure 4 The attack tool code knowledge extraction method provided by the present invention is further described in detail. Figure 4 As shown, the specific process is as follows:
[0151] Step 1: Use the official GitHub API and Python's Requests library to collect repositories related to cyberattacks in the open source community and crawl the repository metadata and full code.
[0152] Step 2: Design an instruction set for the attack tool code knowledge extraction task, and provide a task context example for the model based on the instruction set. The specific process of providing the task context example for the model based on the instruction set is as detailed in the above step S203.
[0153] Step 2.1: Based on expert experience and the characteristics of open source attack tools, the types of code knowledge entities to be extracted are limited to eight categories: operating system, vulnerability, attack payload, sensitive file, sensitive path, sensitive link, network protocol and encoding method.
[0154] Step 2.2: Define the code knowledge extraction instruction set from five aspects: task description, entity definition, output format, sample example, and error prevention. In unsupervised pre-training, the large language model obtains a generalization and pattern recognition ability, which is called contextual learning ability, and can quickly adapt to or identify the tasks to be done based on examples. However, if it is based only on contextual examples, the output generated by the model may not meet expectations. By providing the model with semantically clear instructions to drive and limit the process of model generation, the format and content of the model output can be effectively standardized. Therefore, the present invention defines an instruction set for the attack tool code knowledge extraction task, and adds a sample example field to the instruction set. The specific field description is shown in Table 1 above and will not be repeated here.
[0155] Step 3: Build a structure-optimized generative information extraction model. The model structure example is shown in the figure below. Figure 5 As shown;
[0156] Step 3.1: Add an additional linear layer as a prefix network layer before each Transformer layer in the pre-trained large language model. This effectively increases the number of trainable parameters during downstream task adaptation. These additional adjustable parameters continuously optimize task-specific soft prompts during training. Using these optimized soft prompts to guide the model in generating task outputs effectively mitigates the instability caused by discrete prompts being significantly influenced by the prompt template.
[0157] Step 3.2: Use reparameterization technology to initialize the prefix network layer parameters to avoid training instability caused by random initialization of the prefix network layer parameters.
[0158] Step 3.3: Pretraining the language model: By fixing the parameters of the pretrained model, the model retains the prior knowledge learned during pretraining, helping it understand the task context and input samples. This step allows you to select different pretrained language models based on the specific characteristics of the downstream task.
[0159] Step 3.4: Design a contrastive loss decoding strategy as the output layer decoding strategy. This strategy assumes that the generated sequence should come from the maximum candidate in the model's probability distribution, but when generating a sequence, the current token should remain distinct from the previously generated content. This avoids monotony and repetitive generation issues. It also prevents unstable or unreadable generation results caused by random sampling algorithms. Common generative models use a greedy decoding strategy, directly selecting the word with the highest probability as the output. However, this decoding strategy is relatively conservative and can result in the generated text being overly monotonous and repetitive.
[0160] Specifically, the strategy is for the generated content x0,…,x t-1 , at time step t, label x t The calculation process is shown in formula (1).
[0161]
[0162] Among them, V (k) is the mark x t The candidate set is based on the model p θ (v|x <t ) predicts the k labels with the highest probability; α is the adjustment parameter, s is the similarity function, h v 、 Represents the label v and label x respectively t The embedding vector of The maximum value of the similarity between the tag v and the previously generated sequence tag. If the tag v is very similar to the previous sequence tag, its overall probability value is reduced, thereby reducing the possibility of being decoded.
[0163] The generation process is shown in formula (2):
[0164]
[0165] Among them, W0 is the initial context word sequence, T is the length of the word sequence determined during generation, and when the EOS mark (stop mark) appears at a certain time t, the word sequence generation is stopped, w t represents the tth word, w 1:T Represents the sequence from the 1st word to the Tth word, w 1:0 Indicates that there is no previous word at the beginning of the sequence, that is, the initial state of the sequence is empty, P(w 1:T |W0) is to generate the first word w1 to the Tth word w t The probability of the sequence, P(w t ∣w 1:t-1 , W0) represents the process from t=1 to t-1, each step generates word w t probability.
[0166] Step 4: Fine-tune the model based on the labeled data supervision. The pseudo code implementation of the training process is shown in Table 2 above and will not be repeated here.
[0167] Step 4.1: The instruction set, soft hints, and annotated code files are simultaneously used as model inputs. During forward propagation, the input first passes through the prefix network layer and then into the pre-trained language model. After multiple layers of propagation, it is decoded at the output layer to generate the output. The loss is calculated by maximum likelihood estimation between the output and the label, and the parameters of the prefix network layer are updated to optimize the soft hint vector sequence.
[0168] Specifically, the word embedding is performed on the tag sequence of length n to obtain the embedding representation Where e is the dimension of the embedding space. After applying this embedding table, we get a soft-coded representation P e By concatenating the vectors embedded with the context examples and instruction set text, we can get the complete input [P e ;I e ;X e ], I e Embedding vectors representing context examples and instruction sets. This new input will enter the multi-layer stacked decoder layer and prefix network layer of the pre-trained language model for training and inference.
[0169] Step 4.2: For the output sequence probability distribution T generated by the pre-trained model and the prefix network layer, obtain the hidden layer output of the last layer On this basis, a fully connected layer transformation is used to achieve more stable vector convergence. Based on the output vector of the output layer decoder and previously generated code snippets, the token with the highest score calculated by the decoding algorithm is compared and searched as the next input. A new character or token is sampled from this token. When the EOS token appears at a certain time t, word sequence generation stops. Based on the difference between the generated result and the annotated data, the loss is calculated and the network parameters are back-updated. The multi-layer prefix network layer of the fine-tuned model is then supervised to obtain a fine-tuned attack tool code knowledge extraction model.
[0170] Step 5: Input the open source attack tool code into the model in file units to obtain the code knowledge recognized by the model. The specific operations are:
[0171] The soft hint vector sequence is initialized using the embedding vector of the task description in the instruction set. The model input is constructed in the format of soft hint + instruction + [BOS] + code + [EOS]. The model output is the entity type and the corresponding attribute value.
[0172] The attack tool code knowledge extraction method provided in this embodiment, based on a structurally optimized generative information extraction model structure, can generate a large pre-trained model suitable for downstream tasks with a small number of samples and minimal resource consumption. The designed instruction set, soft prompts, and decoding strategy significantly standardize model generation, guiding the model to understand attack tool code and tasks based on the code and text knowledge learned from large-scale unlabeled data during the pre-training phase. This achieves the goal of extracting structured attack tool code knowledge from massive amounts of unstructured open source code, solving the problem of complex knowledge extraction from massive amounts of unstructured open source tool code.
[0173] This embodiment also provides an attack tool code knowledge extraction device, which is used to implement the above-mentioned embodiments and preferred implementations. Details already described will not be repeated here. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0174] This embodiment provides an attack tool code knowledge extraction device, such as Figure 7 As shown, including:
[0175] The code acquisition and instruction set construction module 701 is used to obtain the open source attack tool code and construct an instruction set for the attack tool code knowledge extraction task based on the open source attack tool code.
[0176] The soft hint vector generation module 702 is configured to add a prefix network layer to the pre-trained large language model and generate a soft hint vector based on the prefix network layer.
[0177] The information extraction model construction module 703 is used to obtain the supervised fine-tuning dataset corresponding to the attack tool code, and build a structure-optimized generative information extraction model based on the instruction set, soft hint vector and supervised fine-tuning dataset, pre-trained large language model and prefix network layer.
[0178] The identification module 704 is used to input the open source attack tool code into the structure-optimized generative information extraction model to identify the attack tool code knowledge.
[0179] In some optional implementations, the code acquisition and instruction set construction module 701 includes:
[0180] The knowledge entity setting unit is used to set the code knowledge entity corresponding to the attack tool code in the extraction task.
[0181] The instruction set construction unit is used to construct an instruction set for the attack tool code knowledge extraction task based on the code knowledge entity. The instruction set includes task description, entity definition, output format, sample example and error prevention.
[0182] In some optional implementations, the pre-trained large language model includes multiple Transformer layers, and the soft hint vector generation module 702 includes:
[0183] The prefix network layer building unit is used to add an additional linear layer as a prefix network layer before each Transformer layer of the pre-trained large language model.
[0184] A soft hint vector generation unit is used to initialize the prefix network layer parameters using a reparameterization technique to generate a corresponding soft hint vector. In some optional embodiments, the attack tool code knowledge extraction device further includes:
[0185] The context example adding module is used to add extraction task context examples to the pre-trained large language model based on the instruction set.
[0186] In some optional implementations, the information extraction model building module 703 includes:
[0187] The supervised fine-tuning dataset determination unit is used to annotate the code knowledge entities corresponding to the attack tool code based on expert experience as the supervised fine-tuning dataset.
[0188] The generation result output and optimization unit is used to input the instruction set, soft prompt vector and supervised fine-tuning dataset into the pre-trained large language model through the prefix network layer to output the generation result, and use the contrast loss decoding strategy to optimize the generation result.
[0189] The soft hint vector sequence optimization unit is used to calculate the maximum likelihood estimation loss between the optimized generation result and the extracted task context example, and update the parameters of the prefix network layer based on the maximum likelihood estimation loss to generate an optimized soft hint vector sequence.
[0190] The structure-optimized generative information extraction model construction unit is used to splice the instruction set, the extracted task context examples, the optimized soft prompt vector sequence and the supervised fine-tuning dataset, and then re-input them as new input to the pre-trained large language model and prefix network layer to obtain a structure-optimized generative information extraction model.
[0191] In some optional implementations, the structure-optimized generative information extraction model building unit includes:
[0192] The output sequence probability distribution generation subunit is used to concatenate the instruction set, extracted task context examples, optimized soft prompt vector sequence and supervised fine-tuning dataset, and then re-input them as new input to the pre-trained large language model and prefix network layer to generate the output sequence probability distribution.
[0193] The prefix network layer fine-tuning subunit is used to calculate the loss based on the difference between the output sequence probability distribution and the supervised fine-tuning dataset, reversely update the prefix network layer parameters based on the loss, and fine-tune the prefix network layer based on the updated prefix network layer parameters to obtain a generative information extraction model with optimized structure.
[0194] In some optional implementations, the identification module 704 includes:
[0195] The model input construction unit is used to organize the open source attack tool code into a file-based format, and construct the model input format by combining the soft prompt vector, instruction set, and preset start and end markers in the structurally optimized generative information extraction model.
[0196] The recognition unit is used to input the soft prompt vector, instruction set, open source attack tool code in file form, and preset start and end marks into the structure-optimized generative information extraction model based on the model input format, and identify the attack tool code knowledge entity and corresponding attribute value.
[0197] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0198] The attack tool code knowledge extraction device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0199] The embodiment of the present invention also provides a computer device having the above Figure 7 The attack tool code knowledge extraction device shown.
[0200] See also Figure 8 , Figure 8 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 8As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 8 A processor 10 is taken as an example.
[0201] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0202] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0203] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0204] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0205] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 8 The bus connection is taken as an example.
[0206] The input device 30 can receive input digital or character information and generate key signal input related to user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0207] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0208] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0209] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for extracting attack tool code knowledge, characterized in that: The method comprises: Obtain open source attack tool code and build an instruction set for the attack tool code knowledge extraction task based on the open source attack tool code; Adding a prefix network layer to the pre-trained large language model and generating a soft hint vector based on the prefix network layer; Obtain the supervised fine-tuning dataset corresponding to the attack tool code, and build a structure-optimized generative information extraction model based on the instruction set, soft hint vectors, supervised fine-tuning dataset, pre-trained large language model, and prefix network layer; The open source attack tool code is input into the structure-optimized generative information extraction model to identify attack tool code knowledge.
2. The method according to claim 1, characterized in that The instruction set for constructing the attack tool code knowledge extraction task includes: Set the code knowledge entity corresponding to the attack tool code in the extraction task; An instruction set for an attack tool code knowledge extraction task is constructed based on the code knowledge entity. The instruction set includes a task description, an entity definition, an output format, a sample example, and error prevention.
3. The method according to claim 1, characterized in that The pre-trained large language model includes multiple Transformer layers, and the prefix network layer is added to the pre-trained large language model, and a soft hint vector is generated based on the prefix network layer, including: Add an additional linear layer as a prefix network layer before each Transformer layer of the pre-trained large language model; The prefix network layer parameters are initialized using reparameterization technology to generate the corresponding soft hint vector.
4. The method according to claim 1, wherein Before obtaining a supervised fine-tuning dataset corresponding to the attack tool code knowledge, the method further includes: Adding examples of extraction task context to pre-trained large language models based on instruction sets.
5. The method according to claim 1, wherein Obtaining a supervised fine-tuning dataset corresponding to the attack tool code includes: The code knowledge entities corresponding to the attack tool codes are annotated based on expert experience as a supervised fine-tuning dataset.
6. The method according to claim 4, characterized in that The structure-optimized generative information extraction model based on the instruction set, soft hint vector, supervised fine-tuning dataset, pre-trained large language model, and prefix network layer is constructed, including: Inputting the instruction set, soft hint vector, and supervised fine-tuning dataset into the pre-trained large language model through a prefix network layer to output a generation result, and optimizing the generation result using a contrastive loss decoding strategy; Calculate the maximum likelihood estimation loss between the optimized generation result and the extracted task context example, and update the parameters of the prefix network layer based on the maximum likelihood estimation loss to generate an optimized soft prompt vector sequence; The instruction set, extracted task context examples, optimized soft prompt vector sequence and supervised fine-tuning dataset are concatenated and re-input into the pre-trained large language model and prefix network layer as new input to obtain a structurally optimized generative information extraction model.
7. The method according to claim 6, characterized in that The instruction set, the extracted task context example, the optimized soft prompt vector sequence, and the supervised fine-tuning dataset are concatenated and re-input as new input to the pre-trained large language model and the prefix network layer to obtain a structurally optimized generative information extraction model, including: The instruction set, the extracted task context examples, the optimized soft prompt vector sequence, and the supervised fine-tuning dataset are concatenated and re-input into the pre-trained large language model and prefix network layer as new input to generate the output sequence probability distribution; The loss is calculated as the difference between the output sequence probability distribution and the supervised fine-tuning dataset. The prefix network layer parameters are updated inversely based on the loss, and the prefix network layer is fine-tuned based on the updated prefix network layer parameters to obtain a generative information extraction model with optimized structure.
8. The method according to claim 1, characterized in that Inputting the open source attack tool code into the structure-optimized generative information extraction model to identify attack tool code knowledge includes: The open source attack tool code is organized into a file-based format, and a model input format is constructed by combining the soft prompt vector, instruction set, and preset start and end markers in the structure-optimized generative information extraction model; Based on the model input format, soft prompt vectors, instruction sets, open source attack tool codes in file form, and preset start and end markers are input into the structure-optimized generative information extraction model to identify the attack tool code knowledge entities and corresponding attribute values.
9. An attack tool code knowledge extraction device, characterized in that: The device comprises: The code acquisition and instruction set construction module is used to obtain the open source attack tool code and construct the instruction set for the attack tool code knowledge extraction task based on the open source attack tool code; A soft hint vector generation module, configured to add a prefix network layer to the pre-trained large language model and generate a soft hint vector based on the prefix network layer; An information extraction model construction module is used to obtain a supervised fine-tuning dataset corresponding to the attack tool code and build a structured generative information extraction model based on the instruction set, soft hint vectors and supervised fine-tuning dataset, a pre-trained large language model, and a prefix network layer; The identification module is used to input the open source attack tool code into the structure-optimized generative information extraction model to identify the attack tool code knowledge.
10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the attack tool code knowledge extraction method according to any one of claims 1 to 8 by executing the computer instructions.