Large model data processing method and device, equipment and storage medium

By constructing a threat entity relationship model and fine-tuning the large language model using a knowledge-guided parameter matrix, the limitations of knowledge and computational overhead in the large language model in the field of cybersecurity are resolved, achieving efficient and accurate data processing.

CN120822524BActive Publication Date: 2025-11-25PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511325336.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-11-25
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Large language models suffer from limitations in knowledge, illusions, and insufficient dynamic adaptability in the field of cybersecurity. Existing technologies also suffer from error accumulation and increased computational overhead in long inference chain scenarios.

Method used

By constructing a threat entity relationship model, using an encoder for semantic encoding and semantic fusion layer enhancement, an enhanced semantic representation is generated. The initial large model is then fine-tuned using a knowledge-guided parameter matrix to form the target large model. This directly integrates cybersecurity knowledge into the model fine-tuning process, avoiding the computational resource consumption caused by additional modules.

Benefits of technology

It improves the data processing performance of large language models in the field of cybersecurity, ensures the accuracy and inference efficiency of complex cybersecurity data processing, and avoids error accumulation and computational delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822524B_ABST
    Figure CN120822524B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a large model data processing method, device and equipment and a storage medium, and relates to the technical field of network security. The method comprises the following steps: constructing a threat entity relationship model according to at least an asset knowledge base, a vulnerability knowledge base and an attack behavior knowledge base; inputting the threat entity relationship model into an encoder to perform semantic coding to obtain an encoding vector corresponding to each node; inputting the encoding vector into a semantic fusion layer to perform semantic enhancement to obtain an enhanced semantic representation corresponding to each node; generating a plurality of entity samples according to the enhanced semantic representation, and constructing a knowledge guide parameter matrix based on the entity samples; inputting the entity samples and the knowledge guide parameter matrix into an initial large model to perform feature representation to obtain a predicted feature representation; and fine-tuning the initial large model according to the predicted feature representation to obtain a target large model. Through the overall design, the accuracy of complex network security data processing is ensured, the reasoning efficiency is taken into account, and the data processing performance of the large language model in the field of network security is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of network security, and particularly relates to a large model data processing method and device, equipment and a storage medium. BACKGROUND

[0002] Large language models have shown excellent capabilities in the field of natural language processing, but in the specific field of network security, they still face problems such as knowledge limitations, illusion phenomena and insufficient dynamic adaptability. Therefore, the model needs to be enhanced to improve the professionalism, security and robustness of large models in the field of network security, and to realize the dynamic adaptability of the model.

[0003] In related technologies, one solution is to use a multi-hop retrieval enhancement framework to solve the problem of incomplete information in solving complex problems through iterative query rewriting. However, this approach has the problem of error accumulation in long reasoning chain scenarios. Another solution is to use a set of plug-and-play modules to enhance the performance of black-box large models, but this will introduce additional computational overhead, resulting in increased inference delay. SUMMARY

[0004] The main purpose of the embodiments of the present application is to propose a large model data processing method, device, equipment and storage medium to improve the data processing performance of large language models in the field of network security.

[0005] To achieve the above purpose, the first aspect of the embodiments of the present application proposes a large model data processing method, comprising:

[0006] constructing a threat entity relationship model according to at least an asset knowledge base, a vulnerability knowledge base and an attack behavior knowledge base obtained based on a network security database, wherein the threat entity relationship model comprises a plurality of nodes;

[0007] inputting the threat entity relationship model into an encoder for semantic coding to obtain an encoding vector corresponding to each node, wherein the encoding vector constitutes encoding embedding data;

[0008] inputting the encoding embedding data into a semantic fusion layer for semantic enhancement to obtain an enhanced semantic representation corresponding to each node;

[0009] generating a plurality of entity samples according to the enhanced semantic representation, constructing a knowledge guide parameter matrix based on the entity samples, inputting the entity samples and the knowledge guide parameter matrix into an initial large model for feature representation to obtain a predicted feature representation, fine-tuning the initial large model according to the predicted feature representation to obtain a target large model, and using the target large model to perform data processing on target data to obtain a target feature representation.

[0010] In some embodiments, the inputting the encoded embedding data into the semantic fusion layer for semantic enhancement to obtain an enhanced semantic representation corresponding to each of the nodes comprises:

[0011] Obtaining a tactic data matrix of the ATT&CK framework, splicing the encoded embedding data and the tactic data matrix to obtain spliced data;

[0012] Inputting the spliced data into the semantic fusion layer for semantic enhancement, and obtaining the enhanced semantic representation corresponding to each of the nodes by using a self-attention mechanism.

[0013] In some embodiments, the training process of the semantic fusion layer at least comprises:

[0014] Extracting a plurality of knowledge triples from the threat entity relationship model, the knowledge triples comprising a head entity node, a relationship, and a tail entity node;

[0015] According to the enhanced semantic representation corresponding to the head entity node and the tail entity node, and the embedding vector of the relationship, a positive sample score corresponding to the knowledge triple is calculated;

[0016] Randomly replacing the head entity node or the tail entity node to obtain a negative sample triple corresponding to the knowledge triple, and calculating a negative sample score corresponding to the negative sample triple;

[0017] According to at least the positive sample score and the negative sample score, a knowledge constraint loss value is obtained, and the semantic fusion layer is trained based on the knowledge constraint loss value to obtain a trained semantic fusion layer.

[0018] In some embodiments, the generating a plurality of entity samples according to the enhanced semantic representation, and constructing a knowledge guide parameter matrix based on the entity samples comprises:

[0019] Selecting the nodes from the threat entity relationship model as the entity samples, and taking the corresponding enhanced semantic representation as a sample feature corresponding to the entity sample;

[0020] If two of the entity samples are associated entities, setting a corresponding knowledge guide parameter as a preset weight value, otherwise the knowledge guide parameter is zero, and the knowledge guide parameter matrix is constructed according to the knowledge guide parameter.

[0021] In some embodiments, the inputting the entity samples and the knowledge guide parameter matrix into an initial large model for feature representation to obtain a predicted feature representation comprises:

[0022] When the corresponding sample features are input into the initial large model for feature extraction, a corresponding attention score matrix is obtained;

[0023] computing a sum of the knowledge guidance parameter matrix and the attention score matrix to obtain an updated score matrix, and obtaining the predicted feature representation based on the updated score matrix.

[0024] In some embodiments, the fine-tuning the initial large model according to the predicted feature representation to obtain a target large model further comprises:

[0025] selecting a plurality of anchor point samples from the threat entity relationship model, and obtaining at least one positive sample and at least one negative sample corresponding to the anchor point samples;

[0026] obtaining an anchor point feature representation corresponding to the anchor point samples, a positive feature representation corresponding to the positive samples respectively, and a negative feature representation corresponding to the negative samples respectively, and computing a contrastive learning loss value according to the anchor point feature representation, the positive feature representation and the negative feature representation;

[0027] computing a task loss value according to the predicted feature representation of the entity sample and corresponding label data;

[0028] computing a total loss value based on the task loss value and the contrastive learning loss value, and fine-tuning the initial large model using the total loss value to obtain the target large model.

[0029] In some embodiments, the inputting the threat entity relationship model into an encoder for semantic encoding to obtain an encoding vector corresponding to each node comprises:

[0030] generating a plurality of semantic paths according to the threat entity relationship model, and each node is located in at least one semantic path;

[0031] In the encoder, for each node, the semantic path in which the node is located is selected as a current path one by one, a neighbor node is determined in the current path, an initial node feature corresponding to the current path is obtained according to the node feature of the node and the neighbor node feature of the neighbor node, and all the initial node features are fused to obtain the corresponding encoding vector.

[0032] To achieve the above object, a second aspect of the embodiment of the present application proposes a large model data processing device, comprising:

[0033] a threat entity relationship model construction module configured to construct a threat entity relationship model according to at least an asset knowledge base, a vulnerability knowledge base and an attack behavior knowledge base obtained based on a network security database, wherein the threat entity relationship model comprises a plurality of nodes;

[0034] A semantic encoding module is configured to input the threat entity relationship model into an encoder to perform semantic encoding, so as to obtain an encoding vector corresponding to each node, and the encoding vector constitutes encoding embedding data.

[0035] A semantic enhancement module is configured to input the encoding embedding data into a semantic fusion layer to perform semantic enhancement, so as to obtain an enhanced semantic representation corresponding to each node.

[0036] A guidance fine-tuning module is configured to generate a plurality of entity samples according to the enhanced semantic representation, construct a knowledge guidance parameter matrix based on the entity samples, input the entity samples and the knowledge guidance parameter matrix into an initial large model to perform feature representation, obtain a predicted feature representation, fine-tune the initial large model according to the predicted feature representation to obtain a target large model, and use the target large model to perform data processing on target data to obtain a target feature representation.

[0037] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the method of the first aspect when executing the computer program.

[0038] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, which is a storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect.

[0039] The method, device, equipment and storage medium for processing data of a large model provided in the embodiments of the present application are as follows: a threat entity relationship model is constructed according to at least an asset knowledge base, a vulnerability knowledge base and an attack behavior knowledge base obtained based on a network security database, the threat entity relationship model includes a plurality of nodes; the threat entity relationship model is input into an encoder for semantic coding to obtain an encoding vector corresponding to each node, and the encoding vector constitutes coding embedding data; the coding embedding data is input into a semantic fusion layer for semantic enhancement to obtain an enhanced semantic representation corresponding to each node; a plurality of entity samples are generated according to the enhanced semantic representation, and a knowledge guide parameter matrix is constructed based on the entity samples, the entity samples and the knowledge guide parameter matrix are input into an initial large model for feature representation to obtain a predicted feature representation, and the initial large model is fine-tuned according to the predicted feature representation to obtain a target large model, and the target large model is used to process target data to obtain a target feature representation. The embodiments of the present application first construct a threat entity relationship model based on the three knowledge bases of assets, vulnerabilities and attack behaviors, convert the dispersed network security knowledge into a structured node correlation network, ensure the integrity and correlation of the initial knowledge, and directly establish a stable correlation between the knowledge to reduce the reasoning error from the source. Secondly, the semantic coding of the nodes by the encoder and the enhancement processing of the semantic fusion layer further strengthen the semantic correlation between the nodes, so that the enhanced semantic representation of each node contains the global knowledge context, and the error propagation caused by the lack of local information in the reasoning process is avoided, and the error accumulation risk of the long reasoning chain is eliminated from the knowledge modeling and semantic understanding level. In addition, the knowledge guide parameter matrix is introduced in the fine-tuning process of the large model to realize the lightweight integration of knowledge enhancement and model performance. It is not an independent module added outside the black box large model, but the network security knowledge is integrated into the fine-tuning process of the initial large model in the form of "entity sample + knowledge guide parameter matrix", and the structured knowledge is converted into a parameter constraint that can be directly learned by the model by using the knowledge guide parameter matrix, so that the target large model obtained by fine-tuning itself has the network security knowledge processing capability, without calling external modules in the reasoning stage, avoiding the calculation resource consumption and data interaction delay caused by introducing additional modules, and ensuring the efficient operation of the reasoning process. Through the overall design process, the accuracy of complex network security data processing can be ensured, the reasoning efficiency can be considered, and the data processing performance of the large language model in the network security field is improved. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 is a flowchart of the method for processing data of a large model provided in the embodiments of the present application.

[0041] Figure 2 is a schematic diagram of the asset knowledge base provided in the embodiments of the present application.

[0042] Figure 3 is a schematic diagram of the vulnerability knowledge base provided in the embodiments of the present application.

[0043] Figure 4 FIG. 1 is a schematic diagram of an attack behavior knowledge base provided by an embodiment of the present application.

[0044] Figure 5 FIG. 2 is a flowchart of inputting a threat entity relationship model into an encoder for semantic encoding to obtain an encoding vector corresponding to each node, provided by an embodiment of the present application.

[0045] Figure 6 FIG. 3 is a flowchart of inputting embedding data into a semantic fusion layer for semantic enhancement to obtain an enhanced semantic representation corresponding to each node, provided by an embodiment of the present application.

[0046] Figure 7 FIG. 4 is a flowchart of a training process of the semantic fusion layer, provided by an embodiment of the present application.

[0047] Figure 8 FIG. 5 is a flowchart of generating a plurality of entity samples according to the enhanced semantic representation, and constructing a knowledge guide parameter matrix based on the entity samples, provided by an embodiment of the present application.

[0048] Figure 9 FIG. 6 is a flowchart of inputting the entity samples and the knowledge guide parameter matrix into an initial large model for feature representation to obtain a predicted feature representation, provided by an embodiment of the present application.

[0049] Figure 10 FIG. 7 is a flowchart of fine-tuning the initial large model according to the predicted feature representation to obtain a target large model, provided by an embodiment of the present application.

[0050] Figure 11 FIG. 8 is a schematic diagram of the overall flow of large model data processing, provided by an embodiment of the present application.

[0051] Figure 12 FIG. 9 is a structure block diagram of a large model data processing apparatus, provided by another embodiment of the present application.

[0052] Figure 13 FIG. 10 is a schematic diagram of the hardware structure of an electronic device, provided by an embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0054] It should be noted that although the functional modules are divided in the apparatus schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a manner different from the module division in the apparatus or the order in the flowchart.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application.

[0056] First, the terms involved in the present application are analyzed:

[0057] Artificial Intelligence (AI): is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; Artificial intelligence is a branch of computer science, and artificial intelligence aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. The research in this field includes robots, language recognition, image recognition, natural language processing and expert systems. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0058] Large language models have shown excellent capabilities in the field of natural language processing, but in the specific field of network security, they still face problems such as knowledge limitations, illusion phenomena, and insufficient dynamic adaptability. Therefore, the model needs to be enhanced to improve the professionalism, security and robustness of large models in the field of network security, and to realize the dynamic adaptability of the model.

[0059] In related technologies, one solution is to use a multi-hop retrieval enhancement framework to solve the problem of incomplete information by iterative query rewriting. However, this approach has the problem of error accumulation in long reasoning chain scenarios. Another solution is to use a set of plug-and-play modules to enhance the performance of black-box large models, but this introduces additional computational overhead, leading to increased reasoning delay. Therefore, large models based on knowledge enhancement have the problem of insufficient coverage of domain knowledge in specific technical fields, and lack of deep cognitive ability in term-intensive scenarios in professional fields such as network security, law and finance. In addition, when the model undergoes complex queries or multiple edits, it will generally experience knowledge interference, and the model perplexity will significantly increase.

[0060] Based on this, the embodiment of the application provides a large model data processing method, device and equipment and a storage medium. First, based on a threat entity relationship model constructed by three knowledge bases of assets, vulnerabilities and attack behaviors, scattered network security knowledge is converted into a structured node correlation network, ensuring the integrity and correlation of initial knowledge and directly establishing stable correlation between knowledge, thereby reducing reasoning errors from the source. Secondly, the use of an encoder for semantic coding of nodes and enhanced processing of the semantic fusion layer further strengthens the semantic correlation between nodes, so that the enhanced semantic representation of each node contains the global knowledge context, avoiding error propagation caused by local information loss in the reasoning process, and eliminating the risk of error accumulation in long reasoning chains from the knowledge modeling and semantic understanding level. In addition, a knowledge guide parameter matrix is introduced in the large model fine-tuning process to realize the lightweight integration of knowledge enhancement and model performance. It is not an independent module added outside the black box large model, but the network security knowledge is integrated into the fine-tuning process of the initial large model in the form of "entity sample + knowledge guide parameter matrix", and the structured knowledge is converted into parameter constraints that can be directly learned by the model using the knowledge guide parameter matrix, so that the target large model obtained by fine-tuning itself has network security knowledge processing capability, without the need to call external modules in the reasoning stage, avoiding the calculation resource consumption and data interaction delay caused by the introduction of additional modules, and ensuring efficient operation in the reasoning process. Through the overall design process, the accuracy of complex network security data processing can be ensured, and the reasoning efficiency can be considered, thereby improving the data processing performance of the large language model in the network security field.

[0061] The embodiment of the application provides a large model data processing method, device, equipment and storage medium, which is specifically described through the following embodiments. First, the large model data processing method in the embodiment of the application is described.

[0062] The embodiment of the application can acquire and process related data based on artificial intelligence technology. Artificial intelligence (AI) is a theory, method, technology and application system for using a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0063] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software level technology. Artificial intelligence basic technology generally includes, such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning and other major directions.

[0064] The large model data processing method provided by the embodiments of the present application relates to the technical field of network security. The large model data processing method provided by the embodiments of the present application can be applied in a terminal, can also be applied in a server end, and can also be a computer program running in a terminal or a server end. For example, the computer program can be a native program or a software module in an operating system; it can be a native application (Application, APP), that is, a program that needs to be installed in an operating system to run, such as a client supporting large model data processing, that is, a program that can run only by being downloaded into a browser environment; it can also be a small program that can be embedded into any APP. In short, the above computer program can be any form of application program, module or plug-in. Among them, the terminal communicates with the server through a network. The large model data processing method can be executed by the terminal or the server, or cooperatively executed by the terminal and the server.

[0065] In some embodiments, the terminal can be a smartphone, a tablet computer, a notebook computer, a desktop computer, or a smart watch, etc. The server can be a standalone server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms; it can also be a service node in a blockchain system, the service nodes in the blockchain system form a peer-to-peer (Peer To Peer, P2P) network among them, and the P2P protocol is an application layer protocol running on the transmission control protocol (Transmission Control Protocol, TCP) protocol. The terminal and the server can be connected through Bluetooth, universal serial bus (Universal Serial Bus, USB) or network communication connection mode, and the embodiments do not limit this.

[0066] The application is operable with numerous general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that can be suitable for use with the application include personal computers, server computers, handheld or laptop devices, tablet devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and the like. The application can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like, that perform particular tasks or implement particular abstract data types. The application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media including memory storage devices.

[0067] The large model data processing method in the embodiment of the application is described below.

[0068] Figure 1 is an optional flowchart of the large model data processing method provided by the embodiment of the application, Figure 1 The method in can include, but is not limited to, steps 110 to 140. It can be understood that the embodiment does not make specific limitations on Figure 1 The order of steps 110 to 140 in is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs.

[0069] Step 110: constructing a threat entity relationship model according to at least an asset knowledge base, a vulnerability knowledge base and an attack behavior knowledge base obtained based on a network security database.

[0070] In an embodiment, network security data is obtained from general CVE vulnerability databases, STIX / TAXII open source threat intelligence, CWE databases, CAPEC databases, CPE databases, MITRE ATT&CK threat frameworks and other network security databases. The network security data at least includes asset, vulnerability, attack behavior and other information. Then, the asset knowledge base, the vulnerability knowledge base and the attack behavior knowledge base are constructed according to the network security data.

[0071] In an embodiment, when constructing the asset knowledge base, first, the operating system, application program and device are named according to a unified naming specification. Then, hostnames, IP addresses, installed software and versions and the like are obtained from the network security data, and are logically associated, such as running relationship, dependency relationship and the like, to construct a topology graph between assets and form a structured asset knowledge base. Referring to Figure 2 , Figure 2 is a schematic diagram of the asset knowledge base provided by the embodiment of the application.Figure 2 The assets in the asset knowledge base and the logical relationship between the assets are schematized, where the assets can be specific software or systems, such as Windows 7, Office 2016, JDK 1.8, etc. shown in the figure, and the assets are abstracted as an "asset" node. These asset nodes are associated with the IP addresses (such as 192.xxx.1.2) on the right side through the connection lines, and can clearly indicate "which software or system is installed on which IP address asset". Accordingly, risk assessment and impact analysis can be performed, for example, when a vulnerability is disclosed, the specific IP asset affected can be quickly located.

[0072] In an embodiment, when building the vulnerability knowledge base, the vulnerability number, description, reference link, CVSS severity score, affected platform CPE list, repair suggestion, and other information related to the vulnerability are obtained from the network security data. Among them, the CWE database classifies the underlying software defect types behind the vulnerability, such as "CWE-78: OS command injection". Then the data records related to the vulnerability are associated through the CVE number, specifically, the vulnerability (CVE) is accurately associated with the specific asset (CPE), and the affected version range is recorded, forming a structured vulnerability knowledge base. Refer to Figure 3 , Figure 3 is a schematic diagram of the vulnerability knowledge base provided by the embodiment of the present application. In the figure, a specific CVE vulnerability number (CNNVD-XXX-1846) is taken as an example, the vulnerability type is "executable command injection", which corresponds to the command injection defect in CWE, and then a plurality of affected hardware device models are listed around the vulnerability, such as NETGATE D7800, and the specific version range affected by each device, such as "≤1.0.1.34". Through the vulnerability knowledge base, the abstract vulnerability can be accurately matched with the specific asset model and version affected.

[0073] In an embodiment, when building the attack behavior knowledge base, the standardized knowledge system of tactics (such as "execution") and techniques (such as "T1059: command and script interpreter") used by global hackers and attack organizations provided by the ATT&CK threat framework can be used, and the attack pattern related information can be obtained from the CAPEC database. In addition, specific intrusion indicators (IOCs) and attack activity reports can also be obtained from the STIX / TAXII open source threat intelligence. Then, the attack behavior in the security device alarm (such as Snort, Topsec) and the log is mapped to the technique ID in the ATT&CK framework using the obtained data. At the same time, the observed attack source IP, destination IP, and CVE vulnerability used are associated with the attack technique, so as to restore the complete attack chain and obtain the attack behavior knowledge base. Refer toFigure 4 , Figure 4 is a schematic diagram of an attack behavior knowledge base provided by an embodiment of the present application. Figure 4 Take a complete attack activity as an example for illustration. The attack target is "Web server A", "CVE-XXXX-2309" is the vulnerability exploited by the attacker, and the attack traffic originates from the external network Extranet-xx1 and the destination address is the server of the internal network Intranet-YY. At the same time, the security devices "Snort-1xxx" and "topsec-7xxx" detect the attack and generate an alarm. Therefore, through the attack behavior knowledge base, the isolated elements such as attack source, destination, exploited vulnerability, triggered alarm, etc. can be connected into a complete attack chain, and the context of attack behavior and the kill chain model can be intuitively reflected.

[0074] In an embodiment, the main entities in the asset knowledge base include asset nodes and asset relationships. The asset nodes include servers, network devices, terminal devices, etc., and contain information such as IP addresses, hostnames, operating systems, software versions, etc. The asset relationships can be "runs on", "belongs to", "connected to", etc., and are used to describe the dependency relationship between assets. The main entities of the vulnerability knowledge base include vulnerability nodes and affected assets. The vulnerability nodes include CVE numbers, CNNVD numbers, etc., and contain vulnerability descriptions, CVSS scores, repair suggestions, etc. The affected assets are used to describe the assets affected by the vulnerability and the version range. The main entities of the attack behavior knowledge base include attack nodes and attack relationships. The attack nodes include attack event IDs, attack types, etc. The attack relationships include "exploit", "trigger", "target", etc., and are used to describe the association between attack behavior and assets, vulnerabilities. Then, the three are integrated, the entity association across knowledge bases is performed, and a threat entity relationship model is obtained.

[0075] It can be understood that the above entities can be nodes, and when multiple knowledge bases are integrated, some entities (such as assets and vulnerabilities) can exist in multiple knowledge bases at the same time, so these nodes need to be reused to realize cross-knowledge base association. For example, a certain asset node exists in the asset knowledge base and the attack behavior knowledge base, and a certain vulnerability node exists in the vulnerability knowledge base and the attack behavior knowledge base. Therefore, when the nodes are reused, the association between assets and vulnerabilities can be established through the Affected-Asset relationship, for example, CVE-2023-1234 affects Web-Server-01. The association between assets and attack behaviors can also be established, and the attack behavior is associated with the asset through the targets relationship. For example, Attack-2023-001 targets Web-Server-01. In addition, the attack behavior can also be associated with the vulnerability through the exploits relationship, for example, Attack-2023-001 exploits CVE-2023-1234.

[0076] Specifically, the number of nodes corresponding to different knowledge bases can be inconsistent, and some nodes can be associated according to actual conditions to finally obtain a global threat entity relationship model, which includes multiple nodes, and the corresponding node types can include assets (Asset), vulnerabilities (Vulnerability), attacks (Attack), etc., and the relationship types can include runs_on (asset dependency relationship), affects (vulnerability affects asset), exploits (attack exploits vulnerability), targets (attack targets asset), triggers (attack triggers alarm), etc. Among them, the threat entity relationship model can at least dynamically associate an attack chain and perform attack path analysis, for example, given a vulnerability (such as CVE-2023-1234), the possible affected assets (Web-Server-01) and the potential attack path (Attack-2023-001) can be automatically derived.

[0077] Step 120: inputting the threat entity relationship model into an encoder for semantic coding to obtain an encoding vector corresponding to each node.

[0078] In an embodiment, after obtaining the threat entity relationship model, first, each node is coded to obtain corresponding node features, and the node features corresponding to nodes of different node types are mapped and aligned. Then, the threat entity relationship model is semantically coded. Referring to Figure 5 , Figure 5 is a flowchart provided by an embodiment of the present application for inputting the threat entity relationship model into an encoder for semantic coding to obtain an encoding vector corresponding to each node, and specifically includes the following steps:

[0079] Step 510: generating a plurality of semantic paths according to the threat entity relationship model.

[0080] In an embodiment, a plurality of semantic paths can be generated based on vulnerability influence, for example: Vulnerability->affects->Software->runs_on->IP, indicating which software is affected by a vulnerability, and which IP these software run on. A plurality of semantic paths can also be generated based on attack exploitation, for example: IP (src)->launches->Attack->exploits->Vulnerability->affects->IP (dst), indicating that an attack launched by a source IP exploits a vulnerability, and this vulnerability eventually affects a target IP. In addition, a plurality of semantic paths can be generated based on alert association, for example: Alert->triggered_by->Attack->targets->IP, indicating that an alert is triggered by which attack, and what asset is the target of this attack. As can be seen, semantic paths are used to define the message passing relationship between different node types, which contains rich semantic information and is no longer divided by simple neighbor relationship in a single knowledge base. Each node is located in at least one semantic path.

[0081] In an embodiment, the threat entity relationship model can be input into a pre-trained large language model, and the above-mentioned generation method of semantic paths can also be input into the large language model in the form of prompt words to guide the large language model to extract a plurality of semantic paths from the threat entity relationship model. Semantic path extraction can also be performed through regular expressions, preset rules, graph neural networks, etc., which are not limited in this embodiment.

[0082] Step 520: in the encoder, for each node, select the semantic path in which the node is located as the current path, determine the neighbor nodes in the current path, obtain the initial node features corresponding to the current path according to the node features of the node and the neighbor node features of the neighbor nodes, and fuse all the initial node features to obtain the corresponding encoding vector.

[0083] In an embodiment, the encoder can be constituted by a graph neural network (GNN) and be pre-trained. For each node, it can be in multiple semantic paths, and thus is processed individually for each semantic path. At this time, the corresponding semantic path is selected as the current path one by one, the encoder is used to determine the corresponding other nodes on the current path as neighbor nodes, and the node features of the nodes and the neighbor node features of the neighbor nodes are aggregated to obtain the initial node features corresponding to the current path. It can be understood that the same node will obtain different initial features in different semantic paths, because it fuses neighbor information from different perspectives. Further, after the node obtains the corresponding initial features in different semantic paths, the initial features are averaged and fused to obtain the encoding vector corresponding to the node, so as to represent the complex semantic information.

[0084] Taking a node Vulnerability in a semantic path Vulnerability->affects->Software->runs_on->IP as an example, the node v=CVE-2010-2309 is found according to the current path, all Software nodes affected by v are found, and then all IP nodes on which the Software nodes run are found. All the nodes passed are taken as neighbor nodes of the node v in the current path. Next, the node features of the node v are obtained, the neighbor node features of each neighbor node are aggregated by averaging, weighting and the like, and the feature vector obtained by aggregation is added to the node features, so that the initial features corresponding to the node in the current path are obtained. The initial features encode the aggregation information of "all IP addresses that can be affected by the vulnerability CVE-2010-2309".

[0085] For each node in the threat entity relationship model, the corresponding encoding vector is obtained in the above manner, and the encoding vectors constitute the encoding embedding data .

[0086] Step 130: input the encoding embedding data into a semantic fusion layer for semantic enhancement to obtain an enhanced semantic representation corresponding to each node.

[0087] In an embodiment, referring to Figure 6 , Figure 6 is a flowchart provided by an embodiment of the present application for inputting the encoding embedding data into the semantic fusion layer for semantic enhancement to obtain an enhanced semantic representation corresponding to each node, and specifically includes the following steps:

[0088] Step 610: Obtain the tactic data matrix of the ATT&CK framework, splice the encoding embedding data and the tactic data matrix to obtain spliced data.

[0089] In an embodiment, the tactic data matrix E is obtained from the ATT&CK framework ATT , which is used to indicate a tactic-attack technique relationship table. The rows in the tactic data matrix can represent different attack techniques, such as T1059 (command and script interpreter), T1190 (exploitation of public-facing applications), etc., and the columns can represent different tactics, such as TA0001 (initial access), TA0002 (execution), etc., and the matrix value E ATT [i,j] can represent whether the ith technique belongs to the jth tactic. If the technique T1059 belongs to the tactic TA0002, then E_ATT["T1059", "TA0002"] = 1, otherwise 0. As can be seen, this tactic data matrix is a binary matrix or a correlation strength matrix. Next, the tactic data matrix and the encoding embedding data are spliced to obtain spliced data .

[0090] Step 620: Input the spliced data into the semantic fusion layer for semantic enhancement, and obtain the enhanced semantic representation corresponding to each node by using a self-attention mechanism.

[0091] In an embodiment, the semantic fusion layer is a Transformer architecture, which uses multi-head self-attention to learn the semantics of the input data for semantic enhancement. After processing by the semantic fusion layer, the enhanced semantic representation corresponding to each entity node in the threat entity relationship model is obtained, represented as:

[0092]

[0093] wherein Z is the enhanced semantic representation, , and W represents the model weight of the semantic fusion layer. After the enhancement processing by the semantic fusion layer, the enhanced semantic representation of each entity node not only contains the context corresponding to the encoding embedding data, but also fuses the semantic information of the corresponding tactic data, and the information amount is much larger than the encoding effect of using only the encoder.

[0094] In an embodiment, referring to Figure 7 , Figure 7 the training process flowchart of the semantic fusion layer provided by the embodiments of the present application, at least includes the following steps:

[0095] Step 710: Extract a plurality of knowledge triples from the threat entity relationship model.

[0096] In an embodiment, the knowledge triple (h, r, t) is two entity nodes and the relationship between the entity nodes extracted from the threat entity relationship model, including the head entity node h, the relationship r and the tail entity node t, for example: (CVE-2020-1234, exploits, Technique-T1059), indicating that a certain vulnerability exploits a certain technique, or (Alert-567, triggered_by, Attack-Attack1), indicating that a certain alert is triggered by a certain attack.

[0097] Step 720: calculating the positive sample score corresponding to the knowledge triple according to the enhanced semantic representation corresponding to the head entity node and the tail entity node respectively and the embedding vector of the relationship.

[0098] In an embodiment, during the training process, the enhanced semantic representation Zh and Zt corresponding to the head entity node and the tail entity node output by the semantic fusion layer respectively and the embedding of the relationship are calculated to obtain the corresponding embedding vector wr, and the positive sample score corresponding to the knowledge triple is calculated according to this: f(h, r, t) = ||(Zh+wr)-Zt||, the meaning of this calculation is that the enhanced semantic representation of the head entity node plus the embedding vector of the relationship should be approximately equal to the enhanced semantic representation of the tail entity node.

[0099] Step 730: randomly replacing the head entity node or the tail entity node to obtain the negative sample triple corresponding to the knowledge triple, and calculating the negative sample score corresponding to the negative sample triple.

[0100] In an embodiment, other entity nodes are randomly selected from the threat entity relationship model to replace the head entity node and / or the tail entity node to construct the negative sample triple corresponding to the knowledge triple, and the negative sample score corresponding to the negative sample triple is calculated in the same way as described above.

[0101] Step 740: obtaining the knowledge constraint loss value according to at least the positive sample score and the negative sample score, training the semantic fusion layer based on the knowledge constraint loss value to obtain the trained semantic fusion layer.

[0102] In an embodiment, the knowledge constraint loss value obtained according to at least the positive sample score and the negative sample score is represented as:

[0103]

[0104] wherein, the negative sample score is represented as f(h, r, t) = ||(Zh+wr)-Zt||, the positive sample score is represented as f(h, r, t) = ||(Zh+wr)-Zt||, the interval hyperparameter is represented as the threat entity relationship model is represented as

[0105] It can be seen that the constraint target of the knowledge constraint loss value is to hope that the positive sample score is at least one interval parameter higher than the negative sample score, and if this interval is not reached, a loss will be generated. The goal of training is to minimize the sum of this loss on all triplets. After the adjustment of the knowledge constraint loss value, the entity nodes with real relationships in the space will be close, and the entity nodes without relationships will be pushed away, so that the semantic fusion layer obtained by training can learn the information-rich feature representation.

[0106] Step 140: generating a plurality of entity samples according to the enhanced semantic representation, constructing a knowledge guide parameter matrix based on the entity samples, inputting the entity samples and the knowledge guide parameter matrix into the initial large model for feature representation, obtaining a predicted feature representation, and fine-tuning the initial large model according to the predicted feature representation to obtain a target large model.

[0107] In an embodiment, referring to Figure 8 , Figure 8 is a flowchart provided by the embodiment of the present application for generating a plurality of entity samples according to the enhanced semantic representation, and constructing a knowledge guide parameter matrix based on the entity samples, specifically comprising the following steps:

[0108] Step 810: selecting a node from the threat entity relationship model as an entity sample, and taking the corresponding enhanced semantic representation as the sample feature corresponding to the entity sample.

[0109] Step 820: if the two entity samples are associated entities, set the corresponding knowledge guide parameter to a preset weight value, otherwise the knowledge guide parameter is zero, and construct a knowledge guide parameter matrix according to the knowledge guide parameter.

[0110] In an embodiment, a node is selected from the threat entity relationship model as an entity sample, and then the enhanced semantic representation corresponding to the entity sample is taken as a sample feature. Next, a knowledge guide parameter matrix is constructed, and whether the entity samples are associated entities is determined. If they are associated entities, the corresponding knowledge guide parameter is set to a preset weight value, otherwise the knowledge guide parameter is zero. The associated entity judgment here can be determined by using a graph neural network in the threat entity relationship model, or by using a graph topology analysis method, and the embodiment of the present application does not limit this.

[0111] wherein the knowledge guide parameter matrix is represented as:

[0112]

[0113] wherein, It can be set according to the actual situation, for example, according to the hop number between two entity nodes to represent the degree of association between them.

[0114] Next, refer to Figure 9 , Figure 9 is the flowchart provided by the embodiment of the application for inputting entity samples and knowledge guide parameter matrix into the initial large model to obtain the predicted feature representation, which specifically includes the following steps:

[0115] Step 910: When the corresponding sample features are input into the initial large model for feature extraction, the corresponding attention score matrix is obtained.

[0116] In an embodiment, the initial large model is a large language model trained by using general data , the target domain data is constituted by the enhanced semantic representation in the network security field obtained in the foregoing , the target domain data is used for transfer learning, so that the initial large model is fine-tuned to obtain a target large model, and the target large model is used for data processing of the target data to obtain a target feature representation, and the target large model is represented as , represents the fine-tuning parameter.

[0117] In an embodiment, a plurality of training batches are set, and in each training batch, a plurality of sample features are input into the initial large model for feature extraction, and the initial large model uses the Transformer mechanism to project the input data through three different linear transformation layers to a query (Query), key (Key), and value (Value) space to obtain a corresponding attention score matrix.

[0118] Specifically, assuming that the plurality of input sample features constitute input data X, the projected query matrix Q, value matrix V, and key matrix K are respectively:

[0119]

[0120]

[0121]

[0122] wherein, , , are trainable weight matrices, and the query matrix Q, the value matrix V, and the key matrix K each contain a query component qi, a value component vi, and a key component ki corresponding to each sample feature.

[0123] Next, the product of the transpose of the query matrix Q and the key matrix K is calculated, that is, the attention score matrix S is obtained, which is represented as: , and the element It can be understood that the element Sij represents a similarity score between the ith sample feature and the jth sample feature.

[0124] Step 920: Calculate the sum of the knowledge guidance parameter matrix and the attention score matrix to obtain an updated score matrix, and obtain a predicted feature representation based on the updated score matrix.

[0125] In an embodiment, assuming there are N sample features, the dimensions of the knowledge guidance parameter matrix and the attention score matrix are both Therefore, the sum of the knowledge guidance parameter matrix and the attention score matrix is calculated, that is, the corresponding knowledge guidance parameter is associated with the similarity score to obtain the updated score matrix S'. When performing the attention mechanism, instead of simply calculating the similarity between the query vector and the key vector, the knowledge guidance parameter is used as a bias term to force the initial large model to focus on logically related entity nodes in the network security field, thereby making the learning process of the initial large model conform to the business logic of the vertical field. That is, even if Sij is small, that is, from a data perspective, sample feature i and sample feature j are not similar, but as long as Bij is large, that is, in the network security field, the two entity features are strongly associated, the final similarity score will also be large. Conversely, even if Sij is large, but if Bij = 0, indicating that the two are unrelated in the network security technology field, the final similarity score will not be excessively magnified.

[0126] In an embodiment, after obtaining the updated score matrix S', a softmax transformation is performed, and then multiplied by the value vector to obtain the predicted feature representation.

[0127] In an embodiment, after obtaining the predicted feature representation, the initial large model can be fine-tuned accordingly. During the fine-tuning process, a contrastive learning loss is also introduced based on the task loss to better improve the performance of the target large model. Referring to Figure 10 , Figure 10 The flowchart provided by the embodiments of the present application for fine-tuning the initial large model to obtain the target large model according to the predicted feature representation, specifically includes the following steps:

[0128] Step 1010: Select a plurality of anchor samples from the threat entity relationship model, and obtain at least one positive sample and at least one negative sample corresponding to the anchor samples.

[0129] In an embodiment, the entity nodes corresponding to a plurality of threat data are selected from the threat entity relationship model as anchor samples, such as attack behaviors. Then, other entity nodes located in the same attack event as the anchor samples are obtained as at least one positive sample corresponding to the anchor samples, and entity nodes located in different attack events from the anchor samples are obtained as at least one negative sample corresponding to the anchor samples. It can be understood that the anchor samples can be a subset of the entity samples.

[0130] Step 1020: Obtain the anchor feature representation corresponding to the anchor sample, the positive feature representation corresponding to the positive sample, and the negative feature representation corresponding to the negative sample. Calculate the contrastive learning loss value based on the anchor feature representation, the positive feature representation, and the negative feature representation.

[0131] In one embodiment, the anchor feature representation corresponding to the anchor sample is obtained from the enhanced semantic representation. Positive feature representations corresponding to positive samples and the negative feature representations corresponding to the negative samples respectively Next, the contrastive learning loss value is calculated based on the anchor point feature representation, positive feature representation, and negative feature representation. Wherein, the contrastive learning loss value corresponding to the i-th anchor point sample... Represented as:

[0132]

[0133] in, Describes the set of positive samples. Let represent the set of negative samples, and s() represent the similarity calculation function. This represents the hyperparameters. It's understandable that, for anchor samples in the same training batch, all the contrastive learning loss values ​​are summed to obtain the total contrastive learning loss value.

[0134] Step 1030: Calculate the task loss value based on the predicted feature representation of the entity sample and the corresponding label data.

[0135] In one embodiment, for entity samples in the same training batch, the predicted feature representation obtained from the initial large model prediction needs to be compared with its corresponding label data to calculate a task loss value. Here, the label data is prior information and can be set according to actual conditions, such as the alarm classification of a certain entity node. The corresponding task loss value can be the cross-entropy loss value, used to represent the difference between the predicted feature representation and the label data; its calculation process is not limited here.

[0136] Step 1040: Calculate the total loss value based on the task loss value and the contrastive learning loss value, and use the total loss value to fine-tune the initial large model to obtain the target large model.

[0137] In one embodiment, based on task loss value and contrastive learning loss value Calculated total loss value Represented as: ,in, is a weight hyperparameter. The initial large model is fine-tuned in the vertical field by using the total loss value, and the target large model is obtained. Through this collaborative mechanism, the large language model not only learns to complete specific network security tasks (such as alarm classification), but its internal representation is also standardized and optimized by specific network security field knowledge, thereby becoming an expert model with both general language ability and deep security field knowledge, and finally improving the discriminant ability of the target large model in the network security task.

[0138] In an embodiment, referring to Figure 11 , Figure 11 is the overall flowchart of the large model data processing provided by the embodiment of the application. First, a threat entity relationship model is constructed based on a network security database, and then semantic fusion and knowledge constraint are performed through a graph calculation process (encoder + semantic fusion layer) to obtain an enhanced semantic representation corresponding to each entity node in the threat entity relationship model, which is used in the subsequent vertical field training process of the initial large model to solve the problem of insufficient network security field knowledge coverage. In the fine-tuning process of the initial large model, the enhanced semantic representation corresponding to each entity node in the threat entity relationship model is introduced as target field data through transfer learning, and the attention weight distribution of the initial large model for the network security knowledge field is improved through the knowledge-guided parameter matrix to solve the problems of knowledge interference and increased model perplexity, and to improve the discriminant accuracy of the target large model for threat features. Moreover, in the training process, the same type of threat samples and different types of threat samples are separated based on contrastive learning to strengthen the sensitivity of the target large model to subtle threat differences, solve the problem of lack of deep understanding of network security tasks by general large models, and significantly improve the discriminant ability of the target large model in the network security task.

[0139] The technical scheme provided by the embodiments of the present application at least constructs a threat entity relationship model according to an asset knowledge base, a vulnerability knowledge base and an attack behavior knowledge base obtained based on a network security database, the threat entity relationship model includes a plurality of nodes; inputs the threat entity relationship model into an encoder for semantic coding to obtain a coding vector corresponding to each node, and the coding vector constitutes coding embedding data; inputs the coding embedding data into a semantic fusion layer for semantic enhancement to obtain an enhanced semantic representation corresponding to each node; generates a plurality of entity samples according to the enhanced semantic representation, and constructs a knowledge guide parameter matrix based on the entity samples, inputs the entity samples and the knowledge guide parameter matrix into an initial large model for feature representation to obtain a predicted feature representation, and fine-tunes the initial large model according to the predicted feature representation to obtain a target large model, and the target large model is used for data processing on target data to obtain a target feature representation. The embodiments of the present application first construct a threat entity relationship model based on the three knowledge bases of assets, vulnerabilities and attack behaviors, convert the dispersed network security knowledge into a structured node association network, ensure the integrity and correlation of the initial knowledge, directly establish the stable association between the knowledge, and reduce the reasoning error from the source. Secondly, the semantic coding of the nodes by the encoder and the enhancement processing of the semantic fusion layer further strengthen the semantic association between the nodes, so that the enhanced semantic representation of each node contains the global knowledge context, avoids the error propagation caused by the lack of local information in the reasoning process, and eliminates the error accumulation risk of the long reasoning chain from the knowledge modeling and semantic understanding level. In addition, the knowledge guide parameter matrix is introduced in the fine-tuning process of the large model, realizing the lightweight integration of knowledge enhancement and model performance. It is not an independent module added outside the black box large model, but the network security knowledge is integrated into the fine-tuning process of the initial large model in the form of “entity sample + knowledge guide parameter matrix”, and the structured knowledge is converted into a parameter constraint that can be directly learned by the model by using the knowledge guide parameter matrix, so that the target large model obtained by fine-tuning itself has the network security knowledge processing capability, without calling external modules in the reasoning stage, avoiding the calculation resource consumption and data interaction delay caused by introducing additional modules, and ensuring the efficient operation of the reasoning process. Through the overall design process, the accuracy of complex network security data processing can be guaranteed, the reasoning efficiency can be considered, and the data processing performance of the large language model in the network security field is improved.

[0140] The embodiments of the present application also provide a large model data processing device, which can implement the large model data processing method. Figure 12 The device comprises:

[0141] The threat entity relationship model construction module 1210 is configured to construct a threat entity relationship model according to an asset knowledge base, a vulnerability knowledge base and an attack behavior knowledge base obtained based on a network security database, and the threat entity relationship model includes a plurality of nodes.

[0142] The semantic encoding module 1220 is configured to perform semantic encoding on the threat entity relationship model input encoder to obtain an encoding vector corresponding to each node, and the encoding vectors constitute the encoding embedding data.

[0143] The semantic enhancement module 1230 is configured to input the encoding embedding data into a semantic fusion layer to perform semantic enhancement to obtain an enhanced semantic representation corresponding to each node.

[0144] The guidance fine-tuning module 1240 is configured to generate a plurality of entity samples according to the enhanced semantic representation, construct a knowledge guidance parameter matrix based on the entity samples, input the entity samples and the knowledge guidance parameter matrix into an initial large model to perform feature representation to obtain a predicted feature representation, fine-tune the initial large model according to the predicted feature representation to obtain a target large model, and use the target large model to perform data processing on target data to obtain a target feature representation.

[0145] The specific implementation of the large model data processing apparatus in this embodiment is basically the same as the specific implementation of the large model data processing method described above, and will not be repeated here.

[0146] The embodiments of the present application also provide an electronic device, which comprises:

[0147] at least one memory;

[0148] at least one processor;

[0149] at least one program;

[0150] The program is stored in the memory, and the processor executes the at least one program to implement the large model data processing method provided in the embodiments of the present application. The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc.

[0151] Please refer to Figure 13 , Figure 13 The hardware structure of the electronic device of another embodiment is illustrated, and the electronic device comprises:

[0152] The processor 1301 can be implemented in the form of a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute related programs to implement the technical solutions provided in the embodiments of the present application.

[0153] The memory 1302 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1302 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 1302 and are called and executed by the processor 1301 to implement the large model data processing method of the embodiments of the present application.

[0154] The input / output interface 1303 is configured to realize information input and output.

[0155] The communication interface 1304 is configured to realize the communication interaction between the device and other devices. The communication can be realized by a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).

[0156] The bus 1305 is configured to transmit information between various components (for example, the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304) of the device.

[0157] The processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304 are connected to each other through the bus 1305 to realize the communication connection between the device.

[0158] The embodiments of the present application also provide a storage medium. The storage medium is a storage medium, and the storage medium stores a computer program. The computer program is executed by the processor to implement the large model data processing method.

[0159] The memory is a non-transitory storage medium, which can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor. These remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0160] The large model data processing method, device, equipment and storage medium provided by the embodiments of the present application are as follows: a threat entity relationship model is constructed according to at least an asset knowledge base, a vulnerability knowledge base and an attack behavior knowledge base obtained based on a network security database, the threat entity relationship model including a plurality of nodes; the threat entity relationship model is input into an encoder for semantic coding to obtain an encoding vector corresponding to each node, the encoding vector constituting encoding embedding data; the encoding embedding data is input into a semantic fusion layer for semantic enhancement to obtain an enhanced semantic representation corresponding to each node; a plurality of entity samples are generated according to the enhanced semantic representation, and a knowledge guide parameter matrix is constructed based on the entity samples, the entity samples and the knowledge guide parameter matrix are input into an initial large model for feature representation to obtain a predicted feature representation, and the initial large model is fine-tuned according to the predicted feature representation to obtain a target large model, the target large model being used for data processing of target data to obtain a target feature representation. The embodiments of the present application first construct a threat entity relationship model based on the three knowledge bases of assets, vulnerabilities and attack behaviors, convert the dispersed network security knowledge into a structured node association network, ensure the integrity and relevance of the initial knowledge, and directly establish a stable association between the knowledge to reduce the reasoning error from the source. Secondly, the semantic coding of the nodes by the encoder and the enhancement processing of the semantic fusion layer further strengthen the semantic association between the nodes, so that the enhanced semantic representation of each node contains the global knowledge context, avoiding the error propagation caused by the lack of local information in the reasoning process, and eliminating the error accumulation risk of the long reasoning chain from the knowledge modeling and semantic understanding level. In addition, the knowledge guide parameter matrix is introduced in the fine-tuning process of the large model to realize the lightweight integration of knowledge enhancement and model performance. It is not an independent module added outside the black box large model, but the network security knowledge is integrated into the fine-tuning process of the initial large model in the form of "entity sample + knowledge guide parameter matrix", and the structured knowledge is converted into a parameter constraint that can be directly learned by the model by using the knowledge guide parameter matrix, so that the target large model obtained by fine-tuning itself has the network security knowledge processing capability, without calling external modules in the reasoning stage, avoiding the calculation resource consumption and data interaction delay caused by introducing additional modules, and ensuring the efficient operation of the reasoning process. Through the overall design process, the accuracy of complex network security data processing can be guaranteed, and the reasoning efficiency can be considered, and the data processing performance of the large language model in the network security field is improved.

[0161] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0162] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and can include more or fewer steps than the figures, or combine certain steps, or different steps.

[0163] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, that is, can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.

[0164] Those skilled in the art can understand that all or some steps in the above disclosed method, functions of the modules / units in the system and the device can be implemented as software, firmware, hardware and appropriate combinations thereof.

[0165] The terms "first", "second", "third", "fourth" and the like in the description of the present application and the above-described figures (if any) are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0166] It should be understood that in the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the association between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0167] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiments described above are merely illustrative, for example, the division of the above units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.

[0168] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place or distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0169] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0170] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of each embodiment of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), magnetic disk or optical disk, and various program storage media.

[0171] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, but this does not limit the scope of the embodiments of the present application. Any modifications, equivalent replacements and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.

Claims

1. A method for processing large model data, characterized in that, include: A threat entity relationship model is constructed based at least on an asset knowledge base, a vulnerability knowledge base, and an attack behavior knowledge base obtained from a cybersecurity database. The threat entity relationship model includes multiple nodes. Multiple semantic paths are generated based on the threat entity relationship model, and each node is located in at least one semantic path. In the encoder, for each node, the semantic path where the node is located is selected one by one as the current path, neighboring nodes are determined in the current path, and the initial node features corresponding to the current path are obtained based on the node features of the node and the neighboring node features of the neighboring nodes. All the initial node features are fused to obtain the corresponding encoding vector, and the encoding vector constitutes the encoded embedding data. The encoded embedded data is input into the semantic fusion layer for semantic enhancement to obtain the enhanced semantic representation corresponding to each node; Nodes are selected from the threat entity relationship model as entity samples, and the corresponding enhanced semantic representations are used as sample features corresponding to the entity samples. If any two entity samples are related entities, the corresponding knowledge guidance parameter is set to a preset weight value; otherwise, the knowledge guidance parameter is zero. A knowledge guidance parameter matrix is ​​constructed based on the knowledge guidance parameter. When the corresponding sample features are input into the initial large model for feature extraction, the corresponding attention score matrix is ​​obtained. The sum of the knowledge guidance parameter matrix and the attention score matrix is ​​calculated to obtain an updated score matrix. A predicted feature representation is obtained based on the updated score matrix. The initial large model is fine-tuned based on the predicted feature representation to obtain a target large model. The target large model is used to process target data to obtain a target feature representation.

2. The large model data processing method according to claim 1, characterized in that, The step of inputting the encoded embedded data into the semantic fusion layer for semantic enhancement to obtain the enhanced semantic representation corresponding to each node includes: Obtain the tactical data matrix of the ATT&CK framework, and concatenate the encoded embedded data and the tactical data matrix to obtain concatenated data; The concatenated data is input into the semantic fusion layer for semantic enhancement, and the enhanced semantic representation corresponding to each node is obtained by using a self-attention mechanism.

3. The large model data processing method according to claim 2, characterized in that, The training process of the semantic fusion layer includes at least the following: Multiple knowledge triples are extracted from the threat entity relationship model, and the knowledge triples include a head entity node, a relationship, and a tail entity node; The positive sample score corresponding to the knowledge triple is calculated based on the enhanced semantic representation and the embedding vector of the relation corresponding to the head entity node and the tail entity node, respectively. Randomly replace the head entity node or the tail entity node to obtain the negative sample triplet corresponding to the knowledge triplet, and calculate the negative sample score corresponding to the negative sample triplet. At least the positive sample score and the negative sample score are used to obtain the knowledge constraint loss value. The semantic fusion layer is then trained based on the knowledge constraint loss value to obtain the trained semantic fusion layer.

4. The large model data processing method according to claim 1, characterized in that, The step of fine-tuning the initial large model based on the predicted feature representation to obtain the target large model further includes: Multiple anchor point samples are selected from the threat entity relationship model, and at least one positive sample and at least one negative sample corresponding to the anchor point sample are obtained; Obtain the anchor feature representation corresponding to the anchor sample, the positive feature representation corresponding to the positive sample, and the negative feature representation corresponding to the negative sample. Calculate the contrastive learning loss value based on the anchor feature representation, the positive feature representation, and the negative feature representation. The task loss value is calculated based on the predicted feature representation of the entity sample and the corresponding label data; The total loss value is calculated based on the task loss value and the contrastive learning loss value, and the target large model is obtained by fine-tuning the initial large model using the total loss value.

5. A large model data processing device, characterized in that, include: Threat Entity Relationship Model Construction Module: Used to construct a threat entity relationship model based at least on an asset knowledge base, vulnerability knowledge base, and attack behavior knowledge base obtained from a cybersecurity database. The threat entity relationship model includes multiple nodes. Semantic encoding module: used to generate multiple semantic paths according to the threat entity relationship model, each node is located in at least one semantic path; in the encoder, for each node, the semantic path where the node is located is selected one by one as the current path, neighboring nodes are determined in the current path, the initial node features corresponding to the current path are obtained according to the node features of the node and the neighboring node features of the neighboring nodes, and all the initial node features are fused to obtain the corresponding encoding vector, the encoding vector constitutes the encoded embedding data; Semantic enhancement module: used to input the encoded embedded data into the semantic fusion layer for semantic enhancement, so as to obtain the enhanced semantic representation corresponding to each node; The guided fine-tuning module is used to select nodes as entity samples from the threat entity relationship model, and use the corresponding enhanced semantic representation as the sample features of the entity samples. If any two entity samples are related entities, the corresponding knowledge guidance parameter is set to a preset weight value; otherwise, the knowledge guidance parameter is zero. A knowledge guidance parameter matrix is ​​constructed based on the knowledge guidance parameter. When the corresponding sample features are input into the initial large model for feature extraction, the corresponding attention score matrix is ​​obtained. The sum of the knowledge guidance parameter matrix and the attention score matrix is ​​calculated to obtain an updated score matrix. A predicted feature representation is obtained based on the updated score matrix. The initial large model is fine-tuned based on the predicted feature representation to obtain a target large model. The target large model is used to process target data to obtain a target feature representation.

6. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the large model data processing method according to any one of claims 1 to 4.

7. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the large model data processing method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Knowledge graph construction and attack path prediction method for network security

    CN120455149A

  • System and method for knowledge graph construction using capsule neural network

    US20220180065A1