A vertical domain large model training method for software-defined deception defense

CN122797656APending Publication Date: 2026-09-22GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611273188.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

大模型处理过程中能够补充外部知识并提高回答的事实依据,但领域知识主要在推理阶段被临时引入,难以转化为模型自身对软件定义欺骗防御结构的稳定理解

Benefits of technology

[0008]本发明提供的面向软件定义欺骗防御的垂域大模型训练方法的有益效果在于:通过构建软件定义欺骗防御异构本体图,将训练样本映射为本体路径,并基于本体路径生成覆盖向量、构造掩码重构任务,进一步设计包含多项损失的联合训练目标;同时根据本体覆盖度和风险等级动态调整训练权重,并可设置本体路由式多适配器,使模型在不同软件定义欺骗防御任务上获得有针对性的参数更新,提升垂域大模型对软件定义欺骗防御知识结构的内化能力、任务推理能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797656A_ABST
    Figure CN122797656A_ABST
Patent Text Reader

Abstract

The application provides a vertical domain large model training method for software-defined deception defense, comprising: mapping training samples into path structures in a heterogeneous ontology graph, generating ontology coverage vectors and calculating ontology coverage; using a vertical domain large model as a basic large model, constructing a multi-objective joint loss function, determining training weights, weighting the multi-objective joint loss function as a joint training target, and learning and updating the basic large model parameters through the joint training target; selecting an adapter set, defining an adapter training target in combination with the multi-objective joint loss function for learning and updating adapter parameters through back propagation; and performing consistency scoring on the input and output of the basic large model and converting the feedback loss or preference samples into training feedback until the consistency score meets a preset passing threshold to complete the training. The method can improve the internalization ability and task reasoning ability of the vertical domain large model for the knowledge structure of software-defined deception defense.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, and in particular to a method for training large-scale vertical models for software-defined deception defense. Background Technology

[0002] With the development of large language models, using large vertical domain models to assist in the generation of deception defense strategies, log interpretation, decoy design, and security operations has become a feasible direction.

[0003] While general-purpose large-scale models possess a certain level of language understanding and reasoning capabilities, software-defined deception defense is not simply about deploying a few decoy resources. It requires a comprehensive consideration of the matching relationships between attack paths, asset relationships, decoy authenticity, observed events, and defensive actions. Furthermore, existing knowledge augmentation or retrieval enhancement schemes typically retrieve relevant content from external knowledge bases, rule bases, or threat intelligence bases after a user asks a question or security event is input, and then feed this information along with the input into the large-scale model to generate an answer. While the large-scale model can supplement external knowledge and improve the factual basis of the answer during processing, domain knowledge is mainly introduced temporarily during the reasoning stage and is difficult to transform into a stable understanding of the software-defined deception defense structure within the model itself.

[0004] Training large models using existing technologies to learn domain knowledge for software-defined deception defense is difficult because it is hard to grasp the domain reasoning relationships of "how attack behavior triggers software-defined deception objects, how software-defined deception objects generate observation events, and how observation events correspond to defensive actions". When large models process software-defined deception defense information, they are prone to problems such as one-sided understanding of concepts, mismatch of policy objects, lack of executable constraints in generated content, or unstable security boundaries.

[0005] Therefore, it is necessary to provide a training method for vertical domain large models for software-defined deception defense, so as to improve the internalization ability and task reasoning ability of vertical domain large models of software-defined deception defense knowledge structure. Summary of the Invention

[0006] The purpose of this invention is to provide a vertical domain large model training method for software-defined deception defense.

[0007] The vertical domain large-scale model training method for software-defined deception defense provided by this invention includes: constructing a heterogeneous ontology graph for software-defined deception defense; obtaining training samples and mapping the training samples to path structures in the heterogeneous ontology graph; generating ontology coverage vectors for the training samples based on the mapping results; calculating ontology coverage by comprehensively considering the scores of each coverage dimension in the ontology coverage vectors; using a vertical domain large-scale model as the basic large-scale model, defining mask reconstruction prediction, language model prediction, relation prediction, path rationality scoring, and task training; constructing corresponding loss functions based on each training element; constructing a multi-objective joint loss function by combining the various losses; and determining the training accuracy by comprehensively considering the ontology coverage, risk level, and path complexity of the ontology paths. The training weights are applied to the multi-objective joint loss function as a joint training objective. The parameters of the base model are learned and updated through backpropagation using the joint training objective. Multiple adapters are set up on the base model to carry different parameter increments. The routing weights of each adapter are calculated by combining the ontology coverage vector and auxiliary features of the training samples. The adapter training objective is defined by combining the routing weights and the multi-objective joint loss function for backpropagation to learn and update the adapter parameters. An ontology consistency discriminator is set up to score the consistency between the input and output of the base model. The consistency score is converted into feedback loss or preference samples as training feedback until the consistency score meets the preset threshold, thus completing the training.

[0008] The beneficial effects of the vertical domain large-scale model training method for software-defined deception defense provided by this invention are as follows: By constructing a heterogeneous ontology graph for software-defined deception defense, training samples are mapped to ontology paths, and coverage vectors are generated based on the ontology paths. A mask reconstruction task is constructed, and a joint training objective including multiple losses is further designed. At the same time, the training weights are dynamically adjusted according to the ontology coverage and risk level, and ontology-routable multi-adaptors can be set up so that the model can obtain targeted parameter updates on different software-defined deception defense tasks, thereby improving the vertical domain large-scale model's internalization ability of the software-defined deception defense knowledge structure and its task reasoning ability.

[0009] In one possible embodiment, entity recognition and relation extraction are performed on the training samples. After aligning the identified entities with the heterogeneous ontology graph, the training samples are converted into ontology paths by combining the extracted relations to obtain an ontology path set. It is then determined whether the node type or relation type of each coverage dimension appears in the ontology path set, and the value of the corresponding coverage dimension is determined to obtain the ontology coverage vector. Finally, the scores of each coverage dimension are weighted and averaged to obtain the ontology coverage.

[0010] In another possible embodiment, mask reconstruction prediction training includes generating mask paths based on the ontology path set according to node type to obtain a mask sample set, performing mask reconstruction prediction based on the mask sample set, and calculating the mask reconstruction loss; language model prediction training includes predicting lexical terms based on the context before the target lexical term and calculating the language model loss; relation prediction training includes predicting the relations between nodes based on the text and nodes of the training samples and calculating the relation prediction loss; path rationality scoring training includes performing rationality prediction on the real path and the inconsistent path constructed based on the real path respectively and calculating the rationality scoring loss; task training includes processing the training samples according to the set task requirements and outputting the processing results, and calculating the task loss by combining the processing results and the task type.

[0011] In other possible embodiments, the risk level of the training sample is the weighted average of the values ​​of various risk factors; the training weight is determined by combining the ontology coverage, risk level and path complexity of the ontology path of the training sample, including: adjusting the basic weight by combining the ontology coverage, risk level and path complexity of the ontology path of the training sample, and limiting the adjusted weight within a preset range to obtain the training weight.

[0012] Multiple adapters are set up on the basic large model to carry parameter increments for different software-defined deception objects or different task types; the software-defined deception objects corresponding to the adapters include honey spots, honey arrays, honey holes, and honey gardens; the task types corresponding to the adapters include attack path reasoning, observation event revelation, defense action generation, and security boundary constraints.

[0013] The routing weights of each adapter are calculated by combining the ontology coverage vector and auxiliary features of the training samples. This includes: concatenating the ontology coverage vector and the auxiliary features, performing a linear transformation, and then performing a bias shift and normalization on the linear transformation result to obtain the routing weights. The auxiliary features include, in order, the one-hot encoding of the task type, the risk level, the path complexity, the normalized value of the number of effective paths, the average confidence of the entity relationship, and the security boundary indication value.

[0014] Consistency scoring is performed on the input and output of the basic large model, including: entity recognition and relation extraction of the output of the basic large model and mapping it to a heterogeneous ontology graph to obtain the output ontology path; the matching degree of the input ontology path and the output ontology path on key nodes and relation edges is evaluated respectively; the degree to which the output of the basic large model satisfies the safety boundary node constraints is evaluated; and the weighted summation of the scoring results is used to obtain the consistency score.

[0015] The consistency score is converted into a feedback loss or a preference sample as training feedback, including: comparing the consistency score with preset pass thresholds and rewrite thresholds; when the consistency score is lower than the preset pass threshold, converting the consistency score into a feedback loss and adding it to the adapter training objective; when the consistency score is greater than the preset rewrite threshold but lower than the preset pass threshold, rewriting is performed based on the input ontology path, and the output of the basic large model and the rewritten result are combined to form a preference sample for retraining. Attached Figure Description

[0016] Figure 1 A flowchart illustrating a vertical domain large model training method for software-defined deception defense provided in an embodiment of the present invention;

[0017] Figure 2 This is a schematic diagram illustrating the implementation structure of a vertical domain large model training method for software-defined deception defense provided in an embodiment of the present invention.

[0018] Figure 3 This is a schematic diagram of an electronic device structure provided in an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention. Unless otherwise defined, the technical or scientific terms used herein should have the ordinary meaning understood by those skilled in the art. The terms "comprising" and similar expressions used herein mean that the element or object preceding the word covers the element or object listed following the word and its equivalents, but do not exclude other elements or objects.

[0020] This embodiment provides a method for training large-scale vertical models for software-defined deception defense.

[0021] See Figure 1 and Figure 2 The method includes:

[0022] S101: Construct a heterogeneous ontology graph for software-defined deception defense, obtain training samples and map the training samples to path structures in the heterogeneous ontology graph, generate ontology coverage vectors of the training samples based on the mapping results, and calculate ontology coverage by combining the scores of each coverage dimension in the ontology coverage vector.

[0023] In one possible embodiment, entity recognition and relation extraction are performed on the training samples. After aligning the identified entities with the heterogeneous ontology graph, the training samples are converted into ontology paths by combining the extracted relations to obtain an ontology path set. It is then determined whether the node type or relation type of each coverage dimension appears in the ontology path set, and the value of the corresponding coverage dimension is determined to obtain the ontology coverage vector. Finally, the scores of each coverage dimension are weighted and averaged to obtain the ontology coverage.

[0024] In one specific embodiment, a heterogeneous ontology graph for software-defined deception defense is constructed to uniformly represent honeypots, honey arrays, honey holes, honey gardens, and their associated attack paths, observation events, defensive actions, and security boundaries. The heterogeneous ontology graph for software-defined deception defense can be represented as follows: ,in, Represents the set of ontology nodes. Represents the set of edges between nodes. This represents the set of relation types corresponding to the edges. A set of attributes representing a node or edge. The graph structure of a heterogeneous ontology graph unifies the modeling of objects, behaviors, events, and constraints scattered throughout a software-defined deception defense system.

[0025] Specifically, the ontology node set can be further divided into asset nodes, attack nodes, software-defined deception object nodes, observation event nodes, defense action nodes, and security boundary nodes, which can be represented as follows: Among them, asset nodes Used to represent asset objects such as hosts, services, accounts, and business systems, and attack nodes. Used to represent attack phases, attack techniques, or attack paths, and to observe event nodes. Used to represent observed events such as login attempts, file access, credential touches, abnormal commands, and network connections, serving as a defense action node. Used to represent defensive actions such as alarms, isolation, source tracing, deception and redirection, and policy hardening; security boundary nodes. Used to represent security boundaries such as authorization restrictions, output granularity, and sensitive information constraints. A collection of software-defined deception object nodes. It can be further expressed as These correspond to honey points, honey arrays, honey holes, and honey gardens, respectively. A honey point represents a single deceptive touchpoint, a honey array represents a collaborative deceptive structure formed by multiple honey points, a honey hole represents a controlled trap space that induces attackers to engage deeply in interaction, and a honey garden represents a relatively complete simulated business environment or deceptive scenario.

[0026] The set of edges between nodes can be represented as ,in, Indicates from node To the node Existence Relationship For example, a relation can represent "an attack path acts on an asset", "a software-defined deception object triggers an observation event", "an observation event indicates an attack phase", or "a security boundary constraint defense action".

[0027] Furthermore, for each relation type Configure the relation schema items. A relation schema item must include at least the type of the starting node. Endpoint node type The schema includes the relationship direction, allowed attribute conditions, and security level. For example, the starting point type of a "trigger" relationship is a software-defined deception object node, and the ending point type is an observed event node; the starting point type of a "constraint" relationship is a security boundary node, and the ending point type is a defensive action node. Each relation schema item is stored in the ontology schema table.

[0028] After constructing the heterogeneous ontology graph for software-defined deception defense, the training samples are mapped to path structures within the heterogeneous ontology graph. Training samples can be software-defined deception defense-related text, question-and-answer samples, log explanation samples, policy analysis samples, or event description samples.

[0029] For any training sample The process of mapping training samples to path structures in a heterogeneous ontology graph includes:

[0030] Entity recognition and relation extraction are performed on the training samples to obtain a candidate entity relation set, which can be represented as: ,in, Represents the set of candidate entity relations. This represents the entity relation extraction function. and This indicates that the entity is mentioned in the sample. This represents the extracted candidate relation. This represents the confidence level of entity relationships. Entity relationship extraction is used to identify elements such as attack phases, attack paths, honeypots, honey squares, honey holes, honey courtyards, observed events, defensive actions, and security boundaries from ordinary text. Specifically, after processing the training samples through the entity relationship extraction function, entity mentions are output separately. , Entity recognition probability , and candidate relationships Relationship classification probability The original confidence level of the candidate relationship This is the geometric mean of the entity recognition probability and the relation classification probability, i.e. Its value range is .

[0031] After obtaining the candidate entity relation set, the obtained entity mentions are aligned with the ontology nodes in the heterogeneous ontology graph. The mapping from entity mentions to ontology nodes can be represented as: , ,in, This indicates the mapping relationship between entity references and ontology nodes. Represents the similarity function. Indicates entity mention The vector representation of , Represents the ontology node The vector representation of , This represents a preset mapping threshold. When the similarity is not lower than the threshold, entity mentions are mapped to the corresponding ontology nodes; when the similarity is lower than the threshold, entity mentions can be retained as unaligned entities or candidate nodes to reduce the impact of incorrect mappings on the training process.

[0032] After completing entity alignment, calibration is performed according to the ontology pattern table: If The node type belongs to and The node type belongs to If the confidence level is positive, retain that confidence level; otherwise, set it to negative. Only those with a post-calibration confidence level not lower than the threshold will be considered. Candidate relationships are used for path construction.

[0033] For example, entity alignment adopts a processing principle of candidate recall, comprehensive scoring, and threshold determination: first, a candidate set is recalled from ontology nodes of the same type based on entity category, standard name, alias dictionary, and asset attributes; then, a comprehensive score is calculated. ,in For vector cosine similarity, For literal similarity of names or aliases, For the degree of consistency in asset type, business domain, port, or account attributes, and Select the entity with the highest overall score that is not lower than the entity threshold. The node as If the difference between the highest and second-highest scores is less than the ambiguity threshold... If so, the entity will be marked as an entity to be confirmed, and alignment will not be forced.

[0034] Relation alignment first maps candidate relations to a relation thesaurus or relation classifier. The standard relation type in the code is then checked, followed by the relation direction. / The type constraint and the existence of the corresponding ontology edge are checked. The corresponding edge exists in the ontology, or the ontology schema allows the relation and the confidence of the candidate relation is not lower than a certain threshold. When a relationship is included in the sample relationship set, it is considered an ambiguous relationship and does not participate in strong constraint training when multiple equally feasible standard relationships exist. Relations that fail type or direction verification are not included in the ontology path.

[0035] After aligning entities and relationships, the training samples are converted into one or more ontology paths based on the edge relationships between ontology nodes. The set of paths corresponding to the samples can be represented as follows: ,in, Indicates sample The resulting set of ontology paths, each ontology path It consists of multiple ontology nodes and relation edges. For example, a sample describing "an attacker accessing a decoy account and triggering an alert during lateral movement" can be mapped to a path of "lateral movement—honeypot or honeycomb—account access event—alert action"; a sample describing "constructing a virtual business environment to induce the attacker to engage deeply" can be mapped to a path of "attack detection—honeycomb—multi-round interaction events—source tracing action". Thus, training samples are transformed from ordinary natural language text into path-based representations with a software-defined deception defense semantic structure.

[0036] Specifically, path transformation is used to align the training samples with the set of ontology nodes after aligning the heterogeneous ontology graph. and relation set Construct sample subgraphs for anchor points For valid relationships that appear directly in the sample, adjacent anchor points are connected first, following the order of events in the text; for anchor points that are not directly connected, the heterogeneous ontology graph is used to connect them. The search length is constrained by relation direction and type to not exceed a set threshold. The shortest reachable path. Candidate paths must not contain loops formed by repeating nodes and should simultaneously satisfy the start and end types and safety boundary constraints in the ontology schema table.

[0037] The paths are scored based on the average confidence score of each entity and relation in the sample subgraph, the anchor point coverage, and the path length, with the best path retained. Analyze the Top-K paths and remove duplicate paths. For example, a path score can be calculated using the following formula: , Indicates path The score obtained from the rating Representing a path The average confidence level, Representing a path Anchor point coverage Representing a path The path length. The ontology path set obtained by the method according to embodiments of the present invention satisfies the set constraints in terms of anchor points, length, type, and score.

[0038] , and path threshold These are preset hyperparameters, which can be determined based on the maximum allowed task chain length of the ontology and the accuracy of the validation set path. If no sample exists, the hyperparameters will be used. If a valid path is found, the sample will only participate in language modeling or task loss, and will not participate in mask reconstruction, relation prediction, or path consistency loss.

[0039] In one possible embodiment, after obtaining the set of ontology paths corresponding to the training samples, an ontology coverage vector is further generated based on the ontology paths to describe the coverage of the samples in the software-defined deception defense knowledge structure. For the training samples... Its ontology coverage vector can be expressed as ,in, Indicates training samples The ontology coverage vector, This indicates the preset number of ontology dimensions. Represents the first in the ontology coverage vector The ontology dimension has several possible values. Ontology dimensions can include types such as honeypots, honey arrays, honey holes, honey gardens, attack phases, attack paths, observed events, defensive actions, and security boundaries. This coverage vector is used to characterize which software-defined deception objects, attack-defense relationships, and task attributes the sample involves.

[0040] Specifically, the process of constructing the ontology coverage vector includes: pre-establishing a coverage dimension dictionary with fixed indices. The coverage dimensions can be divided into node type dimension, relationship type dimension, task type dimension, and security boundary dimension: the node type dimension corresponds to honey spots, honey arrays, honey holes, honey gardens, attack phases, etc.; the relationship type dimension corresponds to "acts on," "triggers," "indicates," "responds," "constraints," etc.; the task type dimension is determined by sample labels such as question answering, log interpretation, and policy analysis. A dictionary is generated for each training sample according to the coverage dimensions. Dimensional covering vector.

[0041] When generating the ontology coverage vector, first... Initialize to all zeros, then iterate through All paths in the table: If a node type or relation type appears in any path, then the corresponding dimension is set to [value]. Then, based on the sample task labels and security boundary labels, set the corresponding dimensions, and merge the coverage results of multiple paths according to a logical OR.

[0042] Specifically, for the node type dimension, the th node in the coverage vector The values ​​of each dimension can be determined based on whether the sample path contains ontology nodes of the corresponding type, and their calculation method can be expressed as follows: ,in, Indicates an indicator function, Indicates the first A collection of class ontology nodes. When any path obtained from the sample mapping contains the first... When class nodes, The value is 1 otherwise 0. This method can be used to determine whether a sample covers honey points, honey arrays, honey holes, or honey gardens, and also whether it involves attack path reasoning, observation event interpretation, defense action selection, or security boundary constraints. For the relation type dimension, the value of the first element in the coverage vector is 1. The values ​​of each dimension can be determined based on whether the sample path contains a relation of the corresponding type, and their calculation method can be expressed as follows: , Indicates the first Edge set, Representing an edge For path The top edge. For the task type dimension, values ​​are assigned directly based on the sample labels. If If empty, the path-related dimension remains unchanged. .

[0043] After obtaining the coverage of each dimension, the overall ontology coverage of the sample can be further calculated. The ontology coverage calculation satisfies the following formula: ,in, Indicates sample Body coverage, Indicates the first Weights based on class ontology. Coverage weights. Take non-negative values ​​and normalize them so that... Equal-weight initialization is allowed by default; however, if the training objective focuses more on software-defined deception targets, attack paths, defensive actions, or security boundaries, the corresponding dimension weights can be increased on the validation set. The weights for each training task are determined... The settings are then kept constant to ensure comparable coverage across different samples. For example, if certain dimensions are more important for training software-defined deception defenses, such as the software-defined deception target, attack path, defense action, or security boundary, then higher values ​​can be set. .

[0044] S102: Using a large vertical model as the base model, define mask reconstruction prediction, language model prediction, relation prediction, path rationality scoring, and task training. Construct corresponding loss functions based on each training step, and combine the various losses to construct a multi-objective joint loss function. Determine training weights by comprehensively considering the ontology coverage, risk level, and path complexity of the ontology path of the training samples. Apply the training weights to weight the multi-objective joint loss function as the joint training objective. Learn and update the parameters of the base model through backpropagation using the joint training objective.

[0045] In one possible embodiment, mask reconstruction prediction training includes generating mask paths based on the ontology path set according to node type to obtain a mask sample set, performing mask reconstruction prediction based on the mask sample set, and calculating the mask reconstruction loss; language model prediction training includes predicting lexical terms based on the context before the target lexical term and calculating the language model loss; relation prediction training includes predicting the relations between nodes based on the text and nodes of the training samples and calculating the relation prediction loss; path rationality scoring training includes predicting the rationality of real paths and inconsistent paths constructed based on real paths and calculating the rationality scoring loss; task training includes processing the training samples according to the set task requirements and outputting the processing results, and calculating the task loss by combining the processing results and the task type.

[0046] The risk level of the training samples is the weighted average of the values ​​of various risk factors; the training weights are determined by combining the ontology coverage, risk level and path complexity of the ontology path of the training samples, including: adjusting the basic weights by combining the ontology coverage, risk level and path complexity of the ontology path of the training samples, and limiting the adjusted weights within a preset range to obtain the training weights.

[0047] In one specific embodiment, the training structure for training the vertical domain large model is designed as follows: a base large model is the vertical domain large model, further configured with an ontology path serialization unit, a node reconstruction prediction head, a relationship prediction head, a path rationality scoring head, a task output head, a router, and adapters for multiple domains. Training is performed using the original text of the training data. As the main input, the ontology path Serialization follows the order of "node type - node identifier - relation type - next node", and is separated by a delimiter. The combined inputs share a common base model. The base model outputs word-level hidden states, which are shared by all prediction heads.

[0048] The output sample records after sequentially serializing the training data are ,in, This represents the target output of the monitoring task. Indicates sample The risk level, Indicates sample The average depth or complexity of the corresponding ontology path This indicates the task type label.

[0049] The risk level of the sample is a weighted average of the values ​​of various risk factors, specifically calculated based on the security sensitivity, attack path complexity, output risk, and security boundary requirements of the sample. The risk level is calculated according to the following formula: ,in, Indicates the number of risk factors. Indicates the first The weights corresponding to risk factors Indicates sample In the Values ​​for risk factors. Risk factors may include whether genuine credentials are involved, whether a highly sensitive attack phase is involved, whether highly operable actions are involved, whether authorization constraints are required, and whether security boundary nodes are involved, etc. Values ​​for each risk factor. Determined and normalized from the preset risk configuration table An example configuration is as follows: whether credentials or sensitive data are involved is determined by "not involved, only descriptive and anonymized, contains highly sensitive semantics". , , The attack phase is categorized into low, medium, and high sensitivity levels. , , Action operability is determined by the concept description, configuration suggestions, and directly executable actions. , , When authorization constraints are required or security boundary nodes are included, the corresponding factors are taken. Otherwise take Authentic plaintext credentials and unauthorized execution details are not used as training objectives. Risk weights. are non-negative numbers and satisfy The weighting can be set initially with equal weights, or it can be set by security experts based on the importance of the factors. Taking five categories of risk factors as an example, the following can be used: As an unrestricted initial value, it is then adjusted on the validation set using the safety boundary violation rate and task availability as indicators.

[0050] The average depth or complexity of an ontology path is the normalized average of the ontology path lengths, and its calculation process is as follows: for a non-empty set of paths... First, calculate the depth of a single path. , then calculate ,in, Indicates the number of path nodes. This indicates the preset maximum allowed path length. If... If empty, then .thus, The range of values ​​is It increases with the increase of average path length.

[0051] The node reconstruction prediction head, relation prediction head, path rationality scoring head, and task output head are each designed with different training tasks, and the construction of the training tasks for each head involves the same... Four types of structured training instances are generated in parallel. Among them, the training of the node reconstruction prediction head involves... Form a mask path set The training of the relationship prediction head involves... The training of the path rationality scoring head involves generating positive paths from real paths and generating inconsistent negative paths, while the training of the task output head involves... This generates language modeling or specific task samples. In the forward computation of training, the task output head calculates the language model loss and the task loss, the node reconstruction prediction head calculates the mask reconstruction loss, the relation prediction head calculates the relation prediction loss, and the path rationality scoring head calculates the rationality score loss.

[0052] In a specific embodiment, the training performed during the forward computation includes mask reconstruction prediction training, language model prediction training, relation prediction training, path rationality scoring, and task training. Mask reconstruction prediction training includes generating masked paths based on the ontology path set according to node type to obtain a masked sample set, performing mask reconstruction prediction based on the masked sample set, and calculating the mask reconstruction loss. Language model prediction training includes predicting lexical terms based on the context preceding the target lexical term and calculating the language model loss. Relation prediction training includes predicting the relationships between nodes based on the text and nodes of the training samples and calculating the relationship prediction loss. Path rationality scoring training includes predicting the rationality of real paths and inconsistent paths constructed based on real paths, and calculating the rationality scoring loss. Task training includes processing the training samples according to the set task requirements and outputting the processing results, and calculating the task loss by combining the processing results and the task type.

[0053] Specifically, a mask reconstruction task is constructed based on the ontology path set: for any training sample, its ontology path can be represented as follows: Based on the node type, the positions requiring masking are selected from the ontology path and masked. The masked ontology path can be represented as follows: , Indicates the type of node to be masked. Selectable masking locations include attack path nodes, software-defined deception object nodes, observation event nodes, defensive action nodes, or security boundary nodes.

[0054] The mask location selection follows the principle of point-by-point masking only for the actually existing allowed types in the path. A pre-configured set of allowed mask node types is configured. For each ontology path First find For each of the node locations, a training instance is generated that masks only that node. If multiple nodes of the same type appear in the path, a single-point mask instance is generated for each. No instance is generated for node types not present in the ontology path. If the number of mask instances generated for a single sample exceeds a preset limit... When sampling, stratified sampling is performed according to node type to ensure that at least one instance of each allowed type that has appeared in the ontology path is retained. Then, the remaining positions are supplemented according to entity relationship confidence or random method. Single-point masking is preferred to maintain the uniqueness of the target; multi-point masking can also be used between non-adjacent nodes, but the type label of each masked node should be retained at the same time.

[0055] For training samples A mask sample set can be generated based on its ontology path set, represented as Each element in the masked sample set represents a training objective, i.e., given the ontology path after masking and the original text context, predicting the masked nodes to complete the mask reconstruction. For example, given "lateral movement — [MASK] — account access event — alarm action", the model should predict that the appropriate software-defined deception object may be a honey spot or honey array; given "honey hole — abnormal command interaction — [MASK]", the model should predict that the corresponding defense action may be isolation, tracing, or alarm.

[0056] The candidate output set during reconstruction is limited to the same range as... Same node type and satisfying adjacency relationship Constrained ontology nodes reduce the prediction scope from the entire vocabulary to a set of legitimate candidate nodes. The mask reconstruction loss is then calculated. At that time, The average cross-entropy of each instance in the dataset is calculated; if If empty, then the corresponding sample Recorded as And it does not generate this gradient. Specifically, for training samples... The mask reconstruction loss can be expressed as , Indicates the model in parameters The predicted probability is as follows. This represents the set of parameters participating in the current forward computation in the basic large language model and the shared prediction head. Its role is to map the text context and path context to the conditional probabilities of words, nodes, or relations. The mask reconstruction loss is used to ensure that the model maintains basic text understanding and generation capabilities during training in the software-defined deception defense vertical domain, avoiding the model only learning structural relations and losing its natural language expression capabilities.

[0057] For training samples The language model prediction loss can be expressed as , Indicates the sample length. Indicates the first in the sample Each word element, express The context preceding each word. The language model prediction loss is used to ensure that the model maintains basic text understanding and generation capabilities during training in the software-defined deception defense vertical domain, preventing the model from learning only structural relationships and losing its natural language expression ability.

[0058] For any triple in the ontology path The relationship prediction loss can be expressed as , and Indicates adjacent body nodes, express and The true type of relationship between them. The relationship prediction loss is designed to make the model predict the correct relationship as much as possible, given sample text and two ontology nodes.

[0059] For the real path Inconsistent paths obtained through node replacement, relation perturbation, or path truncation The loss in rationality scoring can be expressed as , The model represents the path In the sample Reasonableness scoring in context, including path refer to or ; This represents a preset interval threshold. The reasonableness score loss requires that the score of the true path is at least higher than the score of the incorrect path. Otherwise, losses will occur, thus enabling the model to distinguish between legitimate software-defined deception defense links and links that do not conform to domain logic. It should be noted that inconsistent paths... according to Generate, for each Construct at least one of the following candidates: replace a node while maintaining the same node type; replace a relation with a relation that has the same start and end types but is semantically incompatible; perform truncation, deorder concatenation, or deletion of key nodes on the path. For each... Each of the three types of operations described above generates one candidate negative path, calculates its current reasonableness score, and selects the path with the highest score as the difficult negative path. If a certain type of operation cannot generate valid candidates, then stratified random sampling is performed from the same type of path in the same batch. Negative paths can be regenerated in each training round to improve perturbation diversity. (Calculation) The average can be calculated for one or more negative paths corresponding to each positive path. Training samples With serialization path A shared, basic model is used to obtain aggregated hidden vectors of words at the delimiter or end of a word. Then obtain through a linear scoring head , The trainable weight vector representing the path rationality score head. This represents the transpose of the trainable weight vector. This represents a trainable bias term for the path rationality scoring head; a higher score indicates a more consistent path with the sample context. The path rationality scoring head and model parameters are trained together using a rationality scoring loss. (Interval threshold) If it is a positive number, it can be selected on the validation set to achieve a high accuracy in distinguishing between positive and negative paths while maintaining stable loss.

[0060] Mission loss Different methods are used to determine the supervision objective depending on the task type: For generative tasks, let the supervision objective be... , This represents the total number of types of monitored targets. The task loss is calculated based on the word-level cross-entropy of the target sequence. For classification tasks, This represents the negative logarithmic probability of the true class. It should be noted that only the task items with supervised labels are calculated for a single sample; in multi-task batches, the average is first calculated within each task, and then normalized according to the task sampling ratio to avoid tasks with larger sample sizes dominating joint training.

[0061] A multi-objective joint loss function is constructed by combining mask reconstruction loss, language model loss, relation prediction loss, reasonableness score loss, and task loss, which can be specifically expressed as follows: ,in, This represents the weighting coefficient, and it takes non-negative values, which can be normalized to make... The system can be initialized with equal weights by default, and then grid search or item-by-item adjustment can be performed based on language quality, node reconstruction accuracy, relation accuracy, path consistency accuracy, and specific task metrics on the validation set; once determined during a single training run, it remains fixed.

[0062] By designing a mask reconstruction task, the model learns the correspondence between attack paths, software-defined deception objects, observed events, and defensive actions. Furthermore, the aforementioned multi-objective joint loss function is designed to enable the model to simultaneously learn text representation, node completion, relationship judgment, path consistency, and specific task capabilities.

[0063] Because different training samples contribute differently to the training of the vertical domain model for software-defined deception defense, some samples only involve the conceptual interpretation of a single software-defined deception object, some samples simultaneously cover attack paths, software-defined deception objects, observed events, and defensive actions, and some samples involve security boundaries or high-risk output constraints. Therefore, the loss weights of the training samples are dynamically adjusted based on their ontology coverage, risk level, and path complexity, so that the model training process pays more attention to samples with complete structures, complex task chains, or strong security constraints.

[0064] The training weights, determined by considering ontology coverage, risk level, and path complexity, can be expressed as follows: ,in, Indicates the basic weight. , , Represents the adjustment coefficient, function Used to limit training weights and Between these, avoid giving too high or too low a weight to a single sample.

[0065] In one specific embodiment , and Normalized to Basic weights Preferred setting ; , , The adjustment coefficient is non-negative and can be an equal-weighted initial value or searched on the validation set based on task metrics, path consistency metrics, and training stability. and For the pre-defined cutoff boundary, the preferred initial value can be taken as follows: and If gradient fluctuations occur or a small number of sample weights remain at their peak for an extended period, the interval can be narrowed. Alternatively, the original weights within a batch can be divided by the batch mean to achieve an average weight. Execute again Cut off.

[0066] Finally, the joint training objective, weighted by the training weights, can be expressed as: ,in, This represents the training sample set. The design of the joint training objective means that the contribution of each sample to the overall training objective is no longer fixed, but is jointly determined by its ontology coverage, risk level, and path complexity. Through this mechanism, the model can more fully learn key relational samples and security-sensitive samples in software-defined deception defense during the training phase, and provide a training basis for subsequent ontology-based routing adapter activation and output consistency judgment.

[0067] The parameters of the base model are updated through backpropagation and gradient optimizers. There are two specific approaches to updating the base model parameters: updating pre-selected parameters in the base model through full or partial fine-tuning. The system either freezes the base model parameters using efficient parameter adaptation methods, or updates only the domain adapter, routing parameters, and each prediction head. After each training round, the system evaluates task accuracy, path consistency, and safety boundary violation rate on the validation set, and determines the entity threshold, relation threshold, margin threshold, dynamic weight range, and output consistency threshold accordingly.

[0068] S103: Set up multiple adapters on the basic large model to carry different parameter increments. Calculate the routing weights of each adapter by combining the ontology coverage vector and auxiliary features of the training samples. Define the adapter training objective by combining the routing weights and the multi-objective joint loss function for backpropagation learning and updating the adapter parameters.

[0069] In one possible embodiment, multiple adapters set on the basic large model are used to carry parameter increments for different software-defined deception objects or different task types; the software-defined deception objects corresponding to the adapters include honey spots, honey arrays, honey holes, and honey gardens; the task types corresponding to the adapters include attack path reasoning, observation event revelation, defense action generation, and security boundary constraints.

[0070] The routing weights of each adapter are calculated by combining the ontology coverage vector and auxiliary features of the training samples. This includes: performing a linear transformation after concatenating the ontology coverage vector and auxiliary features, performing an offset shift on the linear transformation result, and then normalizing it to obtain the routing weights. The auxiliary features include, in order, the one-hot encoding of the task type, the risk level, the path complexity, the normalized value of the number of effective paths, the average confidence of the entity relationship, and the security boundary indication value.

[0071] In one specific embodiment, multiple domain adapters are set up on the basic large model to carry parameter increments for different software-defined deception objects or different task types. The adapters can correspond to honey spots, honey matrices, honey holes, honey gardens, or task modules such as attack path inference, observation event interpretation, defense action generation, and security boundary constraints. Instead of activating the same adapter for all samples, the adapter routing weights are calculated using ontology coverage vectors and sample risk information. For training samples... The routing weight of its adapter can be expressed as , Indicates auxiliary features, and Represents trainable routing parameters and route weights. Each dimension in the vector represents the weight of the corresponding activated adapter. The role of routing weights is to automatically determine the degree of participation of different adapters based on the software-defined deception object and task attributes involved in the sample. The auxiliary features are fixed-length vectors concatenated from various features such as task type, risk level, or path complexity of the sample.

[0072] In one specific embodiment, In order, these include one-hot encoding of task types. , The normalized value of the number of valid paths, the average confidence of entity relationships, and the indication value of whether a safe boundary node is included. and Learning is achieved through backpropagation using joint loss. It should be noted that the routing weights are jointly determined by the ontology coverage vector and auxiliary features. When the adapter corresponds to a software-defined deception object, the category information of the software-defined deception object is represented by the coverage dimensions corresponding to honey points, honey matrices, honey holes, and honey courtyards in the ontology coverage vector. The one-hot encoding of the task type in the auxiliary features is used to supplement the representation of the task attributes of the training samples.

[0073] In one possible embodiment, during the model forward propagation, let the first... The hidden state of the basic large model is , No. One adapter is The hidden state of each layer of the basic large model after ontology routing can be represented as follows: The hidden states, after ontology routing, serve as input to subsequent base model layers and prediction heads. They are used to calculate the multi-objective joint loss function and further form the adapter training objective. Backpropagation updates the activated adapter parameters and related routing parameters, thus establishing a clear training chain involving adapter selection, hidden state calculation, loss calculation, and parameter updates. During backpropagation, the loss gradient is transmitted via the hidden states to the selected adapter parameters and the routing parameters participating in the activation weight calculation. Unselected adapter parameters do not receive the gradient corresponding to the training sample.

[0074] Among them, adapter set Adapter-based routing weights are calculated using a threshold plus the previous one. The rule selection process prioritizes routes with weights greater than the weight selection threshold. The adapter; if the number of selected is less than Then supplement the route with the highest weight. One adapter; if more than Then only the top-weighted ones are retained. indivual. , and The weights of the selected adapters are preset based on computing resources and validation set performance. Internal renormalization yields the actual activation weights, i.e. , Indicates training samples For the The activation weight of each adapter, Indicates training samples For the The routing weight of each adapter; the weight of unselected adapters is reset to [value missing]. .

[0075] In one specific embodiment, to ensure that the adapter routing is consistent with the joint training objective, the adapter training objective can be represented as: , Indicates the joint loss of multiple objectives. Represents the basic large model parameters. Represents the adapter parameter set. Represents a regular term, This represents the weight of the regularization term. This training method enables different software-defined deception defense tasks to obtain targeted parameter updates while maintaining a continuous relationship with ontology paths, coverage vectors, and risk weights. The regularization term is determined by: based on a pre-defined mapping table between the adapter and ontology dimensions / task types, ... Mapping task type to prior correlation vector and normalized to the target distribution. The prior alignment loss of the route can be expressed as: This is used to suppress adapter activation that is unrelated to sample ontology coverage. Meanwhile, in a training batch... Internal calculation of average load for each adapter and with target load (The distribution can be uniform or determined according to the frequency of training corpus categories) Construct a load balancing loss. This is used to prevent routing from being concentrated on a few adapters for an extended period. Ultimately, it enables... , , The coefficients are non-negative and are determined on the validation set.

[0076] The adapter is used to train the objective, and backpropagation is performed to learn and update the adapter parameters. The update of the adapter parameters can be handled in two ways: freezing the base model in the parameter efficient adaptation mode. Update only , The adapter parameters and prediction head are updated simultaneously in the joint fine-tuning mode, while the learning rate of the pre-selected base large model layer is lower than that of the adapter and router.

[0077] S104: Set the ontology consistency discriminator to score the consistency between the input and output of the basic large model, and convert the consistency score into feedback loss or preference samples as training feedback until the consistency score meets the preset threshold to complete the training.

[0078] In one possible embodiment, consistency scoring of the input and output of the basic large model includes: performing entity recognition and relation extraction on the output of the basic large model and mapping it to a heterogeneous ontology graph to obtain the output ontology path; evaluating the matching degree of the input ontology path and the output ontology path on key nodes and relation edges respectively; evaluating the degree to which the output of the basic large model satisfies the safety boundary node constraints; and obtaining a consistency score by weighted summation of the scoring results.

[0079] The consistency score is converted into a feedback loss or a preference sample as training feedback, including: comparing the consistency score with preset pass thresholds and rewrite thresholds; when the consistency score is lower than the preset pass threshold, converting the consistency score into a feedback loss and adding it to the adapter training objective; when the consistency score is greater than the preset rewrite threshold but lower than the preset pass threshold, rewriting is performed based on the input ontology path, and the output of the basic large model and the rewritten result are combined to form a preference sample for retraining.

[0080] In one possible embodiment, an ontology consistency discriminator is set to constrain the model output. This ontology consistency discriminator is used to determine whether the content generated by the model conforms to the domain relationships in the software-defined deception defense ontology, avoiding situations such as mismatches between software-defined deception objects and attack paths, mismatches between observed events and defense actions, or inconsistent security boundaries in the output.

[0081] Specifically, for the output generated by the model in response to the input, entity recognition and relation extraction are first performed, and the output content is mapped to a heterogeneous ontology graph, which can be represented as follows: , Indicates the output The results obtained from entity relation extraction. This indicates that the extraction results are mapped to a heterogeneous ontology graph. The function, This represents the set of ontology paths corresponding to the output. After mapping the output content to a heterogeneous ontology graph, it is transformed into computable ontology paths for structured validation. Once the output ontology paths are obtained, a consistency score is calculated between the output and the corresponding input ontology paths.

[0082] In one possible embodiment, the consistency score is calculated according to the following formula: ,in. This represents the consistency score between input and output. This indicates the degree of matching between the input and output ontology paths at key nodes. This indicates the degree of matching between the input and output ontology paths on the relation edges. This indicates the degree to which the output satisfies the safety boundary node constraints. , , This indicates the weight of each scoring item. The consistency score not only checks whether the model mentions the correct software-defined deception target, but also whether the relationship between its attack path, observed events, defensive actions, and security boundaries is reasonable.

[0083] Specifically, A weighted Jaccard similarity between the output key node set and the input path key node set can be used; The proportion of triples in the output relation triples that simultaneously satisfy the ontology pattern and are compatible with the input path can be adopted; Press Calculation, where The number of violations of authorization, sensitive information, or action granularity constraints. To apply the safety boundary constraint number, all three components are normalized to... If a certain component does not have an applicable object, the scoring weights are renormalized among the remaining components. , , The coefficients are non-negative and sum to 1.

[0084] In one possible embodiment, the consistency discriminator can be applied to both the model training phase and the inference phase of the trained model. The consistency discriminator uses the same entity relation extraction, ontology mapping, and scoring process when applied to different phases, but the scoring results have different effects on subsequent processing.

[0085] Specifically, after obtaining the consistency score, the consistency score is compared with the preset pass threshold and rewrite threshold.

[0086] During the training phase, structured feedback is provided for parameter updates based on the consistency score: when the consistency score falls below a preset threshold, the consistency score is converted into a feedback loss. And the feedback loss is expressed as a coefficient Add the adapter training objective. Furthermore, when the consistency score is greater than the preset rewrite threshold but lower than the preset pass threshold, rewrite the ontology path based on the input, and combine the output of the base model and the rewritten result to form a preference sample for retraining.

[0087] During the inference phase, constraints are applied to the model output based on the consistency score. The output rules can be expressed as follows: Indicates passing the threshold, Indicates the rewrite threshold. This represents the final output result. Specifically, if the consistency score is high (greater than the pass threshold), the original output is retained; if the score is in the middle range (greater than the rewrite threshold but less than the pass threshold), it is rewritten or downgraded based on the input ontology path and ontology graph; if the score is too low (less than the rewrite threshold), the output is rejected or filtered. and satisfy The criteria are determined on the validation set based on false release rate, rewrite success rate, and task availability; for high-risk samples, improvements can be made. and / or improve This is done by tightening the direct pass conditions and expanding the rejection interval, respectively. Preservation is prioritized during rewriting. The aligned nodes and valid relationships in the dataset are deleted or replaced with content that fails the ontology schema and security boundary checks.

[0088] By constraining the model output through consistency scoring, a technical closed loop is completed, from ontology construction, sample pathing, joint training, adapter routing to output consistency constraints. This enables large-scale vertical models to have more stable domain representation and controllable output capabilities in software-defined deception defense tasks.

[0089] The vertical domain large-scale model training method for software-defined deception defense provided by this invention emphasizes transforming the structured knowledge in software-defined deception defense into constraint signals during the model training process, enabling the model to learn the intrinsic relationship between software-defined deception objects, attack behaviors, and defense responses. The design incorporates a complete technical chain from domain ontology construction, sample structured representation, joint model training, task adaptation to output consistency control, which enhances the vertical domain large-scale model's internalization ability of software-defined deception defense knowledge structures, task reasoning ability, and output controllability.

[0090] The structured representation process, which involves sample parsing, ontology alignment, constraint path generation, and overlay encoding, enables samples to form a computable, mappable, and trainable domain knowledge structure. Compared to existing solutions that rely on manual rules, template configuration, or decentralized knowledge bases, this approach improves the structuring of data representation and domain consistency.

[0091] The task of constructing a mask reconstruction based on ontology path is designed, and a multi-objective joint training method is designed, which includes language model loss, relation prediction loss, path consistency loss and task loss. This allows the model to learn the surface semantics of the text, as well as the chain relationship between attack paths, software-defined deception objects, observed events and defensive actions. At the same time, through dynamic training weight adjustment, complex path samples and security-sensitive samples are trained more fully.

[0092] This invention employs an ontology-based routing multi-adaptor architecture, dynamically activating corresponding adapters based on sample coverage vectors, task types, and risk information. This allows the model to adapt parameters differently for honey spots, honey arrays, honey holes, honey courtyards, and various software-defined deception defense tasks. Compared to existing solutions that rely solely on fine-tuning of a single model or enhancement based solely on retrieval, this invention achieves more granular capability adaptation. Furthermore, by using an ontology consistency discriminator to perform path matching and security boundary checks on the model output, the probability of object mismatches, inconsistent relationships, and output out-of-bounds errors can be reduced.

[0093] In other embodiments of this application, an electronic device is disclosed, such as... Figure 3 As shown, the electronic device 300 may include: one or more processors 301; a memory 302; a display 303; one or more application programs (not shown); and one or more computer programs 304. These devices can be connected via one or more communication buses 305. The one or more computer programs 304 are stored in the memory and configured to be executed by the one or more processors 301. The one or more computer programs 304 include instructions that can be used to perform actions such as... Figure 1 And the steps in the corresponding embodiments.

[0094] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0095] In the embodiments of this application, the functional units can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0096] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as flash memory, portable hard disk, read-only memory, random access memory, magnetic disk, or optical disk.

[0097] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.

Claims

1. A method for training a large-scale vertical model for software-defined deception defense, characterized in that, include: Construct a heterogeneous ontology graph for software-defined deception defense, obtain training samples and map the training samples to path structures in the heterogeneous ontology graph, generate ontology coverage vectors of the training samples based on the mapping results, and calculate ontology coverage by combining the scores of each coverage dimension in the ontology coverage vectors. Using a large vertical model as the base model, mask reconstruction prediction, language model prediction, relation prediction, path rationality scoring, and task training are defined. A corresponding loss function is constructed based on each training step. A multi-objective joint loss function is constructed by combining the various losses. The training weights are determined by comprehensively considering the ontology coverage, risk level, and path complexity of the ontology path of the training samples. The training weights are applied to the multi-objective joint loss function as a joint training objective. The parameters of the base model are learned and updated through backpropagation using the joint training objective. Multiple adapters are set on the basic large model to carry different parameter increments. The routing weights of each adapter are calculated by combining the ontology coverage vector and auxiliary features of the training samples. The adapter training objective is defined by combining the routing weights and the multi-objective joint loss function for backpropagation learning and updating the adapter parameters. An ontology consistency discriminator is set up to score the consistency between the input and output of the basic large model. The consistency score is then converted into feedback loss or preference samples as training feedback until the consistency score meets the preset threshold to complete the training.

2. The method according to claim 1, characterized in that, Entity recognition and relation extraction are performed on the training samples. After aligning the recognized entities with the heterogeneous ontology graph, the training samples are converted into ontology paths by combining the extracted relations to obtain an ontology path set. Determine whether the node type or relation type of each coverage dimension appears in the ontology path set, determine the value of the corresponding coverage dimension to obtain the ontology coverage vector, and perform a weighted average of the scores of each coverage dimension to obtain the ontology coverage.

3. The method according to claim 1, characterized in that, The mask reconstruction prediction training includes generating mask paths based on node types according to the ontology path set to obtain a mask sample set, performing mask reconstruction prediction based on the mask sample set, and calculating the mask reconstruction loss. Language model prediction training includes predicting lexical terms based on the context preceding the target lexical term and calculating the language model loss; Relationship prediction training includes predicting the relationships between nodes based on the text and nodes of the training samples and calculating the relationship prediction loss; The training of path rationality scoring includes predicting the rationality of real paths and inconsistent paths constructed based on real paths, and calculating the rationality score loss. Task training includes processing the training samples according to the set task requirements and outputting the processing results, and calculating the task loss by combining the processing results and the task type.

4. The method according to claim 1, characterized in that, The risk level of the training samples is the weighted average of the values ​​of various risk factors; The training weights are determined by combining the ontology coverage, risk level, and path complexity of the ontology paths of the training samples, including: The basic weights are adjusted by combining the ontology coverage, risk level and path complexity of the ontology path of the training samples, and the adjusted weights are limited to a preset range to obtain the training weights.

5. The method according to claim 1, characterized in that, Multiple adapters are set on the basic large model to carry parameter increments for different software-defined deception objects or different task types; The software-defined deception objects corresponding to the adapter include honey spots, honey arrays, honey holes, and honey gardens; The adapter corresponds to the following task types: attack path reasoning, observation event revelation, defense action generation, and security boundary constraints.

6. The method according to claim 1, characterized in that, The routing weights of each adapter are calculated by combining the ontology coverage vector and auxiliary features of the training samples, including: After concatenating the ontology coverage vector and the auxiliary features, a linear transformation is performed. The result of the linear transformation is then subjected to bias shift and normalization to obtain the routing weight. The auxiliary features include, in order, the one-hot encoding of the task type, the risk level, the path complexity, the normalized value of the number of effective paths, the average confidence level of entity relationships, and the security boundary indication value.

7. The method according to claim 1, characterized in that, A consistency score is applied to the input and output of the basic large model, including: The output of the basic large model is subjected to entity recognition and relation extraction and mapped to a heterogeneous ontology graph to obtain the output ontology path; The matching degree of the input ontology path and the output ontology path on key nodes and relation edges is evaluated separately. The degree to which the output of the basic large model satisfies the safety boundary node constraints is evaluated. The consistency score is obtained by weighted summation of the scores.

8. The method according to claim 1, characterized in that, The consistency score is transformed into a feedback loss or preference samples as training feedback, including: The consistency score is compared with preset pass and rewrite thresholds; When the consistency score is lower than the preset pass threshold, the consistency score is converted into feedback loss and added to the adapter training objective; When the consistency score is greater than the preset rewrite threshold but lower than the preset pass threshold, the ontology path is rewritten based on the input, and the output of the basic large model and the rewritten result are combined to form a preference sample for retraining.