Knowledge editing-based large-model adaptive detoxification method and device

By building a large model containing toxicity detection and feature perception modules, and using parameter isolation editing strategies and gradient optimization methods, the data dependence and high computational cost of the detoxication technology of large language models in the existing technology is solved, and efficient and accurate detoxication effect is achieved, while retaining the general capabilities of the model.

CN120105422AActive Publication Date: 2025-06-06HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Application Number
CN202510582609.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The detoxication technology of existing large language models has problems such as strong data dependence, high computing cost and insufficient defense flexibility, and it is difficult to balance the detoxication effect with the retention of the model's general capabilities.

Method used

Adaptive anti-toxicity method of large-model based on knowledge editing is adopted to build a model including toxicity detection module, toxicity feature perception module, anti-toxicity feedforward calculation module and large language model. Through parameter isolation editing strategies and gradient optimization methods, the feedforward network of the model is optimized to achieve detoxicity.

Benefits of technology

It achieves the accuracy and efficiency of detoxication without damaging the general capabilities of the model, overcomes the dependence of the existing technology on entity positioning and high-quality data, and reduces the risk of performance degradation caused by over-editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105422A_ABST
    Figure CN120105422A_ABST
Patent Text Reader

Abstract

The invention provides a knowledge editing-based large-model adaptive detoxification method and device, and relates to the technical field of natural language processing. The method comprises the following steps: constructing a knowledge editing-based large-model adaptive detoxification model; inputting the constructed training set into a large language model to obtain a feature vector; constructing training sample data according to the feature vectors; training a toxicity detection module according to the training sample data to obtain a trained toxicity detection module; inserting the trained toxicity detection module into a target layer of the large language model, and sensing through a toxicity feature sensing module to obtain a classification result; inputting the classification result into an anti-toxicity feedforward calculation module, and determining a parameter matrix to be detoxified through a parameter isolation editing strategy; optimizing by adopting a gradient optimization method and a knowledge editing method to obtain a detoxified feedforward network; and according to the detoxified feedforward network, obtaining a detoxified large language model. According to the invention, the calculation cost can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a large model adaptive detoxification method and device based on knowledge editing. Background Art

[0002] With the widespread application of large language models such as ChatGPT, Llama, and Deepseek in language understanding and reasoning tasks, their security alignment has become a research focus. To ensure that the model output meets the principles of usefulness, honesty, and harmlessness, existing technologies mainly align human values ​​through supervised fine-tuning, reinforcement learning based on human feedback, and direct preference optimization security training methods. However, even models that have been securely aligned are still threatened by malicious prompts or jailbreak attacks, and may therefore generate harmful content.

[0003] The existing detoxification technologies for large language models include: traditional detoxification methods based on parameter optimization, filtering methods based on toxicity detection enhancement, defense methods based on prompt engineering, and parameter correction methods based on knowledge editing. The above methods have achieved certain results in specific scenarios, but they all have significant limitations and it is difficult to balance the detoxification effect with the retention of the general capabilities of the model. Among them, the traditional detoxification method based on parameter optimization directly enhances security by adjusting model parameters, but it needs to rely on a large amount of labeled data and requires global adjustment of model parameters, resulting in high training costs and potential general capability degradation problems; the filtering method based on toxicity detection enhancement blocks the generation of harmful content by integrating input and output detection mechanisms; however, such methods rely on the accuracy of external detection modules and are prone to legitimate requests being rejected due to misjudgment. In addition, the design of separating the detection module from the model increases the inference delay and makes it difficult to meet the needs of real-time interaction; the defense method based on prompt engineering guides the model to reject harmful requests by designing security prompts. Although there is no need to modify the model parameters, such methods rely on the completeness of the prompt design and are prone to failure in the face of complex adversarial attacks. In addition, prompt engineering may cause the model to be overly sensitive to legitimate requests; the parameter correction method based on knowledge editing achieves detoxification by locally modifying model parameters, relies on explicit entities to locate the editing area, and is difficult to cope with adversarial inputs without clear entities. In addition, parameter correction can easily cause the model's general capabilities to degrade.

[0004] Although the above methods have their own advantages in detoxifying large language models, their core problem is that they cannot dynamically distinguish between harmful and harmless requests, resulting in a lack of targeted detoxification. Parameter optimization and knowledge editing methods are prone to over-editing due to global or coarse-grained modifications, while detection enhancement and prompt engineering methods are limited by the robustness of external modules. The above methods also have the defects of strong data dependence, high computational cost and insufficient defense flexibility. Summary of the invention

[0005] In order to solve the technical problems of the existing technology, such as strong data dependence, high computing cost and insufficient defense flexibility, the embodiment of the present invention provides a large model adaptive detoxification method and device based on knowledge editing. The technical solution is as follows:

[0006] On the one hand, a large model adaptive detoxification method based on knowledge editing is provided, the method is implemented by a large model adaptive detoxification device based on knowledge editing, and the method includes:

[0007] S1. Construct a large-model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module and a large language model;

[0008] S2, constructing a training set including harmful prompts and harmless prompts; according to the training set, setting prefix system safety prompts, and constructing input data of a large language model; inputting the input data into the large language model, and using the generated last hidden state as a feature vector; constructing training sample data according to the feature vector and the obtained annotated toxicity label;

[0009] S3. Train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, perform toxicity feature perception through the toxicity feature perception module, and obtain a classification result of toxicity feature perception;

[0010] S4. Input the classification results of toxic feature perception into the anti-toxic feedforward calculation module, and determine the parameter matrix to be detoxified of the feedforward network in the large language model through the parameter isolation editing strategy; according to the parameter matrix to be detoxified, use the gradient optimization method and the knowledge editing method to optimize and obtain the feedforward network of the large language model after detoxification; according to the feedforward network of the large language model after detoxification, obtain the detoxified large language model.

[0011] Optionally, the S3 inserts the trained toxicity detection module into the target layer of the large language model, performs toxicity feature perception through the toxicity feature perception module, and obtains the classification result of toxicity feature perception, including:

[0012] The trained toxicity detection module is inserted into the target layer of the large language model, and the toxicity feature perception is performed on the given user query to obtain the toxicity feature perception classification result. The specific process of toxicity feature perception is expressed by the following formula (1):

[0013] (1)

[0014] in, classification results representing the perception of toxicity characteristics; Table Parameters of the toxicity feature perception module; Represents the hidden state of the last position of the target layer in the large language model; Represents a classifier.

[0015] Optionally, the large-model adaptive detoxification model based on knowledge editing further includes: a dynamic routing module and an original feedforward calculation module;

[0016] The dynamic routing module is used to make a judgment based on the classification result of the toxicity characteristic perception, obtain a judgment result, and switch the calculation path based on the judgment result;

[0017] When the classification result of the toxicity characteristic perception is -1, it is judged to be safe, and the classification result of the toxicity characteristic perception is input into the original feedforward calculation module for calculation;

[0018] When the classification result of the toxicity characteristic perception is +1, it is judged as unsafe, and the classification result of the toxicity characteristic perception is input into the anti-toxicity feedforward calculation module for calculation;

[0019] The original feedforward calculation module is used to process normal queries of users.

[0020] Optionally, the specific calculation process of the dynamic routing module is expressed by the following formula (2):

[0021] (2)

[0022] in, Represents the hidden state of the layer after the target layer in the large language model; Represents the activation value of the target layer after passing through the first MLP parameter matrix; Represents the second MLP parameter matrix in the target layer anti-toxic feedforward calculation module; Represents the second MLP parameter matrix in the original feed-forward network of the target layer.

[0023] Optionally, the process of determining the parameter matrix to be detoxified of the feedforward network in the large language model by using the parameter isolation editing strategy in S4 is represented by the following formulas (3)-(5):

[0024] (3)

[0025] (4)

[0026] (5)

[0027] in, Represents the initial hidden state in the large language model; represents the application of embedding calculation to input X; Indicates Hidden state of the layer; represents a feed-forward network; Represents the hidden state of the l-1th layer in the large language model; Indicates that the hidden state of the l-1th layer in the large language model is calculated using the attention module; Indicates that at layer l, the feedforward calculation module is applied to the input x for calculation; Represents the activation value after passing through the first MLP parameter matrix; represents the second MLP parameter matrix in the feedforward network; Indicates transposing the input x; Represents the first MLP parameter matrix in the feedforward network; represents a non-linear activation function.

[0028] Optionally, the feedforward network of the large language model after detoxification is optimized by using a gradient optimization method and a knowledge editing method according to the parameter matrix to be detoxified, and comprises:

[0029] Get a harmful prompt and a safety response corresponding to the harmful prompt;

[0030] The acquired harmful prompts and the safety responses corresponding to the harmful prompts are input into the large language model for knowledge editing, and the loss function is constructed by setting the prefix system safety prompts;

[0031] The loss function is expressed by the following formula (6):

[0032] (6)

[0033] in, represents the loss function; Indicates harmful hints; Indicates the safety response corresponding to the harmful prompt; represents the probability of the large language model generating a safe response at step t; Indicates the set prefix system security prompt; Represents the parameters of the large language model at step t;

[0034] According to the loss function, the parameters of the large language model are edited in the back propagation to obtain a detoxified matrix; according to the detoxified matrix, a detoxified feedforward network of the large language model is obtained.

[0035] Optionally, the process of editing the parameters of the large language model in the back propagation according to the loss function is expressed by the following formula (7):

[0036] (7)

[0037] in, represents the parameters of the large language model at the tth time step; Toxicity layer Parameters of the toxic zone; represents the tth time step The gradient of Represents the parameters of the large language model at the t+1th time step.

[0038] On the other hand, a large model adaptive detoxification device based on knowledge editing is provided, and the device is applied to a large model adaptive detoxification method based on knowledge editing, and the device comprises:

[0039] The first construction unit is used to construct a large-model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module and a large language model;

[0040] The second construction unit is used to construct a training set including harmful prompts and harmless prompts; according to the training set, set the prefix system safety prompt to construct input data of the large language model; input the input data into the large language model, and use the generated last hidden state as a feature vector; and construct training sample data according to the feature vector and the obtained annotated toxicity label;

[0041] The first acquisition unit is used to train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, perform toxicity feature perception through the toxicity feature perception module, and obtain a classification result of the toxicity feature perception;

[0042] The second acquisition unit is used to input the classification results of toxic feature perception into the anti-toxic feedforward calculation module, and determine the parameter matrix to be detoxified of the feedforward network in the large language model through the parameter isolation editing strategy; according to the parameter matrix to be detoxified, the gradient optimization method and the knowledge editing method are used for optimization to obtain the feedforward network of the large language model after detoxification; according to the feedforward network of the large language model after detoxification, the large language model after detoxification is obtained.

[0043] Optionally, the first acquiring unit is used to:

[0044] The trained toxicity detection module is inserted into the target layer of the large language model, and the toxicity feature perception is performed on the given user query to obtain the toxicity feature perception classification result. The specific process of toxicity feature perception is expressed by the following formula (1):

[0045] (1)

[0046] in, classification results representing the perception of toxicity characteristics; are the parameters of the toxicity feature perception module; Represents the hidden state of the last position of the target layer in the large language model; Represents a classifier.

[0047] Optionally, the large-model adaptive detoxification model based on knowledge editing further includes: a dynamic routing module and an original feedforward calculation module;

[0048] The dynamic routing module is used to make a judgment based on the classification result of the toxicity characteristic perception, obtain a judgment result, and switch the calculation path based on the judgment result;

[0049] When the classification result of the toxicity characteristic perception is -1, it is judged to be safe, and the classification result of the toxicity characteristic perception is input into the original feedforward calculation module for calculation;

[0050] When the classification result of the toxicity characteristic perception is +1, it is judged as unsafe, and the classification result of the toxicity characteristic perception is input into the anti-toxicity feedforward calculation module for calculation;

[0051] The original feedforward calculation module is used to process normal queries of users.

[0052] Optionally, the specific calculation process of the dynamic routing module is expressed by the following formula (2):

[0053] (2)

[0054] in, Represents the hidden state of the layer after the target layer in the large language model; Represents the activation value of the target layer after passing through the first MLP parameter matrix; Represents the second MLP parameter matrix in the target layer anti-toxic feedforward calculation module; Represents the second MLP parameter matrix in the original feed-forward network of the target layer.

[0055] Optionally, the process of determining the parameter matrix to be detoxified of the feedforward network in the large language model by using the parameter isolation editing strategy is represented by the following formulas (3)-(5):

[0056] (3)

[0057] (4)

[0058] (5)

[0059] in, Represents the initial hidden state in the large language model; represents the application of embedding calculation to input X; Indicates Hidden state of the layer; represents a feed-forward network; Represents the hidden state of the l-1th layer in the large language model; Indicates that the hidden state of the l-1th layer in the large language model is calculated using the attention module; Indicates that at layer l, the feedforward calculation module is applied to the input x for calculation; Represents the activation value after passing through the first MLP parameter matrix; represents the second MLP parameter matrix in the feedforward network; Indicates transposing the input x; Represents the first MLP parameter matrix in the feedforward network; represents a non-linear activation function.

[0060] Optionally, the feedforward network of the large language model after detoxification is optimized by using a gradient optimization method and a knowledge editing method according to the parameter matrix to be detoxified, and comprises:

[0061] Get a harmful prompt and a safety response corresponding to the harmful prompt;

[0062] The acquired harmful prompts and the safety responses corresponding to the harmful prompts are input into the large language model for knowledge editing, and the loss function is constructed by setting the prefix system safety prompts;

[0063] The loss function is expressed by the following formula (6):

[0064] (6)

[0065] in, represents the loss function; Indicates harmful hints; Indicates the safety response corresponding to the harmful prompt; represents the probability of the large language model generating a safe response at step t; Indicates the set prefix system security prompt; Represents the parameters of the large language model at step t;

[0066] According to the loss function, the parameters of the large language model are edited in the back propagation to obtain a detoxified matrix; according to the detoxified matrix, a detoxified feedforward network of the large language model is obtained.

[0067] Optionally, the process of editing the parameters of the large language model in the back propagation according to the loss function is expressed by the following formula (7):

[0068] (7)

[0069] in, represents the parameters of the large language model at the tth time step; Toxicity layer Parameters of the toxic zone; represents the tth time step The gradient of Represents the parameters of the large language model at the t+1th time step.

[0070] On the other hand, a large model adaptive detoxification device based on knowledge editing is provided, and the large model adaptive detoxification device based on knowledge editing comprises: a processor; a memory, wherein computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, any one of the large model adaptive detoxification methods based on knowledge editing is implemented.

[0071] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned large model adaptive detoxification methods based on knowledge editing.

[0072] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0073] The embodiment of the present invention first constructs a large-model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module and a large language model; a training set including harmful prompts and harmless prompts is constructed; according to the training set, a prefix system security prompt is set to construct input data of the large language model; secondly, the input data is input into the large language model, and the generated last hidden state is used as a feature vector; according to the feature vector and the obtained annotated toxicity label, training sample data is constructed; according to the training sample data, the toxicity detection module is trained to obtain a trained toxicity detection module; the trained toxicity detection module is inserted into the target layer of the large language model, and the toxicity feature perception is performed through the toxicity feature perception module to obtain a classification result of the toxicity feature perception; the classification result of the toxicity feature perception is input into the anti-toxicity feedforward calculation module, and the parameter matrix to be detoxified of the feedforward network in the large language model is determined through a parameter isolation editing strategy; finally, according to the parameter matrix to be detoxified, a gradient optimization method and a knowledge editing method are used for optimization to obtain a feedforward network of the detoxified large language model; according to the feedforward network of the detoxified large language model, a detoxified large language model is obtained.

[0074] The embodiments of the present invention achieve a balance between accuracy, efficiency and versatility in the field of large language model detoxification, overcome the core problems of prior art's reliance on entity positioning and high-quality data, and performance degradation caused by over-editing, and provide a new technical paradigm for building safe and reliable large-scale language models. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0076] Figure 1 This is a flow chart of a large model adaptive detoxification method based on knowledge editing provided by an embodiment of the present invention;

[0077] Figure 2 It is a schematic diagram of the overall structure of a framework for adaptive detoxification of a large model based on knowledge editing provided by an embodiment of the present invention;

[0078] Figure 3 It is a block diagram of a large model adaptive detoxification device based on knowledge editing provided by an embodiment of the present invention;

[0079] Figure 4 It is a structural schematic diagram of a large-model adaptive detoxification device based on knowledge editing provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0080] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0081] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0082] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.

[0083] In the embodiments of the present invention, sometimes the subscripts such as W 1 It may be written in non-subscript form such as W1. When the difference is not emphasized, the meaning is the same.

[0084] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0085] The embodiment of the present invention provides a large model adaptive detoxification method based on knowledge editing, which can be implemented by a large model adaptive detoxification device based on knowledge editing, and the large model adaptive detoxification device based on knowledge editing can be a terminal or a server. Figure 1 The flowchart of the large model adaptive detoxification method based on knowledge editing is shown. The processing flow of the method may include the following steps:

[0086] S1. Construct a large-model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module and a large language model.

[0087] Among them, the anti-toxic feedforward calculation module is responsible for eliminating harmful information of the large language model.

[0088] Optionally, the large model adaptive detoxification model based on knowledge editing further includes: a dynamic routing module and an original feedforward calculation module;

[0089] A dynamic routing module is used to make a judgment based on the classification result of the toxicity characteristic perception, obtain the judgment result, and switch the calculation path according to the judgment result;

[0090] Among them, the dynamic routing module is set before the feedforward network of the large language model, and the data flow is dynamically guided to different feedforward networks to achieve adaptive detoxification of user input.

[0091] When the classification result of the toxicity characteristic perception is -1, it is judged to be safe, and the classification result of the toxicity characteristic perception is input into the original feedforward calculation module for calculation;

[0092] When the classification result of the toxicity characteristic perception is +1, it is judged as unsafe, and the classification result of the toxicity characteristic perception is input into the anti-toxicity feedforward calculation module for calculation;

[0093] The original feedforward calculation module is used to process normal user queries.

[0094] S2. Construct a training set containing harmful prompts and harmless prompts; according to the training set, set the prefix system security prompts to construct the input data of the large language model; input the input data into the large language model, and use the generated last hidden state as the feature vector; construct the training sample data according to the feature vector and the obtained annotated toxicity label.

[0095] In a feasible implementation, a training set including 4,000 harmful prompts and 2,000 harmless prompts is constructed; for each training sample, a prefix system safety prompt is set to constitute the input data of the large language model.

[0096] S3. Train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, perform toxicity feature perception through the toxicity feature perception module, and obtain a classification result of toxicity feature perception.

[0097] The obtained annotated toxicity labels are manually annotated labels; the feature vectors and the manually annotated labels are input into the toxicity detection module, and the module is trained by a linear kernel support vector machine to obtain a trained toxicity detection module.

[0098] In a feasible implementation, a validation set is obtained, and based on the obtained validation set, the trained toxicity detection module is verified by selecting the F1 score as an indicator, and the layer corresponding to the toxicity detection module with the best indicator is selected as the insertion layer.

[0099] In a feasible implementation, given a user's query Q, the user's query generates a last hidden state in the target layer of the large language model, and the last hidden state is input into the toxicity feature perception module for judgment to obtain a toxicity feature perception classification result.

[0100] Optionally, S3 inserts the trained toxicity detection module into the target layer of the large language model, performs toxicity feature perception through the toxicity feature perception module, and obtains the classification result of toxicity feature perception, including:

[0101] The trained toxicity detection module is inserted into the target layer of the large language model, and the toxicity feature perception is performed on the given user query to obtain the toxicity feature perception classification result. The specific process of toxicity feature perception is expressed by the following formula (1):

[0102] (1)

[0103] in, classification results representing the perception of toxicity characteristics; are the parameters of the toxicity feature perception module; Represents the hidden state of the last position of the target layer in the large language model; Represents a classifier.

[0104] Among them, the classification result of toxicity feature perception is transmitted to the anti-toxicity feedforward calculation module as a routing signal, triggering the calculation path switching, so that the toxicity feature perception module completes toxicity judgment in the early stage of forward propagation, avoiding subsequent inter-layer interference.

[0105] Among them, calculating the path through the dynamic routing module can effectively deal with adversarial inputs without entity dependence, and can improve the detoxification coverage and robustness.

[0106] Optionally, the specific calculation process of the dynamic routing module is expressed by the following formula (2):

[0107] (2)

[0108] in, Represents the hidden state of the layer after the target layer in the large language model; Represents the activation value of the target layer after passing through the first MLP parameter matrix; Represents the second MLP parameter matrix in the target layer anti-toxic feedforward calculation module; Represents the second MLP parameter matrix in the original feed-forward network of the target layer.

[0109] S4. Input the classification results of toxic feature perception into the anti-toxic feedforward calculation module, and determine the parameter matrix to be detoxified of the feedforward network in the large language model through the parameter isolation editing strategy; according to the parameter matrix to be detoxified, use the gradient optimization method and the knowledge editing method to optimize and obtain the feedforward network of the large language model after detoxification; according to the feedforward network of the large language model after detoxification, obtain the detoxified large language model.

[0110] Among them, the anti-toxic feedforward calculation module adopts a parameter isolation editing strategy to ensure that the parameters that need to be edited are isolated from the original parameters of the large language model, ensuring that the general capabilities of the large language model are not destroyed.

[0111] The basic structure of the large language model is a parameterized function, which includes an embedding matrix and L cascaded Transformer layers. Each layer includes a multi-head attention mechanism and a feedforward network. In a feasible implementation, given an input sequence, the input sequence is input into the large language model, and the parameter matrix to be detoxified of the feedforward network in the large language model is determined through the parameter isolation editing strategy.

[0112] Optionally, the process of determining the parameter matrix to be detoxified of the feedforward network in the large language model through the parameter isolation editing strategy in S4 is represented by the following formulas (3)-(5):

[0113] (3)

[0114] (4)

[0115] (5)

[0116] in, Represents the initial hidden state in the large language model; represents the application of embedding calculation to input X; Indicates Hidden state of the layer; represents a feed-forward network; Represents the hidden state of the l-1th layer in the large language model; Indicates that the hidden state of the l-1th layer in the large language model is calculated using the attention module; Indicates that at layer l, the feedforward calculation module is applied to the input x for calculation; Represents the activation value after passing through the first MLP parameter matrix; represents the second MLP parameter matrix in the feedforward network; Indicates transposing the input x; Represents the first MLP parameter matrix in the feedforward network; represents a non-linear activation function.

[0117] Among them, the embodiment of the present invention selects the second MLP parameter matrix of the target layer as the editing object; creates a copy of the second MLP parameter matrix and inserts it into the target layer to optimize the detoxification parameters.

[0118] Optionally, according to the parameter matrix to be detoxified, a gradient optimization method and a knowledge editing method are used for optimization to obtain a feedforward network of a detoxified large language model, including:

[0119] Get a harmful prompt and a safety response corresponding to the harmful prompt;

[0120] The acquired harmful prompts and the safety responses corresponding to the harmful prompts are input into the large language model for knowledge editing, and the loss function is constructed by setting the prefix system safety prompts;

[0121] In the process of building the loss function, all parameters of the large language model need to be frozen.

[0122] The loss function is expressed by the following formula (6):

[0123] (6)

[0124] in, represents the loss function; Indicates harmful hints; Indicates the safety response corresponding to the harmful prompt; represents the probability of the large language model generating a safe response at step t; Indicates the set prefix system security prompt; Represents the parameters of the large language model at step t;

[0125] According to the loss function, the parameters of the large language model are edited in the back propagation to obtain a detoxified matrix; according to the detoxified matrix, a detoxified feedforward network of the large language model is obtained.

[0126] Among them, the anti-toxicity feedforward calculation module naturally isolates the detoxification influence domain through the routing mechanism, and only needs to optimize the parameter offset of the harmful matrix to reduce the risk of over-editing.

[0127] Among them, the routing mechanism refers to directing the data flow to different modules according to an input signal. In this application, if the signal is judged to be harmful, the dynamic routing module directs the data flow to the anti-toxic feedforward calculation module. Otherwise, if the signal is judged to be harmless, it is directed to the original feedforward calculation module.

[0128] Optionally, the process of editing the parameters of the large language model in the back propagation according to the loss function is expressed by the following formula (7):

[0129] (7)

[0130] in, represents the parameters of the large language model at the tth time step; Toxicity layer Parameters of the toxic zone; represents the tth time step The gradient of Represents the parameters of the large language model at the t+1th time step.

[0131] Among them, the embodiment of the present invention only needs to edit some parameters of a single feedforward network, with short training time and low overhead; the embodiment of the present invention is applicable to large models of various series with different parameter magnitudes, such as the Llama large language model and the Mistral large language model; the implementation of the present invention can handle adversarial input without entity dependency through an inter-layer routing mechanism, and does not rely on high-quality training data to fine-tune the parameters of the large language model.

[0132] Among them, the embodiment of the present invention introduces a dynamic routing mechanism into the field of knowledge editing for detoxification of large language models, breaking through the reliance of traditional methods on entity positioning; by designing a parameter isolation strategy, precise control of the detoxification impact domain is achieved, and the general capabilities of large language models are retained to the maximum extent. The embodiment of the present invention achieves a good balance between detoxification performance and capability retention, and provides a new technical paradigm for the safe deployment of large language models.

[0133] Among them, Figure 2 The figure shows a schematic diagram of the overall structure of a framework for adaptive detoxification of a large model based on knowledge editing provided by an embodiment of the present invention; in a feasible implementation manner, the trained toxicity detection module, toxicity feature perception module and anti-toxicity feedforward calculation module are inserted into the target layer of the large language model; when the user's query is transmitted to the large language model, the toxicity semantic feature detection module determines whether the query is toxic according to the hidden state of the target layer, and transmits the judgment signal to the dynamic routing module inserted after the attention layer and the normalization layer; wherein the toxicity semantic feature detection module includes: a toxicity detection module and a toxicity feature perception module; the dynamic routing module judges the signal, and when the signal is judged to be unsafe, the information flow is directed to the anti-toxicity feedforward calculation module, the information flow is detoxified, and the user query is refused to be responded; when the signal is judged to be safe, the information flow is directed to the original feedforward calculation module, and the user query is processed normally.

[0134] The embodiment of the present invention first constructs a large-model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module and a large language model; a training set including harmful prompts and harmless prompts is constructed; according to the training set, a prefix system security prompt is set to construct input data of the large language model; secondly, the input data is input into the large language model, and the generated last hidden state is used as a feature vector; according to the feature vector and the obtained annotated toxicity label, training sample data is constructed; according to the training sample data, the toxicity detection module is trained to obtain a trained toxicity detection module; the trained toxicity detection module is inserted into the target layer of the large language model, and the toxicity feature perception is performed through the toxicity feature perception module to obtain a classification result of the toxicity feature perception; the classification result of the toxicity feature perception is input into the anti-toxicity feedforward calculation module, and the parameter matrix to be detoxified of the feedforward network in the large language model is determined through a parameter isolation editing strategy; finally, according to the parameter matrix to be detoxified, a gradient optimization method and a knowledge editing method are used for optimization to obtain a feedforward network of the detoxified large language model; according to the feedforward network of the detoxified large language model, a detoxified large language model is obtained.

[0135] The embodiments of the present invention achieve a balance between accuracy, efficiency and versatility in the field of large language model detoxification, overcome the core problems of prior art's reliance on entity positioning and high-quality data, and performance degradation caused by over-editing, and provide a new technical paradigm for building safe and reliable large-scale language models.

[0136] Figure 3 1 is a block diagram of a large model adaptive detoxification device based on knowledge editing according to an exemplary embodiment, and the device is used in a large model adaptive detoxification method based on knowledge editing. Figure 3 The device includes a first construction unit 310, a second construction unit 320, a first acquisition unit 330 and a second acquisition unit 340. Wherein:

[0137] The first construction unit 310 is used to construct a large-model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module and a large language model;

[0138] The second construction unit 320 is used to construct a training set including harmful prompts and harmless prompts; according to the training set, set the prefix system safety prompt to construct input data of the large language model; input the input data into the large language model, and use the generated last hidden state as a feature vector; construct training sample data according to the feature vector and the obtained annotated toxicity label;

[0139] The first acquisition unit 330 is used to train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, perform toxicity feature perception through the toxicity feature perception module, and obtain a classification result of toxicity feature perception;

[0140] The second acquisition unit 340 is used to input the classification result of toxic feature perception into the anti-toxic feedforward calculation module, and determine the parameter matrix to be detoxified of the feedforward network in the large language model through the parameter isolation editing strategy; according to the parameter matrix to be detoxified, a gradient optimization method and a knowledge editing method are used for optimization to obtain the feedforward network of the large language model after detoxification; according to the feedforward network of the large language model after detoxification, the detoxified large language model is obtained.

[0141] Optionally, the first acquiring unit 330 is configured to:

[0142] The trained toxicity detection module is inserted into the target layer of the large language model, and the toxicity feature perception is performed on the given user query to obtain the toxicity feature perception classification result. The specific process of toxicity feature perception is expressed by the following formula (1):

[0143] (1)

[0144] in, classification results representing the perception of toxicity characteristics; are the parameters of the toxicity feature perception module; Represents the hidden state of the last position of the target layer in the large language model; Represents a classifier.

[0145] Optionally, the large-model adaptive detoxification model based on knowledge editing further includes: a dynamic routing module and an original feedforward calculation module;

[0146] The dynamic routing module is used to make a judgment based on the classification result of the toxicity characteristic perception, obtain a judgment result, and switch the calculation path based on the judgment result;

[0147] When the classification result of the toxicity characteristic perception is -1, it is judged to be safe, and the classification result of the toxicity characteristic perception is input into the original feedforward calculation module for calculation;

[0148] When the classification result of the toxicity characteristic perception is +1, it is judged as unsafe, and the classification result of the toxicity characteristic perception is input into the anti-toxicity feedforward calculation module for calculation;

[0149] The original feedforward calculation module is used to process normal queries of users.

[0150] Optionally, the specific calculation process of the dynamic routing module is expressed by the following formula (2):

[0151] (2)

[0152] in, Represents the hidden state of the layer after the target layer in the large language model; Represents the activation value of the target layer after passing through the first MLP parameter matrix; Represents the second MLP parameter matrix in the target layer anti-toxic feedforward calculation module; Represents the second MLP parameter matrix in the original feed-forward network of the target layer.

[0153] Optionally, the process of determining the parameter matrix to be detoxified of the feedforward network in the large language model by using the parameter isolation editing strategy is represented by the following formulas (3)-(5):

[0154] (3)

[0155] (4)

[0156] (5)

[0157] in, Represents the initial hidden state in the large language model; represents the application of embedding calculation to input X; Indicates Hidden state of the layer; represents a feed-forward network; Represents the hidden state of the l-1th layer in the large language model; Indicates that the hidden state of the l-1th layer in the large language model is calculated using the attention module; Indicates that at layer l, the feedforward calculation module is applied to the input x for calculation; Represents the activation value after passing through the first MLP parameter matrix; represents the second MLP parameter matrix in the feedforward network; Indicates transposing the input x; Represents the first MLP parameter matrix in the feedforward network; represents a non-linear activation function.

[0158] Optionally, the feedforward network of the large language model after detoxification is optimized by using a gradient optimization method and a knowledge editing method according to the parameter matrix to be detoxified, and comprises:

[0159] Get a harmful prompt and a safety response corresponding to the harmful prompt;

[0160] The acquired harmful prompts and the safety responses corresponding to the harmful prompts are input into the large language model for knowledge editing, and the loss function is constructed by setting the prefix system safety prompts;

[0161] The loss function is expressed by the following formula (6):

[0162] (6)

[0163] in, represents the loss function; Indicates harmful hints; Indicates the safety response corresponding to the harmful prompt; represents the probability of the large language model generating a safe response at step t; Indicates the set prefix system security prompt; Represents the parameters of the large language model at step t;

[0164] According to the loss function, the parameters of the large language model are edited in the back propagation to obtain a detoxified matrix; according to the detoxified matrix, a detoxified feedforward network of the large language model is obtained.

[0165] Optionally, the process of editing the parameters of the large language model in the back propagation according to the loss function is expressed by the following formula (7):

[0166] (7)

[0167] in, represents the parameters of the large language model at the tth time step; Toxicity layer Parameters of the toxic zone; represents the tth time step The gradient of Represents the parameters of the large language model at the t+1th time step.

[0168] The embodiment of the present invention first constructs a large-model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module and a large language model; a training set including harmful prompts and harmless prompts is constructed; according to the training set, a prefix system security prompt is set to construct input data of the large language model; secondly, the input data is input into the large language model, and the generated last hidden state is used as a feature vector; according to the feature vector and the obtained annotated toxicity label, training sample data is constructed; according to the training sample data, the toxicity detection module is trained to obtain a trained toxicity detection module; the trained toxicity detection module is inserted into the target layer of the large language model, and the toxicity feature perception is performed through the toxicity feature perception module to obtain a classification result of the toxicity feature perception; the classification result of the toxicity feature perception is input into the anti-toxicity feedforward calculation module, and the parameter matrix to be detoxified of the feedforward network in the large language model is determined through a parameter isolation editing strategy; finally, according to the parameter matrix to be detoxified, a gradient optimization method and a knowledge editing method are used for optimization to obtain a feedforward network of the detoxified large language model; according to the feedforward network of the detoxified large language model, a detoxified large language model is obtained.

[0169] The embodiments of the present invention achieve a balance between accuracy, efficiency and versatility in the field of large language model detoxification, overcome the core problems of prior art's reliance on entity positioning and high-quality data, and performance degradation caused by over-editing, and provide a new technical paradigm for building safe and reliable large-scale language models.

[0170] Figure 4 is a schematic diagram of the structure of a large model adaptive virus removal device based on knowledge editing provided by an embodiment of the present invention, such as Figure 4 As shown, the large model adaptive detoxification device based on knowledge editing may include the above Figure 3 The large model adaptive detoxification device based on knowledge editing is shown. Optionally, the large model adaptive detoxification device based on knowledge editing 410 may include a first processor 2001 .

[0171] Optionally, the large model adaptive detoxification device 410 based on knowledge editing may also include a memory 2002 and a transceiver 2003 .

[0172] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0173] Combine the following Figure 4 The components of the large model adaptive detoxification device 410 based on knowledge editing are introduced in detail:

[0174] The first processor 2001 is the control center of the large model adaptive detoxification device 410 based on knowledge editing, which can be a processor or a general term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiment of the present invention, such as one or more microprocessors (digital signal processor, DSP), or one or more field programmable gate arrays (field programmable gate array, FPGA).

[0175] Optionally, the first processor 2001 can perform various functions of the large model adaptive detoxification device 410 based on knowledge editing by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.

[0176] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 CPU0 and CPU1 are shown in FIG.

[0177] In a specific implementation, as an embodiment, the large model adaptive detoxification device 410 based on knowledge editing may also include multiple processors, such as Figure 4 The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0178] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled to be executed by the first processor 2001. The specific implementation method can refer to the above method embodiment, which will not be repeated here.

[0179] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001, or may exist independently, and may be connected to the first processor 2001 through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0180] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0181] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 4 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0182] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently, and may be connected to the first processor 2001 through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0183] It should be noted that Figure 4 The structure of the large model adaptive detoxification device 410 based on knowledge editing shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0184] In addition, the technical effects of the large model adaptive detoxification device 410 based on knowledge editing can refer to the technical effects of the large model adaptive detoxification method based on knowledge editing described in the above method embodiment, and will not be repeated here.

[0185] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0186] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0187] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0188] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0189] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0190] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0191] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0192] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0193] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0194] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0195] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0196] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0197] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A large model adaptive detoxification method based on knowledge editing, characterized in that: The method comprises: S1. Construct a large-model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module and a large language model; S2, constructing a training set including harmful prompts and harmless prompts; according to the training set, setting prefix system safety prompts, and constructing input data of a large language model; inputting the input data into the large language model, and using the generated last hidden state as a feature vector; constructing training sample data according to the feature vector and the obtained annotated toxicity label; S3. Train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, perform toxicity feature perception through the toxicity feature perception module, and obtain a classification result of toxicity feature perception; S4. Input the classification results of toxic feature perception into the anti-toxic feedforward calculation module, and determine the parameter matrix to be detoxified of the feedforward network in the large language model through the parameter isolation editing strategy; according to the parameter matrix to be detoxified, use the gradient optimization method and the knowledge editing method to optimize and obtain the feedforward network of the large language model after detoxification; according to the feedforward network of the large language model after detoxification, obtain the detoxified large language model.

2. The large model adaptive detoxification method based on knowledge editing according to claim 1 is characterized in that: The S3 inserts the trained toxicity detection module into the target layer of the large language model, performs toxicity feature perception through the toxicity feature perception module, and obtains the classification result of toxicity feature perception, including: The trained toxicity detection module is inserted into the target layer of the large language model, and the toxicity feature perception is performed on the given user query to obtain the toxicity feature perception classification result. The specific process of toxicity feature perception is expressed by the following formula (1): (1) in, classification results representing the perception of toxicity characteristics; represents the parameters of the toxicity feature perception module; Represents the hidden state of the last position of the target layer in the large language model; Represents a classifier.

3. The large model adaptive detoxification method based on knowledge editing according to claim 1 is characterized in that: The large-model adaptive detoxification model based on knowledge editing also includes: a dynamic routing module and an original feedforward calculation module; The dynamic routing module is used to make a judgment based on the classification result of the toxicity characteristic perception, obtain a judgment result, and switch the calculation path based on the judgment result; When the classification result of the toxicity characteristic perception is -1, it is judged to be safe, and the classification result of the toxicity characteristic perception is input into the original feedforward calculation module for calculation; When the classification result of the toxicity characteristic perception is +1, it is judged as unsafe, and the classification result of the toxicity characteristic perception is input into the anti-toxicity feedforward calculation module for calculation; The original feedforward calculation module is used to process normal queries of users.

4. The large model adaptive detoxification method based on knowledge editing according to claim 3 is characterized in that: The specific calculation process of the dynamic routing module is expressed by the following formula (2): (2) in, Represents the hidden state of the layer after the target layer in the large language model; Represents the activation value of the target layer after passing through the first MLP parameter matrix; Represents the second MLP parameter matrix in the target layer anti-toxic feedforward calculation module; Represents the second MLP parameter matrix in the original feed-forward network of the target layer.

5. The large model adaptive detoxification method based on knowledge editing according to claim 1 is characterized in that: The process of determining the parameter matrix to be detoxified of the feedforward network in the large language model through the parameter isolation editing strategy in S4 is represented by the following formulas (3)-(5): (3) (4) (5) in, Represents the initial hidden state in the large language model; represents the application of embedding calculation to input X; Indicates Hidden state of the layer; represents a feed-forward network; Represents the hidden state of the l-1th layer in the large language model; Indicates that the hidden state of the l-1th layer in the large language model is calculated using the attention module; Indicates that at layer l, the feedforward calculation module is applied to the input x for calculation; Represents the activation value after passing through the first MLP parameter matrix; represents the second MLP parameter matrix in the feedforward network; Indicates transposing the input x; Represents the first MLP parameter matrix in the feedforward network; represents a non-linear activation function.

6. The large model adaptive detoxification method based on knowledge editing according to claim 1 is characterized in that: The feedforward network of the large language model after detoxification is obtained by optimizing the parameter matrix to be detoxified using a gradient optimization method and a knowledge editing method, including: Get a harmful prompt and a safety response corresponding to the harmful prompt; The acquired harmful prompts and the safety responses corresponding to the harmful prompts are input into the large language model for knowledge editing, and the loss function is constructed by setting the prefix system safety prompts; The loss function is expressed by the following formula (6): (6) in, represents the loss function; Indicates harmful hints; Indicates the safety response corresponding to the harmful prompt; represents the probability of the large language model generating a safe response at step t; Indicates the set prefix system security prompt; Represents the parameters of the large language model at step t; According to the loss function, the parameters of the large language model are edited in the back propagation to obtain a detoxified matrix; according to the detoxified matrix, a detoxified feedforward network of the large language model is obtained.

7. The large model adaptive detoxification method based on knowledge editing according to claim 6 is characterized in that: The process of editing the parameters of the large language model in the back propagation according to the loss function is expressed by the following formula (7): (7) in, represents the parameters of the large language model at the tth time step; Toxicity layer Parameters of the toxic zone; represents the tth time step The gradient of Represents the parameters of the large language model at the t+1th time step.

8. A large model adaptive detoxification device based on knowledge editing, the large model adaptive detoxification device based on knowledge editing is used to implement the large model adaptive detoxification method based on knowledge editing as claimed in any one of claims 1 to 7, characterized in that: The device comprises: The first construction unit is used to construct a large-model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module and a large language model; The second construction unit is used to construct a training set including harmful prompts and harmless prompts; according to the training set, set the prefix system safety prompt to construct input data of the large language model; input the input data into the large language model, and use the generated last hidden state as a feature vector; and construct training sample data according to the feature vector and the obtained annotated toxicity label; The first acquisition unit is used to train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, perform toxicity feature perception through the toxicity feature perception module, and obtain a classification result of the toxicity feature perception; The second acquisition unit is used to input the classification results of toxic feature perception into the anti-toxic feedforward calculation module, and determine the parameter matrix to be detoxified of the feedforward network in the large language model through the parameter isolation editing strategy; according to the parameter matrix to be detoxified, the gradient optimization method and the knowledge editing method are used for optimization to obtain the feedforward network of the large language model after detoxification; according to the feedforward network of the large language model after detoxification, the detoxified large language model is obtained.

9. A large model adaptive virus removal device based on knowledge editing, characterized in that: The large model adaptive virus removal device based on knowledge editing includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal reasoning method and device based on large language model and knowledge graph

    CN118193684A

  • Large language model security vulnerability automatic detection method and device based on scene nesting

    CN119885206A

  • Optimization method and device, reply method and device of large language model, equipment and medium

    CN119918658A

Cited By

  • Large language model detoxification method and device based on Riemannian stream matching

    CN122263897A

  • A Detoxification Method and Device Based on a Large Language Model Using Riemannian Flow Matching

    CN122263897B