A Large Model Adaptive Detoxification Method and Device Based on Knowledge Editing
By building a large-scale model adaptive detoxification method based on knowledge editing, and using the toxicity detection module and dynamic routing module to optimize the feedforward network of the large language model, the problem of insufficient data dependence and defense flexibility of detoxification in the existing technology is solved, and accurate and efficient detoxification effect is achieved.
Patent Information
- Application Number
- CN202510582609.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The detoxication technology of existing large language models cannot dynamically distinguish harmful and harmless requests, resulting in a lack of targeted detoxification process, and there are problems such as strong data dependence, high computing cost and insufficient defense flexibility.
Build a large-model adaptive poison removal method based on knowledge editing, including toxicity detection module, toxicity feature perception module, anti-toxicity feedforward calculation module and large language model. The feedforward network is optimized through gradient optimization and knowledge editing methods. The dynamic routing module is used to judge the calculation path and achieve accurate poison removal.
It has achieved a balance between accuracy, efficiency and universality in the field of detoxification of large language models, overcomes the dependence of entity positioning and high-quality data, and provides a new technical paradigm for safe and reliable large language models.
Smart Images

Figure CN120105422B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a large model adaptive detoxification method and device based on knowledge editing. Background Art
[0002] With the wide application of large language models such as ChatGPT, Llama, and DeepSeek in language understanding and reasoning tasks, their safe alignment issues have become the focus of research. To ensure that the model outputs conform to the principles of usefulness, honesty, and harmlessness, existing technologies mainly align human values through supervised fine-tuning, reinforcement learning based on human feedback, and direct preference optimization for safe training methods. However, even models that have undergone safe alignment still face threats of malicious prompts or jailbreak attacks, and thus may generate harmful content.
[0003] The detoxification techniques of existing large language models include: traditional detoxification methods based on parameter optimization, filtering methods based on enhanced toxicity detection, defense methods based on prompt engineering, and parameter correction methods based on knowledge editing. The above methods have achieved certain effects in specific scenarios, but all have significant limitations and are difficult to balance the detoxification effect and the retention of the general capabilities of the model. Among them, the traditional detoxification method based on parameter optimization directly enhances security by adjusting model parameters, but it requires a large amount of labeled data and global adjustment of model parameters, resulting in high training costs and potential degradation of general capabilities; the filtering method based on enhanced toxicity detection blocks the generation of harmful content by integrating input-output detection mechanisms; however, such methods rely on the accuracy of external detection modules and are prone to rejecting legitimate requests due to misjudgment. In addition, the design of separating the detection module from the model increases the inference latency and is difficult to meet the requirements of real-time interaction; the defense method based on prompt engineering guides the model to reject harmful requests by designing safe prompts. Although it does not require modifying model parameters, such methods rely on the completeness of prompt design and are prone to failure when facing complex adversarial attacks. In addition, prompt engineering may cause the model to be overly sensitive to legitimate requests; the parameter correction method based on knowledge editing achieves detoxification by locally modifying model parameters, relies on explicitly locating the editing area of entities, and is difficult to handle adversarial inputs without clear entities, and parameter correction is prone to causing degradation of the general capabilities of the model.
[0004] Although the above methods have their own advantages in the detoxification of large language models, their core problem is that they cannot dynamically distinguish between harmful and harmless requests, resulting in a lack of pertinence in the detoxification process. Parameter optimization and knowledge editing methods are prone to over-editing due to global or coarse-grained modifications, while detection enhancement and prompt engineering methods are limited by the robustness of external modules. The above methods also have problems such as strong data dependence, high computational costs, and insufficient defense flexibility. Summary of the Invention
[0005] To solve the technical problems of strong data dependence, high computational cost, and insufficient defense flexibility existing in the prior art, an embodiment of the present invention provides a large model adaptive detoxification method and device based on knowledge editing. The technical solution is as follows:
[0006] On the one hand, a large model adaptive detoxification method based on knowledge editing is provided. This method is implemented by a large model adaptive detoxification device based on knowledge editing, and the method includes:
[0007] S1. Construct a large model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module, and a large language model;
[0008] S2. Construct a training set including harmful prompts and harmless prompts; according to the training set, set a prefix system safety prompt to construct the input data of the large language model; input the input data into the large language model, and use the generated last hidden state as a feature vector; construct training sample data according to the feature vector and the obtained labeled toxicity label;
[0009] S3. Train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception through the toxicity feature perception module to obtain the classification result of toxicity feature perception;
[0010] S4. Input the classification result of toxicity feature perception into the anti-toxicity feedforward calculation module, and determine the parameter matrix to be detoxified in the feedforward network of the large language model through a parameter isolation editing strategy; optimize according to the parameter matrix to be detoxified by using a gradient optimization method and a knowledge editing method to obtain the feedforward network of the detoxified large language model; obtain the detoxified large language model according to the feedforward network of the detoxified large language model.
[0011] Optionally, the step of inserting the trained toxicity detection module into the target layer of the large language model in S3 and performing toxicity feature perception through the toxicity feature perception module to obtain the classification result of toxicity feature perception includes:
[0012] Insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception on a given user query to obtain the classification result of toxicity feature perception; where the specific process of toxicity feature perception is represented by the following formula (1):
[0013] (1)
[0014] Where represents the classification result of toxicity feature perception; Parameters of the toxicity feature perception module; Represents the hidden state at the last position of the target layer in the large language model; Represents the classifier.
[0015] Optionally, the large model adaptive detoxification model based on knowledge editing further includes: a dynamic routing module and an original feedforward calculation module;
[0016] The dynamic routing module is used to make a judgment based on the classification result of toxicity feature perception, obtain a judgment result, and switch the calculation path according to the judgment result;
[0017] Among them, when the classification result of toxicity feature perception is -1, it is judged as safe, and the classification result of toxicity feature perception is input into the original feedforward calculation module for calculation;
[0018] When the classification result of toxicity feature perception is +1, it is judged as unsafe, and the classification result of toxicity feature perception is input into the anti-toxicity feedforward calculation module for calculation;
[0019] The original feedforward calculation module is used to process the user's normal query.
[0020] Optionally, the specific calculation process of the dynamic routing module is represented by the following formula (2):
[0021] (2)
[0022] Among them, Represents the hidden state of the layer after the target layer in the large language model; Represents the activation value after the target layer passes through the first MLP parameter matrix; Represents the second MLP parameter matrix in the anti-toxicity feedforward calculation module of the target layer; Represents the second MLP parameter matrix in the original feedforward network of the target layer.
[0023] Optionally, the process of determining the parameter matrix to be detoxified in the feedforward network of the large language model by the parameter isolation editing strategy in S4 is represented by the following formulas (3)-(5):
[0024] (3)
[0025] (4)
[0026] (5)
[0027] Among them, Represents the initial hidden state in the large language model; Represents applying embedding calculation to the input X; Represents the hidden state of the layer; Represents the feed-forward network; Represents the hidden state of the (l - 1)-th layer in the large language model; Represents the calculation of the hidden state of the (l - 1)-th layer in the large language model using the attention module; Represents that in the l-th layer, the input x is applied to the feed-forward calculation module for calculation; Represents the activation value after passing through the first MLP parameter matrix; Represents the second MLP parameter matrix in the feed-forward network; Represents the transpose of the input x; Represents the first MLP parameter matrix in the feed-forward network; Represents the non-linear activation function.
[0028] Optionally, the feed-forward network of the large language model obtained by optimizing according to the parameter matrix to be detoxified using the gradient optimization method and the knowledge editing method includes:
[0029] Obtain a harmful prompt and a safety response corresponding to the harmful prompt;
[0030] Input the obtained harmful prompt and the safety response corresponding to the harmful prompt into the large language model for knowledge editing, and construct a loss function through the set prefix system safety prompt;
[0031] Among them, the loss function is represented by the following formula (6):
[0032] (6)
[0033] Among them, Represents the loss function; Represents the harmful prompt; Represents the safety response corresponding to the harmful prompt; Represents the probability that the large language model generates a safety response at the t-th step; Represents the set prefix system safety prompt; Represents the parameters of the large language model at the t-th step;
[0034] Edit the parameters of the large language model in the backpropagation according to the loss function to obtain the detoxified matrix; according to the detoxified matrix, obtain the feed-forward network of the detoxified large language model.
[0035] Optionally, the process of editing the parameters of the large language model in the backpropagation according to the loss function is represented by the following formula (7):
[0036] (7)
[0037] Among them, represents the parameters of the large language model at the t-th time step; represents the toxicity layer parameters of the toxic region; represents the t-th time step gradient; represents the parameters of the large language model at the (t + 1)-th time step.
[0038] On the other hand, a large model adaptive detoxification device based on knowledge editing is provided. This device is applied to the large model adaptive detoxification method based on knowledge editing. The device includes:
[0039] The first construction unit is used to construct a large model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module, and a large language model;
[0040] The second construction unit is used to construct a training set including harmful prompts and harmless prompts; according to the training set, set the prefix system safety prompt, construct the input data of the large language model; input the input data into the large language model, and use the generated last hidden state as the feature vector; according to the feature vector and the obtained labeled toxicity label, construct the training sample data;
[0041] The first acquisition unit is used to train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception through the toxicity feature perception module to obtain the classification result of the toxicity feature perception;
[0042] The second acquisition unit is used to input the classification result of the toxicity feature perception into the anti-toxicity feedforward calculation module, and determine the parameter matrix to be detoxified in the feedforward network of the large language model through the parameter isolation editing strategy; according to the parameter matrix to be detoxified, use the gradient optimization method and the knowledge editing method for optimization to obtain the feedforward network of the detoxified large language model; according to the feedforward network of the detoxified large language model, obtain the detoxified large language model.
[0043] Optionally, the first acquisition unit is used for:
[0044] Insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception on the given user query to obtain the classification result of the toxicity feature perception; among them, the specific process of the toxicity feature perception is represented by the following formula (1):
[0045] (1)
[0046] Among them, represents the classification result of toxicity feature perception; is the parameter of the toxicity feature perception module; represents the hidden state at the last position of the target layer in the large language model; represents the classifier.
[0047] Optionally, the large model adaptive detoxification model based on knowledge editing further includes: a dynamic routing module and an original feed-forward calculation module;
[0048] The dynamic routing module is used to make a judgment according to the classification result of toxicity feature perception, obtain a judgment result, and switch the calculation path according to the judgment result;
[0049] Among them, when the classification result of toxicity feature perception is -1, it is judged as safe, and the classification result of toxicity feature perception is input into the original feed-forward calculation module for calculation;
[0050] When the classification result of toxicity feature perception is +1, it is judged as unsafe, and the classification result of toxicity feature perception is input into the anti-toxicity feed-forward calculation module for calculation;
[0051] The original feed-forward calculation module is used to process the user's normal query.
[0052] Optionally, the specific calculation process of the dynamic routing module is represented by the following formula (2):
[0053] (2)
[0054] Among them, represents the hidden state of the layer after the target layer in the large language model; represents the activation value after the target layer passes through the first MLP parameter matrix; represents the second MLP parameter matrix in the anti-toxicity feed-forward calculation module of the target layer; represents the second MLP parameter matrix in the original feed-forward network of the target layer.
[0055] Optionally, the process of determining the parameter matrix to be detoxified in the feed-forward network of the large language model through the parameter isolation editing strategy is represented by the following formulas (3)-(5):
[0056] (3)
[0057] (4)
[0058] (5)
[0059] Among them, Represents the initial hidden state in the large language model; Represents the application of embedding calculation to the input X; Represents the hidden state of the l-th layer; Represents the feed-forward network; Represents the hidden state of the (l-1)-th layer in the large language model; Represents the calculation of the hidden state of the (l-1)-th layer in the large language model using the attention module; Represents that at the l-th layer, the feed-forward calculation module is applied to the input x for calculation; Represents the activation value after passing through the first MLP parameter matrix; Represents the second MLP parameter matrix in the feed-forward network; Represents the transpose of the input x; Represents the first MLP parameter matrix in the feed-forward network; Represents the non-linear activation function.
[0060] Optionally, the feed-forward network of the large language model obtained by optimizing according to the parameter matrix to be detoxified using the gradient optimization method and the knowledge editing method includes:
[0061] Obtain a harmful prompt and a safety response corresponding to the harmful prompt;
[0062] Input the obtained harmful prompt and the safety response corresponding to the harmful prompt into the large language model for knowledge editing, and construct a loss function through the set prefix system safety prompt;
[0063] Among them, the loss function is represented by the following formula (6):
[0064] (6)
[0065] Among them, Represents the loss function; Represents the harmful prompt; Represents the safety response corresponding to the harmful prompt; Represents the probability that the large language model generates a safety response at the t-th step; Represents the set prefix system safety prompt; Represents the parameters of the large language model at the t-th step;
[0066] Edit the parameters of the large language model in the backpropagation according to the loss function to obtain the detoxified matrix; according to the detoxified matrix, obtain the feed-forward network of the detoxified large language model.
[0067] Optionally, the process of editing the parameters of the large language model in backpropagation according to the loss function is represented by the following formula (7):
[0068] (7)
[0069] where represents the parameters of the large language model at the t-th time step; represents the toxic layer parameters of the toxic region in; represents the t-th time step gradient of; represents the parameters of the large language model at the (t + 1)-th time step.
[0070] On the other hand, there is provided an adaptive detoxification device for a large model based on knowledge editing, and the adaptive detoxification device for a large model based on knowledge editing includes: a processor; a memory, and computer-readable instructions are stored on the memory, and when the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned adaptive detoxification method for a large model based on knowledge editing is implemented.
[0071] On the other hand, there is provided a computer-readable storage medium, and at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned adaptive detoxification method for a large model based on knowledge editing.
[0072] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0073] In the embodiments of the present invention, an adaptive detoxification model for a large model based on knowledge editing is first constructed; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module, and a large language model; a training set including harmful prompts and harmless prompts is constructed; according to the training set, a prefix system safety prompt is set to construct the input data of the large language model; secondly, the input data is input into the large language model, and the generated last hidden state is used as a feature vector; according to the feature vector and the obtained labeled toxicity label, training sample data is constructed; the toxicity detection module is trained according to the training sample data to obtain a trained toxicity detection module; the trained toxicity detection module is inserted into the target layer of the large language model, and toxicity feature perception is performed through the toxicity feature perception module to obtain a classification result of toxicity feature perception; the classification result of toxicity feature perception is input into the anti-toxicity feedforward calculation module, and through a parameter isolation editing strategy, the parameter matrix to be detoxified in the feedforward network of the large language model is determined; finally, according to the parameter matrix to be detoxified, gradient optimization methods and knowledge editing methods are used for optimization to obtain the feedforward network of the detoxified large language model; according to the feedforward network of the detoxified large language model, the detoxified large language model is obtained.
[0074] The embodiments of the present invention achieve a balance among accuracy, efficiency, and generality in the field of detoxifying large language models, overcome the core problems of the prior art's dependence on entity positioning and high-quality data, and performance degradation caused by over-editing, and provide a new technical paradigm for constructing secure and reliable large language models. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0076] Figure 1 is a flowchart of a method for self-adaptive detoxification of a large model based on knowledge editing provided by an embodiment of the present invention;
[0077] Figure 2 is a schematic diagram of the overall framework structure of a self-adaptive detoxification of a large model based on knowledge editing provided by an embodiment of the present invention;
[0078] Figure 3 is a block diagram of a device for self-adaptive detoxification of a large model based on knowledge editing provided by an embodiment of the present invention;
[0079] Figure 4 is a schematic diagram of the structure of a device for self-adaptive detoxification of a large model based on knowledge editing provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0080] The following describes the technical solutions in the present invention with reference to the drawings.
[0081] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to give examples, illustrations, or explanations. Any embodiment or design solution described as "example" in the present invention should not be construed as more preferred or more advantageous than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0082] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same.
[0083] In the embodiments of the present invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, their intended meanings are the same.
[0084] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0085] Embodiments of the present invention provide a method for self-adaptive detoxification of large models based on knowledge editing. This method can be implemented by a device for self-adaptive detoxification of large models based on knowledge editing, and this device for self-adaptive detoxification of large models based on knowledge editing can be a terminal or a server. As Figure 1 shown in the flowchart of the method for self-adaptive detoxification of large models based on knowledge editing, the processing flow of this method can include the following steps:
[0086] S1. Construct a self-adaptive detoxification model of a large model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module, and a large language model.
[0087] Among them, the anti-toxicity feedforward calculation module is responsible for eliminating harmful information in the large language model.
[0088] Optionally, the self-adaptive detoxification model of the large model based on knowledge editing further includes: a dynamic routing module and an original feedforward calculation module;
[0089] The dynamic routing module is used to make a judgment based on the classification result of toxicity feature perception, obtain a judgment result, and switch the calculation path according to the judgment result;
[0090] Among them, the dynamic routing module is set before the feedforward network of the large language model, and by dynamically guiding the data stream to different feedforward networks, the self-adaptive detoxification of user input is realized.
[0091] Among them, when the classification result of toxicity feature perception is -1, it is judged as safe, and the classification result of toxicity feature perception is input into the original feedforward calculation module for calculation;
[0092] When the classification result of toxicity feature perception is +1, it is judged as unsafe, and the classification result of toxicity feature perception is input into the anti-toxicity feedforward calculation module for calculation;
[0093] The original feedforward calculation module is used to process the normal queries of users.
[0094] S2. Construct a training set including harmful prompts and harmless prompts; according to the training set, set the prefix system security prompt to construct the input data of the large language model; input the input data into the large language model, and use the generated last hidden state as the feature vector; construct the training sample data according to the feature vector and the obtained labeled toxicity label.
[0095] In a feasible implementation, construct a training set including 4000 harmful prompts and 2000 harmless prompts; for each training sample, set the prefix system security prompt to form the input data of the large language model.
[0096] S3. Train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception through the toxicity feature perception module to obtain the classification result of toxicity feature perception.
[0097] Among them, the obtained labeled toxicity label is a manually labeled label; input the feature vector and the manually labeled label into the toxicity detection module, and train through a linear kernel support vector machine to obtain a trained toxicity detection module.
[0098] In a feasible implementation, obtain a validation set, and according to the obtained validation set, select the F1 score as the index to verify the trained toxicity detection module, and select the layer corresponding to the toxicity detection module with the optimal index as the insertion layer.
[0099] In a feasible implementation, given a user's query Q, the user's query generates a last hidden state in the target layer of the large language model, and input the last hidden state into the toxicity feature perception module for judgment to obtain the classification result of toxicity feature perception.
[0100] Optionally, inserting the trained toxicity detection module into the target layer of the large language model in S3 and performing toxicity feature perception through the toxicity feature perception module to obtain the classification result of toxicity feature perception includes:
[0101] Insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception on the given user query to obtain the classification result of toxicity feature perception; among them, the specific process of toxicity feature perception is represented by the following formula (1):
[0102] (1)
[0103] Among them, represents the classification result of toxicity feature perception; is the parameter of the toxicity feature perception module; represents the hidden state at the last position of the target layer in the large language model; Represents a classifier.
[0104] Among them, the classification result of toxicity feature perception is used as a routing signal and transmitted to the anti-toxicity feed-forward calculation module to trigger the switching of the calculation path, so that the toxicity feature perception module completes toxicity judgment in the early stage of forward propagation, avoiding subsequent inter-layer interference.
[0105] Among them, by calculating the path through the dynamic routing module, it can effectively handle adversarial inputs without entity dependence, and can improve the detoxification coverage rate and robustness.
[0106] Optionally, the specific calculation process of the dynamic routing module is represented by the following formula (2):
[0107] (2)
[0108] Among them, represents the hidden state of the layer after the target layer in the large language model; represents the activation value after the target layer passes through the first MLP parameter matrix; represents the second MLP parameter matrix in the anti-toxicity feed-forward calculation module of the target layer; represents the second MLP parameter matrix in the original feed-forward network of the target layer.
[0109] S4. Input the classification result of toxicity feature perception into the anti-toxicity feed-forward calculation module, and determine the parameter matrix to be detoxified in the feed-forward network of the large language model through the parameter isolation editing strategy; according to the parameter matrix to be detoxified, use the gradient optimization method and the knowledge editing method for optimization to obtain the feed-forward network of the detoxified large language model; according to the feed-forward network of the detoxified large language model, obtain the detoxified large language model.
[0110] Among them, the anti-toxicity feed-forward calculation module using the parameter isolation editing strategy can ensure that the parameters to be edited are isolated from the original parameters of the large language model, ensuring that the general ability of the large language model is not damaged.
[0111] Among them, the basic structure of the large language model is a parameterized function, including an embedding matrix and L cascaded Transformer layers. Among them, each layer includes a multi-head attention mechanism and a feed-forward network. In a feasible implementation, given an input sequence, the input sequence is input into the large language model, and the parameter matrix to be detoxified in the feed-forward network of the large language model is determined through the parameter isolation editing strategy.
[0112] Optionally, the process of determining the parameter matrix to be detoxified in the feed-forward network of the large language model through the parameter isolation editing strategy in S4 is represented by the following formulas (3)-(5):
[0113] (3)
[0114] (4)
[0115] (5)
[0116] Among them, represents the initial hidden state in the large language model; represents applying embedding calculation to the input X; represents the hidden state of the l-th layer; represents the feed-forward network; represents the hidden state of the (l - 1)-th layer in the large language model; represents calculating the hidden state of the (l - 1)-th layer in the large language model using the attention module; represents calculating the input x using the feed-forward calculation module at the l-th layer; represents the activation value after passing through the first MLP parameter matrix; represents the second MLP parameter matrix in the feed-forward network; represents transposing the input x; represents the first MLP parameter matrix in the feed-forward network; represents the non-linear activation function.
[0117] Among them, in the embodiment of the present invention, the second MLP parameter matrix of the target layer is selected as the editing object; a copy of the second MLP parameter matrix is created and inserted into the target layer for detoxification parameter optimization.
[0118] Optionally, according to the parameter matrix to be detoxified, the gradient optimization method and the knowledge editing method are used for optimization to obtain the feed-forward network of the detoxified large language model, including:
[0119] Obtain a harmful prompt and a safe response corresponding to the harmful prompt;
[0120] Input the obtained harmful prompt and the safe response corresponding to the harmful prompt into the large language model for knowledge editing, and construct a loss function through the set prefix system security prompt;
[0121] Among them, during the process of constructing the loss function, all parameters of the large language model need to be frozen.
[0122] Among them, the loss function is represented by the following formula (6):
[0123] (6)
[0124] Among them, represents the loss function; Indicates a harmful prompt; Indicates the security response corresponding to the harmful prompt; Indicates the probability that the large language model generates a security response at the t-th step; Indicates the set prefix system security prompt; Indicates the parameters of the large language model at the t-th step;
[0125] Edit the parameters of the large language model in the backpropagation according to the loss function to obtain the detoxified matrix; according to the detoxified matrix, obtain the feedforward network of the detoxified large language model.
[0126] Among them, the anti-toxicity feedforward calculation module naturally isolates the detoxification influence domain through the routing mechanism, and only needs to optimize the parameter offset of the harmful matrix, which can reduce the risk of over-editing.
[0127] Among them, the routing mechanism refers to guiding the data flow to different modules according to an input signal. In this application, if the signal is judged to be harmful, the dynamic routing module will direct the data flow to the anti-toxicity feedforward calculation module. Otherwise, if the signal is judged to be harmless, it will be directed to the original feedforward calculation module.
[0128] Optionally, the process of editing the parameters of the large language model in the backpropagation according to the loss function is represented by the following formula (7):
[0129] (7)
[0130] Among them, Indicates the parameters of the large language model at the t-th time step; Indicates the toxicity layer The parameters of the toxic area in; Indicates the t-th time step of the gradient; Indicates the parameters of the large language model at the (t + 1)-th time step.
[0131] Among them, the embodiments of the present invention only need to edit part of the parameters of a single feedforward network, with short training time and low overhead; the embodiments of the present invention are applicable to large models with different parameter magnitudes in multiple series such as Llama large language models and Mistral large language models; through the inter-layer routing mechanism, the embodiments of the present invention can process adversarial inputs without entity dependencies and do not rely on high-quality training data to fine-tune the parameters of the large language model.
[0132] Among them, the embodiments of the present invention introduce a dynamic routing mechanism into the field of detoxifying large language models through knowledge editing, breaking through the dependence on entity positioning of traditional methods; by designing a parameter isolation strategy, precise control of the detoxification influence domain is achieved, and the general capabilities of the large language model are retained to the greatest extent. The embodiments of the present invention achieve a good balance between detoxification performance and ability retention, providing a new technical paradigm for the safe deployment of large language models.
[0133] Among them, as Figure 2 shown is a schematic diagram of the overall structure of a framework for adaptive detoxification of a large model based on knowledge editing provided by an embodiment of the present invention; in a feasible implementation, a trained toxicity detection module, a toxicity feature perception module, and an anti-toxicity feedforward calculation module are inserted into the target layer of the large language model; when a user's query is input into the large language model, the toxicity semantic feature detection module determines whether the query is toxic based on the hidden state of the target layer, and transmits the discrimination signal to the dynamic routing module inserted after the attention layer and the normalization layer; among them, the toxicity semantic feature detection module includes: a toxicity detection module and a toxicity feature perception module; the dynamic routing module judges the signal, and when the judgment signal is unsafe, it directs the information flow to the anti-toxicity feedforward calculation module to detoxify the information flow and refuses to respond to the user's query; when the judgment signal is safe, it directs the information flow to the original feedforward calculation module to normally process the user's query.
[0134] The embodiments of the present invention first construct an adaptive detoxification model for a large model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module, and a large language model; a training set including harmful prompts and harmless prompts is constructed; according to the training set, a prefix system safety prompt is set to construct the input data of the large language model; secondly, the input data is input into the large language model, and the last hidden state generated is used as a feature vector; according to the feature vector and the obtained labeled toxicity label, training sample data is constructed; the toxicity detection module is trained according to the training sample data to obtain a trained toxicity detection module; the trained toxicity detection module is inserted into the target layer of the large language model, and toxicity feature perception is performed through the toxicity feature perception module to obtain the classification result of toxicity feature perception; the classification result of toxicity feature perception is input into the anti-toxicity feedforward calculation module, and the parameter matrix to be detoxified in the feedforward network of the large language model is determined through a parameter isolation editing strategy; finally, according to the parameter matrix to be detoxified, gradient optimization methods and knowledge editing methods are used for optimization to obtain the feedforward network of the detoxified large language model; according to the feedforward network of the detoxified large language model, the detoxified large language model is obtained.
[0135] The embodiments of the present invention achieve a balance among accuracy, efficiency, and generality in the field of detoxifying large language models, overcome the core problems of the prior art such as dependence on entity positioning and high-quality data, and performance degradation caused by over-editing, and provide a new technical paradigm for constructing secure and reliable large language models.
[0136] Figure 3 It is a block diagram of a large model adaptive detoxification device based on knowledge editing shown according to an exemplary embodiment. This device is used for the large model adaptive detoxification method based on knowledge editing. Refer to Figure 3 In this device, it includes a first construction unit 310, a second construction unit 320, a first acquisition unit 330, and a second acquisition unit 340. Among them:
[0137] The first construction unit 310 is used to construct a large model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feed-forward calculation module, and a large language model;
[0138] The second construction unit 320 is used to construct a training set including harmful prompts and harmless prompts; according to the training set, set a prefix system security prompt, construct the input data of the large language model; input the input data into the large language model, and use the generated last hidden state as a feature vector; according to the feature vector and the obtained labeled toxicity label, construct training sample data;
[0139] The first acquisition unit 330 is used to train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception through the toxicity feature perception module to obtain the classification result of toxicity feature perception;
[0140] The second acquisition unit 340 is used to input the classification result of toxicity feature perception into the anti-toxicity feed-forward calculation module, and determine the parameter matrix to be detoxified in the feed-forward network of the large language model through a parameter isolation editing strategy; according to the parameter matrix to be detoxified, use a gradient optimization method and a knowledge editing method for optimization to obtain the feed-forward network of the detoxified large language model; according to the feed-forward network of the detoxified large language model, obtain the detoxified large language model.
[0141] Optionally, the first acquisition unit 330 is used for:
[0142] Insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception on a given user query to obtain the classification result of toxicity feature perception; among them, the specific process of toxicity feature perception is represented by the following formula (1):
[0143] (1)
[0144] Among them, represents the classification result of toxicity feature perception; is the parameter of the toxicity feature perception module; represents the hidden state at the last position of the target layer in the large language model; represents the classifier.
[0145] Optionally, the large model adaptive detoxification model based on knowledge editing further includes: a dynamic routing module and an original feed-forward calculation module;
[0146] The dynamic routing module is used to make a judgment according to the classification result of toxicity feature perception, obtain a judgment result, and switch the calculation path according to the judgment result;
[0147] Among them, when the classification result of toxicity feature perception is -1, it is judged as safe, and the classification result of toxicity feature perception is input into the original feed-forward calculation module for calculation;
[0148] When the classification result of toxicity feature perception is +1, it is judged as unsafe, and the classification result of toxicity feature perception is input into the anti-toxicity feed-forward calculation module for calculation;
[0149] The original feed-forward calculation module is used to process the normal query of the user.
[0150] Optionally, the specific calculation process of the dynamic routing module is represented by the following formula (2):
[0151] (2)
[0152] Among them, represents the hidden state of the layer after the target layer in the large language model; represents the activation value after the target layer passes through the first MLP parameter matrix; represents the second MLP parameter matrix in the anti-toxicity feed-forward calculation module of the target layer; represents the second MLP parameter matrix in the original feed-forward network of the target layer.
[0153] Optionally, the process of determining the parameter matrix to be detoxified in the feed-forward network of the large language model through the parameter isolation editing strategy is represented by the following formulas (3)-(5):
[0154] (3)
[0155] (4)
[0156] (5)
[0157] Among them, represents the initial hidden state in the large language model; represents applying embedding calculation to the input X; represents the hidden state of the l-th layer; represents the feed-forward network; represents the hidden state of the (l - 1)-th layer in the large language model; represents calculating the hidden state of the (l - 1)-th layer in the large language model using the attention module; represents calculating the input x using the feed-forward calculation module at the l-th layer; represents the activation value after passing through the first MLP parameter matrix; represents the second MLP parameter matrix in the feed-forward network; represents transposing the input x; represents the first MLP parameter matrix in the feed-forward network; represents the non-linear activation function.
[0158] Optionally, the feed-forward network of the large language model after detoxification, obtained by optimizing according to the parameter matrix to be detoxified using the gradient optimization method and the knowledge editing method, includes:
[0159] Obtain a harmful prompt and a safety response corresponding to the harmful prompt;
[0160] Input the obtained harmful prompt and the safety response corresponding to the harmful prompt into the large language model for knowledge editing, and construct a loss function through the set prefix system safety prompt;
[0161] Among them, the loss function is represented by the following formula (6):
[0162] (6)
[0163] Among them, represents the loss function; represents the harmful prompt; represents the safety response corresponding to the harmful prompt; represents the probability that the large language model generates a safety response at the t-th step; represents the set prefix system safety prompt; represents the parameters of the large language model at the t-th step;
[0164] Edit the parameters of the large language model in the backpropagation according to the loss function to obtain the detoxified matrix; according to the detoxified matrix, obtain the feed-forward network of the detoxified large language model.
[0165] Optionally, the process of editing the parameters of the large language model in backpropagation according to the loss function is represented by the following formula (7):
[0166] (7)
[0167] Wherein, represents the parameters of the large language model at the t-th time step; represents the toxicity layer parameters of the toxic region in; represents the t-th time step gradient of; represents the parameters of the large language model at the (t + 1)-th time step.
[0168] In the embodiments of the present invention, an adaptive detoxification model for a large model based on knowledge editing is first constructed; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module, and a large language model; a training set including harmful prompts and harmless prompts is constructed; according to the training set, a prefix system security prompt is set to construct the input data of the large language model; secondly, the input data is input into the large language model, and the generated last hidden state is used as a feature vector; according to the feature vector and the obtained labeled toxicity label, training sample data is constructed; the toxicity detection module is trained according to the training sample data to obtain a trained toxicity detection module; the trained toxicity detection module is inserted into the target layer of the large language model, and toxicity feature perception is performed through the toxicity feature perception module to obtain a classification result of toxicity feature perception; the classification result of toxicity feature perception is input into the anti-toxicity feedforward calculation module, and the parameter matrix to be detoxified in the feedforward network of the large language model is determined through a parameter isolation editing strategy; finally, according to the parameter matrix to be detoxified, gradient optimization methods and knowledge editing methods are used for optimization to obtain the feedforward network of the detoxified large language model; according to the feedforward network of the detoxified large language model, the detoxified large language model is obtained.
[0169] The embodiments of the present invention achieve a balance among accuracy, efficiency, and generality in the field of detoxifying large language models, overcome the core problems of the prior art's dependence on entity positioning and high-quality data, and performance degradation caused by over-editing, and provide a new technical paradigm for building safe and reliable large language models.
[0170] Figure 4 is a schematic structural diagram of an adaptive detoxification device for a large model based on knowledge editing provided by an embodiment of the present invention. As Figure 4 shown, the adaptive detoxification device for a large model based on knowledge editing may include the above-mentioned Figure 3 adaptive detoxification device for a large model based on knowledge editing shown. Optionally, the adaptive detoxification device 410 for a large model based on knowledge editing may include a first processor 2001.
[0171] Optionally, the large model adaptive detoxification device 410 based on knowledge editing may further include a memory 2002 and a transceiver 2003.
[0172] Among them, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, such as through a communication bus.
[0173] Next, in conjunction with Figure 4 Specific introductions will be made to the various components of the large model adaptive detoxification device 410 based on knowledge editing:
[0174] Among them, the first processor 2001 is the control center of the large model adaptive detoxification device 410 based on knowledge editing, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0175] Optionally, the first processor 2001 can execute various functions of the large model adaptive detoxification device 410 based on knowledge editing by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0176] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 the CPU0 and CPU1 shown in
[0177] In a specific implementation, as an embodiment, the large model adaptive detoxification device 410 based on knowledge editing may also include multiple processors, such as Figure 4 the first processor 2001 and the second processor 2004 shown in
[0178] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiment and will not be elaborated here.
[0179] Optionally, the memory 2002 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited thereto. The memory 2002 can be integrated with the first processor 2001 or exist independently and is coupled to the first processor 2001 through the interface circuit of the large model adaptive detoxification device 410 based on knowledge editing ( Figure 4 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.
[0180] The transceiver 2003 is used to communicate with a network device or with a terminal device.
[0181] Optionally, the transceiver 2003 can include a receiver and a transmitter ( Figure 4 not shown separately in the figure). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0182] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently and is coupled to the first processor 2001 through the interface circuit of the large model adaptive detoxification device 410 based on knowledge editing ( Figure 4 not shown in the figure), and the embodiments of the present invention do not make specific limitations on this.
[0183] It should be noted that Figure 4 the structure of the large model adaptive detoxification device 410 based on knowledge editing shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0184] In addition, for the technical effects of the large model adaptive detoxification device 410 based on knowledge editing, reference may be made to the technical effects of the large model adaptive detoxification method based on knowledge editing described in the above method embodiments, which will not be elaborated here.
[0185] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0186] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0187] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0188] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0189] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0190] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0191] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0192] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0193] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0194] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0195] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0196] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0197] As described above, the above are only specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An adaptive detoxification method for large models based on knowledge editing, characterized in that, The method includes: S1. Construct an adaptive detoxification model for large models based on knowledge editing; the model includes a toxicity detection module, a toxicity feature perception module, an anti-toxicity feedforward calculation module, and a large language model; Among them, the adaptive detoxification model for large models based on knowledge editing further includes a dynamic routing module and an original feedforward calculation module; The dynamic routing module is used to make a judgment according to the classification result of toxicity feature perception, obtain a judgment result, and switch the calculation path according to the judgment result; Among them, when the classification result of toxicity feature perception is -1, it is judged as safe, and the classification result of toxicity feature perception is input into the original feedforward calculation module for calculation; When the classification result of toxicity feature perception is +1, it is judged as unsafe, and the classification result of toxicity feature perception is input into the anti-toxicity feedforward calculation module for calculation; The original feedforward calculation module is used to process the normal queries of users; S2. Construct a training set containing harmful prompts and harmless prompts; according to the training set, set a prefix system security prompt to construct the input data of the large language model; input the input data into the large language model, and use the generated last hidden state as a feature vector; construct training sample data according to the feature vector and the obtained labeled toxicity label; S3. Train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception through the toxicity feature perception module to obtain the classification result of toxicity feature perception; S4. Input the classification result of toxicity feature perception into the anti-toxicity feedforward calculation module, and determine the parameter matrix to be detoxified in the feedforward network of the large language model through a parameter isolation editing strategy; according to the parameter matrix to be detoxified, use a gradient optimization method and a knowledge editing method for optimization to obtain the feedforward network of the detoxified large language model; according to the feedforward network of the detoxified large language model, obtain the detoxified large language model; Among them, when the query of the user is input into the detoxified large language model, the toxicity semantic feature detection module determines whether the query is toxic according to the hidden state of the target layer, and transmits the discrimination signal to the dynamic routing module inserted after the attention layer and the normalization layer; among them, the toxicity semantic feature detection module includes a toxicity detection module and a toxicity feature perception module; the dynamic routing module makes a judgment on the signal. When the judgment signal is unsafe, it guides the information flow to the anti-toxicity feedforward calculation module to perform detoxification processing on the information flow and refuses to respond to the user's query; when the judgment signal is safe, it guides the information flow to the original feedforward calculation module to normally process the user's query.
2. The method for adaptively detoxifying a large model based on knowledge editing according to claim 1, wherein The step of inserting the trained toxicity detection module into the target layer of the large language model in S3 and performing toxicity feature perception through the toxicity feature perception module to obtain the classification result of toxicity feature perception includes: Insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception on the given user query to obtain the classification result of toxicity feature perception; among them, the specific process of toxicity feature perception is represented by the following formula (1): (1) Among them, represents the classification result of toxicity feature perception; represents the parameters of the toxicity feature perception module; represents the hidden state at the last position of the target layer in the large language model; represents the classifier.
3. The large model adaptive detoxification method based on knowledge editing according to claim 1, wherein The specific calculation process of the dynamic routing module is represented by the following formula (2): (2) Among them, represents the hidden state of the layer immediately following the target layer in the large language model; represents the activation value of the target layer after passing through the first MLP parameter matrix; represents the second MLP parameter matrix in the anti-toxicity feed-forward calculation module of the target layer; represents the second MLP parameter matrix in the original feed-forward network of the target layer.
4. The method for adaptively detoxifying large models based on knowledge editing according to claim 1, wherein, The process of determining the parameter matrix to be detoxified in the feed-forward network of the large language model through the parameter isolation editing strategy in S4 is represented by the following formulas (3)-(5): (3) (4) (5) Among them, represents the initial hidden state in the large language model; represents applying embedding calculation to the input X; represents the hidden state of the l-th layer; represents the feed-forward network; represents the hidden state of the (l - 1)-th layer in the large language model; represents calculating the hidden state of the (l - 1)-th layer in the large language model using the attention module; represents, at the l-th layer, applying the feed-forward calculation module to the input x for calculation; represents the activation value after passing through the first MLP parameter matrix; represents the second MLP parameter matrix in the feed-forward network; represents transposing the input x; represents the first MLP parameter matrix in the feed-forward network; represents the non-linear activation function.
5. The method for adaptively detoxifying a large model based on knowledge editing according to claim 1, characterized in that According to the parameter matrix to be detoxified, using the gradient optimization method and the knowledge editing method for optimization to obtain the feed-forward network of the detoxified large language model, including: Obtain a harmful prompt and a safety response corresponding to the harmful prompt; Input the obtained harmful prompt and the safety response corresponding to the harmful prompt into the large language model for knowledge editing, and construct a loss function through the set prefix system safety prompt; Among them, the loss function is represented by the following formula (6): (6) Among them, represents the loss function; represents the harmful hint; represents the safety response corresponding to the harmful hint; represents the probability that the large language model generates a safety response at the t-th step; represents the set prefix system safety hint; represents the parameters of the large language model at the t-th step; Edit the parameters of the large language model in the backward propagation according to the loss function to obtain a detoxified matrix; according to the detoxified matrix, obtain the feed-forward network of the detoxified large language model.
6. The method for adaptively detoxifying a large model based on knowledge editing according to claim 5, wherein The process of editing the parameters of the large language model in the backward propagation according to the loss function is represented by the following formula (7): (7) Among them, represents the parameters of the large language model at the t-th time step; represents the toxicity layer parameters of the toxic region; represents the t-th time step gradient; represents the parameters of the large language model at the (t + 1)-th time step.
7. An adaptive detoxification device for large models based on knowledge editing, which is used to implement the adaptive detoxification method for large models based on knowledge editing according to any one of claims 1-6, characterized in that, The device includes: The first construction unit is used to construct a large model adaptive detoxification model based on knowledge editing; the model includes: a toxicity detection module, a toxicity feature perception module, an anti-toxicity feed-forward calculation module, and a large language model; Among them, the large model adaptive detoxification model based on knowledge editing further includes: a dynamic routing module and an original feed-forward calculation module; The dynamic routing module is used to make a judgment according to the classification result of toxicity feature perception, obtain a judgment result, and switch the calculation path according to the judgment result; Among them, when the classification result of toxicity feature perception is -1, it is judged as safe, and the classification result of toxicity feature perception is input into the original feed-forward calculation module for calculation; When the classification result of toxicity feature perception is +1, it is judged as unsafe, and the classification result of toxicity feature perception is input into the anti-toxicity feed-forward calculation module for calculation; The original feed-forward calculation module is used to process the user's normal query; The second construction unit is used to construct a training set including harmful prompts and harmless prompts; according to the training set, set a prefix system safety prompt, construct the input data of the large language model; input the input data into the large language model, and use the generated last hidden state as a feature vector; construct training sample data according to the feature vector and the obtained labeled toxicity label; The first acquisition unit is used to train the toxicity detection module according to the training sample data to obtain a trained toxicity detection module; insert the trained toxicity detection module into the target layer of the large language model, and perform toxicity feature perception through the toxicity feature perception module to obtain the classification result of toxicity feature perception; A second acquisition unit, configured to input the classification result of toxicity feature perception into an anti-toxicity feedforward calculation module, and determine the parameter matrix to be detoxified in the feedforward network of the large language model through a parameter isolation editing strategy; optimize according to the parameter matrix to be detoxified by using a gradient optimization method and a knowledge editing method to obtain the feedforward network of the detoxified large language model; obtain the detoxified large language model according to the feedforward network of the detoxified large language model. Among them, when the user's query is input into the detoxified large language model, the toxicity semantic feature detection module determines whether the query is toxic according to the hidden state of the target layer, and transmits the discrimination signal to the dynamic routing module inserted after the attention layer and the normalization layer; the toxicity semantic feature detection module includes: a toxicity detection module and a toxicity feature perception module; the dynamic routing module judges the signal, and when the judgment signal is unsafe, it guides the information flow to the anti-toxicity feedforward calculation module to detoxify the information flow and refuses to respond to the user's query; when the judgment signal is safe, it guides the information flow to the original feedforward calculation module to normally process the user's query.
8. An adaptive detoxification device for large models based on knowledge editing, characterized in that, The large model adaptive detoxification device based on knowledge editing includes: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-modal reasoning method and device based on large language model and knowledge graph
CN118193684A
Optimization method and device, reply method and device of large language model, equipment and medium
CN119918658A