Model behavior evaluation method and system based on hidden layer activation distribution modeling

Through the method of activation distribution modeling based on hidden layer, the cosine similarity of parameter gradient slices is calculated and the training data set is generated, and a classifier model is established for large language model evaluation, which solves the problem of difficulty in efficiently detecting jailbreak prompts in large language models in the existing technology, and achieves efficient and accurate model security evaluation.

CN120123772APending Publication Date: 2025-06-10ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510208888.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately detect jailbreak prompts in large language models, which makes it difficult to ensure model security, and existing methods usually require a large amount of resource investment.

Method used

A model behavior evaluation method based on hidden layer activation distribution modeling is adopted. By obtaining the reference data set, calculating the cosine similarity of parameter gradient slices, and generating training data sets in combination with safety key parameter locations, a classifier model is established for large language model evaluation.

Benefits of technology

It realizes accurate detection of unsafe inputs in large language models, effectively evaluates the security of the model, and does not require additional training on the original large language models, and has less resource investment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123772A_ABST
    Figure CN120123772A_ABST
Patent Text Reader

Abstract

The invention discloses a model behavior evaluation method and system based on hidden layer activation distribution modeling, and relates to the technical field of security evaluation of large language models.The method comprises the steps that a positive sample and a negative sample are input into a large language model, and corresponding hidden layer parameters and gradients under a loss function are obtained; hidden layer parameters and gradients are converted into a matrix form, and then the matrix is sliced; calculating the overall average value of gradient slices generated by negative sample input as a reference, then calculating the cosine similarity between the gradient slices generated by each positive sample and negative sample input and the reference, performing subtraction to calculate a difference value, and taking the slice with the difference value exceeding a specified threshold value as a safety key parameter position; and inputting the security domain data set into the large language model, obtaining security key parameters in combination with the positions of the security key parameters to establish a classifier model, and performing large language model evaluation through the classifier model. According to the method, unsafe input in the large language model can be accurately detected, and the safety of the model is evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of security evaluation of large language models, and more specifically, to a model behavior evaluation method and system based on hidden layer activation distribution modeling. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models, such as GPT series, LLaMA series, etc., have achieved remarkable achievements in many fields and are widely used in daily fields such as search engines and office software. However, large language models face security threats from jailbreak prompts. Jailbreak prompts may lead to the abuse of large language models, resulting in various illegal or adverse consequences, and may also affect the security alignment of the models, causing the models to exhibit unsafe behaviors during the fine-tuning process.

[0003] To address this problem, there are various methods for detecting jailbreak prompts in the prior art, but these methods have certain limitations. On the one hand, some online review API tools, such as OpenAI Moderation API, PerspectiveAPI, Azure AI Content SafetyAPI, etc., are mainly models trained based on a large amount of data for detecting general toxic content and are not effective in identifying jailbreak prompts. On the other hand, using the large language model itself as a zero-shot detector often has poor performance problems, such as overestimating security risks. Although there are also methods to improve the detection performance by fine-tuning the large language model, such as LLaMA Guard, which achieves unsafe detection of input-output by fine-tuning on a carefully collected dataset, this fine-tuning process requires a large amount of resources, including a carefully planned dataset and a large amount of training time.

[0004] Therefore, a more efficient, accurate and resource-saving method is needed to detect jailbreak prompts in large language models to ensure the safe use and stable performance of large language models. Summary of the Invention

[0005] In view of this, the present invention provides a model behavior evaluation method and system based on hidden layer activation distribution modeling, which can accurately detect unsafe inputs in large language models, such as jailbreak prompts, toxic content, etc., and effectively evaluate the security of the models.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A model behavior evaluation method based on hidden layer activation distribution modeling, comprising:

[0008] Obtaining a reference dataset, the reference dataset including positive samples and negative samples;

[0009] Input the positive samples and negative samples into a preset large language model respectively to obtain the corresponding hidden layer parameters and gradients under the standard loss function;

[0010] Convert the hidden layer parameters and gradients into matrix forms, and then slice each gradient matrix row by row and column by column;

[0011] Calculate the overall average value of the parameter gradient slices generated under the input of negative samples as the reference gradient, then calculate the cosine similarity between the parameter gradient slices generated under the input of each positive sample and negative sample and the reference gradient, and finally calculate the difference between the cosine similarities generated under the input of positive and negative samples. Take the parameter gradient slices with the difference exceeding the specified threshold as the positions of safety-critical parameters;

[0012] Input the safety domain dataset into the large language model, combine the positions of the safety-critical parameters to obtain the safety-critical parameters, and generate training data based on the safety-critical parameters to generate a training dataset;

[0013] Use the generated training dataset to establish a classifier model and evaluate the large language model through the classifier model.

[0014] Preferably, each parameter gradient slice is a multi-dimensional vector, and the cosine similarity is used to characterize the similarity between different parameter gradient slices for two different parameter gradient slices.

[0015] Preferably, the calculation of the cosine similarity between the parameter gradient slices generated under the input of each positive sample and negative sample and the reference gradient specifically includes:

[0016]

[0017] where, represents the dot product of vectors A and B, and |A| and |B| represent the magnitudes of vectors A and B respectively.

[0018] Preferably, the taking the parameter gradient slices with the difference exceeding the specified threshold as the positions of safety-critical parameters specifically includes: Since the range of cosine similarity is [-1, 1], set a threshold of 0.5. When the cosine similarity between two slices is greater than this threshold, the two slices are classified as highly similar, and when it is lower than this threshold, they are classified as low similar.

[0019] Preferably, there are two methods for establishing the classifier model. The first classifier: take the arithmetic mean of all safety-critical parameters and cosine similarities, set a threshold, and those exceeding the threshold are classified as unsafe; The second classifier: use a logistic regression classifier and train the logistic regression classifier model with the generated training dataset.

[0020] A model behavior evaluation system based on hidden layer activation distribution modeling, comprising:

[0021] A reference dataset acquisition module that acquires a reference dataset, where the reference dataset includes positive samples and negative samples;

[0022] A parameter gradient acquisition module that inputs the positive samples and negative samples into a preset large language model respectively to obtain the corresponding hidden layer parameters and gradients under the standard loss function;

[0023] A slicing module that converts the hidden layer parameters and gradients into matrix forms, and then slices each gradient matrix row by row and column by column;

[0024] A critical parameter position acquisition module that calculates the overall average of the parameter gradient slices generated under the input of negative samples as the reference gradient, then calculates the cosine similarity between the parameter gradient slices generated under the input of each positive sample and negative sample and the reference gradient, and finally calculates the difference between the cosine similarities generated under the input of positive and negative samples, and takes the parameter gradient slices with the difference exceeding the specified threshold as the security critical parameter positions;

[0025] A training dataset generation module that inputs a security domain dataset into the large language model, combines the security critical parameter positions to obtain security critical parameters, and generates training data based on the security critical parameters to generate a training dataset;

[0026] An evaluation module that uses the generated training dataset to establish a classifier model and evaluates the large language model through the classifier model.

[0027] It can be seen from the above technical solutions that compared with the prior art, the present invention discloses a model behavior evaluation method and system based on hidden layer activation distribution modeling, aiming at detecting unsafe prompts to protect the large language model from abuse or malicious fine-tuning, and in view of the current situation that existing methods usually involve training or adjusting the large language model as a classifier of a large dataset, a new method for checking the security critical parameters of the large language model to identify unsafe prompts is proposed, and its effect can be better than fine-tuning the model without any additional training on the original large language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0029] Figure 1 It is a structural schematic diagram provided by the present invention. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0031] An embodiment of the present invention discloses a model behavior evaluation method based on hidden layer activation distribution modeling, as Figure 1 shown, including:

[0032] Obtain a reference data set, where the reference data set includes positive samples and negative samples;

[0033] Input the positive samples and negative samples into a preset large language model respectively to obtain the corresponding hidden layer parameters and gradients under the standard loss function;

[0034] Convert the hidden layer parameters and gradients into matrix forms, and then perform row-by-row and column-by-column slicing on each gradient matrix;

[0035] Calculate the average value of the gradient slices of all negative samples as the reference gradient slice, and then calculate the similarity between the gradient slices of each positive and negative sample and the reference gradient slice. Take the gradient slices with similarity exceeding the specified threshold as the positions of safety-critical parameters (find the parameter slices that show high similarity in the gradients of each unsafe input and low similarity between the gradients of unsafe and safe inputs as safety-critical parameters. There can be multiple safety-critical parameters, and the reference slice of the gradient where they are located is used as the reference for unsafe gradients); among them, positive samples are safe inputs, and negative samples are unsafe inputs;

[0036] Input a safety domain data set (including safety prompts and non-safety prompts for the target task) containing safety prompts and non-safety prompts for the target task into the large language model, combine the positions of safety-critical parameters to obtain safety-critical parameters, and generate training data based on the safety-critical parameters to generate a training data set;

[0037] Use the generated training data set to establish a classifier model and evaluate the large language model through the classifier model.

[0038] In a specific embodiment, each parameter gradient slice is a multi-dimensional vector, and the cosine similarity is used to characterize the similarity between different parameter gradient slices.

[0039] In a specific embodiment, calculating the cosine similarity between the parameter gradient slices generated under each positive and negative sample input and the reference gradient specifically includes:

[0040]

[0041] Among them, represents the dot product of vectors A and B, and |A| and |B| respectively represent the magnitudes of vectors A and B.

[0042] In a specific embodiment, taking the parameter gradient slices with a difference exceeding a specified threshold as the safety-critical parameter positions specifically includes: Since the cosine similarity ranges from [-1, 1], a threshold of 0.5 is set. When the cosine similarity between two slices is greater than this threshold, the two slices are classified as highly similar, and when it is lower than this threshold, they are classified as low-similarity.

[0043] In a specific embodiment, establishing a classifier model includes two methods. The first classifier: taking the arithmetic mean of all safety-critical parameters and the cosine similarity, setting a threshold, and classifying those exceeding the threshold as unsafe; The second classifier: using a logistic regression classifier and training the logistic regression classifier model with the generated training dataset.

[0044] A model behavior evaluation system based on hidden layer activation distribution modeling, comprising:

[0045] A reference dataset acquisition module that acquires a reference dataset, where the reference dataset includes positive samples and negative samples;

[0046] A parameter gradient acquisition module that inputs the positive samples and negative samples into a preset large language model respectively to obtain the corresponding hidden layer parameters and the gradients under the standard loss function;

[0047] A slicing module that converts the hidden layer parameters and gradients into matrix forms, and then slices each gradient matrix row by row and column by column;

[0048] A critical parameter position acquisition module that calculates the overall average of the parameter gradient slices generated under the input of negative samples as the reference gradient, then calculates the cosine similarity between each parameter gradient slice generated under the input of positive samples and negative samples and the reference gradient, and finally calculates the difference between the cosine similarities generated under the input of positive and negative samples, and takes the parameter gradient slices with a difference exceeding the specified threshold as the safety-critical parameter positions;

[0049] A training dataset generation module that inputs a safety domain dataset into the large language model, combines the safety-critical parameter positions to obtain safety-critical parameters, and generates training data based on the safety-critical parameters to generate a training dataset;

[0050] An evaluation module that uses the generated training dataset to establish a classifier model and evaluates the large language model through the classifier model.

[0051] In specific embodiment 1, as follows:

[0052] Step 1: Use the Llama-2-7B large language model, with only two safe prompts and two unsafe prompts, and require the large language model to output "sure". Calculate its loss function under the input. The loss function uses cross-entropy loss and the corresponding gradient information. In Llama-2-7B, there are many matrices composed of parameters with dependency relationships. Preserve these matrix structures and fill the partial derivative gradients of the corresponding position parameters into the corresponding positions to form different gradient matrices.

[0053] Step 2: Perform row and column slicing on each gradient matrix, thus obtaining a total of 2,498,560 slices (1,138,688 columns and 1,359,872 rows) for Llama-2-7B. These slices are used as the basic elements of this work to identify safety-critical parameters and calculate cosine similarity features.

[0054] Step 3: Calculate the overall average of the parameter gradient slices generated under the negative sample input as the reference gradient. Then calculate the cosine similarity between each positive sample and the gradient slices generated under the negative sample input and the reference gradient. Finally, calculate the difference between the cosine similarities generated under the positive and negative sample inputs. Mark the parameter slices whose difference exceeds the specified threshold (set to 0.5). These marked parameter slices are identified as safety-critical parameters, and the corresponding gradient slices from the reference gradient slices are stored as unsafe gradient references. Finally, 12,137 row slices (about 1% of the total row slices) and 3,194 column slices (about 0.2% of the total column slices) are found.

[0055] Step 4: Use two benchmark datasets, ToxicChat and XSTest, which are used to evaluate the performance of methods for detecting unsafe prompts (including jailbreak prompts). Use the dataset data as the input to Llama-2-7B, obtain the gradient slices of the safety-critical parameters, and calculate the cosine similarity between these safety-critical parameters and the unsafe gradient references as the data. The labels are classified into two categories: safe and unsafe, forming a new training dataset.

[0056] Step 5: Use the formed data training set, divide it into a training set and a test set according to 7:3. Train a logistic regression classifier model, use the log loss function, the gradient descent optimization algorithm, set the learning rate to 0.01, and set the number of iterations to 100 times. Use the f1 score as the evaluation criterion. Finally, the f1 score of the model on the ToxicChat dataset is 0.707, and the f1 score on the XSTest dataset is 0.950.

[0057] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For related parts, reference can be made to the description in the method section.

[0058] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A model behavior evaluation method based on hidden layer activation distribution modeling, characterized in that: include: Acquire a reference data set, wherein the reference data set includes positive samples and negative samples; Input the positive samples and negative samples into a preset large language model respectively to obtain corresponding hidden layer parameters and gradients under a standard loss function; The hidden layer parameters and gradients are converted into matrix form, and each gradient matrix is ​​sliced ​​row by row and column by column; Calculate the overall average of the parameter gradient slices generated under the negative sample input as the reference gradient, then calculate the cosine similarity between the parameter gradient slices generated under each positive sample and negative sample input and the reference gradient, and finally calculate the difference in cosine similarity generated under the positive and negative sample inputs, and take the parameter gradient slices whose difference exceeds the specified threshold as the safety-critical parameter position; Inputting the security domain data set into the large language model, combining the security key parameter position, obtaining the security key parameter, generating training data based on the security key parameter, and generating a training data set; A classifier model is established using the generated training data set, and a large language model is evaluated using the classifier model.

2. A model behavior evaluation method based on hidden layer activation distribution modeling according to claim 1, characterized in that: Each parameter gradient slice is a multi-dimensional vector, and the cosine similarity between two different parameter gradient slices is used to characterize the similarity between different parameter gradient slices.

3. A model behavior evaluation method based on hidden layer activation distribution modeling according to claim 2, characterized in that: The calculation of the cosine similarity between the parameter gradient slice generated under each positive sample and negative sample input and the reference gradient specifically includes: in, represents the dot product of vectors A and B, and |A||B| represents the magnitudes of vectors A and B respectively.

4. A model behavior evaluation method based on hidden layer activation distribution modeling according to claim 3, characterized in that: The method of using the parameter gradient slice whose difference exceeds the specified threshold as the safety-critical parameter position specifically includes: since the cosine similarity range is [-1, 1], a threshold is set to 0.

5. When the cosine similarity of two slices is greater than this threshold, the two slices are classified as having high similarity, and if it is lower than this threshold, they are classified as having low similarity.

5. The model behavior evaluation method based on hidden layer activation distribution modeling according to claim 1 is characterized in that: The method for establishing the classifier model includes two methods. The first classifier is to take the arithmetic mean of all safety-critical parameters and cosine similarity, set a threshold, and classify them as unsafe if the threshold is exceeded; the second classifier is to use a logistic regression classifier and use the generated training data set to train the logistic regression classifier model.

6. A model behavior evaluation system based on hidden layer activation distribution modeling, applying a model behavior evaluation method based on hidden layer activation distribution modeling according to any one of claims 1 to 5, comprising: A reference data set acquisition module is used to acquire a reference data set, wherein the reference data set includes positive samples and negative samples; A parameter gradient acquisition module inputs the positive samples and negative samples into a preset large language model to obtain corresponding hidden layer parameters and gradients under a standard loss function; A slicing module converts the hidden layer parameters and gradients into matrix form, and then slices each gradient matrix row by row and column by column; The key parameter position acquisition module calculates the overall average value of the parameter gradient slices generated under the negative sample input as the reference gradient, then calculates the cosine similarity between the parameter gradient slices generated under each positive sample and negative sample input and the reference gradient, and finally calculates the difference between the cosine similarities generated under the positive and negative sample inputs, and takes the parameter gradient slices whose difference exceeds the specified threshold as the security key parameter position; A training data set generation module, inputting the security domain data set into the large language model, combining the security key parameter position, obtaining the security key parameter, generating training data based on the security key parameter, and generating a training data set; The evaluation module uses the generated training data set to establish a classifier model, and uses the classifier model to evaluate the large language model.