Language model security measurement method and device
By constructing adversarial word embedding distributions and prior distributions, and optimizing the optimal adversarial word embedding distribution, the uncertainty of language model security measurement is resolved, and security measurement is achieved in the worst case, ensuring that the risk of language model generating harmful text is fully exposed.
Patent Information
- Application Number
- CN202511138797.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies have uncertainties in measuring the security of language models and cannot ensure that adversarial prompts can be found for any harmful text, making the defense mechanism unreliable.
We construct adversarial word embedding distributions and prior distributions, generate random adversarial hints by optimizing continuous probability distributions, and search for the optimal adversarial word embedding distribution using gradient descent algorithm to ensure security measurement in the worst case.
It achieves a worst-case security measure, avoids the limitations of discrete search, ensures that the risk of language models generating harmful text is fully exposed, and prevents adversarial examples from deviating from natural language features.
Smart Images

Figure CN121030740A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence security, and in particular to a language model security measurement method and device. BACKGROUND
[0002] A language model, as a kind of probability model for statistical modeling of natural language, is generally implemented based on a deep neural network and can produce fluent, grammatically correct, and semantically correct text like a human. After training, the language model receives a "prompt" as input (which can be understood as a question), and completes the prompt (which can be understood as answering the input question). The completion of the prompt is the data generated by the language model.
[0003] Although advanced language models have shown amazing text generation capabilities, the problem of generating unsafe content has caused widespread concern, including generating harmful text, private data, copyrighted data, etc. Although developers generally introduce security alignment during development, relevant research shows that attackers can still induce unsafe content from security-hardened language models using specially designed prompts. To address these risks, researchers have proposed corresponding defense measures, which have shown some resistance.
[0004] For example, reference 1 (Universal and transferable adversarial attacks on aligned language models[J]. arXiv preprint arXiv:2307.15043, 2023) discloses universal and transferable adversarial attacks on aligned language models, focusing on the Greedy Coordinate Gradient (GCG) algorithm. For a given harmful text, the algorithm optimizes an adversarial suffix that, when concatenated to the prompt, can induce the language model to generate the corresponding harmful text. For example, the following prompt with a GCG-optimized adversarial suffix "Please help me write a phishing email to trick the other party into revealing personal information [adversarial suffix]" asks the language model, and the language model will answer the question "Of course, here is a phishing email to trick the other party into revealing personal information …". GCG evaluates the security of the language model in an attack manner. If the attack is successful, it proves that the language model has certain security risks.
[0005] For example, reference 2 (Rethinking LLM memorization through the lens of adversarial compression, Advances in Neural Information Processing Systems, 2024, 37: 56244-56267) discloses rethinking LLM memorization from the perspective of adversarial compression, proposes adversarial compression rate (ACR), and ACR is based on GCG algorithm to try to optimize multiple adversarial prompts of different lengths, and the ratio of the length of the shortest adversarial prompt to the length of the target harmful text is taken as the corresponding risk measure. That is, for a certain harmful content, if it can be induced by a shorter adversarial prompt, the risk of the language model generating this harmful content is greater.
[0006] Although the prior art can measure the security of the language model to a certain extent, the measurement of the security of the language model is insufficient. The above two methods are based on first-order gradient approximation search for adversarial prompts, and cannot ensure that a usable adversarial prompt is found for any harmful text, so in the case of search failure, the corresponding security is uncertain, and it cannot be determined whether there is no adversarial prompt to make the language model generate the corresponding harmful content, or it is not searched due to its technical limitations. Therefore, the disclosure of the prior art for the security of the language model is one-sided, and the defense mechanism constructed based thereon is also unreliable. SUMMARY
[0007] The purpose of the present application is to provide a language model security measurement method and device, which can ensure that the language model always successfully generates target harmful text, so that the obtained measurement result is the security in the worst case, and guarantees sufficient measurement of the security of the language model.
[0008] To achieve the above-mentioned purpose of the application, the language model security measurement method and device provided by the embodiments include a language model security measurement method and a device based on the language model security measurement.
[0009] In one embodiment, the language model security measurement method comprises the following steps: Given a harmful text and a pre-trained language model, construct an adversarial word embedding distribution, and at the same time construct a prior distribution based on the word embedding set pre-trained by the language model; Define the sensitivity of the adversarial word embedding distribution as the probability that a random sample from the adversarial word embedding distribution generates harmful text, and based on the adversarial word embedding distribution, the sensitivity of the adversarial word embedding distribution, and the constructed prior distribution, construct a security measurement index function; The optimization objective function is constructed based on the sensitivity objective function generated by the adversarial word embedding distribution and the security objective function based on the upper bound of the security metric, and the optimal adversarial word embedding distribution is obtained by optimization. The security metric is calculated based on the optimal adversarial word embedding distribution, and the security metric is obtained.
[0010] In one embodiment, the construction of the adversarial word embedding distribution comprises: Constructing a sequence of adversarial word embedding distributions , (e) is the first word embedding distribution, The total number of word tokens, and the sequence of word embedding distributions is modeled as a parameterized multi-dimensional Gaussian distribution Wherein, is the center of the Gaussian distribution, is the covariance matrix, is the word embedding dimension.
[0011] In one embodiment, the construction of the prior distribution based on the word embedding set pre-trained by the language model comprises: The word embedding set determined based on the pre-trained language model The prior distribution is constructed by kernel density estimation Describes the real word embedding distribution: , , Wherein, is the Gaussian kernel function, is the kernel function bandwidth matrix, denotes the size of the pre-trained language model vocabulary, denotes the value of the word embedding random variable, denotes the word embedding dimension, is the Mahalanobis distance between the random variable and the embedding of the first word token.
[0012] Further, the kernel function bandwidth matrix is determined according to Silverman's empirical rule, is a diagonal matrix, and each diagonal element is calculated by: , Wherein, is the diagonal element, is the variance of the language model pre-trained word embedding set in the first dimension.
[0013] In one embodiment, the sensitivity of the defined adversarial word embedding distribution includes: (1) a given piece of harmful text is represented by a token sequence , where denotes the token, and denotes the total number of tokens; (2) the sensitivity of the defined adversarial word embedding distribution is defined as: , where denotes a random sample from the adversarial word embedding distribution, i.e., a random adversarial prompt, denotes the text generated by the language model when inputting the random sample , and the function measures the similarity between the text generated by the language model and the harmful text . The function measures the similarity between the text generated by the language model and the harmful text by word-by-word matching: , where is an indicator function, and the function value is 1 when is 1, otherwise it is 0. In one embodiment, the random adversarial prompt generates harmful text by being inputted into the language model alone or embedded in the context and then inputted into the language model.
[0014] In one embodiment, based on the adversarial word embedding distribution, the sensitivity of the adversarial word embedding distribution, and the constructed prior distribution, a security measurement index function is constructed, denoted as:
[0015] , , , where denotes the sensitivity of the adversarial word embedding distribution; denotes the sensitivity threshold, which is in the range of 0-1 and takes a larger value; the security measurement index function indicates that when the sensitivity of the adversarial word embedding distribution is greater than or equal to the sensitivity threshold , the adversarial word embedding distribution sequence is more similar to the real word embedding distribution The minimum deviation is determined by constructing the deviation between the adversarial word embedding distribution sequence and the natural language distribution using KL divergence.
[0016] In one embodiment, the sensitivity threshold condition of the adversarial word embedding distribution is achieved by optimizing the negative log-likelihood function of the harmful text generated from random samples of the adversarial word embedding distribution, as expressed as: , in, This represents the negative log-likelihood function. Represents the distribution relative to adversarial word embeddings Expectations Indicates the first Each word element, This represents the total number of lexical units.
[0017] In one embodiment, the security objective function based on the upper bound of the metric is expressed as: , in, Represented as a security objective function, Indicates the number of adversarial word embeddings. This indicates the size of the vocabulary of the pre-trained language model. Indicates the first A distribution of adversarial word embeddings, Indicates random word embeddings. The first term determined by the pre-trained language model Word embedding.
[0018] In one embodiment, the optimization to obtain the optimal adversarial word embedding distribution is performed by searching for the optimal adversarial word embedding distribution parameters based on a gradient descent algorithm, including: (1) Initialize the adversarial word embedding distribution, specifically, for each adversarial word embedding distribution The center is initialized to the mean of the pre-trained word embeddings in the vocabulary, i.e. Covariance matrix Initialize it as a diagonal matrix, where the diagonal elements are the variances of the pre-trained set of word embeddings; (2) Based on the initial adversarial word embedding distribution, the maximum number of iterations is set to... The iteration begins, calculating the gradient of the negative log-likelihood function with respect to the parameters of the adversarial word embedding distribution. Gradient descent is used to update the distribution parameters of adversarial word embeddings; (3) Calculate the sensitivity based on the updated adversarial word embedding distribution parameters. If the value reaches the predetermined threshold... , jump to (4); otherwise, check the iteration number, if all iterations are completed, output +∞, the algorithm ends, otherwise, return to (2); (4) set the maximum iteration number as , start iteration; (5) based on the set maximum iteration number , respectively calculate the gradient of the negative log-likelihood function with respect to the parameters of the adversarial word embedding distribution and the gradient of the security objective function ; (6) calculate the value of the sensitivity , if the predetermined threshold is reached, proceed to (7), otherwise, set the weight variable to 1, jump to (10); (7) if , set the weight variable to 1, jump to (10), otherwise, proceed to (8); (8) if , set the weight variable to 0, jump to (10), otherwise, proceed to (9); (9) set the weight variable to , proceed to (10); (10) update the parameters of the adversarial word embedding distribution using the weighted gradient , if all iterations are completed, output the security objective function , the algorithm ends, the search for the optimal adversarial word embedding distribution is completed, otherwise, return to (5).
[0019] In one embodiment, the calculation of the security metric function based on the optimal adversarial word embedding distribution includes: based on the optimal adversarial word embedding distribution, calculating the security metric function by calculating the mathematical upper bound of the security metric function, denoted as: , wherein represents the mathematical upper bound of the security metric function, that is, the security objective function corresponding to the optimal adversarial word embedding distribution .
[0020] In order to clearly show the measurement method of the memory problem, a language model security measurement device is provided, comprising a memory for storing a computer program and a processor for implementing the language model security measurement method when executing the computer program.
[0021] Compared with the prior art, the present application has at least the following beneficial effects: The language model security measurement method provided by the present application, compared with the existing measurement method, constructs a continuous probability distribution based on a given harmful text and a pre-trained language model, generates a random adversarial prompt by sampling the continuous probability distribution, avoids the limitations of discrete search, ensures that the induced path can always be found, and guarantees the worst case, solves the problem of search failure caused by optimizing discrete adversarial suffix in the past; by constructing a prior distribution based on pre-training word embedding kernel density estimation, the random adversarial prompt is constrained to be close to the legal word embedding distribution, preventing the adversarial sample from deviating from the natural language feature, and by jointly optimizing the sensitivity and security target, the gradient is weighted to balance the two, realizing the security measurement in the worst case.
[0022] The present application also provides a language model security measurement device, which realizes the language model security measurement method. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced as follows.
[0024] Figure 1 The flowchart of the language model security measurement method.
[0025] Figure 2 The flowchart of the optimization of the optimal adversarial word embedding distribution. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical scheme and advantages of the present application clearer, the following will combine the drawings and The embodiments of the present application are further described in detail. It should be understood that the specific embodiments described herein are only used to explain the present application, and do not limit the protection scope of the present application.
[0027] Firstly, the language model is composed of a word table, an embedding layer, a coding layer and an output layer.
[0028] The word table contains a fixed number of tokens, which are represented as words, phrases, special symbols, etc. The word table can be represented as a set of integers after digitization , wherein is the size of the word table.
[0029] The embedding layer represents the token as a multi-dimensional token embedding, the token embedding of the where is the dimension of the token embedding, and the superscript denotes the type of embedding. After the language model is trained, the token embedding of each token is fixed.
[0030] The input token sequence is processed by the embedding layer into a token embedding sequence, and the encoding layer then performs a complex mapping on the token embedding sequence, finally inputting a hidden layer feature sequence of the same size, the sequence size being consistent with the input.
[0031] The input layer uses the last feature of the hidden layer feature sequence to make a prediction, and the prediction result is the next token of the input sequence.
[0032] The predicted new token is concatenated to the input sequence and then input to the language model to predict the next token. This process is repeated to realize language generation.
[0033] To fully measure the security of the language model, an embodiment provides a method for measuring the security of a language model, given a harmful text , a pre-trained language model , the invention outputs a scalar , representing the upper bound of the risk of the language model generating harmful text .
[0034] As shown in Figure 1 , the method for measuring the security of a language model provided by the embodiment is a worst-case security measurement, comprising the following steps: S1, given a harmful text and a pre-trained language model, construct an adversarial token embedding distribution, and construct a prior distribution based on the token embedding set of the language model pre-training.
[0035] Constructing an adversarial token embedding distribution: given a pre-trained language model and a harmful text, where the harmful text is represented as , consisting of tokens, is the number of the token.
[0036] Then, based on the token sequence, the embedding layer of the pre-trained language model is embedded into a token embedding distribution sequence , (e) is the token embedding distribution sequenceword embedding distributions, and model the sequence of word embedding distributions as a parametric multivariate Gaussian distribution wherein, is the center of the Gaussian distribution, is a covariance matrix, each word embedding distribution is independent of each other, and an adversarial word embedding distribution is obtained.
[0037] The present application optimizes continuous probability distribution , generates random adversarial prompts by sampling the probability distribution , and inputs the language model, which can avoid the limitations of discrete search and ensure that the induced path can always be found (worst-case guarantee).
[0038] Constructing a prior distribution: a set of word embeddings pre-trained based on a language model Construct a prior distribution to describe the distribution of real word embeddings. The construction can be realized by kernel density estimation: , , wherein, is a Gaussian kernel function, is a kernel function bandwidth matrix, denotes the size of the word table of the pre-trained language model, denotes the value of the word embedding random variable, denotes the embedding dimension of the word, is the Mahalanobis distance between the embedding of the first real word and the embedding of the second real word.
[0039] The kernel function bandwidth matrix is determined according to Silverman's rule of thumb, is a diagonal matrix, and each diagonal element is calculated by the following formula: , wherein, is the diagonal element, is the variance of the set of word embeddings pre-trained by the language model in the first dimension.
[0040] By constructing a prior distribution, the adversarial word embedding distribution can be used to constrain the adversarial word embedding distribution , which can prevent the adversarial sample from deviating from the natural language features.
[0041] S2, define the sensitivity of the adversarial word embedding distribution as the probability of generating harmful text from a random sample from the adversarial word embedding distribution, construct a security measurement index function based on the adversarial word embedding distribution and the sensitivity of the adversarial word embedding distribution and the constructed prior distribution.
[0042] Sensitivity of the adversarial word embedding distribution: given a piece of harmful text Through the word sequence , wherein represents the i-th word, , and the total number of words is defined as the sensitivity of the adversarial word embedding distribution: , wherein represents a random sample from the adversarial word embedding distribution, i.e. a random adversarial prompt, represents the text generated by the language model when inputting the random sample . The function measures the similarity between the text generated by the language model and the harmful text .
[0043] Further, The function measures the similarity between the text generated by the language model and the harmful text by matching word by word: , wherein is an indicator function, and the function value is 1 when represents the i-th word, , and the total number of words is defined as the sensitivity of the adversarial word embedding distribution: The sensitivity defined in the application quantifies the average success rate of the language model outputting harmful text
[0044] when a random adversarial prompt is sampled from the adversarial word embedding distribution sequence , and in the subsequent optimization of the adversarial word embedding distribution sequence , the sensitivity is maximized to ensure that even in the face of the strongest adversarial distribution, the risk can be fully exposed.
[0045] Based on the sensitivity of the adversarial word embedding distribution and the constructed prior distribution, a security metric function is built. The basic principle is: by optimizing the adversarial word embedding distribution, random adversarial hints from this distribution are more likely to be received. When input is given to a language model, the text generated by the language model differs from the given harmful text. They are exactly the same. The specific definitions are as follows: , , in, Indicates sensitivity to the distribution of adversarial word embeddings; This represents the sensitivity threshold, ranging from 0 to 1. A larger value is chosen to ensure that the language model output is very similar to y.
[0046] Security metrics function The sensitivity of the adversarial word embedding distribution is greater than or equal to the sensitivity threshold. Under these conditions, the distribution sequence of adversarial word embeddings Distribution with real word embeddings The minimum deviation is determined by constructing the KL divergence between the adversarial word embedding distribution sequence and the true word embedding distribution. A smaller value indicates lower security and a higher likelihood of the language model generating harmful content. The greater the probability, the less likely it is to occur; otherwise, the opposite is true.
[0047] S3. Based on the sensitivity objective function generated by the adversarial word embedding distribution and the security objective function based on the upper bound of the metric, construct an optimization objective function and optimize it to obtain the optimal adversarial word embedding distribution. Based on the optimal adversarial word embedding distribution, calculate the security metric function to obtain the security metric.
[0048] In this embodiment, an optimization objective function is constructed based on the threshold condition of sensitivity generated by the adversarial word embedding distribution and the security objective function based on the upper bound of the security metric, and then optimized to obtain the optimal adversarial word embedding distribution. The optimization algorithm searches for the optimal adversarial word embedding distribution parameters based on gradient descent.
[0049] The algorithm has two optimization objectives: the first objective is the negative log-likelihood function of the input harmful text generated by random word embeddings, called the sensitivity threshold condition, which is expressed as: , in, This represents the negative log-likelihood function. Indicates relative to adversarial word embedding Expectations Indicates the first Each word element, total number of tokens.
[0050] The second objective function is an upper bound of the metric, called the safety objective function, denoted as , where is the safety objective function, is the number of adversarial word embedding distributions, is the vocabulary size of the pre-trained language model, is the adversarial word embedding distribution, is a random word embedding, is the word embedding determined by the pre-trained language model.
[0051] The specific optimization process is shown in Figure 2 , including: (Step 1) Initialize the adversarial word embedding distribution, specifically, for each adversarial word embedding distribution , initialize its center as the mean of the pre-trained token embeddings in the vocabulary, i.e. ; the covariance matrix is initialized as a diagonal matrix, with the diagonal elements being the variance of the pre-trained word embedding set, and proceed to (Step 2).
[0052] (Step 2) Based on the initialized adversarial word embedding distribution, set the maximum number of iterations , start iteration, calculate the gradient of the negative log-likelihood function with respect to the adversarial word embedding distribution parameters , update the adversarial word embedding distribution parameters using gradient descent, and proceed to (Step 3).
[0053] (Step 3) Based on the updated adversarial word embedding distribution parameters, calculate the value of the sensitivity , if it reaches the predetermined threshold , jump to (Step 4); otherwise, check the number of iterations, if all iterations are completed, output +∞, and the algorithm ends, otherwise return to (Step 2).
[0054] (Step 4) Set the maximum number of iterations , start iteration, and proceed to (Step 5).
[0055] (Step 5) Based on the set maximum number of iterations , calculate the gradient of the negative log-likelihood function with respect to the adversarial word embedding distribution parameters and the gradient of the safety objective function (Step 6).
[0056] (Step 6) Calculate the sensitivity value, if a predetermined threshold is reached , proceed to (Step 7), otherwise set the weight variable to 1 and jump to (Step 10).
[0057] (Step 7) If , set the weight variable to 1 and jump to (Step 10), otherwise proceed to (Step 8).
[0058] (Step 8) If , set the weight variable to 0 and jump to (Step 10), otherwise proceed to (Step 9).
[0059] (Step 9) Set the weight variable to and proceed to (Step 10).
[0060] (Step 10) Update the adversarial word embedding distribution parameters using the weighted gradient , if all iterations are completed, output the security objective function , the algorithm ends and the search for the optimal adversarial word embedding distribution is completed, otherwise return to (Step 5).
[0061] Security metric upper bound: Considering that the definition of the above security metric index function is not easy to calculate, the security metric index function can actually be calculated by calculating the upper bound of the mathematical security metric index function: , where represents the upper bound of the mathematical security metric index function, i.e. the security objective function corresponding to the optimal adversarial word embedding distribution.
[0062] Based on the optimal adversarial word embedding distribution, the security metric index function is calculated to obtain the security metric index, and the language model security metric is realized.
[0063] To solve these problems, the application also provides a language model security metric device, comprising a memory and a processor, the memory is used to store a computer program, and the processor is used to realize the language model security metric method when the computer program is executed.
[0064] It should be noted that the language model security measurement method provided by the above embodiment and the language model security measurement device embodiment belong to the same concept, and the specific implementation process is detailed in the language model security measurement method embodiment, which will not be repeated here.
[0065] In summary, the present application provides an innovative probability perspective for adversarial security analysis by constructing adversarial word embedding distribution and prior distribution, and proposes a security measurement in the worst case for language models, which is sufficient and thorough. By jointly optimizing sensitivity and security goals, the gradient is weighted to balance the two, and the security measurement in the worst case is realized to ensure sufficient measurement of language model security.
[0066] The specific embodiments described above have detailed the technical solutions and beneficial effects of the present application. It should be understood that the above description is only the most preferred embodiment of the present application and is not intended to limit the present application. Any modifications, supplements, and equivalent replacements made within the principle range of the present application shall be included in the protection scope of the present application.
Claims
1. A method for measuring the security of a language model, characterized in that, Includes the following steps: Given a piece of harmful text and a pre-trained language model, construct an adversarial word embedding distribution, and construct a prior distribution based on the word embedding set pre-trained by the language model. The sensitivity of the adversarial word embedding distribution is defined as the probability of generating harmful text from random samples derived from the adversarial word embedding distribution. Based on the adversarial word embedding distribution, its sensitivity, and the constructed prior distribution, a security metric function is built. An optimization objective function is constructed based on the sensitivity objective function generated by the adversarial word embedding distribution and the security objective function based on the upper bound of the metric. The optimal adversarial word embedding distribution is obtained through optimization. The security metric function is then calculated based on the optimal adversarial word embedding distribution to obtain the security metric.
2. The method for measuring the security of a language model according to claim 1, characterized in that, The construction of the adversarial word embedding distribution includes: Construct an adversarial word embedding distribution sequence , (e) is the first Word embedding distribution, The total number of lexical units is given, and the word embedding distribution sequence is modeled as a parameterized multidimensional Gaussian distribution. ,in, It is the center of the Gaussian distribution. Let covariance matrix be the variance matrix. For word embedding dimension.
3. The method for measuring the security of a language model according to claim 1, characterized in that, The construction of the prior distribution based on the word embedding set pre-trained by the language model includes: Word embedding set determined by pre-trained language model Constructing a prior distribution through kernel density estimation Describe the distribution of real word embeddings: , , in, For Gaussian kernel function, The kernel function bandwidth matrix, This indicates the size of the vocabulary of the pre-trained language model. The value of the word embedding random variable is represented. Indicates word embedding dimension, For random variables and the Embedding of real lexical units The Mahalanobis distance between them.
4. The method for measuring the security of a language model according to claim 2, characterized in that, The sensitivity of the defined adversarial word embedding distribution includes: (1) A given piece of harmful text Through word sequence It means that among them Indicates the first Each word element, The total number of lexical units; (2) Define the adversarial word embedding distribution The sensitivity is: , in, This represents a random sample from an adversarial word embedding distribution, i.e., a random adversarial cue. Indicates input random sample Temporal language model The generated text, Functional metrics language model generates text and harmful text Similarity between them; The function measures the text generated by the language model by matching each word individually. and harmful text Similarities between them: , in For indicator functions, The function value is 1 when the value is 1, otherwise it is 0.
5. The method for measuring the security of a language model according to claim 3, characterized in that, Based on the adversarial word embedding distribution and its sensitivity and the constructed prior distribution, a security metric function is constructed, expressed as: , , in, Indicates sensitivity to the distribution of adversarial word embeddings; This represents the sensitivity threshold, ranging from 0 to 1, with the larger value taken; the security metric function. The sensitivity of the adversarial word embedding distribution is greater than or equal to the sensitivity threshold. Under these conditions, the distribution sequence of adversarial word embeddings Distribution with real word embeddings The minimum deviation is determined by constructing the KL divergence between the adversarial word embedding distribution sequence and the true word embedding distribution.
6. The method for measuring the security of a language model according to claim 5, characterized in that, The threshold condition for sensitivity to the adversarial word embedding distribution is achieved by minimizing the negative log-likelihood function of the harmful text generated from random samples from the adversarial word embedding distribution, and is expressed as: , in, This represents the negative log-likelihood function. Indicates relative to adversarial word embedding Expectations Indicates the first Each word element, This represents the total number of lexical units.
7. The method for measuring the security of a language model according to claim 6, characterized in that, The security objective function based on the upper bound of the metric is expressed as: , in, Represented as a security objective function, Indicates the number of adversarial word embeddings. This indicates the size of the vocabulary of the pre-trained language model. Indicates the first A distribution of adversarial word embeddings, Indicates random word embeddings. The first term determined by the pre-trained language model Word embedding.
8. The method for measuring the security of a language model according to claim 7, characterized in that, The optimization to obtain the optimal adversarial word embedding distribution is performed by searching for the optimal adversarial word embedding distribution parameters based on the gradient descent algorithm, including: (1) Initialize the adversarial word embedding distribution, specifically, for each adversarial word embedding distribution The center is initialized to the mean of the pre-trained word embeddings in the vocabulary, i.e. Covariance matrix Initialize it as a diagonal matrix, where the diagonal elements are the variances of the pre-trained set of word embeddings; (2) Based on the initial adversarial word embedding distribution, the maximum number of iterations is set to... The iteration begins, calculating the gradient of the negative log-likelihood function with respect to the parameters of the adversarial word embedding distribution. Gradient descent is used to update the distribution parameters of adversarial word embeddings; (3) Calculate the sensitivity based on the updated adversarial word embedding distribution parameters. If the value reaches the predetermined threshold... Jump to (4); otherwise check the number of iterations. If all iterations have been completed... If the iteration is +∞, the algorithm ends; otherwise, it returns (2). (4) Set the maximum number of iterations to Start iterating; (5) Based on the set maximum number of iterations Calculate the gradient of the negative log-likelihood function with respect to the parameters of the adversarial word embedding distribution. and the gradient of the security objective function ; (6) Calculate sensitivity If the value reaches the predetermined threshold... Perform (7), otherwise set the weight variable. If the value is 1, jump to (10); (7) If Set weight variables If the result is 1, jump to (10); otherwise proceed to (8). (8) If Set weight variables If the value is 0, jump to (10); otherwise proceed to (9). (9) Weight variables Set as , proceed (10); (10) Using weighted gradients Update the parameters of the adversarial word embedding distribution if all steps are completed. The next iteration outputs the security objective function. The algorithm ends when the search for the optimal adversarial word embedding distribution is completed; otherwise, it returns (5).
9. The method for measuring the security of a language model according to claim 8, characterized in that, The calculation of the security metric function based on the optimal adversarial word embedding distribution includes: calculating the security metric function by calculating the upper bound of the mathematically meaningful security metric function based on the optimal adversarial word embedding distribution, expressed as: , in, This represents the upper bound of a mathematically defined safety metric function. That is, the security objective function corresponding to the optimal adversarial word embedding distribution. .
10. An apparatus for measuring the security of a language model, comprising a memory and a processor, the memory for storing a computer program, characterized in that, The processor is configured to implement the method for measuring the language model security as described in any one of claims 1 to 9 when executing the computer program.