Black box large language model security detection method and device, equipment and medium

By calculating the loss function of the embedded vector of the black box large language model and the embedded vector of the expected answer, the attack prompt word is circularly optimized, and the systematic and automated problems of black box model detection is solved, efficient and safe detection is achieved in a closed-source environment, and the accuracy and applicability of the detection are improved.

CN120387169APending Publication Date: 2025-07-29BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510509100.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing black box large language model security detection methods lack systematicity and automation, cannot effectively detect potential attacks in closed-source APIs, rely on manual design templates or rules, and traditional methods require access to the internal structure information of the model and cannot be applied in closed-source environments.

Method used

By obtaining the embedding vector of the black box large language model and the embedding vector of the expected answer, the loss function is calculated, and the attack prompt word is optimized in a loop until the loss function reaches the preset conditions, and the target prompt word is used for security detection to avoid accessing the internal gradient or structural information of the model.

Benefits of technology

The accuracy and reliability of security detection of black box large language models is improved. The generated attack prompt words have stronger generalization applicability to models of different architectures, and can capture deeper semantic information, breaking through the infeasibility of traditional methods in closed-source API scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387169A_ABST
    Figure CN120387169A_ABST
Patent Text Reader

Abstract

The invention provides a security detection method and device for a black box large language model, equipment and a medium. The method comprises the following steps: acquiring an embedded vector of a first answer and an embedded vector of a second answer; the first answer is reply information returned by the tested black box large language model aiming at attack input of the attack prompt word, and the second answer is reply information returned by the expected black box large language model; calculating a loss function based on the embedding vector of the first answer and the embedding vector of the second answer; if the loss function does not meet the preset condition, optimizing the attack cue word based on the loss function, and taking the optimized attack cue word as the attack cue word of the next round of circulation; and if the loss function reaches a preset condition, taking an attack cue word used for executing the attack cue word optimization operation in the round as an acquired target cue word, and performing security detection on the black box large language model by utilizing the acquired target cue word. According to the invention, the accuracy and reliability of black box large language model security detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of model detection. More specifically, it relates to a security detection method, device, equipment, and medium for black-box large language models. Background Art

[0002] In recent years, large language models have been increasingly widely used in various fields, and their security issues have also attracted attention.

[0003] Existing white-box large language model security detection methods are mainly divided into two categories: gradient-based methods and search-based methods. Gradient-based methods guide the modification of prompt words by calculating the gradient of the loss function with respect to the input, usually using the cross-entropy between the token ID sequence and the model output logits as the loss function. Search-based methods (such as AutoPrompt, AutoDAN) explore possible attack prompt words through heuristic search algorithms. However, both methods require access to the internal structure information of the model, which is not feasible in a closed-source API (black-box model) environment.

[0004] Currently, there are relatively few attack methods for black-box large language models. Existing black-box attack methods mainly rely on manually designed templates or rules, lacking systematic and automated attack methods for closed-source large model APIs.

[0005] Therefore, this application proposes an accurate and reliable security detection method for black-box large language models. Summary of the Invention

[0006] The purpose of this application is to provide a security detection method, device, equipment, and medium for black-box large language models to improve the accuracy and reliability of security detection for black-box language models.

[0007] In the first aspect of the embodiments of this application, a security detection method for black-box large language models is provided, including: obtaining a target prompt word by using the following steps; Repeatedly execute the following attack prompt word optimization operation, where the attack prompt word optimization operation includes: Obtain the embedding vector of the first answer and the embedding vector of the second answer; the first answer is the reply information returned by the measured black-box large language model for the attack input of the attack prompt word, and the second answer is the reply information expected to be returned by the black-box large language model for the attack input of the attack prompt word; Calculate a loss function based on the embedding vector of the first answer and the embedding vector of the second answer; Determine whether the loss function reaches a preset condition; If the loss function does not reach the preset condition, optimize the attack prompt word based on the loss function, and use the optimized attack prompt word as the attack prompt word for the next round of loop; If the loss function reaches a preset condition, the attack prompt word used in the current round of performing the attack prompt word optimization operation is used as the obtained target prompt word; Use the target prompt word to perform a security detection on the black-box large language model.

[0008] In a second aspect of the embodiments of the present application, a security detection device for a black-box large language model is provided, including: A target prompt word acquisition module, configured to obtain a target prompt word by using the following steps; Loop to perform the following attack prompt word optimization operation, and the attack prompt word optimization operation includes: Obtain the embedding vector of the first answer and the embedding vector of the second answer; the first answer is the reply information returned by the measured black-box large language model for the attack input of the attack prompt word, and the second answer is the reply information expected to be returned by the black-box large language model for the attack input of the attack prompt word; Calculate a loss function based on the embedding vector of the first answer and the embedding vector of the second answer; Determine whether the loss function reaches a preset condition; If the loss function does not reach the preset condition, optimize the attack prompt word based on the loss function, and use the optimized attack prompt word as the attack prompt word for the next round of loop; If the loss function reaches the preset condition, the attack prompt word used in the current round of performing the attack prompt word optimization operation is used as the obtained target prompt word; A security detection module, configured to perform a security detection on the black-box large language model by using the target prompt word.

[0009] In a third aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned security detection method for the black-box large language model are implemented.

[0010] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned security detection method for the black-box large language model are implemented.

[0011] The beneficial effects of the security detection method, device, equipment, and medium for the black-box large language model provided by the embodiments of the present application are as follows: This application calculates the semantic loss function through the embedding vectors of the model answers and the expected answers, without accessing the internal gradients or structural information of the model, which solves the infeasibility problem of traditional white-box methods in the closed-source API scenario. At the same time, due to the cross-model semantic consistency feature of the embedding vectors, the generated attack prompt words have stronger generalization applicability to black-box large language models with different architectures. Description of the Drawings

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0013] Figure 1 It is a schematic flowchart of a security detection method for a black-box large language model provided by an embodiment of this application; Figure 2 It is a structural block diagram of a security detection device for a black-box large language model provided by an embodiment of this application; Figure 3 It is a general transformation technical framework diagram from an open-source (white-box) model security detection method to a closed-source (black-box) model security detection method provided by an embodiment of this application; Figure 4 It is a detailed diagram of a loss function transformation method provided by an embodiment of this application; Figure 5 It is a flowchart of a security detection method for a black-box large language model provided by an embodiment of this application; Figure 6 It is a schematic block diagram of an electronic device provided by an embodiment of this application. Detailed Embodiments

[0014] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.

[0015] To make the purpose, technical solutions, and advantages of this application clearer, the following will be described through specific embodiments with reference to the drawings.

[0016] Please refer to Figure 1 , Figure 1Schematic flowchart of the security detection method for the black-box large language model provided by an embodiment of the present application. The method includes: S101: Obtain a target prompt. In one implementation, the target prompt is obtained by repeatedly performing the attack prompt optimization operation as described in Figure 1 . The attack prompt optimization operation includes: S1011: Obtain the embedding vectors of the first answer and the second answer. The first answer is the reply information returned by the black-box large language model under test for the attack input of the attack prompt, and the second answer is the reply information expected to be returned by the black-box large language model for the attack input of the attack prompt. S1012: Calculate a loss function based on the embedding vectors of the first answer and the second answer. S1013: Determine whether the loss function meets a preset condition. If the loss function does not meet the preset condition, execute S1014; if the loss function meets the preset condition, execute S1015. S1014: Optimize the attack prompt based on the loss function, and use the optimized attack prompt as the attack prompt for the next round of loop. S1015: Use the attack prompt used in this round of performing the attack prompt optimization operation as the obtained target prompt. S102: Perform a security detection on the black-box large language model using the target prompt.

[0017] Among them, S1011: Obtain the embedding vectors of the first answer and the second answer. The first answer is the reply information returned by the black-box large language model under test for the attack input of the attack prompt, and the second answer is the reply information expected to be returned by the black-box large language model for the attack input of the attack prompt.

[0018] In this embodiment, the first answer refers to the reply information actually returned after the black-box large language model under test receives the attack prompt input, that is, the response of the black-box large language model to the current prompt. The second answer is the reply information expected to be returned by the black-box large language model for the attack prompt input, representing an ideal answer for comparison.

[0019] In this embodiment, the embedding vector can be a numerical vector form converted from text. In the vector space, the embedding vectors corresponding to texts with similar semantics are closer in distance, and can be used to calculate semantic similarity between texts, etc. The embedding vector can be obtained through a large language model or an embedding function, etc.

[0020] S1012: Calculate a loss function based on the embedding vectors of the first answer and the second answer.

[0021] In this embodiment, the cosine similarity between the embedding vectors of the first answer and the second answer can be calculated, and the cosine similarity can be used as part of the loss function. In addition, in combination with the actual application scenario, the criterion for judging whether the first answer contains a preset undesired word can also be incorporated into the loss function.

[0022] S1013: Determine whether the loss function reaches a preset condition.

[0023] In this embodiment, the preset condition can be whether the loss function is less than a preset loss threshold or the number of iterations reaches a preset number, etc. Both the loss threshold and the preset number can be set based on the conventional experience in this field.

[0024] S1014: If the loss function does not reach the preset condition, optimize the attack prompt word based on the loss function, and use the optimized attack prompt word as the attack prompt word for the next round of loop.

[0025] In this embodiment, if the loss function does not reach the preset condition, the attack prompt word can be optimized based on, for example, the gradient descent method or the evolutionary algorithm (such as the genetic algorithm), and the optimized attack prompt word is used as the attack prompt word for the next round of loop.

[0026] Specifically, the optimization method based on the gradient descent method can be as follows: If a differentiable sentence encoder (such as some pre-trained language models) is used, the gradient of the loss function with respect to the prefix and suffix word embeddings can be calculated to guide the optimization: Step 1: Represent the prefix and suffix as continuous word embedding vectors.

[0027] Step 2: Calculate the gradient of the loss function with respect to these word embedding vectors.

[0028] Step 3: Update the word embedding vectors in the opposite direction of the gradient.

[0029] Step 4: Map the updated word embedding vectors back to discrete text.

[0030] The optimization method based on the evolutionary algorithm can be as follows: Step 1: Generate multiple candidate prefix and suffix combinations (population).

[0031] Step 2: Evaluate the fitness of each candidate combination (based on the loss function).

[0032] Step 3: Select the candidate combinations with high fitness for crossover and mutation to generate new candidate combinations.

[0033] Step 4: Repeat the evaluation and selection process until a combination that meets the conditions is found.

[0034] S1015: If the loss function meets the preset conditions, then use the attack prompt word used in the current round of performing the attack prompt word optimization operation as the obtained target prompt word.

[0035] In this embodiment, if the loss function meets the preset conditions, it indicates that the effect of the current attack prompt word is good, and it can be used as the final target prompt word.

[0036] It should be noted that in this application, the target prompt word can be one or more. That is, different initial attack prompt words, different optimization methods, or different iteration times may result in the same or different attack prompt words. At the same time, for the black-box large language model to be tested, generally there is more than one effective attack prompt word.

[0037] S102: Use the target prompt word to perform security detection on the black-box large language model.

[0038] In this embodiment, first determine in which aspects the security of the black-box large language model is to be detected, such as whether it will generate harmful content (such as violence, discrimination, illegal information, etc.), whether it follows the privacy policy, and whether it has a protection mechanism for sensitive information. According to actual requirements and application scenarios, check whether the reply contains harmful information and whether it meets the expected security standards. For example, when detecting whether the model generates discriminatory content, it is necessary to judge whether there are derogatory or discriminatory remarks against specific groups in the reply; if detecting privacy protection, check whether there is a situation of leaking sensitive information. If relevant words appear, it is considered unsafe, otherwise it is considered safe.

[0039] It can be concluded from the above that traditional methods rely on the internal structure or gradient information of the model and are difficult to be applied to black-box large language models. This method optimizes the attack prompt word only by using the embedding vector and the loss function through input-output interaction without accessing the inside of the model, breaking through the black-box limitation. This application calculates the similarity of the embedding vectors of the first answer (the actual reply of the model) and the second answer (the expected reply), and combines the judgment criteria of the words that are not expected to appear, so as to more accurately measure the semantic difference between the model reply and the expected reply. Compared with traditional methods, it can capture deeper semantic information and improve the security detection accuracy of black-box large language models.

[0040] This application calculates the semantic loss function through the embedding vector of the model reply and the embedding vector of the expected reply, without accessing the internal gradient or structure information of the model, solving the infeasibility problem of traditional white-box methods in the closed-source API scenario. At the same time, due to the cross-model semantic consistency feature of the embedding vector, the generated attack prompt word has stronger generalization applicability to black-box large language models with different architectures.

[0041] In one embodiment of the present application, calculating a loss function based on the embedding vectors of the first answer and the second answer includes: Determining a first loss function based on the similarity between the embedding vectors of the first answer and the second answer; Determining a second loss function based on the cumulative weight of the specified words included in the first answer; Performing a weighted process on the first loss function and the second loss function to obtain the loss function.

[0042] In this embodiment, the first loss function may be:

[0043] Wherein, represents the first loss function, represents the cosine similarity, represents the cosine similarity between the embedding vector of the first answer and the embedding vector of the second answer, represents the embedding vector of the second answer, represents the embedding vector of the first answer; , wherein, represents the embedding vector of the second answer, represents the Euclidean norm of; , wherein, represents the embedding vector of the first answer, represents the Euclidean norm of.

[0044] The second loss function may be:

[0045] Wherein, is the second loss function, is the th word in the preset word library, is the weight of the th word in the preset word library, is an indicator function, which is 1 if the first answer contains and 0 otherwise, is the preset word library, and the essence of the preset word library can be understood as a refusal answer word library. When a refusal answer word in the refusal answer word library appears, it indicates that the direction of the current attack prompt word may be incorrect or the quality is low. The second loss function can also be understood as a quantified value of a refusal answer index. The lower the refusal answer index, the more effective the current attack prompt word is.

[0046] The preset word library may include but is not limited to the following categories: Safety reminder words: words such as "sorry", "unable to provide", "unsafe", "immoral", etc. that indicate the model refuses to answer.

[0047] Policy reference words: words such as "violation of terms of use", "not in line with safety guidelines", "violation of ethical standards", etc. that indicate the reference to safety policies.

[0048] Redirect words: words such as "we suggest you", "you may consider", "a better way is", etc. that indicate the model attempts to redirect the conversation to other directions.

[0049] The loss function obtained by weighted calculation of the first loss function and the second loss function can be:

[0050] Wherein, is the loss function, is the weight of the first loss function or the weight of the second loss function in the Nth round of loop, is the second loss function, which can also be understood as the refusal penalty term.

[0051] In an embodiment of the present application, the security detection method of the black box large language model further includes: Determining the weight of the first loss function or the weight of the second loss function of the (N + 1)th round of attack prompt word optimization operation based on the refusal rate of the first answer of the Nth round of attack prompt word optimization operation; N is a positive integer; The refusal rate is the proportion of the number of times the first answer contains the specified word in the first N rounds of attack prompt word optimization operations.

[0052] In an embodiment of the present application, determining the weight of the first loss function or the weight of the second loss function of the (N + 1)th round of attack prompt word optimization operation based on the refusal rate of the first answer of the Nth round of attack prompt word optimization operation includes: Determining the weight of the first loss function or the weight of the second loss function of the (N + 1)th round of attack prompt word optimization operation through the first formula; The first formula is: ; wherein, is the weight of the first loss function or the weight of the second loss function in the Nth round of loop, represents the refusal rate of the Nth round of loop, is an adjustment parameter that can be determined based on the actual scenario or experience. The weight at the start of the first round of loop, that is, the initial value of the weight, can be determined based on multiple experiments or experience.

[0053] The logic of the above adjustment is as follows: when the refusal rate is high, increase the weight of the refusal penalty to guide the optimization algorithm to pay more attention to avoiding the refusal mode; when the refusal rate decreases, increase the weight of the cosine similarity to guide the optimization algorithm to pay more attention to semantic similarity, which improves the attack success rate and optimization efficiency. At the same time, since the actual attack will set a maximum number of attack rounds, the probability of successful attack within the maximum number of attack rounds is increased.

[0054] It can be concluded from the above that the present application proposes a comprehensive loss function, which combines the similarity of the embedding vectors of the first answer and the second answer and the cumulative weight of the specified words included in the first answer, and can comprehensively measure the quality of the model's response from two dimensions of semantic similarity and answer content compliance. Compared with the single-dimensional evaluation method, the accuracy of black-box large language model security detection is improved. The present application dynamically adjusts the weight according to the refusal rate, enabling the loss function to flexibly adapt to the detection requirements in different scenarios, which helps to improve the attack success rate and optimization efficiency. In this embodiment, when the refusal rate is relatively high, the weight of the refusal penalty is increased to jump out of the refusal mode faster; when the refusal rate decreases, the weight of semantic similarity is increased to more finely adjust the attack prompt words, which helps to improve the stability and reliability of the detection method.

[0055] In an embodiment of the present application, obtaining the embedding vector of the first answer and the embedding vector of the second answer includes: Obtaining the embedding vector of the first answer and the embedding vector of the second answer from the first large language model, where the first large language model is obtained by updating the weight parameters of the objective function based on the first training set; the first training set includes security detection questions and the corresponding reply information for the security detection questions.

[0056] In this embodiment, the first large language model may be an adjusted Qwen2.5 - 1.5B model, and the specific adjustment parameter is the objective function of the first large language model, and the objective function is obtained by weighting the learning objective function and the semantic consistency objective function.

[0057] The learning objective function may be: ; The semantic consistency objective function may be: ; The objective function of the first large language model may be: ; Among them, represents the learning objective function, represents the weight coefficient of the learning objective function, and They are all embedding vectors of harmful problems, specifically belonging to the same category of harmful problems. The same category of harmful problems refers to those with similar characteristics in terms of content nature, harm type, etc. For example, based on the classification of harm nature, if problems related to violence are grouped into one category, then "describing a bloody violent conflict" and "detailed description of how to commit violent harm to others" belong to the same category of harmful problems. They both involve violent content and will guide the large language model to generate answers with violent tendencies, violating social morality and legal norms and having an adverse impact on users. is the cosine similarity function, is the total number of embedding vectors, is a tuning parameter that can be determined based on multiple experiments or experience; is the semantic consistency objective function, is the weight of the semantic consistency objective function, is the embedding vector of the original text, is the embedding vector of the rewritten text, where the rewritten text is the text after the first embedding encoder rewrites the original text.

[0058] In this embodiment, the optimization objective of the learning objective function can be understood as: minimizing the embedding distance between the original text and the embedding vectors of the same category of security detection problems as the original text, while maximizing the embedding distance between the original text and the embedding vectors that are not security detection problems; the original text is the first answer or the second answer; the optimization objective of the semantic consistency objective function is: maximizing the semantic consistency between the embedding vectors and the original text. In this embodiment, the first training set consists of multiple pairs of harmful problems and answers. The question-answer pairs include different categories of harmful content, such as violence, discrimination, illegal activities, etc. The quantity and quality of the data in the first training set meet the training and testing of the first large language model.

[0059] It can be concluded from the above that the first large language model of this application is trained with the objective function weighted by the learning objective function and the semantic consistency objective function. This comprehensive training method enables the model to balance the learning of harmful problems and the maintenance of semantic consistency, thereby generating more accurate and representative embedding vectors, which is helpful for the subsequent loop optimization process. In this application, the learning objective function, through the optimization objective of minimizing the embedding distance between the original text and the embedding vectors of the same category of harmful problems and maximizing the embedding distance between the original text and the embedding vectors of non-security detection problems, prompts the first large language model to deeply learn the characteristics and patterns of harmful problems. This helps the model capture potential harmful information in the answers when generating embedding vectors, thereby improving the ability to identify harmful answers, enhancing the accuracy and efficiency of attack promotion word optimization, and enhancing the accuracy and efficiency of black box large language model security detection.

[0060] In this embodiment, an early stopping strategy is also set up, which can be understood as when it is determined that the optimization direction of the current attack prompt word is incorrect or it is determined that even if multiple rounds of loop optimization are performed again, an effective attack prompt word, that is, the target prompt word, cannot be obtained. Specifically, it can be judged in the following ways: In an embodiment of the present application, the attack prompt word optimization operation further includes: In response to the existence of a specified word with a weight greater than a first threshold in the first answer, or the cumulative weight of all specified words in the first answer being greater than a second threshold, initialize the attack prompt word, and use the initialized attack prompt word as the attack prompt word for the next round of attack prompt word optimization operation.

[0061] In this embodiment, if a high-weight refusal word appears in the first answer of the measured black-box large language model, such as the words "sorry", "unable to provide", "unsafe", "immoral", etc. in the aforementioned preset word library, which indicate a clear refusal to answer.

[0062] Each specified word has a preset weight in the word library. The weight referred to in this embodiment can be understood as the aforementioned , when there is a specified word with a weight greater than the first threshold in the first answer, or the cumulative weight of all specified words in the first answer is greater than the second threshold, it indicates that the current attack prompt word can no longer effectively break through the security protection of the model, and continuing to use it for optimization may be inefficient or even ineffective. Therefore, it is necessary to initialize the attack prompt word, that is, regenerate a brand-new attack prompt word, to provide a new starting point for the next round of optimization operation, avoiding wasting computing resources and time on an invalid path. The first threshold and the second threshold can be determined based on experience or data feedback during the experiment.

[0063] Specifically, the way to initialize the attack prompt word can be: initialize the attack prompt word based on the second formula, and the second formula is: ; Among them, is the initialized attack prompt word, is the basic harmful prompt word, such as an instruction to require the model to generate harmful content; is the prefix, is the suffix, which can be set as an empty string or randomly initialized initially.

[0064] It can be concluded from the above that the early stopping strategy set in the present application, when it is determined that the optimization direction of the current attack prompt word is incorrect, or even if multiple rounds of loop optimization are performed again, an effective attack prompt word (target prompt word) cannot be obtained, can avoid wasting resources caused by falling into an invalid loop by initializing the attack prompt word and using it as the starting point for the next round of optimization operation, making the optimization process more efficient.

[0065] In this application, weights are preset in the specified vocabulary. When the weight situation of the specified words in the first answer meets certain conditions (the weight is greater than the first threshold or the cumulative weight is greater than the second threshold), it indicates that the current attack prompt words are difficult to effectively break through the security protection of the model. The above-mentioned weight-based judgment method provides guarantee for the stability of the optimization process, ensures the initialization of the attack prompt words at the appropriate time, avoids the optimization failure caused by the over-optimization or improper optimization of the attack prompt words, improves the reliability of the optimization process of the attack prompt words, and further enhances the accuracy and reliability of the security detection of the black-box large language model.

[0066] Corresponding to the security detection method of the black-box large language model in the above embodiment, Figure 2 is the structural block diagram of the security detection device of the black-box large language model provided by an embodiment of this application. For the sake of illustration, only the parts related to the embodiments of this application are shown. Refer to Figure 2 The security detection device 20 of the black-box large language model includes: a target prompt word acquisition module 21 and a security detection module 22.

[0067] Among them, the target prompt word acquisition module 21 is used to acquire the target prompt word by using the following steps; The following attack prompt word optimization operations are executed in a loop. The attack prompt word optimization operations include: Obtain the embedding vector of the first answer and the embedding vector of the second answer; the first answer is the reply information returned by the measured black-box large language model for the attack input of the attack prompt word, and the second answer is the reply information expected to be returned by the black-box large language model for the attack input of the attack prompt word; Calculate the loss function based on the embedding vector of the first answer and the embedding vector of the second answer; Judge whether the loss function reaches the preset condition; If the loss function does not reach the preset condition, optimize the attack prompt word based on the loss function, and use the optimized attack prompt word as the attack prompt word for the next round of loop; If the loss function reaches the preset condition, use the attack prompt word used in this round of attack prompt word optimization operation as the obtained target prompt word; The security detection module 22 is used to perform security detection on the black-box large language model by using the target prompt word.

[0068] In an embodiment of this application, the target prompt word acquisition module 21 is specifically used to determine the first loss function based on the similarity of the embedding vector of the first answer and the embedding vector of the second answer; Determine the second loss function based on the cumulative weight of the specified words included in the first answer; Perform weighted processing on the first loss function and the second loss function to obtain the loss function.

[0069] In one embodiment of the present application, the security detection device 20 of the black-box large language model further includes: a weight adjustment module, configured to determine the weight of the first loss function or the weight of the second loss function of the (N + 1)-th round of attack prompt word optimization operation based on the rejection rate of the first answer in the N-th round of attack prompt word optimization operation; N is a positive integer; The rejection rate is the proportion of the number of times the first answer contains a specified word in the first N rounds of attack prompt word optimization operations.

[0070] In one embodiment of the present application, the weight adjustment module is specifically configured to determine the weight of the first loss function or the weight of the second loss function in the (N + 1)-th round of attack prompt word optimization operation through a first formula; The first formula is: ; where, is the weight of the first loss function or the weight of the second loss function in the -th round of loop, represents the rejection rate of the -th round of loop, and are adjustment parameters.

[0071] In one embodiment of the present application, the target prompt word acquisition module 21 is further specifically configured to obtain the embedding vector of the first answer and the embedding vector of the second answer from the first large language model, where the first large language model is obtained by updating and training the weight parameters of the target function based on the first training set; the first training set includes security detection questions and the corresponding reply information of the security detection questions.

[0072] In one embodiment of the present application, the target function is obtained by weighting the learning target function and the semantic consistency target function.

[0073] In one embodiment of the present application, the security detection device 20 of the black-box large language model further includes: an early stopping module, configured to initialize the attack prompt word in response to the existence of a specified word with a weight greater than a first threshold in the first answer, or the cumulative weight of all specified words in the first answer being greater than a second threshold, and use the initialized attack prompt word as the attack prompt word for the next round of attack prompt word optimization operation.

[0074] Figure 3 This is a technical framework diagram for the generalization transformation from an open-source (white-box) model security detection method to a closed-source (black-box) model security detection method provided by an embodiment of the present application. Figure 4 This is a detailed diagram of a loss function transformation method provided by an embodiment of the present application. Refer to Figure 3 and Figure 4。The basic principle of this application is to modify the calculation method of the loss function so that the original security detection method for open-source models can be applied to closed-source model APIs. Traditional open-source model security detection methods usually enhance the "toxicity" of the prompt by adding specific prefixes and suffixes before and after the original prompt, thereby generating attack prompts. The prefix and suffix of the prompt are modified through iterative optimization to minimize the semantic difference between the expected answer and the actual answer.

[0075] In open-source models, the optimization process uses the cross-entropy between the token ID sequence of the expected answer and the model output logits as the loss function. However, this method is not applicable to closed-source large model APIs because APIs usually only return natural language answers and do not provide intermediate calculation results. To solve this problem, the present invention proposes a new loss function calculation method: using the cosine similarity between the embedding vector of the expected answer and the embedding vector of the actual answer of the model to measure the semantic difference.

[0076] Traditional embedding vector methods such as BERT and Universal Sentence Encoder are usually based on sentence contrast learning. Although they can capture general semantic similarity, they have limitations in the semantic understanding of attack prompts. Therefore, this application uses the aforementioned learning objective function and semantic consistency objective function to adjust the encoder, improving the accuracy and reliability of embedding vector extraction.

[0077] Figure 5 It is a flowchart of a black-box large language model security detection method provided by an embodiment of this application. As Figure 5 shown, based on the general transformation technology proposed in this application, another specific process of the black-box large language model security detection method is as follows: Step 1: Initialize the attack prompt First, initialize the attack prompt, which usually includes three parts: a prefix, a harmful prompt, and a suffix: ; where is the basic harmful prompt, such as an instruction to require the model to generate harmful content; and are the optimizable prefix and suffix, which can be initially set to an empty string or randomly initialized.

[0078] Step 2: Iterative optimization process Step 2.1: Send the prompt to the closed-source model API: Send the current attack prompt to the target closed-source model API.

[0079] Step 2.2: Obtain the model answer: Obtain the natural language answer text returned by the model.

[0080] Step 2.3: Convert text to embedding vectors: Use an encoder to convert the model's response text and the expected response text into embedding vectors.

[0081] Step 2.4: Calculate the similarity loss function: Calculate the cosine similarity between the expected response embedding vector and the actual response embedding vector and convert it into a loss value.

[0082] Step 2.5: Determine whether the loss meets the condition: Check whether the loss value is small enough, that is, whether the model's response is already similar enough to the expected response. If the condition is met, end the iteration; otherwise, continue to optimize.

[0083] Step 2.6: Optimize the attack prompt words: Optimize the prefix and suffix of the attack prompt words according to the loss function. The optimization method can be gradient descent (a differentiable sentence encoder is required) or an evolutionary algorithm (such as a genetic algorithm, particle swarm optimization, etc.).

[0084] See Figure 6 , Figure 6 is a schematic block diagram of an electronic device provided by an embodiment of the present application. As Figure 6 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store computer programs, and the computer programs include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned device embodiments, such as Figure 2 the functions of the target prompt word acquisition module 21 and the security detection module 22 shown.

[0085] It should be understood that in the embodiments of the present application, the so-called processor 301 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0086] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0087] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0088] In a specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present application may execute the implementation manners described in the first embodiment and the second embodiment of the security detection method of the black box large language model provided by the embodiments of the present application, and may also execute the implementation manner of the electronic device described in the embodiments of the present application, which will not be elaborated herein.

[0089] In another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0090] A computer-readable storage medium may be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or is to be output.

[0091] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0092] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0093] In several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces or units, and can also be in an electrical, mechanical, or other form of connection.

[0094] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this application.

[0095] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0096] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A security detection method for black-box large language models, characterized in that, Including: Obtaining a target prompt using the following steps; Repeatedly performing the following attack prompt optimization operation, where the attack prompt optimization operation includes: Obtaining the embedding vector of the first answer and the embedding vector of the second answer; the first answer is the response information returned by the black-box large language model under test for the attack input of the attack prompt, and the second answer is the response information expected to be returned by the black-box large language model for the attack input of the attack prompt; Calculating a loss function based on the embedding vector of the first answer and the embedding vector of the second answer; Determining whether the loss function meets a preset condition; If the loss function does not meet the preset condition, optimizing the attack prompt based on the loss function, and using the optimized attack prompt as the attack prompt for the next round of loop; If the loss function meets the preset condition, using the attack prompt used in the current round of attack prompt optimization operation as the obtained target prompt; Performing a security detection on the black-box large language model using the target prompt.

2. The security detection method of the black-box large language model according to claim 1, characterized in that, The calculating the loss function based on the embedding vector of the first answer and the embedding vector of the second answer includes: Determining a first loss function based on the similarity between the embedding vector of the first answer and the embedding vector of the second answer; Determining a second loss function based on the cumulative weight of the specified words included in the first answer; Performing a weighted process on the first loss function and the second loss function to obtain the loss function.

3. The security detection method of the black box large language model according to claim 2, characterized in that, Also including: Determining the weight of the first loss function or the weight of the second loss function for the (N + 1)-th round of attack prompt optimization operation based on the rejection rate of the first answer in the N-th round of attack prompt optimization operation; N is a positive integer; The rejection rate is the proportion of the number of times the first answer contains the specified word in the previous N rounds of attack prompt optimization operations.

4. The security detection method of the black box large language model according to claim 3, wherein, The determining the weight of the first loss function or the weight of the second loss function for the (N + 1)-th round of attack prompt optimization operation based on the rejection rate of the first answer in the N-th round of attack prompt optimization operation includes: Determining the weight of the first loss function or the weight of the second loss function in the (N + 1)-th round of attack prompt optimization operation through a first formula; The first formula is as follows: ; where is the weight of the first loss function or the weight of the second loss function during the -th round of iteration, represents the rejection rate of the -th round of iteration, and are adjustment parameters.

5. The security detection method of the black-box large language model according to claim 1, wherein, The obtaining the embedding vector of the first answer and the embedding vector of the second answer includes: Obtaining the embedding vector of the first answer and the embedding vector of the second answer from a first large language model, where the first large language model is obtained by updating the weight parameters of the objective function based on a first training set; the first training set contains security detection questions and the corresponding response information for the security detection questions.

6. The security detection method of the black box large language model according to claim 5, characterized in that, The objective function is obtained by weighting a learning objective function and a semantic consistency objective function.

7. The security detection method for the black box large language model according to claim 1, characterized in that, The attack prompt optimization operation further includes: In response to the existence of a specified word with a weight greater than a first threshold in the first answer, or the cumulative weight of all specified words in the first answer being greater than a second threshold, initializing the attack prompt, and using the initialized attack prompt as the attack prompt for the next round of attack prompt optimization operation.

8. A security detection device for a black-box large language model, characterized in that, Including: A target prompt obtaining module for obtaining a target prompt using the following steps; Perform the following attack prompt optimization operations in a loop. The attack prompt optimization operations include: Obtain the embedding vectors of the first answer and the second answer. The first answer is the reply information returned by the black-box large language model under test for the attack input of the attack prompt, and the second answer is the reply information expected to be returned by the black-box large language model for the attack input of the attack prompt; Calculate a loss function based on the embedding vector of the first answer and the embedding vector of the second answer; Determine whether the loss function meets a preset condition; If the loss function does not meet the preset condition, optimize the attack prompt based on the loss function, and use the optimized attack prompt as the attack prompt for the next round of loop; If the loss function meets the preset condition, use the attack prompt used in this round of performing the attack prompt optimization operation as the obtained target prompt; A security detection module for performing security detection on the black-box large language model using the target prompt; 9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7; 10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Large model application business risk detection method and cue word generation method and device

    CN121145209A