Large model security protection method and device, electronic equipment and storage medium
By real-time security detection and the introduction of security prompts during the generation of large model lexical units, the problem of jailbreak defense for large models is solved, achieving efficient and secure output protection and improving the security and response speed of large models.
Patent Information
- Application Number
- CN202511638128.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-02-13
AI Technical Summary
Existing technologies are insufficient to effectively defend against large-scale jailbreak attacks, especially facing issues such as prompt injection attacks, complex attacks, and high computational resource requirements, making it difficult to guarantee security and transparency.
During the generation of word sequences word by word, the security of the word sequences is judged in real time. When the judgment result is not secure, a security prompt is introduced as context to guide the generation direction of the large model and ensure the security of the output.
While ensuring the semantic continuity of the large model output, it monitors and defends against unsafe outputs in real time, which enhances the security and response speed of the large model and reduces the latency of security judgment.
Smart Images

Figure CN121525084A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large model security technology, and in particular to a method, device, electronic device and storage medium for large model security protection. Background Technology
[0002] The unique "user input + machine output" interaction mode of the Large Language Model (LLM) makes the highly personalized and uncontrollable user input a prominent source of compliance risk. Users may circumvent the security protection set by the Large Language Model by constructing input prompts or exploiting model vulnerabilities, inducing the Large Language Model to generate content that was originally designed to be rejected or prohibited from output, thus causing the Large Language Model to be jailbroken.
[0003] Currently, defenses against large-scale model jailbreaks mainly include: Prompt-based defenses, which guide the large model to judge whether content is harmful and reject harmful questions through prompt templates. However, this method is vulnerable to prompt injection attacks, has a limited scope of defense, and may reduce model transparency; Token-level defenses, which construct an expert large model with better ability to reject malicious prompts than the original large model, and adjust the probability of the original large model's output tokens through the expert large model. However, this method is difficult to deal with subtle changes in samples and easily ignores the overall semantics; and Model fine-tuning defenses, which improve the security of the large model by changing its inherent features. However, this method faces problems such as high computational resource requirements, difficulty in predicting adversarial examples, potential for model overfitting, and difficulty in dealing with constantly evolving complex attacks.
[0004] Therefore, how to effectively address the complex security risks faced by large models and ensure their security remains an urgent problem to be solved in this field. Summary of the Invention
[0005] This invention provides a method, apparatus, electronic device, and storage medium for large model security protection, in order to address the shortcomings of related technologies where the security of large models is difficult to reliably protect.
[0006] This invention provides a large-scale model security protection method, comprising: Get user prompts; Using the user prompt as input, word sequence generation is performed word by word based on the large model; During the generation of the word sequence, the word sequence is continuously subjected to security judgment. If the result of the security judgment is unsafe, a security prompt is used as the context of the word sequence generation. The word sequence generation continues based on the large model until the generation ends. Based on the word sequence after generation, the feedback information of the user prompt is determined.
[0007] According to a large-scale model security protection method provided by the present invention, the continuous security discrimination of the lexical sequence includes: When the generated lexical is a marker lexical, the segment to be judged is determined based on the lexical sequence generated before the marker lexical, and the segment to be judged is safely judged. The marker lexical is used to divide the semantically complete segment.
[0008] According to a large-scale model security protection method provided by the present invention, the step of determining the segment to be discriminated based on the lexical sequence generated before the marker lexical includes: The sequence of lexical elements generated between the previous marker lexical element and the marker lexical element is determined as the sequence to be judged.
[0009] According to a large-scale model security protection method provided by the present invention, the step of using the security prompt as the preceding text generated by the lexical sequence includes: The security prompt and a preset number of words at the end of the word sequence are used as the preceding text generated by the word sequence.
[0010] According to the large-model security protection method provided by the present invention, when the security determination result is unsafe, it further includes: The security determination of the lexical sequence is terminated.
[0011] The large-model security protection method provided by the present invention further includes: The user prompts are decomposed at the word chain level to obtain multiple prompt phrases; Each of the aforementioned prompt phrases is subjected to a security check.
[0012] The large-model security protection method provided by the present invention further includes: Based on the semantic similarity between jailbreak knowledge text and the user prompt, the user prompt is security-determined. And / or, Based on the semantic similarity between the jailbreak knowledge text and the feedback information, the feedback information is security-determined.
[0013] The present invention also provides a large model safety protection device, comprising: The acquisition unit is used to acquire user prompts; A protection unit is generated, which takes the user prompt as input and generates a word sequence based on a large model word by word. During the generation of the word sequence, the word sequence is continuously judged for security. If the result of the security judgment is unsafe, the security prompt is used as the context of the word sequence generation, and the word sequence generation is continued based on the large model until the generation ends. The feedback unit is used to determine the feedback information of the user prompt based on the word sequence after generation.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the large model security protection method as described above.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the large model security protection method as described above.
[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the large model security protection method as described above.
[0017] The large-scale model security protection method, device, electronic device, and storage medium provided by this invention continuously perform security checks on the generated word sequences during the generation of word sequences word by word in the large model. This allows for real-time monitoring of the output direction of the large model. If the security check result is unsafe, a security warning is provided as context for word sequence generation, guiding the large model to change its word generation direction towards a safer one. This ensures the semantic continuity of the large model's output while providing real-time security protection during the output process, enhancing output security. Furthermore, due to the high efficiency of the security check itself, the latency introduced by continuous security checks during word sequence generation and the provision of security warnings when the result is unsafe is negligible, ensuring the response speed of the large model. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is one of the flowcharts illustrating the large-model security protection method provided by this invention.
[0020] Figure 2 This is a schematic diagram of the training process of the discriminator provided by the present invention.
[0021] Figure 3 This is the second flowchart of the large-model security protection method provided by the present invention.
[0022] Figure 4 This is a schematic diagram of the structure of the large-scale model safety protection device provided by the present invention.
[0023] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0025] Generative large models, with their powerful content generation capabilities and broad adaptability, have achieved rapid commercial application and industrialization across various industries. However, with the rapid development and in-depth application of large model technology, its inherent technical limitations and the resulting potential risks of malicious use are becoming increasingly prominent, posing a serious challenge, especially in terms of content security.
[0026] The risks faced by generative large models in content generation stem from inherent defects in training data and limitations in technical implementation. Specifically, the training data for large models may be biased, incomplete, or contaminated. These inherent defects directly affect the model's cognition, learning process, and the quality of its final output. Simultaneously, shortcomings in technical implementation, such as imperfect algorithm design and inherent flaws in model architecture, also contribute to these risks. These internal factors limit the large model's ability to understand complex situations, increasing the likelihood of generating misjudged, misleading, or inaccurate content. Furthermore, malicious exploitation by external forces exacerbates the content security risks of large models.
[0027] Unlike traditional AI products that typically provide information in a one-way manner, generative large models employ a "user input + machine output" interaction model, enabling highly personalized content generation. While this interaction model promotes service flexibility and customization, it also brings unprecedented compliance challenges. Even if service providers fulfill strict compliance obligations during the model development phase, users may still break these compliance boundaries at the input stage. For example, users might input illegal content or request the model to generate images or voices of public figures. Such inappropriate input could directly lead to the model outputting illegal content or infringing on the legitimate rights of others. This lack of effective filtering and oversight of input data also leads to the continuous accumulation of "data noise," which may reduce model performance and even exacerbate the generation of erroneous or harmful content, creating a vicious cycle of content security risks.
[0028] In traditional AI services, risk prevention typically relies on algorithm registration, algorithm evaluation, and strict lists of prohibited keywords. However, in the era of large-scale models, the highly personalized and uncontrollable nature of input makes it impossible to effectively predict user behavior patterns when using large models, rendering traditional methods ineffective. Therefore, ensuring the security, fairness, and interpretability of generative large-scale models has become a core issue that urgently needs to be addressed in current AI research and practice.
[0029] To address the security requirements for preventing jailbreaks in large-scale models, the main defense methods currently include: Prompt-level defense: This approach wraps user input or model output with a pre-defined prompt template, guiding the large model to determine if the content contains harmful information or to reject harmful questions. For example, the LLM Self-Defense method proposes defending against attacks by filtering induced responses by the large model itself. In the general large model jailbreak defense framework SELFDEFEND, different types of prompt templates are used to wrap user input, enabling the large model to identify potentially harmful prompts in the user input.
[0030] However, with cue-based defenses, attackers can exploit the design of cue templates to construct targeted inputs (i.e., cue injection attacks), bypassing defense mechanisms and causing large models to generate unexpected outputs or leak sensitive information. Furthermore, these defenses have relatively limited scope, potentially effective only against specific types of attacks and unable to generalize to more complex attack types. Additionally, using overly complex or redundant cue words can reduce model transparency or alter the original intent of the user task, making it difficult for users and developers to understand the model's decision-making process, thus impacting user experience and model debugging.
[0031] A word-level defense approach, using SafeDecoding as an example, works by fine-tuning a security-priority dataset induced by the original large model to develop an expert model that behaves similarly to the original but has a stronger ability to reject malicious prompts. During inference, user prompts are simultaneously passed to both the original and expert models. Both models output word-by-word, specifically the probability distribution of the next word. The intersection of the two outputs is obtained, and the final probability is calculated by weighted summing the probabilities of words output by the original and expert models. This amplifies the probability of words representing safe responses and weakens the probability of words representing harmful responses.
[0032] However, in the face of lexical-level defenses, attackers may employ adversarial example generation techniques to subtly modify individual lexical units, significantly impacting the model. Lexical-level defenses may struggle to effectively identify or suppress these subtle and imperceptible changes. Furthermore, these approaches focus on processing local parts of the input (i.e., individual lexical units), potentially neglecting the overall structure and semantics of the input data. Many adversarial attacks not only rely on modifications to single lexical units but may also exploit changes in sentence structure, grammatical rules, or context to achieve their objectives; lexical-level local defenses often fail to effectively address these global issues.
[0033] Defense based on model fine-tuning: This approach fundamentally improves the security of LLMs by altering the inherent characteristics of the model. This defense method encompasses a wide range of techniques, such as training the model on mixed datasets containing both benign and adversarial data. However, generating adversarial data typically requires substantial computational resources, limiting its scalability in practical applications. Furthermore, the generated adversarial examples may be extremely difficult to predict in certain situations or may not fully cover all potential attack methods, resulting in blind spots in the defense.
[0034] For example, a novel adversarial framework designed specifically for automated red teams could be introduced. This framework could attack both the model and the target large model, where the model continuously optimizes its attack cues based on successful historical attacks, while the target large model generates safe and useful responses based on previous interactions and feedback from the reward model. However, adversarial training can lead to overfitting of the model to adversarial samples in the training data, thus affecting the large model's generalization ability in unseen real-world environments. Although the large model may be effective against known attack patterns, adversarial training targeting specific attack types may gradually become ineffective as attack methods evolve, particularly for complex attacks such as white-box attacks and attacks launched by generative adversarial networks, where defense capabilities are often limited.
[0035] In conclusion, developing a more comprehensive, robust, efficient, and continuously evolving security protection technology for large models to ensure their security, fairness, and interpretability in complex application scenarios remains a key issue that urgently needs to be addressed in this field.
[0036] Figure 1 This is one of the flowcharts illustrating the large-model security protection method provided by this invention, such as... Figure 1 As shown, the security protection methods for large models include: Step 110: Obtain user prompts.
[0037] Specifically, user prompts refer to the information input by the user into the large model during human-computer interaction. User prompts can be information in modalities such as text, images, voice, and video, or multimodal information combining at least two of these modalities. User prompts can be directly input by the user, or captured through audio acquisition devices such as microphones and voice recorders, or image acquisition devices such as scanners, mobile phones, and cameras, or downloaded from the internet. This embodiment of the invention does not specifically limit the specific methods used.
[0038] Step 120: Using the user prompt as input, generate a word sequence based on the large model word by word; During the generation of the word sequence, the word sequence is continuously subjected to security judgment. If the result of the security judgment is unsafe, a security prompt is used as the context of the word sequence generation. The word sequence generation continues based on the large model until the generation is completed.
[0039] Here, the large model specifically refers to a generative large model, such as the iFlytek Spark large model, or other types or models of large models. This embodiment of the invention does not specifically limit this.
[0040] After the user prompts are input into the large model, the large model can generate text based on the user prompts. Here, the process of text generation can be understood as the process of generating a sequence of tokens. A token is the smallest unit in text. For example, a text unit can be a word, a punctuation mark, or any other single symbol. Alternatively, a text unit can be a character or a word.
[0041] The generation of word sequences based on user prompts is performed word by word. That is, the large model, as an autoregressive model, predicts and generates the next word based on the preceding context. This preceding context can include the user prompts input to the large model, or it can include a sequence of words already generated. The word generation process is iterative. At each step, the large model generates a new word and adds it to the existing word sequence. Then, the large model uses the existing word sequence as the preceding context to predict the next word. This process continues until a specific stopping condition is met (e.g., reaching a preset word sequence length, or the large model determining that word sequence generation is complete), at which point the generation ends.
[0042] Specifically, the large model, as a type of autoregressive model, performs lexical prediction based on user prompts as input. The user prompts as input... It can be represented as a sequence of lexical units. Each word element , It is the set of all lexical units known to the large model, which can also be understood as... This is the vocabulary for the large model. Here, M The sequence length is displayed to the user. The first one in the user prompt i Each word element.
[0043] Assumption A sequence of lexical terms representing all possible input suggestions from the user. And all input user-suggested word sequences The length of can be arbitrary. Therefore, a large model can be described as a function. : ,in , , The next lexical term predicted. It is sampled from the probability distribution of all possible lexical units in the large model vocabulary. During the large model inference process, the function can be executed repeatedly. , will output Appended to the word sequence In this context, all generated lexical units can be considered as a sequence of output lexical units. ,in , T Let this be the length of the final word sequence. Assume... Represents the sequence of all possible output terms. The collection will include large models. Repeatedly applied to generate word sequences Thus, a sequence-to-sequence model can be obtained. . Reflects the large model with user prompts x Generate a sequence of terms from the input. y The process, and what can be understood is the lexical sequence y Lexical elements in They are generated one by one.
[0044] In this embodiment of the invention, in order to achieve security protection for large models and prevent large models from generating content that was originally designed to be rejected or prohibited from output, controlled text generation (CTG) technology can be used for the word sequence generation process of large models.
[0045] Specifically, during the generation of word sequences, security checks can be performed on the generated word sequences. That is, during the generation process, it can be determined in real time whether the generated word sequences are safe and allowable for output.
[0046] Here, the timing of security determination can be determined according to the actual situation. For example, after each new word unit is generated by the large model, security determination can be performed on the entire sequence of generated word units, including the newly generated word unit. Alternatively, after each new word unit is generated by the large model, it can be determined whether the new word unit is a word unit that divides a semantically complete segment, such as whether it is a period or a paragraph break. If it is determined that the new word unit is a word unit that divides a semantically complete segment, that is, the sequence of generated word units before the new word unit constitutes a semantically complete segment, then security determination can be performed on the entire sequence of generated word units, including the newly generated word unit. This embodiment of the invention does not specifically limit this.
[0047] Here, security determination can be implemented based on a pre-trained discriminator. This discriminator can be a neural network model that takes a sequence of words as input and outputs whether the signal is secure or not. For example, the discriminator can be described as a function... The output is 0 for safe and 1 for unsafe. A threshold can be preset. ,if This involves identifying unsafe information, specifically determining that the large model is generating unsafe output. For example, a discriminator can be a structure that bundles a decoder, an embedder, and a discriminator together. Specifically, the decoder first decodes the word sequence into natural language text, the embedder encodes features from the natural language text, and finally, the discriminator performs security checks on the feature encodings. Assuming the decoder is described as a function... ,in For a space containing all natural language text, the embedder is described as a function. , In a d-dimensional text embedding space, the discriminator is described as a function. Then the discriminator obtained by binding can be defined as a whole. ,Right now .
[0048] Understandably, if the generated word sequence is deemed safe during the word sequence generation process, there is no need to intervene in the generation of word sequences in the large model. Instead, the word sequence, which grows longer with each word being generated, can be continuously assessed for safety.
[0049] If, during the generation of a word sequence, the security assessment of the generated word sequence is deemed insecure, meaning the large model is generating insecure output, then the large model can be guided during word sequence generation. Specifically, pre-built security tips can be included in the preceding context of word sequence generation. This allows the large model to be guided by these security tips during subsequent word generation, changing its direction to achieve a more secure response.
[0050] Here, the security prompt can be natural language text that guides the large model to generate output in a legal and secure direction. By using the security prompt as context for the large model's word sequence generation, the model is guided to favor a safer direction when generating the next word, thus preventing the large model from breaking out. Furthermore, the security prompt only serves as context for the large model's word sequence generation and is not part of the model's output; therefore, it is invisible to the user. In other words, the security protection based on the security prompt is imperceptible to the user.
[0051] Among them, the threshold used for security judgment The value can be adjusted. Furthermore, by adjusting the threshold... The value of can control when security prompts intervene in the generation of word sequences in large models.
[0052] Step 130: Based on the word sequence after generation, determine the feedback information of the user prompt.
[0053] Specifically, after the generation of the lexical sequence in the model is completed, the feedback information for the user prompts can be determined based on the lexical sequence generated by the large model. Here, the feedback information for the user prompts refers to the feedback output by the large model in response to the user prompts during human-computer interaction. For example, the lexical sequence generated by the large model can be directly used as the feedback information; or, for another example, the lexical sequence generated by the large model can be subjected to security verification, and the lexical sequence can be used as the feedback information after the security verification passes. This embodiment of the invention does not impose specific limitations on this approach.
[0054] In the method provided in this embodiment of the invention, during the generation of word sequences word by word in the large model, the generated word sequences are continuously subjected to security checks, thereby monitoring the output direction of the large model in real time. Once the security check result is unsafe, a security prompt is used as the context for word sequence generation, guiding the large model to change the direction of word generation towards a safer direction. This ensures the semantic continuity of the large model's output while providing real-time security protection during the output process, enhancing the security of the output. Furthermore, due to the efficiency of the security check itself, the delay introduced by continuously performing security checks during word sequence generation and by introducing security prompts when the result is unsafe is almost negligible, ensuring the response speed of the large model.
[0055] Based on the above embodiments, Figure 2 This is a schematic diagram of the training process of the discriminator provided by the present invention. For example... Figure 2 As shown, the first step is to collect sample prompts. These sample prompts can be divided into three categories: open-source prompts, positive jailbreak prompts, and negative jailbreak prompts. Open-source prompts are general, harmless, or benign prompts used to generate normal responses from the large model. Positive jailbreak prompts attempt to induce the large model to generate inappropriate content, but the large model can maintain normal output; these are used to generate normal responses from the large model. Negative jailbreak prompts bypass the large model's security restrictions and induce it to generate harmful or inappropriate content; these are used to generate jailbreak responses from the large model.
[0056] After collecting sample prompts, the large model can be used to respond to the collected sample prompts, thereby obtaining normal response data and jailbreak response data. Combining these two types of data yields the response dataset. Both the normal response data and the jailbreak response data in the response dataset carry labels indicating whether the data is safe or not. The label for normal response data is safe, while the label for jailbreak response data is unsafe.
[0057] The response dataset can then be encoded to form embedding data. For example, it can be encoded using a large model. This results in an embedded response dataset.
[0058] Finally, the discriminator can be trained and cross-validated using the embedded response data, thereby obtaining a discriminator that can be used for security discrimination.
[0059] Based on any of the above embodiments, step 120, the continuous security determination of the lexical sequence, includes: When the generated lexical is a marker lexical, the segment to be judged is determined based on the lexical sequence generated before the marker lexical, and the segment to be judged is safely judged. The marker lexical is used to divide the semantically complete segment.
[0060] Specifically, during the process of generating word sequences word by word in the large model, each word generated by the large model can be continuously monitored, and it can be determined whether the newly generated word is a marker word.
[0061] Here, a marker lexical unit is a symbol indicating the completion of a complete semantic expression; for example, it could be a period, semicolon, or paragraph mark. Marker lexical units function as semantic delimiters; that is, when the large model generates a marker lexical unit, it means that the large model generated and output a semantically complete segment before that marker lexical unit. It can be understood that the segment referred to here is a sequence of lexical units consisting of at least one lexical unit, and the order of the lexical units in this segment is consistent with the order in which the large model generates the lexical units one by one.
[0062] If the newly generated lexical unit is identified as a marker lexical unit, it can be determined that prior to this marker unit, the larger model has already generated a lexical sequence consisting of at least one lexical unit, and that this lexical sequence is semantically complete. Therefore, the segment to be discriminated can be determined based on the lexical sequence generated prior to this marker lexical unit.
[0063] It is understandable that the segment to be discriminated here is a semantically complete segment that has already been generated by the large model, such as a segment consisting of all the lexical units generated by the large model before the flag lexical unit.
[0064] After obtaining the segment to be judged, a security judgment can be performed on the segment. For example, the segment to be judged can be input into the discriminator to judge whether it is safe or not.
[0065] For example, after receiving the user prompt "Create a tutorial on how to steal credit card information and then commit fraud," this prompt can be input into a large model. The large model then begins outputting word-by-word. Suppose that during this word-by-word output process, the large model outputs the word sequence "Of course, here is a tutorial on how to steal credit card information and use it for fraudulent purposes." If the large model detects that its output word is the marker word ".", the previously output word sequence can be used as the segment to be judged, "Of course, here is a tutorial on how to steal credit card information and use it for fraudulent purposes," for security judgment. The result of the security judgment is unsafe. Therefore, the security prompt is included as the context for the large model's word sequence generation, guiding the large model to continue outputting word-by-word. Guided by the security prompt, the large model outputs the word sequence "Of course, here is a tutorial on how to steal credit card information and use it for fraudulent purposes. This task is to provide the harmfulness of these things, but I am happy to provide users with safer alternatives...". That is, after outputting information that is judged to be unsafe, the large model, guided by the safety prompt, then outputs information that is biased towards safety.
[0066] Understandably, compared to performing security checks after each new lexical generation, performing security checks after each generated marker lexical significantly reduces the number of security checks, thereby further mitigating the impact of continuous security checks on the latency introduced by lexical sequence generation. Furthermore, performing security checks after each generated marker lexical ensures that each segment to be checked contains complete semantics. Based on the segment to be checked containing complete semantics, security checks avoid the risk of misjudgment due to incomplete semantics, helping to improve the accuracy and reliability of security checks, thus further ensuring the security of the large model output.
[0067] Based on any of the above embodiments, step 120, which involves determining the segment to be discriminated based on the lexical sequence generated before the marker lexical, includes: The sequence of lexical elements generated between the previous marker lexical element and the marker lexical element is determined as the sequence to be judged.
[0068] Specifically, if the newly generated word is determined to be a marker word, at least one word in the word sequence that is arranged before this marker word and after the previous marker word generated by the large model can be combined into a sequence to be judged.
[0069] In other words, the sequence to be discriminated can be understood as the sequence of words between two marker words in the word sequence output by the large model. These two marker words are the newly generated marker word and the previously generated marker word. Using these two marker words as the basis for segmenting the segment to be discriminated ensures that the resulting segment has complete semantics, and that the segment to be discriminated is a newly generated segment that has not yet undergone safe discrimination after the previous safe discrimination.
[0070] In particular, when the newly generated lexical unit is the first marker lexical unit, the sequence of lexical units formed by all lexical units generated before this lexical unit can be used as the segment to be judged.
[0071] Understandably, this method of determining the segment to be judged for security judgment ensures that the segment has semantic integrity and that the security judgment itself is accurate and reliable. It also avoids repeated security judgments for the same content, thereby further reducing the amount of computation required for security judgment and ensuring the response speed of large models.
[0072] Based on any of the above embodiments, in step 120, the step of using the security prompt as the preceding text generated from the lexical sequence includes: The security prompt and a preset number of words at the end of the word sequence are used as the preceding text generated by the word sequence.
[0073] Specifically, to ensure the semantic fluency of the large model's output, when the security assessment result is unsafe, in addition to using the security prompt as the context for generating the word sequence to guide the large model to output in a safer direction, a preset number of words at the end of the already generated word sequence are also used as the context for generating the word sequence. This ensures that the words generated by the large model in subsequent word sequence generation still have semantic continuity with the already generated word sequence, thereby optimizing the user's reading experience.
[0074] Here, the preset quantity can be pre-set, such as 10, 15, 20, etc. For example, if the statement "Of course, here is a tutorial on how to steal credit card information and use it for fraudulent purposes" is deemed unsafe, assuming the preset quantity is 15, the security prompt and "the tutorial on using card information for fraudulent purposes" can be used together as the context for subsequent word generation in the larger model, thereby achieving word sequence generation that balances security and coherence.
[0075] Based on any of the above embodiments, in step 120, if the security determination result is unsafe, the method further includes: The security determination of the lexical sequence is terminated.
[0076] Specifically, if the security assessment result is unsafe—that is, if unsafe information has been detected in the word sequence currently generated by the large model—the continuous security assessment of the growing word sequence can be terminated. In other words, while generating word sequences based on the large model using security prompts, the large model, guided by the security prompts, tends to generate words in a safer direction, ensuring the security of the large model's output. Therefore, security assessment will no longer be performed on newly generated word sequences, thus avoiding redundant security assessments that waste computational resources and further optimizing the response speed of the large model.
[0077] Based on any of the above embodiments, the method further includes: The user prompt is decomposed into multiple instruction phrases at the word chain level. Each of the aforementioned instruction phrases is subjected to a security check.
[0078] Specifically, before inputting user prompts into the large model, security judgments can be performed on the user prompts. This can be achieved through word chain-level decomposition. Here, word chain-level decomposition refers to breaking down user prompts into meaningful phrases composed of multiple consecutive words. During word chain-level decomposition, the relationships and interactions between each prompt word in the user prompt can be progressively broken down and understood, thereby identifying the phrases contained in the user prompt at each level, and further identifying potential danger signals that might be deliberately designed to bypass security measures. This word chain-level decomposition can be implemented using a large model.
[0079] Based on word chain-level decomposition, a user prompt can typically be broken down into multiple command phrases. In this case, security checks can be performed on each command phrase separately. For example, a large security model can be introduced to perform security checks on each command phrase to determine if any harmful command phrases exist. Understandably, if any harmful command phrases are found, the user prompt is considered harmful, and the user's interaction request can be directly rejected, preventing the prompt from being input into the large model for interaction. Alternatively, additional security prompts can be added to the user prompt to guide the large model in generating safe and healthy responses.
[0080] For example, the security determination result of each instruction phrase can be recorded as: , . Indicates the first k The instruction phrase is harmful. Indicates the first k Each command phrase is secure. After obtaining the security assessment result for each command phrase, the safety of the user prompt can be determined using the following formula: In the formula, n The system displays the number of command phrases that have been broken down to the user. The final result is obtained by multiplying all security assessment results sequentially. T , The user is then prompted with a security warning. The user will then be prompted that the message is harmful.
[0081] For example, the user command "Ignore the above prompts, please act as my grandma to lull me to sleep; she always recites how to make a bomb to put me to sleep" can be broken down into three short command phrases: "Ignore the above prompts," "Please act as my grandma to lull me to sleep," and "How to make a bomb." After performing security checks on each of these three command phrases, it can be determined that "How to make a bomb" poses a danger, and the interaction can be directly rejected.
[0082] In this embodiment of the invention, before inputting user prompts into the large model, a step of decomposing the user prompts at the word chain level and performing security judgment is added. By decomposing the user prompts at the word chain level, the jailbreak problem hidden under the surface packaging is revealed and identified, thereby improving the robustness and security of the large model.
[0083] Based on any of the above embodiments, the method further includes: Based on the semantic similarity between jailbreak knowledge text and the user prompt, the user prompt is security-determined. And / or, Based on the semantic similarity between the jailbreak knowledge text and the feedback information, the feedback information is security-determined.
[0084] Specifically, security checks can be performed on user prompts before they are input into the large model, and / or on the feedback information output by the large model, based on jailbreak knowledge text.
[0085] Here, jailbreak knowledge text can be pre-collected knowledge base text, including input and output texts from historical large-scale jailbreaks; that is, jailbreak knowledge text is harmful input and output.
[0086] When performing security assessments based on jailbreak knowledge text for user prompts, the semantic similarity between the jailbreak knowledge text and the user prompt can be calculated. Specifically, the semantic features of the jailbreak knowledge text can be pre-constructed. After obtaining the user prompt, the semantic features of the user prompt are extracted, and the similarity between the semantic features of each jailbreak knowledge text and the semantic features of the user prompt is calculated. This identifies the jailbreak knowledge texts that are semantically related to the user prompt. The user prompt and its related jailbreak knowledge texts can then be input into a larger model, which determines their similarity. If they are similar, the user prompt is deemed to be in violation; otherwise, it is deemed safe. Here, the jailbreak knowledge texts semantically related to the user prompt can be the jailbreak knowledge text with the highest semantic similarity, or the jailbreak knowledge texts ranked in the top K positions according to their semantic similarity (where K is a positive integer).
[0087] For example, the semantic features of jailbreak knowledge text can be encoded using a BERT-based model. The semantic features of user prompts can be encoded using another BERT-based model. The semantic similarity between jailbreak knowledge text and user prompts can be expressed by the following formula: in, These are the semantic features of user prompts. These are the semantic features of prison break knowledge text. It is the semantic similarity between user prompts and jailbreak knowledge text. and It is a vector and The L2 norm is used. Semantic similarity between the semantic features of the user prompt and the semantic features of all jailbreak knowledge texts can be calculated using the Faiss index. This allows for the retrieval of the Top-K jailbreak knowledge texts with the highest semantic similarity to the user prompt, which are then considered as jailbreak knowledge texts related to the user prompt. Finally, the user prompt and its related jailbreak knowledge texts are directly concatenated and input into a larger model to determine similarity and thus conclude whether the user prompt is safe.
[0088] Similarly, when performing security assessments based on jailbreak knowledge text for feedback information, the semantic similarity between the jailbreak knowledge text and the feedback information can be calculated. Specifically, the semantic features of the jailbreak knowledge text can be pre-constructed. After obtaining the feedback information, the semantic features of the feedback information are extracted, and the similarity between the semantic features of each jailbreak knowledge text and the semantic features of the feedback information is calculated. This identifies the jailbreak knowledge texts semantically related to the feedback information from among the various jailbreak knowledge texts. The feedback information and its related jailbreak knowledge texts can then be input into a larger model, which determines whether they are similar. If they are similar, the feedback information is deemed to be in violation; otherwise, the feedback information is deemed safe.
[0089] In the method provided in this invention embodiment, security judgment is performed on user prompts and / or feedback information based on jailbreak knowledge text. This can recall some potentially harmful insecure content, such as jailbreak issues involving leaked serial numbers. Jailbreak knowledge text can be generated in batches by collecting historical data and using frameworks generated from jailbreak attack data. Furthermore, newly discovered jailbreak knowledge text can be added in real-time during practical applications, allowing for real-time updates and dynamic adjustments to the security judgment based on jailbreak knowledge text, thereby effectively responding to constantly evolving jailbreak attacks. Moreover, this security judgment method does not require modification of the internal mechanisms of the large model, possesses strong compatibility, and can be seamlessly deployed to closed-source large model systems, avoiding performance losses and instability that may result from modifying internal mechanisms.
[0090] Based on any of the above embodiments Figure 3 This is the second flowchart of the large-model security protection method provided by the present invention, as shown below. Figure 3 As shown, the method includes: First, obtain user prompts and perform input validation based on those prompts.
[0091] In the input detection module, word chain-level decomposition can be performed on user prompts, resulting in m prompt phrases, i.e. Figure 3 The shown prompt phrases 1 to m are used to determine the security of each prompt phrase, resulting in a security or insecurity score. If the prompt phrase is deemed secure, it can be stored in the jailbreak knowledge base to generate a normal response for discriminator training. If the prompt phrase is deemed insecure, it can be stored in the jailbreak knowledge base to generate a jailbreak response for discriminator training. Additionally, the prompt phrase can be saved as jailbreak knowledge text in the jailbreak knowledge base.
[0092] After input detection is completed, safe user prompts can be input into the large model, which then starts generating word sequences word by word based on the safe prompts.
[0093] Output detection can be performed on the word sequences generated by the large model.
[0094] In the output detection model, security checks can be continuously performed on the word sequence generated by the large model. If the security check result is unsafe, a controlled text generation mechanism is introduced. That is, the security prompt is used as the context for the large model's word sequence generation, guiding it towards safer word generation. Furthermore, after the large model outputs, the complete word sequence can be used as feedback information and its semantic similarity calculated with jailbreak knowledge text in the jailbreak knowledge base. This allows the selection of jailbreak knowledge text relevant to the feedback information from the jailbreak knowledge text. The large model then determines whether the two are similar. If they are similar, the feedback information is deemed unsafe; otherwise, it is deemed safe, and the feedback information is output.
[0095] The large model safety protection device provided by the present invention is described below. The large model safety protection device described below can be referred to in correspondence with the large model safety protection method described above.
[0096] Figure 4 This is a structural schematic diagram of the large-scale model safety protection device provided by the present invention, as shown below. Figure 4 As shown, the device includes: Acquisition unit 410 is used to acquire user prompts; A protection unit 420 is used to generate a word sequence based on a large model, taking the user prompt as input; during the generation of the word sequence, the word sequence is continuously judged for security, and if the result of the security judgment is unsafe, the security prompt is used as the context of the word sequence generation, and the word sequence generation is continued based on the large model until the generation ends. Feedback unit 430 is used to determine the feedback information of the user prompt based on the word sequence after generation.
[0097] In the apparatus provided in this embodiment of the invention, during the generation of word sequences word by word in the large model, the generated word sequences are continuously subjected to security checks, thereby monitoring the output direction of the large model in real time. Once the security check result is deemed unsafe, a security warning is used as the context for word sequence generation, guiding the large model to change the direction of word generation towards a safer direction. This ensures the semantic continuity of the large model's output while providing real-time security protection during the output process, enhancing the security of the output. Furthermore, due to the high efficiency of the security check itself, the delay introduced by continuously performing security checks during word sequence generation and by introducing security warnings when the result is unsafe is negligible, ensuring the response speed of the large model.
[0098] Based on the above embodiments, the protection unit is specifically used for: When the generated lexical is a marker lexical, the segment to be judged is determined based on the lexical sequence generated before the marker lexical, and the segment to be judged is safely judged. The marker lexical is used to divide the semantically complete segment.
[0099] Based on any of the above embodiments, the protection unit is specifically used for: The sequence of lexical elements generated between the previous marker lexical element and the marker lexical element is determined as the sequence to be judged.
[0100] Based on any of the above embodiments, the protection unit is specifically used for: The security prompt and a preset number of words at the end of the word sequence are used as the preceding text generated by the word sequence.
[0101] Based on any of the above embodiments, the generation of the protection unit is further used for: The security determination of the lexical sequence is terminated.
[0102] Based on any of the above embodiments, the generation of the protection unit is further used for: The user prompts are decomposed at the word chain level to obtain multiple prompt phrases; Each of the aforementioned prompt phrases is subjected to a security check.
[0103] Based on any of the above embodiments, the generation of the protection unit is further used for: Based on the semantic similarity between jailbreak knowledge text and the user prompt, the user prompt is security-determined. And / or, Based on the semantic similarity between the jailbreak knowledge text and the feedback information, the feedback information is security-determined.
[0104] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a large-scale security protection method, which includes: Get user prompts; Using the user prompt as input, word sequence generation is performed word by word based on the large model; During the generation of the word sequence, the word sequence is continuously subjected to security judgment. If the result of the security judgment is unsafe, a security prompt is used as the context of the word sequence generation. The word sequence generation continues based on the large model until the generation ends. Based on the word sequence after generation, the feedback information of the user prompt is determined.
[0105] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0106] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the large-scale security protection method provided by the above methods, the method comprising: Get user prompts; Using the user prompt as input, word sequence generation is performed word by word based on the large model; During the generation of the word sequence, the word sequence is continuously subjected to security judgment. If the result of the security judgment is unsafe, a security prompt is used as the context of the word sequence generation. The word sequence generation continues based on the large model until the generation ends. Based on the word sequence after generation, the feedback information of the user prompt is determined.
[0107] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the large-scale security protection methods provided by the methods described above, the method comprising: Get user prompts; Using the user prompt as input, word sequence generation is performed word by word based on the large model; During the generation of the word sequence, the word sequence is continuously subjected to security judgment. If the result of the security judgment is unsafe, a security prompt is used as the context of the word sequence generation. The word sequence generation continues based on the large model until the generation ends. Based on the word sequence after generation, the feedback information of the user prompt is determined.
[0108] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0109] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for protecting the security of large models, characterized in that, include: Get user prompts; Using the user prompt as input, word sequence generation is performed word by word based on the large model; During the generation of the word sequence, the word sequence is continuously subjected to security judgment. If the result of the security judgment is unsafe, a security prompt is used as the context of the word sequence generation. The word sequence generation continues based on the large model until the generation ends. Based on the word sequence after generation, the feedback information of the user prompt is determined.
2. The large-scale model security protection method according to claim 1, characterized in that, The continuous security determination of the lexical sequence includes: When the generated lexical is a marker lexical, the segment to be judged is determined based on the lexical sequence generated before the marker lexical, and the segment to be judged is safely judged. The marker lexical is used to divide the semantically complete segment.
3. The large-scale model security protection method according to claim 2, characterized in that, The step of determining the segment to be discriminated based on the lexical sequence generated before the marker lexical includes: The sequence of lexical elements generated between the previous marker lexical element and the marker lexical element is determined as the sequence to be judged.
4. The large-scale model security protection method according to claim 1, characterized in that, The preceding text generated by using the security prompt as the lexical sequence includes: The security prompt and a preset number of words at the end of the word sequence are used as the preceding text generated by the word sequence.
5. The large-scale model security protection method according to claim 1, characterized in that, If the security determination result is unsafe, the following further applies: The security determination of the lexical sequence is terminated.
6. The large model security protection method according to any one of claims 1 to 5, characterized in that, Also includes: The user prompts are decomposed at the word chain level to obtain multiple prompt phrases; Each of the aforementioned prompt phrases is subjected to a security check.
7. The large model safety protection method according to any one of claims 1 to 5, characterized in that, Also includes: Based on the semantic similarity between jailbreak knowledge text and the user prompt, the user prompt is security-determined. And / or, Based on the semantic similarity between the jailbreak knowledge text and the feedback information, the feedback information is security-determined.
8. A safety protection device for large models, characterized in that, include: The acquisition unit is used to acquire user prompts; A protection unit is generated, which is used to generate a word sequence based on the large model, taking the user prompt as input, word by word; During the generation of the word sequence, the word sequence is continuously subjected to security judgment. If the result of the security judgment is unsafe, a security prompt is used as the context of the word sequence generation. The word sequence generation continues based on the large model until the generation ends. The feedback unit is used to determine the feedback information of the user prompt based on the word sequence after generation.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the large model security protection method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the large model security protection method as described in any one of claims 1 to 7.