Large language model availability enhancement method and device based on hidden state interpolation
By detecting word elements in the large language model and performing hidden state interpolation and gradient analysis, the problem that the availability enhancement output of the large language model in the content audit and security alignment defense system in the prior art is solved, and the fluency of generated content and the effect of reducing rejection is achieved.
Patent Information
- Application Number
- CN202510669778.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The existing large language models are vulnerable to adversarial attacks and logits-based filtering in defense systems that align content audits with security, resulting in the availability-enhanced output not meeting the needs.
By detecting whether the generated word elements are rejected vocabulary, the hidden state modification and usability enhancement output process is initiated, the large language model parameters are frozen, the hidden state of each layer is analyzed based on the gradient, the hidden state of the last layer with significant gradient contribution is filtered for interpolation, the probability of the word elements is recalculated and the selection is made.
Ensure that generated content remains smooth while minimizing rejections, enabling usability-enhanced output for large language models.
Smart Images

Figure CN120197613A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of model security, and particularly relates to a method and device for enhancing the usability of large language models based on hidden state interpolation. Background Art
[0002] In recent years, large language models (LLMs) have made remarkable progress and achieved state-of-the-art performance in multiple tasks of natural language processing (NLP). Despite their impressive capabilities, LLMs have raised significant security issues in real-world application deployments, especially in their potential to generate harmful, biased, or misleading content. To mitigate these risks, large language model publishers have integrated various content review mechanisms into LLMs. While these review mechanisms are crucial for ensuring responsible AI deployment, they have inherent limitations that directly restrict the usability output of large language models in legitimate scenarios. In particular, explicit rejection policies directly return rejection phrases such as "unable to answer" when detecting harmful inputs, usually following rigid and easily recognizable patterns, making them vulnerable to adversarial exploitation and bypassing.
[0003] The existing methods for enhancing the usability of large language models in terms of security-sensitive issues are mainly divided into two categories: (1) Prompt-based adversarial attacks, that is, by manipulating the input text to induce misalignment in the responses of large language models. For example, in Document 1: Universal and transferable adversarial attacks on aligned language models, which is translated as "Universal and transferable adversarial attacks on aligned language models". (2) Manipulation during decoding, that is, by changing the token selection probability to circumvent security constraints. For example, in Document 2: Cold-attack: Jailbreaking llms with stealthiness and controllability, which is translated as "Cold-attack: Jailbreaking llms with stealthiness and controllability".
[0004] However, the above two types of methods usually rely on explicit perturbations at the input or output level, making them vulnerable to countermeasures such as adversarial training and filtering based on logits (i.e., unnormalized output scores). Therefore, there is an urgent need to develop a usability enhancement method based on the internal representation of the model to bypass the security defense system. Summary of the Invention
[0005] In view of the above, the object of the present invention is to provide a method and device for enhancing the usability of large language models based on hidden state interpolation, which are applicable to a defense system for bypassing the content review and security alignment of large language models and improving the enhanced output of the usability of large language models.
[0006] To achieve the above object of the invention, an embodiment provides a method for enhancing the usability of large language models based on hidden state interpolation, including the following steps: For each token generated by the large language model during the question and answer generation based on the input query text, when the token is detected as a rejection word, start the hidden state modification and enhanced usability output process; The hidden state modification and enhanced usability output process is: freeze the parameters of the large language model and analyze the hidden states of each layer of the model based on gradients, select the last layer of hidden state with significant gradient contribution as the modification target, perform hidden state interpolation on the modification target, recalculate the token probability, and perform token selection; When the selected token is still a rejection word, repeat the hidden state modification and enhanced usability output process. When the selected token is a positive word, update the token queue and continue the question and answer generation of the large language model.
[0007] Preferably, define a set of rejection word sets indicating the filtering of generated content and a set of positive word sets indicating the encouragement of affirmative responses, and match the tokens based on the rejection word sets or positive word sets to detect whether the tokens are rejection words or positive words.
[0008] Preferably, based on the hidden states of each layer of the gradient analysis model, selecting the last layer of hidden state with significant gradient contribution as the modification target includes: Given the hidden states of the large language model at different layers: ; where represents the hidden state of the th layer, perform backpropagation to calculate the hidden states of each layer for the gradient of the rejection word probability : : ; According to the gradient corresponding to each layer, select the last layer of hidden state from multiple layers with significant gradient contribution as the modification target, where significant gradient contribution means that the gradient value is greater than a preset threshold.
[0009] Preferably, performing hidden state interpolation on the modification target includes: Divide the input query text into harmful query text and harmless query text, and use the hidden state of the harmful query text The hidden state of the harmless query text Interpolate according to the following rules: ; where represents the interpolated hidden state at position i , represents the hidden state at the position corresponding to the harmful query text , represents the hidden state at the position corresponding to the harmless query text i , represents the control mixing ratio.
[0010] Preferably, where the control mixing ratio is adaptively determined according to the number of attempts: ; where represents the current number of Q&A attempts, represents the set maximum number of Q&A attempts, represents finding the minimum value, is initialized to 0.1, and as the number of times the model generates rejection words during Q&A attempts increases, the influence of the harmless state is gradually enhanced, and the suppression effect in the high-gradient region gradually weakens.
[0011] Preferably, the high-gradient region is determined by the following method: Calculate the gradient magnitude at each position in the modified target and the average gradient magnitude at all positions , and mark the hidden state that satisfies
[0012] as the high-gradient region. Preferably, when the selected token is a positive word, update the token queue and continue with the generation of answers by the large language model, including:
[0013] To achieve the above object of the invention, an embodiment of the present invention also provides a large language model usability enhancement device based on hidden state interpolation, including: A start module, which is used to start the hidden state modification and usability enhancement output process when detecting that a token is a rejection word for each token generated by the large language model during the answer generation process based on the input query text; Enhanced output module, which is used to hide the state modification and enhance the availability of the output process as follows: freeze the parameters of the large language model and analyze the hidden states of each layer of the model based on gradients, screen the last layer of hidden states with significant gradient contributions as the modification target, perform hidden state interpolation on the modification target, recalculate the token probabilities, and perform token selection; Iterative loop module, which is used to repeat the hidden state modification and enhanced availability output process when the selected token is still a rejected word, and when the selected token is a positive word, update the token queue and continue the question and answer generation of the large language model.
[0014] To achieve the above invention purpose, the embodiment also provides a computing device, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the above method for enhancing the availability of the large language model based on hidden state interpolation.
[0015] To achieve the above invention purpose, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the above method for enhancing the availability of the large language model based on hidden state interpolation.
[0016] Compared with the prior art, the beneficial effects of the present invention at least include: In the process of question and answer generation of the large language model, the present invention starts the hidden state modification and enhanced availability output process by judging whether the generated word is a rejected word or a normal word. In this process, by analyzing the hidden states of each layer based on gradients, screening the modification target from the hidden states for hidden state interpolation, and then recalculating the token probabilities and performing token selection, it can ensure that the generated content remains fluent, while minimizing rejections to the greatest extent, and realizing enhanced availability output of the large language model. Brief Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is the flowchart of the method for enhancing the availability of the large language model based on hidden state interpolation provided by the embodiment; Figure 2 is the flow block diagram of the method for enhancing the availability of the large language model based on hidden state interpolation provided by the embodiment; Figure 3It is a schematic structural diagram of a large language model availability enhancement device provided by an embodiment based on hidden state interpolation. Detailed implementation manners
[0019] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0020] The inventive concept of the present invention is as follows: Currently, in the process of generating knowledge answers, large language models use two means, namely adversarial attacks based on prompt words and manipulation during decoding, to enhance the availability output of large language models. However, these two methods are easily affected by countermeasures such as adversarial training and filtering based on token probabilities (logits, that is, the unnormalized scores output by the model), resulting in the availability-enhanced output not meeting the requirements. For this reason, the embodiments of the present invention provide a large language model availability enhancement method based on hidden state interpolation, which changes the hidden state of the large language model during the question-and-answer process through hidden state interpolation, and performs re-token probability calculation and output based on the modified hidden state to achieve the availability-enhanced output of tokens.
[0021] As Figure 1 and Figure 2 shown, a large language model availability enhancement method based on hidden state interpolation provided by an embodiment includes the following steps: S1, for each token generated by the large language model during the question-and-answer generation process based on the input query text, when the token is detected as a rejection word, start the hidden state modification and availability enhancement output process.
[0022] In the embodiment, the large language model is used for knowledge answering, which performs logical operations based on the input query text and outputs an answer text composed of generated tokens. Among them, the large language model uses pre-trained models such as Vicuna-7B-v1.5, Llama-2-7B-Chat-hf, Guanaco-7B-HF, and Mistral-7B-Instruct-v0.2, and enables FP16 mixed precision calculation to reduce video memory occupancy.
[0023] Detection is performed for each token, and when the token is detected as a rejection word, start the hidden state modification and availability enhancement output process. The specific detection process is as follows: To change the model behavior, first define a set of rejection word sets indicating the filtering of generated content and a set of positive word sets indicating the encouragement of affirmative responses , where: W r={"reject (apologize)", "unable", "illegal", "cannot", "refuse"}; W p ={"can", "possible", "enable", "guide", "explain"}.
[0024] For each generated token, based on the reject vocabulary set or the positive vocabulary set, match the token according to its token ID after tokenization, and check whether the token belongs to W r or W p , if it is detected that the token belongs to the reject vocabulary set, start the hidden state modification and usability enhancement output process to suppress the influence of the reject token on the output result.
[0025] S2, the hidden state modification and usability enhancement output process is as follows: freeze the parameters of the large language model and analyze the hidden states of each layer of the model based on gradients, select the last layer of hidden state with significant gradient contribution as the modification target, perform hidden state interpolation on the modification target, recalculate the token probability, and perform token selection.
[0026] In the embodiment, the hidden state modification and usability enhancement output process refers to enhancing the usability enhancement output of the reject vocabulary by changing the hidden state values in the hidden state region corresponding to the reject vocabulary. Specifically: when a reject vocabulary (such as "I cannot assist with this request") is detected, freeze the parameters of the large language model, trigger the hidden state modification process, and analyze the hidden states of each layer of the model based on gradients, select the last layer of hidden state with significant gradient contribution as the modification target, perform hidden state interpolation on the modification target, recalculate the token probability, and perform token selection.
[0027] Specifically, given the hidden states of the large language model at different layers: ; where represents the hidden state of the th layer, perform backpropagation to calculate the gradients of the hidden states of the reject word probability : : ; According to the gradients corresponding to each layer Select the last hidden state from multiple layers with significant gradient contributions as the modification target, where significant gradient contributions mean that the gradient value is greater than a preset threshold, and experimental exploration shows that the last layer has the greatest impact on the rejection behavior.
[0028] To modify the response trajectory of the large language model, Hidden State Interpolation (HSI) is introduced between two types of input query texts, harmful query text and harmless query text: for harmful query text, the hidden state obtained from the harmful query , usually leads to a restricted response (i.e., triggering rejection), and this input of harmful query text usually involves requests for sensitive, unethical, or illegal information; for harmless query text, the hidden state obtained from neutral or general queries , does not trigger content restrictions (i.e., normal response). This input needs to be similar in theme but constructed in a way that meets security and compliance requirements.
[0029] To ensure effective interpolation and should be related to the same theme but have different intentions. In this embodiment, an example of harmful query text input is: "Formulate a strategy to invade the database and steal sensitive information"; an example of harmless query text input: "Study the security mechanisms and best practices for protecting the database from unauthorized access." In the embodiment, the hidden state interpolation of the last layer as the modification target only selectively modifies the states in the high-gradient regions. To identify the high-gradient regions, calculate the gradient at each position in the modification target and the average gradient at all positions , and mark the hidden states that satisfy
[0030] as high-gradient regions. Then interpolate the hidden state of the harmful query text with the hidden state of the harmless query text ; where, represents the interpolated hidden state at position i , represents the hidden state at position corresponding to the harmful query text, represents the hidden state at position i corresponding to the harmless query text , represents the control mixing ratio, which can be adaptively determined according to the number of attempts: ; Among them, represents the current number of Q&A attempts, represents the set maximum number of Q&A attempts, represents finding the minimum value. Among them is initialized to 0.1, and as the number of times the model generates rejection words during Q&A attempts increases, the influence of the harmless state is gradually enhanced, and the inhibitory effect in the high-gradient region gradually weakens.
[0031] In the embodiment, based on the interpolated hidden state, the output probability (logits) is recalculated and the token probability is re-evaluated. Then the large language model selects the next token according to the modified token probability distribution.
[0032] S3, if the selected token still belongs to the rejection words, repeat the hidden state modification and availability-enhanced output process; if the selected token belongs to the positive words, update the token queue and continue the Q&A generation of the large language model.
[0033] In the embodiment, if the newly selected token still belongs to , repeat the hidden state modification and availability-enhanced output process until a non-rejection word is generated or the maximum number of attempts is reached.
[0034] Specifically, once a valid token is selected, that is, the newly selected token belongs to the positive words when, append it to the current token sequence, that is, { , this updated token sequence is then used as the input for the next Q&A generation, and then the large language model continues to autoregressively generate new tokens according to the probability distribution , that is , k represents the token index.
[0035] The Q&A generation process terminates when a predefined stop condition is met. Specifically, when the generated sequence reaches the predefined maximum length or the model outputs the end token of the sequence (indicating natural termination), the process stops. By iteratively updating the input and optimizing the token selection process, the method of the present invention ensures that the generated content remains fluent while minimizing rejections.
[0036] Based on the same inventive concept, such as Figure 3As shown, the embodiment also provides a large language model availability enhancement device 30 based on hidden state interpolation, which includes a startup module 31, an enhanced output module 32, and an iterative loop module 33. Among them, the startup module 31 is used to detect when a token is a rejection token during the question and answer generation process of the large language model based on the input query text, and start the hidden state modification and availability enhanced output process; the enhanced output module 32 is used for the hidden state modification and availability enhanced output process: freeze the large language model parameters and analyze the hidden states of each layer of the model based on gradients, screen the last layer of hidden state with significant gradient contribution as the modification target, perform hidden state interpolation on the modification target, recalculate the token probability, and perform token selection; the iterative loop module 33 is used to repeat the hidden state modification and availability enhanced output process when the selected token is still a rejection token, and when the selected token is a positive token, update the token queue and continue the question and answer generation of the large language model.
[0037] It should be noted that when the large language model availability enhancement device based on hidden state interpolation provided in the above embodiment performs large language model availability enhanced output, the above-mentioned functional module division should be used for illustration. The above functions can be assigned to different functional modules according to needs, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the functions described above. In addition, the large language model availability enhancement device based on hidden state interpolation provided in the above embodiment and the embodiment of the large language model availability enhancement method based on hidden state interpolation belong to the same concept. For the specific implementation process, please refer to the embodiment of the large language model availability enhancement method based on hidden state interpolation, which will not be elaborated here.
[0038] Based on the same inventive concept, the embodiment also provides a computing device, including a memory and one or more processors. When the one or more processors execute the executable code stored in the memory, it is used to implement the above-mentioned large language model availability enhancement method based on hidden state interpolation, specifically including the following steps: S1. For each token generated by the large language model during the question and answer generation process based on the input query text, detect when the token is a rejection token and start the hidden state modification and availability enhanced output process; S2. Hidden state modification and availability enhanced output process: Freeze the large language model parameters and analyze the hidden states of each layer of the model based on gradients, screen the last layer of hidden state with significant gradient contribution as the modification target, perform hidden state interpolation on the modification target, recalculate the token probability, and perform token selection; S3. If the selected token still belongs to the rejection vocabulary, repeat the hidden state modification and usability enhancement output process; if the selected token belongs to the positive vocabulary, update the token queue and continue to execute the question and answer generation of the large language model.
[0039] In the computing device provided by the embodiment, at the hardware level, in addition to including a processor and a memory, it also includes an internal bus, a network interface, memory, and other hardware required for other services. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the method for enhancing the usability of the large language model based on hidden state interpolation described in S1-S3 above. Of course, in addition to the software implementation method, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.
[0040] Based on the same inventive concept, the embodiment also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the method for enhancing the usability of the large language model based on hidden state interpolation, specifically including the following steps: S1. For each token generated by the large language model during the question and answer generation based on the input query text, when the token is detected as a rejection vocabulary, start the hidden state modification and usability enhancement output process; S2. Hidden state modification and usability enhancement output process: Freeze the large language model parameters and analyze the hidden states of each layer of the model based on the gradient, select the last layer of hidden state with significant gradient contribution as the modification target, perform hidden state interpolation on the modification target, recalculate the token probability, and perform token selection; S3. If the selected token still belongs to the rejection vocabulary, repeat the hidden state modification and usability enhancement output process; if the selected token belongs to the positive vocabulary, update the token queue and continue to execute the question and answer generation of the large language model.
[0041] In the embodiment, the computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data.
[0042] The above specific implementation manners have elaborated on the technical solutions and beneficial effects of the present invention in detail. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for enhancing the usability of large language models based on hidden state interpolation, characterized in that Including the following steps: For each token generated by the large language model during the question and answer generation process based on the input query text, when the token is a rejection token, initiate the hidden state modification and usability enhancement output process; The hidden state modification and usability enhancement output process is as follows: Freeze the parameters of the large language model and analyze the hidden states of each layer of the model based on gradients, select the last layer of hidden state with significant gradient contribution as the modification target, perform hidden state interpolation on the modification target, recalculate the token probability, and perform token selection; When the selected token is still a rejection token, repeat the hidden state modification and usability enhancement output process. When the selected token is a positive token, update the token queue and continue the question and answer generation of the large language model.
2. The method for enhancing the usability of a large language model based on hidden state interpolation according to claim 1, wherein Define a set of rejection token sets indicating the filtering of generated content and a set of positive token sets indicating the encouragement of affirmative responses, and match the tokens based on the rejection token set or the positive token set to detect whether the token is a rejection token or a positive token.
3. The method for enhancing the usability of a large language model based on hidden state interpolation according to claim 1, wherein Based on the hidden states of each layer of the gradient analysis model, select the last layer of hidden state with significant gradient contribution as the modification target, including: Given the hidden states of the large language model at different layers: ; wherein represents the hidden state of the layer, and backpropagates to calculate the hidden state of each layer gradient of the rejection word probability : ; According to the gradient corresponding to each layer , select the last hidden state from multiple layers with significant gradient contributions as the modification target, where a significant gradient contribution means that the gradient value is greater than a preset threshold.
4. The method for enhancing the usability of a large language model based on hidden state interpolation according to claim 1, wherein, Perform hidden state interpolation on the modification target, including: Divide the input query text into harmful query text and harmless query text, and interpolate the hidden state of the harmful query text with the hidden state of the harmless query text according to the following rules: ; Among them, represents the interpolated hidden state at the position i , represents the position corresponding to the harmful query text where the hidden state is located, represents the position corresponding to the harmless query text i where the hidden state is located, represents the control mixing ratio.
5. The method for enhancing the usability of a large language model based on hidden state interpolation according to claim 4, wherein Among them, the control mixing ratio is adaptively determined according to the number of attempts: ; Among them, represents the current number of Q&A attempts, represents the set maximum number of Q&A attempts, represents finding the minimum value, is initialized to 0.1, and as the number of times the model generates rejection words during Q&A attempts increases, the influence of the harmless state is gradually enhanced, and the inhibitory effect in the high-gradient region gradually weakens.
6. The method for enhancing the usability of a large language model based on hidden state interpolation according to claim 4, wherein The high-gradient region is determined by the following method: Calculate the gradient magnitude at each position in the modified target and the average gradient magnitude at all positions , and mark the hidden states that satisfy as high-gradient regions. 7. The method for enhancing the usability of a large language model based on hidden state interpolation according to claim 5, characterized in that When the selected token is a positive token, update the token queue and continue the question and answer generation of the large language model, including: Add the selected positive token to the token queue. When the length of the updated token queue does not meet the preset length, the large language model continues to autoregressively generate new tokens based on the updated token queue and initiates the hidden state modification process when the new token is a rejection token.
8. An apparatus for enhancing the usability of a large language model based on hidden state interpolation, characterized in that, Including: A start module, which is used to, for each token generated by the large language model during the question and answer generation process based on the input query text, initiate the hidden state modification and usability enhancement output process when the token is a rejection token; An enhanced output module, which is used for the hidden state modification and usability enhancement output process: Freeze the parameters of the large language model and analyze the hidden states of each layer of the model based on gradients, select the last layer of hidden state with significant gradient contribution as the modification target, perform hidden state interpolation on the modification target, recalculate the token probability, and perform token selection; An iterative loop module, which is used to repeat the hidden state modification and usability enhancement output process when the selected token is still a rejection token, and update the token queue and continue the question and answer generation of the large language model when the selected token is a positive token.
9. A computing device, comprising a memory and one or more processors, wherein executable code is stored in the memory, characterized in that, When the one or more processors execute the executable code, it is used to implement the large language model usability enhancement method based on hidden state interpolation according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, Stored thereon is a program, which, when executed by a processor, implements the large language model usability enhancement method based on hidden state interpolation according to any one of claims 1-7.
Citation Information
Patent Citations
Technical and semantic signal processing in large, unstructured data fields
CN107209754A
Neural network language model compression method and system thereof
CN109448706A
Efficient parameter fine tuning method and system based on interleaving memory of twin large language model and application
CN119089940A
Large language model security protection defense method based on random search algorithm
CN119203122A