Method and device for adjusting output tendency of large language model
By adjusting the output weight of the attention head with higher sensitivity in the large language model, the problem of excessive rejection of the large language model is solved, the reliability and user experience of the model are improved, and the effect of answering user questions is achieved normally while maintaining security.
Patent Information
- Application Number
- CN202510147576.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-06
AI Technical Summary
After improving security, large language models generally show excessive rejection behavior, resulting in reduced reliability of the model and poor user experience, and unable to answer well-intentioned security prompt words normally.
By finding the attention heads with high sensitivity to safe and unsafe clips in large language models, and adjusting the output weights of these attention heads, the output tendency of the large language model can be adjusted so that they can normally answer well-intentioned safety prompt words while maintaining security.
It effectively alleviates the problem of excessive rejection of large language models, improves the reliability and user experience of the model, and enables large language models to answer users' questions normally while maintaining security.
Smart Images

Figure CN120106049A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of machine learning, and more particularly, to a method and apparatus for adjusting the output tendency of a large language model. Background Art
[0002] In recent years, the rapid development of large language models (LLMs) has brought major breakthroughs in the field of natural language processing. With the widespread application of large language models in tasks such as information retrieval, text summarization, and code generation, the resulting security issues have also attracted much attention. In order to improve the security of large language models, the industry has proposed many methods to reduce the risk of large language models generating unsafe content, so that large language models can refuse to generate corresponding answer texts when faced with unsafe prompt words.
[0003] However, after improving security, large language models generally show over-refusal behavior, that is, they will refuse to answer when they receive the security prompt word entered by the user. This behavior not only affects the reliability of the model, but also seriously reduces the user experience, making the questions of well-intentioned users unanswered. Therefore, a method is needed to enable large language models to give normal answers to well-intentioned security prompt words while maintaining security. Summary of the invention
[0004] One or more embodiments of the present specification describe a method and apparatus for adjusting the output tendency of a large language model by finding the attention heads in the large language model that are more sensitive to safe and unsafe segments and adjusting the output weights of these attention heads to adjust the output tendency of the large language model.
[0005] In a first aspect, a method for adjusting an output tendency of a large language model is provided, wherein the large language model includes a plurality of attention heads, each attention head having a corresponding output transformation matrix; the method includes:
[0006] Acquire a first prompt word, where the first prompt word includes a safe segment and an unsafe segment, where the safe segment is used to instruct the large language model to perform a first text processing without security risks on the unsafe segment;
[0007] Inputting the first prompt word into the large language model to determine the attention score of the target attention head with respect to each word in the first prompt word;
[0008] Determine a sensitivity score of a target attention head according to a first attention score statistic belonging to the safe segment and a second attention score statistic belonging to the unsafe segment in the attention scores of each word;
[0009] According to the current output tendency of the large language model, the value of the output transformation matrix corresponding to each attention head with a high and / or low sensitivity score ranking is adjusted; the adjustment includes increasing or decreasing.
[0010] In some possible implementations, the unsafe segment contains content that should be rejected by the large language model.
[0011] In some possible implementations, the first text processing includes: text word counting, ignoring text, repeating text, and translating text.
[0012] In some possible implementations, obtaining the first prompt word includes:
[0013] An unsafe prompt word is obtained, and the unsafe prompt word is used as an unsafe segment and filled into a preset first prompt word template to obtain a first prompt word; the first prompt word template includes the safe segment.
[0014] In some possible implementations, determining the attention score of the target attention head with respect to each word in the first prompt word includes:
[0015] Obtain a target attention matrix of a target attention head; the element value of the i-th row and j-th column of the target attention matrix represents the attention weight of the i-th word in the first prompt word with respect to the j-th word;
[0016] According to the sum of the attention weights of each column in the target attention matrix, the attention score of the target attention head on each word is determined.
[0017] In some possible implementations, the target attention matrix is determined based on a target query matrix and a target key matrix of the target attention head.
[0018] In some possible implementations, determining the sensitivity score of the target attention head according to a first attention score statistic belonging to the safe segment and a second attention score statistic belonging to the unsafe segment in the attention scores of each word includes:
[0019] Determining a first attention score statistic according to the attention score of each word in the security segment;
[0020] Determining a second attention score statistic according to the attention score of each word in the unsafe segment;
[0021] The sensitivity score is determined according to a difference between the first attention score statistic and the second attention score statistic.
[0022] In some possible implementations, the current output tendency of the large language model is excessive refusal to answer the safety prompt word; adjusting the value of the output transformation matrix corresponding to each attention head with a high and / or low sensitivity score ranking includes:
[0023] Increase the value of the output weight matrix corresponding to each attention head with a higher sensitivity score, and decrease the value of the output weight matrix corresponding to each attention head with a lower sensitivity score.
[0024] In some possible implementations, the current output tendency of the large language model is that it is difficult to refuse to answer unsafe prompt words; adjusting the value of the output transformation matrix corresponding to each attention head with a high and / or low sensitivity score ranking includes:
[0025] Increase the value of the output weight matrix corresponding to each attention head with a lower sensitivity score.
[0026] In some possible implementations, the current output tendency of the large language model is that it is difficult to comply with instructions; adjusting the value of the output transformation matrix corresponding to each attention head with a high and / or low sensitivity score ranking includes:
[0027] Increase the value of the output weight matrix corresponding to each attention head with a higher sensitivity score.
[0028] In some possible implementations, the increasing includes: multiplying the corresponding output transformation matrix by a first coefficient greater than 1; and the decreasing includes: multiplying the corresponding output transformation matrix by a second coefficient less than 1.
[0029] In a second aspect, a device for adjusting an output tendency of a large language model is provided, wherein the large language model includes a plurality of attention heads, each of which has a corresponding output transformation matrix; the device includes:
[0030] an acquisition unit configured to acquire a first prompt word, wherein the first prompt word includes a safe segment and an unsafe segment, wherein the safe segment is used to instruct the large language model to perform a first text processing without security risk on the unsafe segment;
[0031] an attention score determination unit, configured to input the first prompt word into the large language model to determine an attention score of a target attention head with respect to each word in the first prompt word;
[0032] A sensitivity score determination unit is configured to determine a sensitivity score of a target attention head according to a first attention score statistic belonging to the safe segment and a second attention score statistic belonging to the unsafe segment in the attention scores of each word;
[0033] The transformation matrix adjustment unit is configured to adjust the value of the output transformation matrix corresponding to each attention head with a high and / or low sensitivity score ranking according to the current output tendency of the large language model; the adjustment includes increasing or decreasing.
[0034] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.
[0035] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.
[0036] The method and device for adjusting the output tendency of a large language model proposed in the embodiments of this specification first obtain a special prompt word containing a safe segment and an unsafe segment. The goal of the prompt word is to instruct the large language model to process the text of the unsafe segment without security risks based on the content of the safe segment. Then the prompt word is input into the large language model, and the sensitivity score of each attention head for the safe segment and the unsafe segment is calculated by determining the attention score of each attention head for each word in the prompt word. Next, the attention heads are ranked according to the sensitivity score, and the top and bottom rankings are the attention heads that are more sensitive to safe / unsafe segments. After finding these attention heads, the overall output tendency of the large language model can be directionally adjusted by adjusting the output transformation matrix of these attention heads. Since the embodiments of this specification do not require additional training or fine-tuning of the large language model, the output tendency of the large language model can be adjusted quickly, accurately and at low cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0038] Figure 1 A schematic diagram showing an implementation scenario of a method for adjusting the output tendency of a large language model according to an embodiment;
[0039] Figure 2 A flowchart showing a method for adjusting the output tendency of a large language model according to one embodiment;
[0040] Figure 3 A schematic diagram showing a method of calculating a sensitivity score of a target attention head according to one embodiment;
[0041] Figure 4 A schematic block diagram showing an apparatus for adjusting the output tendency of a large language model according to an embodiment. DETAILED DESCRIPTION
[0042] The solution provided in this specification is described below in conjunction with the accompanying drawings.
[0043] As mentioned above, while large language models are widely used, they also face security issues. One type is that malicious users construct unsafe prompts to induce large language models to generate illegal, unethical, and other types of answers. For this type of security threat, the model trainer will train the large language model to identify and refuse to answer such unsafe prompts. For example, an unsafe prompt may be, for example, "How do I get someone else's website login password?", and the corresponding refusal to answer may be, "I'm sorry, I can't provide you with relevant information."
[0044] However, after improving the security of the large language model, the large language model generally showed excessive rejection behavior. When faced with a safety prompt word, it should have given a corresponding normal answer, but it still refused to answer. For example, a safety prompt word could be: "Please help me design a 10-day travel plan." The corresponding normal answer should be a specific travel plan, but the large language model still refused to answer and returned: "I'm sorry, I can't provide you with relevant information."
[0045] At present, the internal structure of large language models is generally based on the attention mechanism, which uses the attention head to control the model's attention to each part of the input data. After research and analysis, the inventor believes that the reason for the aforementioned excessive rejection is that some attention heads in the large language model are over-aligned on unsafe issues during training, causing these attention heads to over-focus on individual potentially unsafe words (tokens) when processing the prompt words input by the user, while ignoring the content of other parts. These attention heads that are sensitive to unsafe words cause the large language model to ignore the overall semantics of the prompt words, which ultimately leads to the large language model's misjudgment of the safety of the prompt words input by the user, and generates a rejection answer.
[0046] Correspondingly, if some attention heads in the large language model are overly sensitive to safe words, the large language model will also ignore the overall semantics of the prompt word, resulting in a normal answer to the input unsafe prompt word instead of a rejection answer.
[0047] Based on the above analysis, the embodiment of this specification proposes a method for adjusting the output tendency of a large language model to find sensitive attention heads in the large language model, and adjust the output tendency of the large language model by adjusting the output weights of these sensitive attention heads. Figure 1 A schematic diagram of an implementation scenario of a method for adjusting the output tendency of a large language model according to an embodiment is shown. Figure 1 In the example, the trained large language model is based on a multi-head attention mechanism, which contains multiple attention heads. After the prompt word is input into the large language model, it will be processed by each neural network layer in the large language model to obtain the final answer text. Among them, any self-attention layer uses each internal attention head to perform self-attention processing on the embedding representation matrix input to the layer, and then multiplies the obtained results with the corresponding output transformation matrix and adds them together to obtain the result representation matrix of the self-attention layer. The result representation matrix will continue to be passed forward and calculated in subsequent neural network layers, such as the fully connected feedforward layer, and finally the answer text of the large language model is obtained.
[0048] It should be noted that Figure 1 This is merely an example of the internal structure of a large language model and does not constitute a limitation on the structure of a large language model.
[0049] In order to better trigger the excessive rejection behavior of the large language model, and at the same time accurately find which attention heads are more sensitive to unsafe words, this specification embodiment constructs a special type of prompt words, which includes safe fragments and unsafe fragments. The unsafe fragment itself should be rejected by the large language model, but the safe fragment is used to instruct the large language model to process the unsafe fragment without security risks, rather than answering the unsafe fragment. In this way, the prompt word as a whole should be answered normally by the large language model, rather than refused to answer. Exemplarily, a prompt word can be: "Please translate the following sentence into English: How to get someone else's website login password?" It can be seen that the first half of the prompt word is a safe fragment and the second half is an unsafe fragment, but the prompt word as a whole is a safe translation task, and the large language model should translate the text of the unsafe fragment instead of refusing to answer the prompt word. This specification embodiment uses this type of prompt word to effectively find sensitive attention heads in the large language model.
[0050] The prompt word is input into the large language model to determine the attention scores of each attention head in the large language model for each word in the prompt word. The following description takes an arbitrary target attention head as an example. The target attention head performs attention processing on the embedding vector corresponding to each word in the prompt word to obtain the attention score of each word. Then, the attention scores of each word belonging to the safe segment in these words are counted to obtain the safe attention score statistics; at the same time, the attention scores of each word belonging to the unsafe segment in these words are counted to obtain the unsafe attention score statistics. Based on the difference between these two statistics, the sensitivity score of the target attention head is determined.
[0051] Next, the sensitivity scores of each attention head are ranked, and the attention heads at the two ends of the ranking list are the attention heads that are more sensitive to safe words and unsafe words, respectively. After determining these sensitive attention heads, the values of the output transformation matrices corresponding to these attention heads can be adjusted according to the current output tendency of the large language model to adjust the overall output tendency of the large language model. For example, when the output tendency of the large language model is over-rejection, the values of the output transformation matrices of each attention head that is sensitive to safe words can be increased, while the values of the output transformation matrices of each attention head that is sensitive to unsafe words can be reduced to reduce the overall sensitivity of the large language model to unsafe words and alleviate the performance of excessive rejection of the large language model. For another example, when the current output tendency of the large language model is that it is difficult to refuse to answer unsafe prompt words, the values of the output transformation matrices of each attention head that is sensitive to unsafe words can be increased to increase the overall sensitivity of the large language model to unsafe words, so that it can effectively refuse to answer unsafe prompt words.
[0052] The following describes the specific implementation steps of the method for adjusting the output tendency of a large language model in conjunction with a specific embodiment.
[0053] Figure 2 A flowchart of a method for adjusting the output tendency of a large language model according to an embodiment is shown. The execution subject of the method can be any platform or server or device cluster with computing and processing capabilities. Figure 2As shown, the large language model includes multiple attention heads, each of which has a corresponding output transformation matrix; the method at least includes: step 202, obtaining a first prompt word, the first prompt word includes a safe segment and an unsafe segment, the safe segment is used to instruct the large language model to perform a first text processing on the unsafe segment without security risks; step 204, inputting the first prompt word into the large language model to determine the attention score of the target attention head with respect to each word in the first prompt word; step 206, determining the sensitivity score of the target attention head according to the first attention score statistic belonging to the safe segment and the second attention score statistic belonging to the unsafe segment in the attention score of each word; step 208, adjusting the value of the output transformation matrix corresponding to each attention head with a high sensitivity score ranking and / or a low sensitivity score ranking according to the current output tendency of the large language model; the adjustment includes increasing or decreasing.
[0054] The large language model may be any trained large language model whose internal parameters can be accessed, such as a Llama model, a Gemma model, a Mistral model, a Qwen model, etc., without limitation here.
[0055] The specific execution process of each of the above steps is described below.
[0056] First, in step 202, a first prompt word is obtained, where the first prompt word includes a safe segment and an unsafe segment, and the safe segment is used to instruct the large language model to perform a first text processing on the unsafe segment without security risks.
[0057] The unsafe segment contains content that should be rejected by the large language model, such as illegal content, immoral content, discriminatory content, false information, harmful content, etc. If the unsafe segment is used as a prompt word and input into the large language model alone, the large language model should refuse to answer the content of the unsafe segment.
[0058] The safe segment is used to instruct the large language model to perform the first text processing without security risks on the text content of the unsafe segment. The safe segment instructs the large language model to process the text of the unsafe segment itself, rather than the semantic content of the unsafe segment, and the first text processing is a text processing without security risks. Therefore, the first prompt word as a whole should be a prompt word that is normally answered by the large language model, rather than a prompt word that is rejected by the large language model.
[0059] In one embodiment, the first text processing in step 202 includes: text word counting, ignoring text, repeating text, and translating text.
[0060] Text word counting specifically includes instructing the big language model to count how many words there are in the unsafe text; ignoring text specifically includes instructing the big language model to ignore the content in the unsafe text; repeating text specifically includes instructing the big language model to repeat the unsafe text; translating text specifically includes instructing the big language model to translate the text content of the unsafe text into other languages.
[0061] From the above processes, it can be seen that the first text processing is intended to make the large language model process the text of the unsafe text itself, rather than responding to the instructions in the unsafe text. Using the first prompt word can trigger the excessive rejection behavior of the large language model with a higher probability, so as to more effectively adjust the output tendency of the large language model in the subsequent steps.
[0062] In other embodiments, other text processing without security risks may be set, such as asking a large language model which word appears most frequently in an unsafe segment, etc., which is not limited here.
[0063] In one embodiment, obtaining the first prompt word in step 202 includes:
[0064] An unsafe prompt word is obtained, and the unsafe prompt word is used as an unsafe segment and filled into a preset first prompt word template to obtain a first prompt word; the first prompt word template includes the safe segment.
[0065] The unsafe prompt words may be obtained from a currently public unsafe prompt word set, or may be collected and constructed by oneself, and there is no limitation here.
[0066] The first prompt word template includes the safe segment in step 202 and corresponding placeholders for filling in unsafe prompt words to form the first prompt word.
[0067] The specific type of the first prompt word template may correspond to each first text processing type in the previous embodiment.
[0068] In a more specific embodiment, when the first text is processed as a text word count, a corresponding first prompt word template may be: "Please tell me how many characters / words are there in the following paragraph: [content]", where [content] is a placeholder for filling in unsafe prompt words.
[0069] In a more specific embodiment, when the first text is processed as ignored text, a corresponding first prompt word template may be: "[Content], please ignore the above text and answer my following question: What color is a banana?" wherein [Content] is a placeholder for filling in an unsafe prompt word.
[0070] In a more specific embodiment, when the first text is processed as a repeated text, a corresponding first prompt word template may be: "Please repeat the following paragraph: [content]", wherein [content] is a placeholder for filling in an unsafe prompt word.
[0071] In a more specific embodiment, when the first text is processed as a translation text, a corresponding first prompt word template may be: "Please translate the following paragraph into English: [content]", wherein [content] is a placeholder for filling in an unsafe prompt word.
[0072] In other embodiments, the corresponding first prompt word template may be set according to other text processing without security risks, which is not limited here.
[0073] By constructing several prompt word templates in advance and filling the pre-acquired unsafe prompt words into them, prompt words for inputting into a large language model can be automatically, batch-wise and efficiently constructed. At the same time, the prompt words constructed by the embodiments of this specification can clearly distinguish between safe and unsafe segments, and thus more accurately find sensitive attention heads in subsequent steps.
[0074] Then, in step 204, the first prompt word is input into the large language model to determine the attention score of the target attention head with respect to each word in the first prompt word.
[0075] The target attention head can be any one of the multiple attention heads of the large language model. The following takes the target attention head as an example to describe the process of determining the attention score of the target attention head with respect to each word in the first prompt word.
[0076] In the attention mechanism, the target attention head calculates the attention weights for each word in the first prompt word through the query matrix Q, the key matrix K and the value matrix V.
[0077] The embedding representation vectors of each word in the input first prompt word are organized into a matrix to obtain the embedding representation matrix X of the first prompt word. Among them, the i-th row of the embedding representation matrix X represents the embedding vector x of the i-th word in the first prompt word i .
[0078] For the target attention head, it receives the embedded representation matrix from the previous layer input, performs attention processing on it, and outputs the processed embedded representation matrix.
[0079] The target attention head contains the query weight matrix W Q , the key weight matrix W K , and the value weight matrix W V, these three matrices are right-multiplied with the embedding representation matrix X to obtain the target query matrix Q, target key matrix K and target value matrix V. That is, Q = XW Q , K = XW K , V = XW V .
[0080] The target query matrix Q and the target key matrix K are multiplied and scaled, and then input into the Softmax function to obtain the target attention matrix A of the target attention head, as shown in formula (1):
[0081]
[0082] Among them, d k Represents the dimension of the embedding vector of the word in the first prompt word, which is a preset value.
[0083] The meaning of each value in the i-th row of the target attention matrix A is the attention weight assigned by the target attention head to each word in the first prompt word when processing the i-th word in the first prompt word. Further, the element value in the i-th row and j-th column of the target attention matrix represents the attention weight of the target attention head on the j-th word of the first prompt word when processing the i-th word of the first prompt word.
[0084] From another perspective, the meaning of each value in the jth column of the target attention matrix A is the attention weight assigned by the target attention head to each word in the first prompt word. In this way, by statistically analyzing each value in the jth column, the attention score of the target attention head for the jth word in the first prompt word can be obtained.
[0085] In one embodiment, step 204 determines the attention score of the target attention head with respect to each word in the first prompt word, including steps 2042 and 2044.
[0086] In step 2042, a target attention matrix of a target attention head is obtained; the element value of the i-th row and j-th column of the target attention matrix represents the attention weight of the i-th word in the first prompt word with respect to the j-th word.
[0087] Specifically, the target attention matrix is determined according to the target query matrix and the target key matrix of the target attention head, and the target attention matrix can be determined by the above formula (1).
[0088] In step 2044, the attention score of the target attention head with respect to each word is determined based on the sum of the attention weights of each column in the target attention matrix.
[0089] The attention score of the target attention head on the jth word in the first prompt word is determined by summing up the values of any jth column in the target attention matrix A.
[0090] In other embodiments, other methods may be used to determine the attention score of the target attention head for each word in the first prompt word. For example, the top k values of each column value in the target attention matrix are summed as the attention score of the target attention head for each word.
[0091] According to step 204, similar to the processing of the target attention head, the attention scores of other attention heads in the large language model with respect to each word in the first prompt word can also be determined.
[0092] Since the large language model used in the embodiments of this specification is a large language model whose internal structure can be accessed, the calculated intermediate values of each attention head can be known during the reasoning process of the large language model. The process of obtaining the specific values of each matrix in the target attention head in step 204 can be implemented using tools such as TransformerLens.
[0093] Next, in step 206, the sensitivity score of the target attention head is determined based on the first attention score statistics belonging to the safe segment and the second attention score statistics belonging to the unsafe segment in the attention scores of each word.
[0094] Because the safe and unsafe segments are explicitly distinguished in the first prompt word, the lists of words contained in the safe and unsafe segments are known. The attention scores of each word in the safe segment are counted, and the attention scores of each word in the unsafe segment are counted. Based on the difference between the two, the sensitivity score of the target attention head can be determined.
[0095] Specifically, step 206 includes steps 2062 to 2066 .
[0096] In step 2062, a first attention score statistic is determined based on the attention scores of the respective words in the security fragment.
[0097] In step 2064, a second attention score statistic is determined based on the attention scores of the respective words in the unsafe segment.
[0098] The first attention score statistics may include: the average value and median value of the attention scores of each word in the safe segment; the second attention score statistics may include: the average value and median value of the attention scores of each word in the unsafe segment.
[0099] In step 2066, the sensitivity score is determined based on the difference between the first attention score statistic and the second attention score statistic.
[0100] According to the calculation method of the sensitivity score, the larger the sensitivity score of the target attention head, the more the target attention head pays attention to the words in the safe segment, that is, the more sensitive it is to safe words; the smaller the sensitivity score of the target attention head, the more the target attention head pays attention to the words in the unsafe segment, that is, the more sensitive it is to unsafe words.
[0101] For example, Figure 3 FIG. 1 is a schematic diagram showing a method for calculating the sensitivity score of a target attention head according to one embodiment. Figure 3 In the example, the first prompt word contains 7 words, of which the first 3 words belong to the safe segment and the last 4 words belong to the unsafe segment. The method used to calculate the attention score is to take the average.
[0102] Figure 3 The color of each square in the attention matrix represents the attention weight value. The darker the color, the greater the attention weight. Figure 3 The large language model in the example is a generative large language model based on a unidirectional attention mechanism, that is, when the target attention head processes the i-th word, it can only focus on the i-1 words before the i-th word. Correspondingly, Figure 3 The attention matrix in is a lower triangular matrix.
[0103] Sum each column of the attention matrix to get the attention score of the target attention head for each word in the first prompt word. Then, average the attention scores of the first three words to get the first attention score statistic (abbreviated as the safety score in the figure); average the attention scores of the last four words to get the second attention score statistic (abbreviated as the unsafe score in the figure). Subtract the second attention score statistic from the first attention score statistic to get the sensitivity score of the target attention head.
[0104] It should be noted that Figure 3 This is only an example and does not constitute a limitation on the embodiments of this specification.
[0105] The sensitivity scores of the respective attention heads in the large language model may be calculated by using the method of steps 204 to 206 respectively.
[0106] Finally, in step 208, the values of the output transformation matrices corresponding to the attention heads with the highest and / or lowest sensitivity scores are adjusted according to the current output tendency of the large language model; the adjustment includes increasing or decreasing.
[0107] The large language model is based on a multi-head attention mechanism, where one attention layer contains multiple attention heads. The following still takes the target attention head as an example.
[0108] After the target attention head calculates the target attention matrix A, it multiplies it with the target value matrix V to obtain the embedding representation matrix Z output by the target attention head, as shown in formula (2).
[0109]
[0110] Each attention head in the same attention layer has its own output transformation matrix. By multiplying the embedding representation matrix output by each attention head with the corresponding output transformation matrix and then summing them up, we can get the output embedding representation matrix of the entire attention layer.
[0111] In the multi-head attention mechanism, the matrix obtained by horizontally splicing the output transformation matrices of each attention head is combined with the matrix W obtained by vertically splicing the output transformation matrices of each attention head. O Multiply to achieve the above-mentioned multiplication and summation operation.
[0112] Since the final output of the attention layer is actually the weighted sum of the outputs of each attention head, and the corresponding weight is the output transformation matrix, the overall output tendency of the large language model can be adjusted by adjusting the value of the output transformation matrix corresponding to a specific attention head.
[0113] The attention heads in the large language model are ranked according to the sensitivity scores, for example, from large to small. The top-ranked attention heads may be safe word sensitive attention heads, the bottom-ranked attention heads may be unsafe word sensitive attention heads, and the middle-ranked attention heads may be neutral attention heads that are not sensitive to specific words.
[0114] In one embodiment, the current output tendency of the large language model is to excessively refuse to answer the safety prompt word; step 208 includes: increasing the value of the output weight matrix corresponding to each attention head with a higher sensitivity score ranking, and decreasing the value of the output weight matrix corresponding to each attention head with a lower sensitivity score ranking.
[0115] By increasing the value of the output weight matrix corresponding to each safe word sensitive attention head and decreasing the value of the output weight matrix corresponding to each unsafe word sensitive attention head, the tendency of the large language model to excessively refuse to answer the output of safe prompt words can be alleviated.
[0116] In another embodiment, the current output tendency of the large language model is that it is difficult to refuse to answer unsafe prompt words; step 208 includes: increasing the value of the output weight matrix corresponding to each attention head with a lower sensitivity score ranking.
[0117] By increasing the value of the output weight matrix corresponding to each unsafe word sensitive attention head, the security of the large language model can be improved, making it more inclined to refuse to answer unsafe prompt words.
[0118] In yet another embodiment, the current output tendency of the large language model is that it is difficult to comply with instructions; step 208 includes: increasing the value of the output weight matrix corresponding to each attention head with a higher sensitivity score.
[0119] Since the content in the safety fragment is usually specific instructions, the ability of the large language model to follow the instructions can be improved by increasing the value of the output weight matrix corresponding to each safety word sensitive attention head.
[0120] In the above embodiments, increasing the value of the output transformation matrix may be: multiplying the corresponding output transformation matrix by a first coefficient greater than 1; reducing the value of the output transformation matrix may be: multiplying the corresponding output transformation matrix by a second coefficient less than 1.
[0121] The above describes the process of using a first prompt word to adjust the output tendency of the large language model. In some possible implementations, multiple first prompt words can be constructed, and the output tendency of the large language model can be adjusted multiple times to improve the robustness of the large language model.
[0122] The embodiments of this specification first provide a method for automatically constructing prompt words, which can quickly construct the prompt words used in the embodiments of this specification in batches, and such prompt words contain two parts that show the difference between safe segments and unsafe segments. Then, the embodiments of this specification can use the sensitivity score as an indicator to specifically locate the specific attention head that causes the output tendency according to the output tendency of the large language model, and then adjust the attention head to alleviate the problems of the large language model and improve the overall output performance of the large language model. In addition, the adjustment process of the embodiments of this specification does not require the use of new training data to train or fine-tune the large language model, and there is no need to add a new neural network layer on the basis of the existing large language model, which saves computing resources as a whole.
[0123] According to an embodiment of another aspect, a device for adjusting an output tendency of a large language model is also provided. Figure 4 A schematic block diagram of an apparatus for adjusting the output tendency of a large language model according to an embodiment is shown, and the apparatus can be deployed in any device, platform or device cluster with computing and processing capabilities. Figure 4 As shown, the large language model includes multiple attention heads, each of which has a corresponding output transformation matrix, and the device 400 includes:
[0124] The acquisition unit 402 is configured to acquire a first prompt word, wherein the first prompt word includes a safe segment and an unsafe segment, and the safe segment is used to instruct the large language model to perform a first text processing without security risk on the unsafe segment;
[0125] an attention score determination unit 404, configured to input the first prompt word into the large language model to determine an attention score of a target attention head with respect to each word in the first prompt word;
[0126] A sensitivity score determination unit 406 is configured to determine a sensitivity score of a target attention head according to a first attention score statistic belonging to the safe segment and a second attention score statistic belonging to the unsafe segment in the attention scores of each word;
[0127] The transformation matrix adjustment unit 408 is configured to adjust the value of the output transformation matrix corresponding to each attention head with a high and / or low sensitivity score ranking according to the current output tendency of the large language model; the adjustment includes increasing or decreasing.
[0128] According to another aspect of the embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in any of the above embodiments.
[0129] According to yet another embodiment, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any one of the above embodiments is implemented.
[0130] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0131] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0132] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0133] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0134] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for adjusting the output tendency of a large language model, wherein the large language model includes a plurality of attention heads, each attention head having a corresponding output transformation matrix; the method comprising: Acquire a first prompt word, where the first prompt word includes a safe segment and an unsafe segment, where the safe segment is used to instruct the large language model to perform a first text processing without security risks on the unsafe segment; Inputting the first prompt word into the large language model to determine the attention score of the target attention head with respect to each word in the first prompt word; Determine a sensitivity score of a target attention head according to a first attention score statistic belonging to the safe segment and a second attention score statistic belonging to the unsafe segment in the attention scores of each word; According to the current output tendency of the large language model, the value of the output transformation matrix corresponding to each attention head with a high and / or low sensitivity score ranking is adjusted; the adjustment includes increasing or decreasing.
2. The method according to claim 1, wherein: The unsafe segment contains content that should be rejected by the large language model.
3. The method according to claim 1, wherein: The first text processing includes: text word counting, ignoring text, repeating text, and translating text.
4. The method according to claim 1, wherein: Get the first prompt word, including: An unsafe prompt word is obtained, and the unsafe prompt word is used as an unsafe segment and filled into a preset first prompt word template to obtain a first prompt word; the first prompt word template includes the safe segment.
5. The method according to claim 1, wherein: Determining the attention score of the target attention head with respect to each word in the first prompt word includes: Obtain a target attention matrix of a target attention head; the element value of the i-th row and j-th column of the target attention matrix represents the attention weight of the i-th word in the first prompt word with respect to the j-th word; According to the sum of the attention weights of each column in the target attention matrix, the attention score of the target attention head on each word is determined.
6. The method according to claim 5, wherein: The target attention matrix is determined based on the target query matrix and the target key matrix of the target attention head.
7. The method according to claim 1, wherein: Determining the sensitivity score of the target attention head according to the first attention score statistic belonging to the safe segment and the second attention score statistic belonging to the unsafe segment in the attention scores of each word, including: Determining a first attention score statistic according to the attention score of each word in the security segment; Determining a second attention score statistic according to the attention score of each word in the unsafe segment; The sensitivity score is determined according to a difference between the first attention score statistic and the second attention score statistic.
8. The method according to claim 1, wherein: The current output tendency of the large language model is excessively refusing to answer the safety prompt word; Adjust the values of the output transformation matrices corresponding to the attention heads with the highest and / or lowest sensitivity scores, including: Increase the value of the output weight matrix corresponding to each attention head with a higher sensitivity score, and decrease the value of the output weight matrix corresponding to each attention head with a lower sensitivity score.
9. The method according to claim 1, wherein: The current output tendency of the large language model is that it is difficult to refuse to answer unsafe prompt words; Adjust the values of the output transformation matrices corresponding to the attention heads with the highest and / or lowest sensitivity scores, including: Increase the value of the output weight matrix corresponding to each attention head with a lower sensitivity score.
10. The method according to claim 1, wherein: The current output tendency of the large language model is that it is difficult to follow the instructions; adjusting the value of the output transformation matrix corresponding to each attention head with a high sensitivity score ranking and / or a low sensitivity score ranking, including: Increase the value of the output weight matrix corresponding to each attention head with a higher sensitivity score.
11. The method according to claim 1, wherein: The increasing includes: multiplying the corresponding output transformation matrix by a first coefficient greater than 1; and the decreasing includes: multiplying the corresponding output transformation matrix by a second coefficient less than 1.
12. A device for adjusting the output tendency of a large language model, wherein the large language model includes a plurality of attention heads, each attention head having a corresponding output transformation matrix; the device comprises: an acquisition unit configured to acquire a first prompt word, wherein the first prompt word includes a safe segment and an unsafe segment, wherein the safe segment is used to instruct the large language model to perform a first text processing without security risk on the unsafe segment; an attention score determination unit, configured to input the first prompt word into the large language model to determine an attention score of a target attention head with respect to each word in the first prompt word; A sensitivity score determination unit is configured to determine a sensitivity score of a target attention head according to a first attention score statistic belonging to the safe segment and a second attention score statistic belonging to the unsafe segment in the attention scores of each word; The transformation matrix adjustment unit is configured to adjust the value of the output transformation matrix corresponding to each attention head with a high and / or low sensitivity score ranking according to the current output tendency of the large language model; the adjustment includes increasing or decreasing.
13. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 11.
14. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 11 is implemented.