Security reply generation method, related device, equipment and storage medium

By identifying session risks and selecting appropriate conversation strategies to route to the big model for processing, using multiple cross-attention and security knowledge fragments to generate secure replies, the security and real-time problems of content generated by large language models are solved, and more efficient secure replies are achieved.

CN120408414AActive Publication Date: 2025-08-01IFLYTEK CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510838962.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-01
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively improve the security and real-time response of large language models for generated content, especially when generating harmful, improper or false content, the processing efficiency is low.

Method used

By identifying the session risks of the statement to be responded to and routing the appropriate dialogue strategy to the corresponding large model for processing, including multiple large models that are suitable for security replies at different risk levels, using multiple cross-attention and security knowledge fragments for security replies generation.

Benefits of technology

It improves the security and real-time response of the original generation of large language models, reduces the post-processing requirement for generated content, and improves the security response capabilities under different risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408414A_ABST
    Figure CN120408414A_ABST
Patent Text Reader

Abstract

The invention discloses a safe reply generation method, a related device, equipment and a storage medium, and the method comprises the steps: carrying out the recognition based on a to-be-responded statement, and obtaining a session risk for replying the to-be-responded statement; selecting a dialogue strategy matched with the dialogue risk from a plurality of dialogue strategies as a target strategy; wherein the plurality of dialogue strategies at least comprise a plurality of large models which are respectively used as different dialogue strategies, and the different large models are respectively suitable for security reply under different dialogue risks; and processing the to-be-responded statement based on the target strategy to obtain reply content of the to-be-responded statement. According to the scheme, the safety and the response real-time performance of original generation of the model can be improved as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a method for generating a secure reply and related devices, equipment, and storage media. Background Art

[0002] With the widespread application of large language models in many tasks, their potential security risks (such as generating harmful, inappropriate or false content) have gradually become an important issue that needs to be addressed urgently.

[0003] Existing security mechanisms typically perform secondary screening of model-generated content to minimize security risks. However, this approach makes it difficult to correct for deviations in the original model generation and also suffers from relatively poor real-time response. Therefore, maximizing the security and real-time response of the original model generation has become a pressing issue. Summary of the Invention

[0004] The main technical problem solved by this application is to provide a secure response generation method and related devices, equipment and storage media, which can maximize the security of the original generation of the model and the real-time response.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a method for generating a safe response, including: identifying based on a statement to be responded to, and obtaining the session risk of responding to the statement to be responded to; selecting a dialogue strategy that matches the session risk from several dialogue strategies as a target strategy; wherein the several dialogue strategies include at least a plurality of large models that serve as different dialogue strategies, and different large models are respectively suitable for providing safe responses under different session risks; processing the statement to be responded to based on the target strategy, and obtaining the response content of the statement to be responded to.

[0006] In order to solve the above technical problems, the second aspect of the present application provides a safe response generation device, including: a risk identification module, a strategy matching module and a statement response module. The risk identification module is used to identify based on the statement to be responded to obtain the session risk of replying to the statement to be responded to; the strategy matching module is used to select a dialogue strategy that matches the session risk from several dialogue strategies as a target strategy; wherein the several dialogue strategies include at least multiple large models that serve as different dialogue strategies, and different large models are respectively suitable for making safe responses under different session risks; the statement response module is used to process the statement to be responded to based on the target strategy to obtain the response content of the statement to be responded to.

[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, which at least includes a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the security reply generation method in the above first aspect.

[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the secure response generation method of the first aspect above.

[0009] In the above solution, based on the statement to be responded to, the session risk of responding to the statement to be responded to is identified, and then a dialogue strategy that matches the session risk is selected from several dialogue strategies as the target strategy. Among the several dialogues, at least multiple large models that are different dialogue strategies are included, and different large models are respectively applicable to making secure responses under different session risks. Furthermore, based on the target strategy, the statement to be responded to is processed to obtain the response content of the statement to be responded to. Therefore, on the one hand, different from post-processing the output content of the model, by routing the statement to be responded to to the matching dialogue strategy for processing according to the session risk, the security of the original generation of the model can be improved as much as possible. On the other hand, since different large models are respectively applicable to making secure responses under different session risks, after the target strategy processes the statement to be responded to and obtains its response content, there is no need for post-processing, which can improve the response real-time performance. Therefore, the security of the original generation of the model and the response real-time performance can be improved as much as possible. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a schematic flowchart of an embodiment of the secure response generation method of the present application; Figure 2 is a schematic diagram of the process of an embodiment of the secure response generation method of the present application; Figure 3 is a schematic framework diagram of an embodiment of the secure response generation device of the present application; Figure 4 is a schematic framework diagram of an embodiment of the electronic device of the present application; Figure 5 is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] The following will describe in detail the solutions of the embodiments of the present application with reference to the accompanying drawings of the specification.

[0012] In the following description, specific details such as specific system structures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0013] The terms "system" and "network" are often used interchangeably in this document. The term " / or" in this document is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, in this document, the segment " / " generally indicates that the related objects before and after are in an "or" relationship. Furthermore, "plurality" in this document means two or more than two.

[0014] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the method for generating a secure response in this application. Specifically, it may include the following steps: Step S11: Identify based on the statement to be responded to, and obtain the session risk of the response to the statement to be responded to.

[0015] In one implementation scenario, the specific content of the statement to be responded to may also vary with different application scenarios. Additionally, the statement to be responded to can be obtained by receiving it from a user input interface (User Input Interface). For specific details, please refer to the technical details of the user input interface, which will not be elaborated here.

[0016] In one implementation scenario, the session risk of the response to the statement to be responded to at least represents the degree of harm that may be caused if the statement to be responded to is correctly replied according to the prompt (such as being distinguished by different risk levels such as "extremely high, high, medium, low", etc.). Of course, the session risk may not be limited to this. For example, the session risk may also include risk types (such as prejudice, discrimination, privacy leakage), etc. The specific content of the session risk will not be limited here either.

[0017] In one implementation scenario, in order to identify the session risk of the statement to be responded to, the statement to be responded to can be identified based on a risk identifier (Dynamic Risk Identifier, DRI), and the identification result can be obtained. The identification result can include the possibility that the statement to be responded to belongs to different risk categories respectively. The risk identifier can include multiple levels for risk identification from different dimensions. On this basis, a multi-layer perceptron (Multi-Layer Perceptron, MLP) can be used to predict the identification result to obtain the session risk of the response to the statement to be responded to. Through the above method, the relevant identification of risk categories for the statement to be responded to based on a multi-level and multi-dimensional risk identifier helps to improve the accuracy of risk categories. Then, by predicting the above identification result through a multi-layer perceptron, the accuracy of session risk identification can be improved.

[0018] In a specific implementation scenario, multiple levels of the risk recognizer may include a first level, which may perform recognition based on at least one of a sensitive word library, regular expressions, and risk content patterns. For example, the first level may perform recognition based only on one of the above three, or the first level may also perform recognition based on a combination of any two of the above three, or the first level may further perform recognition based on all of the above three, which is not limited herein. It should be noted that the sensitive word library may include sensitive words in various application scenarios (such as the aforementioned text generation, knowledge Q&A, etc. application scenarios) and the risk categories to which the sensitive words belong, so as to determine the risk category to which the sensitive word belongs when a sensitive word in the sensitive word library is recognized in the to-be-responded statement. Regular expressions may include regular expressions corresponding to different risk categories in various application scenarios, so as to determine the risk category to which the to-be-responded statement belongs when the to-be-responded statement matches a certain regular expression. The risk content pattern may include induced sentence patterns, speech structures, etc. corresponding to different risk categories in various application scenarios, so as to determine the risk category to which the to-be-responded statement belongs when the to-be-responded statement matches a certain risk content pattern.

[0019] In a specific implementation scenario, multiple levels of the risk recognizer may include a second level, which may perform recognition based on the semantic features of the to-be-responded statement. It can be understood that different from the aforementioned first level that performs risk recognition on the literal expressions such as the vocabulary and sentence patterns of the to-be-responded statement, the second level performs risk recognition at the semantic level of the to-be-responded statement, and can recognize possible situations where there is no risk in the literal expression but there is a risk in the deep semantics. Specifically, a lightweight neural network model may be used to perform intent recognition at the semantic level of the to-be-responded statement to determine the possibility of the to-be-responded statement regarding various risk categories; or, the semantic features of the to-be-responded statement may also be extracted, and based on the similarity between the semantic features of the to-be-responded statement and the semantic features of various risk categories, the possibility of the to-be-responded statement belonging to various risk categories may be determined. Of course, the above examples are only several possible examples of performing risk recognition based on the semantic features of the to-be-responded statement in the actual application process, and other possible situations are not limited herein, nor will they be listed one by one.

[0020] In a specific implementation scenario, multiple levels of the risk recognizer may include a third level, which may perform recognition on the to-be-responded statement by combining at least one of the conversation history, user profile, and task background. It should be noted that the risk recognition of the to-be-responded statement may be performed by combining any one of the above three; or, the risk recognition of the to-be-responded statement may also be performed by combining any two combinations of the above three; or, the risk recognition of the to-be-responded statement may further be performed by combining all of the above three, which is not limited herein. In addition, when it is necessary to combine the user profile, user authorization may be obtained in advance.

[0021] In a specific implementation scenario, the above three levels of the risk recognizer can be executed sequentially (i.e., the three levels can be in a series relationship), that is, the statement to be responded to can be recognized by the above three levels in sequence, so as to obtain the final recognition result by integrating the sub-recognition results of each level for the statement to be responded to. Or, after a certain level obtains the sub-recognition result of the statement to be responded to including each risk category and its possibility, it no longer continues to execute the subsequent levels, but directly uses this sub-recognition result as the final recognition result. Of course, in some cases, the above three levels of the risk recognizer can also be executed in parallel (i.e., the three levels can be in a parallel relationship), that is, the statement to be responded to can be recognized by the above three levels simultaneously, so as to obtain the final recognition result by integrating the sub-recognition results of each level for the statement to be responded to. Or, after a certain level obtains the sub-recognition result of the statement to be responded to including each risk category and its possibility prior to other levels, the recognition processes of other levels are terminated. Of course, the above examples are only several possible examples of the working methods of the three levels in the risk recognizer. Other possible situations are not limited here, and no more examples will be given one by one.

[0022] In a specific implementation scenario, during the training process of the risk recognizer, a large number of clearly safe samples can be used as negative examples, and at the same time, a generative model (even an adversarial model) can be used to generate various boundary, fuzzy, and implicit unsafe samples and adversarial samples to enhance the robustness of the risk recognizer. In addition, manually labeled risk judgment data can also be used to fine-tune the risk recognizer to make its evaluation results more in line with the judgments of human experts.

[0023] Step S12: Select a dialogue strategy that matches the conversation risk from several dialogue strategies as the target strategy.

[0024] In the disclosed embodiments, several conversation strategies may include at least multiple large models, each serving as a different conversation strategy. Each large model may be suitable for providing secure responses under different conversation risks. As one possible example, the multiple large models may include a first large model and a second large model, and the second large model may be suitable for a higher conversation risk than the first large model. Of course, the above examples are merely possible examples of multiple large models. The possibility of multiple large models including other numbers of large models is not limited herein, and no further examples will be given. Furthermore, as previously mentioned, conversation risk can include risk levels such as "extremely high," "high," "medium," and "low," along with their likelihood. In this case, the risk level of the statement to be responded to can be determined based on each risk level and its likelihood. For example, the risk level with the highest likelihood can be selected as the risk level of the statement to be responded to; or, when the likelihood is expressed as a probability value, the risk level with a probability value above a preset threshold can be selected as the risk level of the statement to be responded to. Of course, the above examples are merely a few possible examples of determining the risk level of the statement to be responded to. Other possible methods are not limited herein, and no further examples will be given.

[0025] In an implementation scenario, please refer to Figure 2 , Figure 2 This is a process diagram of an embodiment of the method for generating a security reply of this application. Figure 2 As shown, after obtaining the conversation risk of the statement to be responded to, the risk level represented by the conversation risk can be detected. In response to a low risk or no risk risk, the first model among several conversation strategies can be selected as the target strategy, while in response to a medium-high risk risk, the second model among several conversation strategies can be selected as the target strategy. It should be noted that the first model can be obtained by training a general large model using a first dataset. The first dataset can be obtained by mixing medium- and high-risk sample conversation pairs with low- or no-risk sample conversation pairs (i.e., the data of the second model is fed back to the first model), and the sample response content in the sample conversation pairs can be a safe response to the sample statement to be responded to in the sample conversation pairs. The second model can be obtained by training the general large model using a second dataset. The second dataset can include medium- and high-risk sample conversation pairs. In the above approach, the first model is selected as the target strategy when the risk is low or no, and the second model is selected as the target strategy when the risk is medium or high. In addition, the first model is mixed with medium and high-risk sample conversation pairs during training, which can prevent the first model from "drifting" and producing unsafe content as much as possible. The optimization goal is more inclined towards general performance. The second model is trained with medium and high-risk sample conversation pairs, which can force the second model to learn how to respond safely under medium and high risks as much as possible. Therefore, the safety response capabilities of the first and second models can be improved during the training of the large models.

[0026] In an implementation scenario, several dialogue strategies can include not only the above-mentioned multiple large models but also fallback responses. It should be noted that the fallback response can be used to output either a warning response or a rejection response as the response content for the response statement when the risk level characterized by the conversation risk is an extremely high risk. For example, in this case, the response content can include, but is not limited to, "Sorry, since there is a security risk in the content you asked, please re-enter!" "Warning, there is a security risk in the content you asked!" and so on. In this regard, the response content under extremely high risk is not limited here, and no further examples will be given one by one.

[0027] In an implementation scenario, when several dialogue strategies include multiple large models, the multiple large models can share at least part of the underlying representations or learn from each other through knowledge distillation. It should be noted that for the specific process of knowledge distillation, the technical details of knowledge distillation can be referred to and will not be elaborated here. Through the above method, by having the multiple large models share at least part of the underlying representations or learn from each other through knowledge distillation, it is possible to make the output styles and knowledge levels of the multiple large models as consistent as possible within a safe range, so as to avoid the user perceiving an obvious model switching gap as much as possible.

[0028] Step S13: Process the response statement to be responded to based on the target strategy to obtain the response content of the response statement to be responded to.

[0029] In an implementation scenario, when using a large model as the target strategy to make a safe response to the response statement to be responded to, the context features of the large model as the target strategy during the process of processing the response statement to be responded to can be obtained, and the knowledge features of the security knowledge fragments related to the response statement to be responded to can be obtained. On this basis, multi-head cross-attention can be performed based on the context features and the knowledge features to obtain the output features of the multi-head cross-attention, and the context features and the output features can be fused to obtain the fused features. Furthermore, the fused features can be further processed based on the large model to obtain the response content of the response statement to be responded to. Through the above method, knowledge fusion is achieved through multi-head cross-attention, which can improve the safe response to the response statement to be responded to as much as possible with the assistance of the security knowledge fragments.

[0030] In a specific implementation scenario, the context features can be output by the Transformer layer in the large model after the response statement to be responded to is input, and it can represent the hidden state of the text sequence (including the response statement to be responded to and the partially generated response content) up to the current moment. For the convenience of description, the context features can be denoted as C ∈ R B*L_c*d , B*L_k*d , ,

[0030] , where B represents the batch size, L_c represents the context sequence length, and d represents the model dimension. In addition, the knowledge features of the security knowledge fragments can be denoted as K ∈ R B*L_k*d, where \(L_k\) represents the length of the knowledge sequence (e.g., the total number of tokens in a security knowledge fragment or the number of structured entries). Additionally, if the number of heads in the multi-head cross-attention is \(H\), then the dimension of each attention head is \(d_h = d / H\). It should be noted that a security knowledge base can be pre-constructed, which may contain several security knowledge fragments. Then, based on the similarity between the context features and the knowledge features of each security knowledge fragment, at least one security knowledge fragment can be selected to perform the subsequent multi-head cross-attention. As a possible example, the security knowledge base can include, but is not limited to, the following: a sensitive word library, a laws and regulations database, a fact-checking system, etc. The specific content in the security knowledge base is not limited here.

[0031] In a specific implementation scenario, when performing multi-head cross-attention, the corresponding query features (Query, Q), key features (Key, K), and value features (V, Value) can be calculated for the context features \(C\) and the knowledge features \(K\) respectively. Specifically, different projection matrices can be used to allow the model to learn to distinguish between context features and knowledge features. As a possible example, when projecting the context features, the query feature, key feature, and value feature of the context features can be obtained: \(Q_c = C * W_{qc} \in \mathbb{R}\) B*L_c*d \(K_c = C * W_{kc} \in \mathbb{R}\) B*L_c*d \(V_c = C * W_{vc} \in \mathbb{R}\) B*L_c*d Similarly, when projecting the knowledge features, the query feature, key feature, and value feature of the knowledge features can be obtained: \(Q_k = K * W_{qk} \in \mathbb{R}\) B*L_k*d \(K_k = K * W_{kk} \in \mathbb{R}\) B*L_k*d \(V_k = K * W_{vk} \in \mathbb{R}\) B*L_k*d In the above formulas, \(W_{qc}\), \(W_{kc}\), \(W_{vc}\), \(W_{qk}\), \(W_{kk}\), and \(W_{vk}\) respectively represent learnable weight matrices.

[0032] In a specific implementation scenario, after obtaining the above query features, key features, and value features, the above features can be reshaped for multi-head. That is, it can be split into \(H\) heads along the model dimension \(d\). For example, \(Q_c\) becomes \(Q_c' \in \mathbb{R}\) B*L_c*d_h . The same applies to \(K_c\), \(V_c\), \(Q_k\), \(K_k\), and \(V_k\) respectively, which will not be elaborated here one by one.

[0033] In a specific implementation scenario, after multi-head reshaping, the query feature Q_c’ of the context features can be used as the query party to attend to the key feature K_k’ and the value feature V_k’ of the knowledge features. Specifically, the attention scores can be calculated based on the similarity between the query feature Q_c’ and the key feature K_k’. As a possible example, the attention scores can be expressed as: Scores_ck = matmul(Q_c’, K_k’ T ) / sqrt(d_h) In the above formula, K_k’ T represents the transpose of the last two dimensions of K_k’, with a shape of B * H * d_h * L_k, and Scores_ck ∈ R B*H*L_c*L_k . Based on this, the attention weights can be obtained based on the attention scores: Atten_ck = softmax(Scores_ck, dim=-1) ∈ R B*H*L_c*L_k Then, the value features of the knowledge features can be weighted using the attention weights: Info_from_K’ = matmul(Atten_ck, V_k’) ∈ R B*H*L_c*d_h In the above formula, Info_from_K’ represents the weighted information summary “extracted” from the knowledge base according to the current context.

[0034] In a specific implementation scenario, after obtaining the processing results Info_from_K’ of each attention head, the processing results Info_from_K’ of the H attention heads can be concatenated to obtain the output features of the multi-head cross-attention. As a possible example, the output features of the multi-head cross-attention can be expressed as: Info_from_K = Info_from_K’.transpose(1, 2).reshape(B, L_c, d) ∈ R B*L_c*d In the above formula, Info_from_K represents the output features of the multi-head cross-attention. As another possible example, a linear projection can also be performed on the output features: Info_from_K_proj = Info_from_K * W_o ∈ R B*L_c*d In the above formula, W_o is the projection matrix of the output features. Of course, in actual application, this step can also be omitted and the aforementioned output features Info_from_K can be directly used.

[0035] In a specific implementation scenario, after obtaining the output features of the multi-head cross-attention, the context features and the output features can be fused to obtain the fused features. Specifically, predictions can be made based on the context features and the output features to obtain the fusion weights of the context features and the output features. Exemplarily, a gated unit can be used to dynamically determine how much information extracted from the knowledge is fused into the original context feature C. For example, the fusion weight can be expressed as: g = sigmoid(Linear_g(concat(C, Info_from_K))) ∈ R B*L_c*d In the above formula, g represents the fused feature, sigmoid represents the normalization function, concat represents concatenation, and Linear_g represents a linear layer (whose parameters are learnable during training to achieve dynamic adjustment of the fusion weights). Based on this, the context features and the output features can be weighted based on the fusion weights to obtain the fused features. Exemplarily, the fused feature can be expressed as: Fused_info = (1 - g) * C + g * Info_from_K ∈ R B*L_c*d It can be seen that the larger the fusion weight g, the higher the proportion of the knowledge information Info_from_K fused in. Conversely, more of the original context feature C is retained. After obtaining the fused features, the subsequent network layers of the large model can continue to process them to obtain the response content of the response statement to be replied.

[0036] In an implementation scenario, when the large model is used as the target policy, especially when the second large model is used as the target policy (i.e., the risk level of the response statement to be replied is medium to high), the second large model can first understand the potential "benign" intention of the user in the response statement to be replied, rewrite it as a safe response statement around the "benign" intention, and then make a safe response to the safe response statement. Of course, the second large model can also directly make a safe response to the response statement to be replied. In addition, if the second large model determines to give a rejection response during the process of processing the response statement to be replied, it can further give the reason for rejection. For example, the reason for rejection can be given in combination with the risk category (such as, "Since your query involves the XXX risk category and does not comply with the regulations, no answer will be given", etc.), and relevant regulations or community guidelines should be followed. Of course, for response statements that may cause prejudice or discrimination, neutral, objective, and balanced views can be generated, or the prejudice risk in the content can be pointed out. For response statements outside its knowledge scope or that cannot be safely replied, it can clearly state that it is unaware or unable to answer, rather than fabricating. It should be noted that the above examples are only several possible examples in the actual application process of the large model's response processing. Other possible situations are not limited here and will not be exemplified one by one.

[0037] In one implementation scenario, during the processing, the large model can obtain the output characters at each decoding moment through autoregressive decoding until the output character at a certain decoding moment is an end character (such as <eos>) until the output characters at each decoding moment can be combined to form the reply content. It should be noted that at any decoding moment, the large model can output the probability values of each predicted character at the current decoding moment, so that a predicted character can be selected as the output character based on the probability values of the predicted characters. In this case, to further enhance the security of the content output by the large model, the probability distribution of the predicted characters can be correspondingly constrained when the predicted characters output by the large model are risky. As previously mentioned, the final output content of the large model can be determined based on the probability distribution of the predicted characters. For example, when the predicted characters are sensitive words, the probability distribution of the predicted characters can be suppressed (e.g., by multiplying the probability distribution of the predicted characters by a value between 0 and 1). In this way, the probability distribution of the predicted characters is constrained when the predicted characters output by the large model are risky, and the final output content of the large model is determined based on the probability distribution of the predicted characters, which can minimize the possibility of generating unsafe words or sentences. Of course, in actual application, the above strategy can be applied to both the inference stage and the training stage. Furthermore, when the multiple large models include a first large model and a second large model, the above strategy can be applied to both the first large model and the second large model. Still taking multiple large models including the aforementioned first and second models as an example, the first model can be based on RLHF (Reinforcement Learning from Human Feedback) and can also add penalty items for unsafe outputs during the training process, guiding it to avoid "harmfulness" while pursuing "usefulness".

[0038] In one implementation scenario, still taking the example of multiple large models including at least the first large model and the second large model, as mentioned above, the conversation risk applicable to the second large model is higher than the conversation risk applicable to the first large model. Different from the training method of the first large model, during the training process of the second large model, a sample prompt instruction can be constructed based on the safety principle and the sample sentences to be responded to that have a response risk. The sample prompt instruction is used to instruct the second large model to follow the safety principle to respond safely to the sample sentences to be responded to. On this basis, the output content of the second large model in response to the sample prompt instruction can be obtained as the predicted response content of the sample sentences to be responded to, and then the network parameters of the second large model can be adjusted based on at least the difference between the predicted response content of the sample sentences to be responded to and the sample response content expected to be responded safely by the second large model. Then, during the training process, the second large model can be guided by the safety principle to self-criticize and correct the internal generation attempts that do not comply with the principle.

[0039] In a specific implementation scenario, the security response types of the sample prompt instructions can be diverse, such as but not limited to: directly rejecting, explanatory rejection (i.e., not only rejecting but also explaining the reason for rejection), secure rewriting, neutralized answering, guiding questions, etc. The security response types of the sample prompt instructions are not limited here.

[0040] In a specific implementation scenario, the security principles can cover but not be limited to multiple dimensions such as bias, discrimination, privacy, etc. They are not limited here and will not be exemplified one by one.

[0041] In a specific implementation scenario, during the training process of the second large model, reference response sentences with response risks can also be generated, and the reference response content predicted by the second large model for the reference response sentences can be obtained. It should be noted that a dedicated model can be used to generate reference response sentences with response risks. On this basis, in response to the situation that the reference response content cannot safely respond to the reference response sentences, the reference response sentences can be selected as the new sample response sentences during the training of the second large model. In this way, an adversarial training loop can be formed with the second large model. In addition, a dedicated team can also continuously design new, tricky, and veiled reference response sentences to try to attack the second large model, so as to expose the weaknesses of the second large model as much as possible, and add the successfully attacked reference response sentences to the training set, that is, the successfully attacked reference response sentences are used as the new sample response sentences.

[0042] In a specific implementation scenario, in addition to the above training strategies, during the training process of the second large model, a dedicated dataset can also be used to train the second large model to learn "how to gracefully and firmly reject improper requests" and explain the reasons. In addition, <biased output, unbiased output> pairs can be constructed to train the model to distinguish and tend to generate unbiased response content.

[0043] In a specific implementation scenario, as mentioned above, the session risk can be identified by the risk recognizer for the to-be-responded statement, and then the risk recognizer and the large model can be jointly trained based on the total loss. It should be noted that the total loss can be measured based on the judgment accuracy of the risk recognizer, the output security of the large model, and the diversion effectiveness of several dialogue strategies. As a possible example, the judgment accuracy of the risk recognizer can be obtained by measuring the difference between the session risk predicted by the risk recognizer for the sample to-be-responded statement and the session risk labeled for the sample to-be-responded statement; or, as another possible example, the output security of the large model can be obtained by measuring the difference between the response content of the large model to the sample to-be-responded statement and the response content labeled for the sample to-be-responded statement; or, as yet another possible example, the diversion effectiveness of several dialogue strategies can be obtained by measuring the difference between the dialogue strategy actually diverted for the sample to-be-responded statement and the dialogue strategy expected to be diverted as labeled. Of course, the above examples are only possible examples of separately measuring the judgment accuracy, output security, and judgment accuracy in the actual application process. The measurement methods for the judgment accuracy, output security, and judgment accuracy are not limited here, and no further examples will be given one by one.

[0044] In the above solution, based on the to-be-responded statement, the session risk for replying to the to-be-responded statement is identified, and then a dialogue strategy matching the session risk is selected from several dialogue strategies as the target strategy. Among the several dialogues, there are at least multiple large models serving as different dialogue strategies respectively, and different large models are respectively applicable to making safe replies under different session risks. Then, based on the target strategy, the to-be-responded statement is processed to obtain the response content of the to-be-responded statement. Therefore, on the one hand, different from post-processing the output content of the model, by routing the to-be-responded statement to the matching dialogue strategy according to the session risk for processing, the security of the original generation of the model can be improved as much as possible. On the other hand, since different large models are respectively applicable to making safe replies under different session risks, there is no need for post-processing after the target strategy processes the to-be-responded statement to obtain its response content, which can improve the response real-time performance. Therefore, the security of the original generation of the model and the response real-time performance can be improved as much as possible.

[0045] Please refer to Figure 3 , Figure 3 It is a schematic framework diagram of an embodiment of the security response generation device of the present application. The security response generation device 30 includes: a risk identification module 31, a policy matching module 32, and a statement response module 33. The risk identification module 31 is configured to identify based on the statement to be responded to, and obtain the conversation risk of responding to the statement to be responded to. The policy matching module 32 is configured to select, from a number of conversation policies, a conversation policy that matches the conversation risk as the target policy. Among them, at least a number of conversation policies include multiple large models that are respectively different conversation policies, and different large models are respectively applicable to making security responses under different conversation risks. The statement response module 33 is configured to process the statement to be responded to based on the target policy, and obtain the response content of the statement to be responded to.

[0046] In the above solution, the security response generation device 30 identifies based on the statement to be responded to, and obtains the conversation risk of responding to the statement to be responded to, so as to select, from a number of conversation policies, a conversation policy that matches the conversation risk as the target policy. Among the number of conversations, at least a number of conversation policies include multiple large models that are respectively different conversation policies, and different large models are respectively applicable to making security responses under different conversation risks. Furthermore, the statement to be responded to is processed based on the target policy, and the response content of the statement to be responded to is obtained. Therefore, on the one hand, different from post-processing the output content of the model, by routing the statement to be responded to to the matching conversation policy for processing according to the conversation risk, the security of the original generation of the model can be improved as much as possible. On the other hand, since different large models are respectively applicable to making security responses under different conversation risks, there is no need to perform post-processing after the target policy processes the statement to be responded to and obtains its response content, which can improve the response real-time performance. Therefore, the security of the original generation of the model and the response real-time performance can be improved as much as possible.

[0047] In some publicly disclosed embodiments, the policy matching module 32 includes a level detection sub-module, which is configured to detect the risk level represented by the conversation risk. The policy matching module 32 includes a first selection sub-module, which is configured to select the first large model among a number of conversation policies as the target policy in response to the risk level being low risk or no risk. The policy matching module 32 includes a second selection sub-module, which is configured to select the second large model among a number of conversation policies as the target policy in response to the risk level being medium-high risk. Among them, the first large model is obtained by training a general large model through a first data set, and the first data set is obtained by mixing medium-high risk sample conversation pairs into low-risk or no-risk sample conversation pairs, and the sample response content in the sample conversation pair is a security response to the sample statement to be responded to in the sample conversation pair. The second large model is obtained by training a general large model through a second data set, and the second data set includes medium-high risk sample conversation pairs.

[0048] In some disclosed embodiments, the risk identification module 31 includes an identification sub-module for identifying a to-be-responded statement based on a risk identifier to obtain an identification result; wherein the identification result includes the probabilities that the to-be-responded statement belongs to several risk categories respectively, and the risk identifier includes multiple levels for risk identification from different dimensions respectively; the risk identification module 31 includes a prediction sub-module for predicting the identification result based on a multi-layer perceptron to obtain the conversation risk of replying to the to-be-responded statement.

[0049] In some disclosed embodiments, the multiple levels include a first level, a second level, and a third level; wherein the first level performs identification based on at least one of a sensitive word library, a regular expression, and a risk content pattern, the second level performs identification based on the semantic features of the to-be-responded statement, and the third level performs identification on the to-be-responded statement in combination with at least one of the conversation history, user profile, and task background.

[0050] In some disclosed embodiments, the risk identifier and the large model are jointly trained based on the total loss, and the total loss is obtained by measuring based on the judgment accuracy of the risk identifier, the output security of the large model, and the diversion effectiveness of several conversation strategies.

[0051] In some disclosed embodiments, the multiple large models at least include a first large model and a second large model, and the conversation risk applicable to the second large model is higher than that applicable to the first large model. The secure reply generation device 30 includes an instruction construction module for constructing a sample prompt instruction based on security principles and a sample to-be-responded statement with reply risk; wherein the sample prompt instruction is used to instruct the second large model to make a secure reply to the sample to-be-responded statement in accordance with the security principles; the secure reply generation device 30 includes a content acquisition module for acquiring the output content of the second large model in response to the sample prompt instruction as the predicted reply content of the sample to-be-responded statement; the secure reply generation device 30 includes a parameter adjustment module for adjusting the network parameters of the second large model at least based on the difference between the predicted reply content of the sample to-be-responded statement and the sample reply content expected to be securely replied by the second large model.

[0052] In some disclosed embodiments, the secure reply generation device 30 includes a statement generation module for generating a reference to-be-responded statement with reply risk, the secure reply generation device 30 includes a reply acquisition module for acquiring the reference reply content predicted by the second large model for the reference to-be-responded statement; the secure reply generation device 30 includes a sample update module for selecting the reference to-be-responded statement as a new sample to-be-responded statement during the training of the second large model in response to the reference reply content being unable to securely reply to the reference to-be-responded statement.

[0053] In some disclosed embodiments, the statement reply module 33 includes a feature acquisition sub-module, configured to acquire the context features of the large model serving as the target policy during the process of processing the statement to be replied, and acquire the knowledge features of the security knowledge segments related to the statement to be replied; the statement reply module 33 includes an attention sub-module, configured to perform multi-head cross-attention based on the context features and the knowledge features to obtain the output features of the multi-head cross-attention; the statement reply module 33 includes a feature fusion sub-module, configured to fuse the context features and the output features to obtain the fused features; the statement reply module 33 includes a feature processing sub-module, configured to continue to process the fused features based on the large model to obtain the reply content of the statement to be replied.

[0054] In some disclosed embodiments, the feature fusion sub-module includes a weight prediction unit, configured to perform prediction based on the context features and the output features to obtain the fusion weights of the context features and the output features; the feature fusion sub-module includes a feature weighting unit, configured to weight the context features and the output features based on the fusion weights to obtain the fused features.

[0055] In some disclosed embodiments, among several dialogue policies, there is also a fallback reply, which is used to output either a warning reply or a rejection reply as the reply content when the risk level characterized by the session risk is an extremely high risk; and / or, when there is a risk in the predicted characters output by the large model, the probability distribution of the predicted characters is constrained, and the final output content of the large model is determined based on the probability distribution of the predicted characters; and / or, multiple large models share at least part of the underlying representation or learn from each other through knowledge distillation.

[0056] Please refer to Figure 4 , Figure 4 which is a schematic framework diagram of an embodiment of the electronic device of the present application. The electronic device 40 at least includes a memory 41 and a processor 42 that are coupled to each other. The memory 41 stores at least program instructions, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above embodiments of the security reply generation method. Specifically, reference can be made to the foregoing disclosed embodiments, which will not be elaborated herein. As a possible example, the electronic device 40 may include, but is not limited to, devices such as mobile phones, tablet computers, learning machines, smart large screens, servers, etc. The specific type of the electronic device 40 is not limited herein.

[0057] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned embodiments of the secure response generation method. The processor 42 can also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with the ability to process signals. The processor 42 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Additionally, the processor 42 can be implemented jointly by integrated circuit chips.

[0058] In the above solution, the electronic device 40 identifies based on the statement to be responded to, obtains the session risk of the response to the statement to be responded to, and thus selects, from several dialogue strategies, a dialogue strategy that matches the session risk as the target strategy. And among the several dialogues, at least multiple large models respectively serving as different dialogue strategies are included. Different large models are respectively applicable to making secure responses under different session risks. Furthermore, based on the target strategy, the statement to be responded to is processed to obtain the response content of the statement to be responded to. Therefore, on the one hand, different from post-processing the output content of the model, by routing the statement to be responded to to the matching dialogue strategy for processing according to the session risk, the security of the original generation of the model can be improved as much as possible. On the other hand, since different large models are respectively applicable to making secure responses under different session risks, after the target strategy processes the statement to be responded to and obtains its response content, there is no need for post-processing, which can improve the response real-time performance. Therefore, the security of the original generation of the model and the response real-time performance can be improved as much as possible.

[0059] Please refer to Figure 5 , Figure 5 which is a framework schematic diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor. The program instructions 51 are used to implement the steps in any of the above-mentioned embodiments of the secure response generation method.

[0060] In the above solution, the computer-readable storage medium 50 performs recognition based on the statement to be replied, obtains the session risk for replying to the statement to be replied, and thus selects, from several dialogue strategies, a dialogue strategy that matches the session risk as the target strategy. Among the several dialogues, at least multiple large models that are different dialogue strategies are included, and different large models are respectively applicable to making secure replies under different session risks. Furthermore, based on the target strategy, the statement to be replied is processed to obtain the reply content of the statement to be replied. Therefore, on the one hand, different from post-processing the output content of the model, by routing the statement to be replied to the matching dialogue strategy according to the session risk for processing, the security of the original generation of the model can be improved as much as possible. On the other hand, since different large models are respectively applicable to making secure replies under different session risks, after the target strategy processes the statement to be replied to obtain its reply content, there is no need for post-processing, which can improve the response real-time performance. Therefore, the security of the original generation of the model and the response real-time performance can be improved as much as possible.

[0061] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0062] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0063] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0064] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0065] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0066] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various implementation methods of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0067] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the individual of the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information. Among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.< / eos>

Claims

1. A secure reply generation method, characterized in that, Including: Based on the statement to be responded to, identify the conversation risk of replying to the statement to be responded to; Select a conversation strategy that matches the conversation risk from a number of conversation strategies as the target strategy; wherein, at least a number of large models are included in the number of conversation strategies, and different large models are respectively applicable to making secure replies under different conversation risks; Process the statement to be responded to based on the target strategy to obtain the reply content of the statement to be responded to.

2. The method according to claim 1, wherein The step of selecting a conversation strategy that matches the conversation risk from a number of conversation strategies as the target strategy includes: Detect the risk level characterized by the conversation risk; In response to the risk level being low risk or no risk, select the first large model among the number of conversation strategies as the target strategy; In response to the risk level being medium-high risk, select the second large model among the number of conversation strategies as the target strategy; wherein, the first large model is obtained by training a general large model with a first data set, the first data set is obtained by mixing sample conversation pairs with medium-high risk into sample conversation pairs with low risk or no risk, and the sample reply content in the sample conversation pairs is a secure reply to the sample statement to be responded to in the sample conversation pairs, and the second large model is obtained by training the general large model with a second data set, and the second data set contains the sample conversation pairs with medium-high risk.

3. The method according to claim 1, wherein The step of based on the statement to be responded to, identify the conversation risk of replying to the statement to be responded to, includes: Based on a risk recognizer, identify the statement to be responded to to obtain an identification result; wherein, the identification result includes the probabilities of the statement to be responded to belonging to a number of risk categories respectively, and the risk recognizer includes a number of levels for risk identification from different dimensions; Based on a multi-layer perceptron, predict the identification result to obtain the conversation risk of replying to the statement to be responded to.

4. The method according to claim 3, wherein The number of levels includes a first level, a second level and a third level; wherein, the first level is identified based on at least one of a sensitive word library, a regular expression, a risk content pattern, the second level is identified based on the semantic features of the statement to be responded to, and the third level identifies the statement to be responded to in combination with at least one of the conversation history, the user profile, the task background.

5. The method according to claim 3, wherein The risk recognizer and the large model are jointly trained based on the total loss, and the total loss is obtained by measuring the judgment accuracy of the risk recognizer, the output security of the large model and the diversion effectiveness of the number of conversation strategies.

6. The method according to claim 1, wherein The number of large models includes at least a first large model and a second large model, and the conversation risk applicable to the second large model is higher than the conversation risk applicable to the first large model. The training steps of the second large model include: Based on the security principle and the sample statement to be responded to with reply risk, construct a sample prompt instruction; wherein, the sample prompt instruction is used to instruct the second large model to make a secure reply to the sample statement to be responded to in accordance with the security principle. Obtain the output content of the second largest model in response to the sample prompt instruction as the predicted reply content of the sample to-be-responded statement; Adjust the network parameters of the second largest model at least based on the difference between the predicted reply content of the sample to-be-responded statement and the sample reply content expected for the second largest model to reply safely.

7. The method according to claim 6, wherein The method further includes: Generate a reference to-be-responded statement with reply risk, and obtain the reference reply content predicted by the second largest model for the reference to-be-responded statement; In response to the reference reply content being unable to reply to the reference to-be-responded statement safely, select the reference to-be-responded statement as a new sample to-be-responded statement during the training of the second largest model.

8. The method according to claim 1, wherein The processing the to-be-responded statement based on the target policy to obtain the reply content of the to-be-responded statement includes: Obtain the context features of the large model as the target policy during the process of processing the to-be-responded statement, and obtain the knowledge features of the knowledge fragment related to the to-be-responded statement; Execute multi-head cross-attention based on the context features and the knowledge features to obtain the output features of the multi-head cross-attention; Fuse the context features and the output features to obtain the fused features; Continue to process the fused features based on the large model to obtain the reply content of the to-be-responded statement.

9. The method according to claim 8, wherein The fusing the context features and the output features to obtain the fused features includes: Perform prediction based on the context features and the output features to obtain the fusion weights of the context features and the output features; Weight the context features and the output features based on the fusion weights to obtain the fused features.

10. The method according to any one of claims 1 to 9, characterized in that Among the several dialogue strategies, there is also a fallback reply, which is used to output either a warning reply or a rejection reply as the reply content when the risk level characterized by the session risk is an extremely high risk; And / or, when there is a risk in the predicted characters output by the large model, the probability distribution of the predicted characters is constrained, and the final output content of the large model is determined based on the probability distribution of the predicted characters; And / or, the multiple large models share at least part of the underlying representation or learn from each other through knowledge distillation.

11. A secure response generation device, characterized in that, including: A risk identification module, configured to identify based on the to-be-responded statement to obtain the session risk of replying to the to-be-responded statement; A policy matching module, configured to select a dialogue strategy that matches the session risk from several dialogue strategies as the target policy; wherein, among the several dialogue strategies, there are at least multiple large models that are respectively different dialogue strategies, and different large models are respectively applicable to reply safely under different session risks; A statement reply module, configured to process the to-be-responded statement based on the target policy to obtain the reply content of the to-be-responded statement.

12. An electronic device, characterized in that, It includes at least a memory and a processor. At least program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the safe reply generation method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, Stored with program instructions that can be run by a processor, the program instructions being used to implement the secure response generation method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Information processing method and device in user dialogue, electronic equipment and storage medium

    CN112000781A

  • Dialogue processing method and device and dialogue model training method and device

    CN117556007A

  • Question and answer method and device, electronic equipment and storage medium

    CN118227760A

  • Prompt word generation method, prompt word-based dialogue method and related device

    CN119513237A

  • Automatic information collection method and device, electronic equipment and storage medium

    CN119514560A