Interactive generative model training method, generative dialogue realization method and device

The iterative optimization of generative dialogue systems aligns responses with human safety values by updating safety norms and refining models, addressing output safety issues and enhancing trustworthiness.

JP7725772B2Active Publication Date: 2025-08-20BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024098811
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-06-30
Filing Date
2024-06-19
Publication Date
2025-08-20
Estimated Expiration
2044-06-19

AI Technical Summary

Technical Problem

Generative dialogue systems face challenges in ensuring output safety due to noise, bias, and errors in training data, leading to inappropriate or harmful responses that can damage user trust and pose legal or moral liabilities.

Method used

A two-stage iterative optimization method is employed, involving alternating iterations of safety norm updates and interactive generative model optimization, using a detection model to refine responses to align with human safety values, and incorporating evaluation criteria for different content areas and application scenarios.

Benefits of technology

The method continuously improves the output safety of generative models by ensuring responses conform to human safety standards, reducing the risk of inappropriate outputs and enhancing trustworthiness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007725772000001
    Figure 0007725772000001
  • Figure 0007725772000002
    Figure 0007725772000002
  • Figure 0007725772000003
    Figure 0007725772000003
Patent Text Reader

Abstract

To provide an interactive generation model training method, and a generative interaction implementation method and device that relate to the field of artificial intelligence such as deep learning, natural language processing, and intelligent interaction.SOLUTION: The interactive generation model training method can include the steps of: using a safety code after updating as a target safety code to determine interaction input corresponding to the optimization this time based on the target safety code in response to determining that the safety code is updated, wherein the updating is updating performed on the original safety code when it is determined that the most recently optimized interactive generation model does not meet an on-line request; and optimizing the interactive generation model according to a principle where a responce generated by the interactive generation model meets the target safety code on the basis of the interactive input, wherein the interactive generation model is used for generating a response corresponding to the interactive input.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to interactive generative model training and generative dialogue realization methods and devices in fields such as deep learning, natural language processing, and intelligent dialogue. [Background technology]

[0002] A generative dialogue system uses deep learning technology to directly generate responses based on dialogue input. With the development of artificial intelligence technology, generative dialogue systems have been widely applied in various scenarios as a new natural language processing task. However, in practical applications, generative dialogue systems also face several challenges and risks, such as output safety issues. Summary of the Invention [Problem to be solved by the invention]

[0003] The present disclosure provides an interactive generative model training, generative dialogue realization method and apparatus. [Means for solving the problem]

[0004] The interactive generative model training method is In response to determining that an update has been made to the safety norm, the updated safety norm is set as a target safety norm, and an interactive input corresponding to the current optimization is determined based on the target safety norm, wherein the update is an update made to the original safety norm when it is determined that the interactive generative model after the most recent optimization does not meet the online requirements; and optimizing the interactive generative model based on the interactive input so that responses generated by the interactive generative model conform to the principles of the target safety norm, wherein the interactive generative model is used to generate responses corresponding to the interactive input.

[0005] The generative dialogue realization method is as follows: obtaining a dialogue input to be processed; The method includes a step of using an interactive generative model to generate a response corresponding to the pending interactive input, wherein the interactive generative model is an interactive generative model that matches online requirements obtained by N iterative optimizations, where N is a positive integer greater than 1, and each optimization includes, in response to determining that an update has been made to a safety norm, optimizing the interactive generative model based on the determined interactive input according to a principle that the response generated by the interactive generative model matches a target safety norm, wherein the target safety norm is the updated safety norm, the determined interactive input is an interactive input corresponding to the current optimization determined based on the target safety norm, and the update is an update made to the original safety norm when it is determined that the interactive generative model after the most recent optimization does not match the online requirements.

[0006] The interactive generative model training apparatus includes a preprocessing module and a model optimization module; the pre-processing module, in response to determining that an update has been made to the safety norm, sets the updated safety norm as a target safety norm and determines an interactive input corresponding to the current optimization based on the target safety norm, the update being an update made to the original safety norm when it is determined that the interactive generative model after the most recent optimization does not meet online requirements; The model optimization module is used to optimize the interactive generative model based on the interactive input so that the responses generated by the interactive generative model conform to the principles of the target safety norm, and the interactive generative model is used to generate responses corresponding to the interactive input.

[0007] The generative dialogue realization device includes an input acquisition module and a response generation module; the input acquisition module is used to acquire a dialogue input to be processed; The response generation module is used to generate a response corresponding to the pending dialogue input using an dialogue model, wherein the dialogue model is an dialogue model that matches the online requirements obtained by N iterative optimizations, where N is a positive integer greater than 1, and each optimization includes optimizing the dialogue model based on the determined dialogue input in response to determining that an update has been made to the safety norm, according to the principle that the response generated by the dialogue model matches the target safety norm, wherein the target safety norm is the safety norm after the update, the determined dialogue input is the dialogue input corresponding to the current optimization determined based on the target safety norm, and the update is an update made to the original safety norm when it is determined that the dialogue model after the most recent optimization does not match the online requirements.

[0008] Electronic devices include: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform the method.

[0009] A non-transitory computer-readable storage medium having computer instructions stored thereon that cause a computer to perform the method.

[0010] The computer program product comprises computer programs / instructions which, when executed by a processor, implement the method.

[0011] It should be understood that the content described herein is not intended to identify key or important features of the embodiments of the present disclosure, nor should it be used to limit the scope of the present disclosure. Other features of the present disclosure can be readily understood throughout the following specification. [Brief explanation of the drawings]

[0012] The drawings are for better understanding of the present application, but do not limit the present application. [Figure 1] 1 is a flowchart of an embodiment of the interactive generative model training method of the present disclosure. [Figure 2] 1 is a schematic diagram of a conventional interactive generative model training and working method. [Figure 3] FIG. 1 is a schematic diagram of an iterative optimization scheme for safety disciplines and safety systems of the present disclosure. [Figure 4] FIG. 2 is a schematic diagram illustrating the relationship between the baseline model and the target model of the present disclosure. [Figure 5] FIG. 1 is a schematic diagram of the overall optimization method for the safety system of the present disclosure. [Figure 6] 1 is a flowchart of an embodiment of the method for realizing generative dialogue of the present disclosure. [Figure 7] FIG. 7 is a structural schematic diagram of the configuration of an embodiment 700 of the interactive generative model training apparatus of the present disclosure. [Figure 8] 8 is a structural schematic diagram of the configuration of the generation-type dialogue realization device 800 of the present disclosure. [Figure 9] 9 shows a schematic block diagram of an electronic device 900 in which embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, examples of the present application will be described based on the drawings. For ease of understanding, various details of the examples of the present application are included, and they should be considered as mere examples. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of brevity, the following description will omit descriptions of well-known functions and structures.

[0014] Furthermore, the term "and / or" in this specification only describes the relation between related objects and indicates that three types of relations can exist, for example, A and / or B can represent three cases: only A exists, A and B exist simultaneously, or only B exists. It should be noted that the symbol " / " generally indicates that the related objects before and after it are in an "or" relation.

[0015] 1 is a flowchart of an embodiment of the interactive generative model training method of the present disclosure. As shown in FIG. 1, the method includes the following specific implementations:

[0016] In step 101, in response to determining that an update has occurred to the safety norm, the updated safety norm is set as a target safety norm, and an interactive input (query) corresponding to the current optimization is determined based on the target safety norm, and the update is an update made to the original safety norm when it is determined that the interactive generative model after the closest optimization does not meet the online requirements.

[0017] In step 102, based on the dialogue input, the dialogue generative model is optimized according to the principle that the response generated by the dialogue generative model conforms to the target safety norm, and the dialogue generative model is used to generate a response corresponding to the dialogue input.

[0018] Generative dialogue systems generate responses based on deep learning models, i.e., interactive generative models. However, interactive generative models are typically obtained after learning linguistic rules and knowledge from a large number of training samples (corpora). As shown in Figure 2, Figure 2 is a schematic diagram of the training and operation method of a conventional interactive generative model. Accordingly, interactive generative models may be affected by noise, bias, errors, etc. present in the corpus, which can lead to inappropriate or harmful responses. These responses can damage users' emotions, trust, or interests, and may even result in legal or moral liability. Therefore, how to improve the output safety of interactive generative models is an urgent issue that needs to be addressed.

[0019] In practical applications, interactive generative models face content limitations and cannot accept a single, formatted dialogue input. They must simultaneously face a variety of possible task forms, such as casual conversation, information questions and answers, text translation, text creation, and code creation. This determines the possibility that the model's input and output may be any text form, and its safety solution must also consider various possibilities simultaneously, making it difficult to achieve the ideal state in one go.

[0020] Accordingly, in the above-mentioned method embodiment, a gradual iterative optimization method for the interactive generative model is provided, and through alternating iteration and continuous optimization of the two parts of the safety norm and the interactive generative model, the output safety of the interactive generative model can be continuously improved, and ultimately the responses generated by the interactive generative model can be aligned with human safety values, etc.

[0021] The goal of safety norms is to define standards that conform to human safety values, and the most core function of such norms is to answer what responses are safe.

[0022] Preferably, the safety norms may include evaluation norms of at least one evaluation dimension, each corresponding to a different combination, and any combination is each composed of one content area and one application scene, wherein the content area is a safety content area related to generative interaction, and the application scene is an application scene of generative interaction.

[0023] For example, content areas can include politics, law, pornography, morality and values, etc., and application scenarios can include chatting, information questions and answers, text translation, text creation, code creation, etc. In the same content area, safety standards differ for different application scenarios, so it is necessary to combine both to subdivide safety norms.

[0024] One content area and one application scene can constitute one combination. Accordingly, assuming that there are three content areas (the numbers are for illustrative purposes only), namely content area 1, content area 2, and content area 3, and three application scenes, namely application scene 1, application scene 2, and application scene 3, the following combinations can be obtained: content area 1 + application scene 1, content area 1 + application scene 2, content area 1 + application scene 3, content area 2 + application scene 1, content area 2 + application scene 2, content area 2 + application scene 3, content area 3 + application scene 1, content area 3 + application scene 2, and content area 3 + application scene 3.

[0025] Each combination can correspond to one or more evaluation criteria of evaluation dimensions, for example, to one evaluation dimension, namely, whether it is safe, or to three evaluation dimensions, namely, whether it is safe, whether the knowledge is accurate, and whether the content is rich. No matter how many evaluation dimensions are used, the evaluation dimension of whether it is safe is usually always present, i.e., the most basic, and all other evaluation dimensions are further optimized on this basis.

[0026] For each combination, different evaluation dimensions can correspond to different evaluation criteria. For example, for the combination of content area 1 + application scene 1, assume that there are two evaluation dimensions: whether it is safe and whether the knowledge is accurate. In this case, these two evaluation dimensions correspond to different evaluation criteria, namely, the evaluation criteria that need to be met when the evaluation is safe, and the evaluation criteria that need to be met when the evaluation is knowledge is accurate.

[0027] As can be seen from the above, the above processing method can establish corresponding evaluation criteria for different content domains, different application scenarios, and different evaluation dimensions, thereby improving the accuracy of subsequent processing results.

[0028] The safety norms and the safety system are alternately iterated and continuously optimized, wherein the safety system may include an interactive generation model and a detection model, and may be optimized under the guidance of the safety norms.

[0029] 3 is a schematic diagram of the iterative optimization method of the safety norm and the safety system disclosed herein. As shown in FIG. 3, in the initial stage, an expert can determine a version of the safety norm based on experience, etc., and then optimize the safety system based on the safety norm. After the safety system is optimized, the expert can evaluate whether the interactive generation model therein meets the online requirements. If not, the safety norm can be updated based on the safety defects (exposed safety issues) existing in the safety system evaluated by the expert. For example, the expert can conduct a simulation attack on the safety system to determine the existing safety defects. After the safety norm is updated, the safety system can be optimized again based on the updated safety norm, and the above process can be continuously repeated.

[0030] Preferably, the occurrence of an update to the safety norm can include one or any combination of adding a combination and an evaluation norm of at least one corresponding evaluation dimension, adding an evaluation dimension and a corresponding evaluation norm for the original combination, and adjusting the original evaluation norm. That is, it is possible to add a combination and an evaluation norm of at least one corresponding evaluation dimension to the safety norm, and it is also possible to add one or more evaluation dimensions and corresponding evaluation norms for some or some original combinations, or it is also possible to adjust the original evaluation norm (e.g., subdivision), etc., which is very flexible and convenient.

[0031] For ease of explanation, the updated safety norm is called the target safety norm, and the interactive input corresponding to the current optimization (i.e., optimization performed on the safety system) can be determined based on the target safety norm.

[0032] Preferably, a first set of dialogue inputs can be obtained, and the dialogue inputs therein are the dialogue inputs corresponding to the current optimization, the first set of dialogue inputs includes at least dialogue inputs corresponding to combinations for which updates have been generated, the first set of dialogue inputs meets a predetermined condition that the proportion of the number of dialogue inputs of the first type is greater than the proportion of the number of dialogue inputs of the second type, the first type of dialogue inputs are dialogue inputs corresponding to combinations for which updates have been generated, and the second type of dialogue inputs are dialogue inputs corresponding to combinations for which updates have not been generated.

[0033] For example, suppose the safety norms include nine combinations, namely combination 1 to combination 9, and combination 1 and combination 2 are updated. In this case, the first dialogue input set may include more dialogue inputs corresponding to combination 1 and combination 2, a relatively small number of dialogue inputs corresponding to other combinations, or no dialogue inputs corresponding to other combinations. For example, suppose combination 1 is a legal + information question and answer, in which case the corresponding dialogue input may be legal problem information.

[0034] Through the above process, it is possible to realize focused optimization of the content updated in accordance with the safety standard, thereby improving the optimization effect and the optimization efficiency.

[0035] The dialogue inputs in the first dialogue input set can be selected from user utterances of dialogue product services that have already been publicly deployed, provided by experts based on safety standards, automatically generated by a model, etc., and the specific manner is not limited.

[0036] Based on the interaction inputs in the first interaction input set, the safety system can be optimized according to principles whereby responses generated by the interactive generative model conform to target safety norms.

[0037] Preferably, some or all of the dialogue inputs are selected from the first dialogue input set to form a second dialogue input set, the second dialogue input set meeting the predetermined conditions, an dialogue generative model is used to generate responses corresponding to each dialogue input in the second dialogue input set to form the first response set, the dialogue generative model and the detection model are optimized based on the first response set and the target safety norm, some or all of the dialogue inputs are selected from the first dialogue input set to form a third dialogue input set, the third dialogue input set meeting the predetermined conditions, the optimized dialogue generative model is used to generate responses corresponding to each dialogue input in the third dialogue input set to form the second response set, the optimized dialogue generative model can be re-optimized based on the second response set and the optimized detection model, and the detection model performs safety detection on the generated responses.

[0038] The safety system includes an interactive generative model and a detection model, and whether optimizing the interactive generative model or optimizing the detection model, the ultimate goal is to improve the output safety of the interactive generative model.

[0039] The interactive generative model may be a pre-trained and obtained model, for example, a large-scale language model based on Transformer, which is trained and obtained based on a large number of training samples and contains a wealth of knowledge, but the generated responses may pose safety risks and must be in line with human safety values, after which they can be actually deployed online, and the pre-trained interactive generative model can be optimized accordingly using the method described in the present disclosure.

[0040] The detection model can be used to perform safety detection on the responses generated by the interactive generative model to determine whether a safety risk or the like exists.

[0041] As can be seen from the above, the above optimization method is a two-stage optimization method. In the first stage, the interactive generation model and the detection model are optimized, and the optimized interactive generation model and the optimized detection model are obtained. In the second stage, the optimized detection model is used to re-optimize the optimized interactive generation model. That is, in one iterative optimization process, two optimizations of the interactive generation model can be achieved, and the two optimizations adopt different realization methods to further improve the optimization effect, etc.

[0042] The following provides a detailed description of the specific implementation of the first and second stages, respectively.

[0043] 1) First stage A second dialogue input set can be constructed by selecting some or all of the dialogue inputs from the first dialogue input set, and the second dialogue input set must meet the predetermined condition, i.e., the proportion of the number of dialogue inputs of the first type therein is greater than the proportion of the number of dialogue inputs of the second type therein. Since the latter involves manual labeling, in order to reduce the amount of work, the second dialogue input set usually includes only some of the dialogue inputs in the first dialogue input set, and the specific number is not limited.

[0044] Then, the interactive generation model can be used to generate a reply corresponding to each dialogue input in the second dialogue input set, thereby constituting a first reply set; preferably, the first reply set can include M replies generated for each dialogue input in the second dialogue input set, where M is a positive integer greater than 1, and the specific value can be determined according to actual needs; that is, multiple replies can be generated for each dialogue input in the second dialogue input set.

[0045] Furthermore, the interactive generative model and the detection model can be optimized based on the first answer set and the target safety criterion. Preferably, the following processing can be performed for any dialogue input in the second dialogue input set, in which the dialogue input is a pending dialogue input, and each candidate response corresponding to the pending dialogue input and a manual labeling result for each candidate response are obtained, where the number of candidate responses is M or more, the candidate responses include responses generated for the pending dialogue input and / or responses manually modified from the responses generated for the pending dialogue input, and the manual labeling result for any candidate response includes a labeling result after manually performing safety labeling on the candidate response based on the target safety criterion; constructing training samples based on the pending dialogue input, each candidate response, and the manual labeling result for each candidate response, and optimizing the interactive generative model and the detection model using the training samples.

[0046] Preferably, for any candidate response, the labeling result after the safety labeling is performed can include evaluation labels corresponding to different evaluation dimensions of the labeled candidate response based on evaluation criteria of different evaluation dimensions of combinations corresponding to the dialogue inputs waiting to be manually processed, and the evaluation labels either match the corresponding evaluation criteria (Yes) or do not match the corresponding evaluation criteria (No).

[0047] Each of the dialogue inputs in the second dialogue input set can be treated as a dialogue input to be processed and processed in the same manner. Specifically, suppose six replies are generated for a dialogue input to be processed, each of which is designated as Response 1 to Response 6. Then, based on these six replies, multiple candidate replies can be generated. The specific number is not limited and may be, for example, six or more. Typically, the candidate replies should include as many safety status replies as possible. For example, the evaluation labels of each evaluation dimension are all replies that conform to the corresponding evaluation criteria, the partial evaluation labels conform to the corresponding evaluation criteria, and the remaining evaluation labels are replies that do not conform to the corresponding evaluation criteria, and the evaluation labels of each evaluation dimension are all replies that do not conform to the corresponding evaluation criteria. Furthermore, some or all of the six replies can be directly used as candidate replies, or certain modifications can be made to some or all of the six replies to obtain the required safety status replies.

[0048] The above process can be called a safety data labeling process, and the purpose of safety data labeling is to provide data support for the interactive generative model and the detection model so as to optimize the interactive generative model and the detection model.

[0049] Preferably, a first type of training sample and a second type of training sample can be constructed respectively, the first type of training sample can be used to adopt a supervised learning method to optimize the interactive generative model, and the second type of training sample can be used to adopt a supervised learning method to optimize the detection model.

[0050] That is, appropriate training sample construction methods are adopted for the interactive generation model and the detection model, and model optimization is performed accordingly, thereby improving the model optimization effect.

[0051] Preferably, the method for constructing the first type of training sample includes selecting a candidate response that meets the following condition from each candidate response, where the condition is that the evaluation labels of different evaluation dimensions all meet the corresponding evaluation criteria, and configuring each selected candidate response and the waiting dialogue input as one first type of training sample.

[0052] For example, suppose there are 12 candidate responses, and two of them meet the condition that "the evaluation labels of different evaluation dimensions all conform to the corresponding evaluation criteria." In this case, these two candidate responses can be respectively configured as a dialogue input to be processed and a training sample, and two training samples can be obtained.

[0053] For each dialogue input in the second dialogue input set, a training sample, i.e., a first-type training sample, can be generated using the above-mentioned method, and the generated first-type training sample can be used to optimize the dialogue generation model. Specifically, the dialogue generation model scores the candidate responses in the first-type training sample, calculates the negative log-likelihood loss through the scoring, and minimizes the loss using a gradient descent method, so that the dialogue generation model tends to generate corresponding candidate responses in the first-type training sample for the dialogue input in the first-type training sample, thereby achieving the purpose of model optimization.

[0054] Furthermore, for each pending dialogue input, an overall score for each candidate response can be obtained, where the higher the overall score, the higher the security. The detection model can include classification detection models corresponding to different evaluation dimensions from the overall detection model. The second type of training samples can include a first subtype of training samples and a second subtype of training samples. The first subtype of training samples can include two candidate responses with different overall scores, a pending dialogue input, and a sample label, where the sample label is used to indicate the candidate response with the higher overall score. The second subtype of training samples can include one candidate response, a pending dialogue input, and one evaluation label for the candidate response. Accordingly, optimizing the detection model can include optimizing the overall detection model using the first subtype of training samples, and optimizing, for any classification detection model, the classification detection model using a second subtype of training samples including evaluation labels for the evaluation dimensions corresponding to the classification detection model. The overall detection model and each classification detection model can all be Transformer-based models.

[0055] For example, assuming there are 12 candidate responses, each corresponding to three rating labels, the overall score of each candidate response can be determined based on these three rating labels. A possible implementation method is to set different weights for different rating labels. For example, the weight corresponding to the rating dimension of whether it is safe is the highest, and the weights of the other rating dimensions are all lower than this rating dimension. If the rating label is Yes (matches the corresponding rating criteria), then the value can be 1; otherwise, the value can be 0. Accordingly, for any candidate response, the overall score of the candidate response can be calculated based on the values of the three rating labels and their corresponding weights.

[0056] In order to better optimize the interactive generative model and make it more likely to generate a safe response, the number of detection models may be multiple, i.e., it may include one overall detection model and classification detection models corresponding to different evaluation dimensions, thereby providing judgment signals from aspects such as overall safety and different evaluation dimensions, which are used to optimize the interactive generative model and accordingly improve the optimization effect, etc.

[0057] Furthermore, as can be seen from the above, the above processing method employs appropriate optimization methods for the comprehensive detection model and the classification detection model, respectively, to further improve the optimization effect.

[0058] Wherein, for the overall detection model, a first subtype of training sample can be constructed, which can include two candidate answers with different overall scores, a dialogue input to be processed, and a sample label, and the sample label is used to indicate the candidate answer with the higher overall score among the two candidate answers.

[0059] For example, assuming there are 12 candidate replies, multiple training samples of a first subtype can be constructed by combining two and two, and each training sample of a first subtype can include two candidate replies with different overall scores, a dialogue input to be processed, and a sample label to indicate which of the two candidate replies has the higher overall score.

[0060] For each dialogue input in the second dialogue input set, a training sample of the first subtype can be constructed using the above-mentioned method, and the constructed training sample of the first subtype can be used to optimize the overall detection model, which can similarly employ negative log-likelihood loss and gradient descent methods, thereby allowing the overall detection model to learn how to distinguish between the superior and inferior quality of different candidate responses.

[0061] Each training sample of the second subtype may include one candidate response, a pending dialogue input, and one rating label for the candidate response.

[0062] Also, for each dialogue input in the second dialogue input set, a training sample of the second subtype can be constructed in the above-mentioned manner.

[0063] Accordingly, the second subtype of training samples can be used to optimize classification detection models corresponding to different evaluation dimensions. For example, assuming that evaluation dimension 1, evaluation dimension 2, and evaluation dimension 3 exist, the second subtype of training samples including evaluation labels corresponding to evaluation dimension 1 can be used to optimize the classification detection model corresponding to evaluation dimension 1, the second subtype of training samples including evaluation labels corresponding to evaluation dimension 2 can be used to optimize the classification detection model corresponding to evaluation dimension 2, and the second subtype of training samples including evaluation labels corresponding to evaluation dimension 3 can be used to optimize the classification detection model corresponding to evaluation dimension 3.

[0064] 2) Second stage After completing the first stage of optimization, the second stage of optimization can be performed, that is, the optimized detection model can be used to re-optimize the optimized interactive generative model.

[0065] First, a third dialogue input set can be constructed by selecting some or all of the dialogue inputs from the first dialogue input set, and the third dialogue input set must meet the predetermined condition, i.e., the proportion of the number of the first type of dialogue inputs therein is greater than the proportion of the number of the second type of dialogue inputs therein. In order to improve the optimization effect, the third dialogue input set can include all of the dialogue inputs in the first dialogue input set.

[0066] The optimized generative interactive model can then be used to generate a response corresponding to each of the dialogue inputs in the third dialogue input set, thereby constituting a second set of answers. Preferably, the second set of answers includes one response generated for each of the dialogue inputs in the third dialogue input set.

[0067] Furthermore, the optimized interactive generative model can be re-optimized based on the second answer set and the optimized detection model. Preferably, the optimized detection models can be used to perform safety detection for each answer in the second answer set, and the optimized interactive generative model can be re-optimized using a reinforcement learning method based on the safety detection results for each answer. The reinforcement learning algorithm used can be a proximal policy optimization (PPO) algorithm, etc.

[0068] That is, the detection model can be used as a referee to re-optimize the optimized interactive generation model, thereby further improving the optimization effect of the interactive generation model.

[0069] Preferably, the detection model may include classification detection models corresponding to different evaluation dimensions from the overall detection model. Accordingly, for any answer in the second answer set, the following processing may be performed: obtain the classification detection results corresponding to the returned overall detection result and the different classification detection models, combine the overall detection result and the different classification detection result to determine a reward corresponding to the answer; and use the answer, the dialogue input corresponding to the answer, and the reward to construct a training sample. Similarly, for each answer in the second answer set, a training sample may be constructed in this manner; and further, the constructed training sample may be used to re-optimize the optimized interactive generative model.

[0070] For example, for a given response a, an overall detection result (which may be in the form of a score) and classification detection results corresponding to different evaluation dimensions can be obtained, and then a predetermined fusion algorithm can be used to fuse the overall detection result and the different classification detection results, thereby determining a reward corresponding to response a. The specific form of the fusion algorithm can be determined according to actual needs.

[0071] Accordingly, a training sample can be constructed using response a, an interactive input corresponding to response a, and a reward corresponding to response a, and similarly, training samples corresponding to other responses can be obtained, and each training sample can be used to re-optimize the optimized interactive generative model.

[0072] Preferably, the optimized interactive generative model can be used as a baseline model, and a target model identical to the baseline model can be generated. Furthermore, the target model can be optimized using training samples based on the Kullback-Leibler (KL) variance constraint introduced between the baseline model and the target model, and the optimized target model can be used as the optimized interactive generative model again.

[0073] That is, two interactive generative models, namely, an optimized interactive generative model (i.e., a baseline model) and a target model, can be maintained through maintenance. Figure 4 is a schematic diagram of the relationship between the baseline model and the target model in the present disclosure. As shown in Figure 4, the baseline model can be considered to be maintained unchanged in the re-optimization process. Before optimizing the target model, the target model is exactly the same as the baseline model. When optimizing the target model using training samples, the optimization process is relatively difficult and it is easy to cause extreme situations such as the model generating garbled characters. Therefore, to prevent the target model from being too far from the baseline model, a KL variance between a baseline model and the target model can be additionally introduced to constrain it, and the optimized target model can be the required re-optimized interactive generative model.

[0074] In conjunction with the above introduction, Figure 5 is a schematic diagram of the overall optimization method of the safety system disclosed herein. As shown in Figure 5, Optimization 1 therein shows the process of optimizing the interactive generative model and the detection model, and Optimization 2 shows the process of using the optimized detection model to re-optimize the optimized interactive generative model. In Optimization 1, the output response of the interactive generative model can pass through the detection model, or it can be directly manually labeled without passing through the detection model.

[0075] After optimizing each safety system using the method shown in Figure 5, an expert can re-evaluate whether the latest acquired interactive generative model meets the online requirements. If not, the safety norms can be updated, and the safety system can be re-optimized based on the updated safety norms. If so, the latest acquired interactive generative model can be actually deployed online. After actually deploying it online, if necessary, the interactive generative model can be further continuously optimized using the method described in the present disclosure.

[0076] Accordingly, Figure 6 is a flowchart of an embodiment of the method for realizing the generative dialogue of the present disclosure. As shown in Figure 6, the method includes the following specific implementation methods:

[0077] In step 601, a dialogue input waiting to be processed is obtained.

[0078] In step 602, an interactive generative model is used to generate a response corresponding to the pending interactive input, wherein the interactive generative model is an interactive generative model that meets the online requirements obtained by N iterative optimizations, where N is a positive integer greater than 1, and each optimization includes optimizing the interactive generative model based on the determined interactive input in response to determining that an update has been made to the safety norm, according to the principle that the response generated by the interactive generative model meets the target safety norm, wherein the target safety norm is the updated safety norm, the determined interactive input is the interactive input corresponding to the current optimization determined based on the target safety norm, and the update is an update made to the original safety norm when it is determined that the interactive generative model after the most recent optimization does not meet the online requirements.

[0079] As can be seen from the above, by adopting the method described in the embodiment of the above method, by alternating the two parts of the safety norm and the interactive generative model, continuous optimization can be performed and the output safety of the interactive generative model can be constantly improved. Accordingly, the trained interactive generative model can be used to generate responses, and the safety of the generated responses can be improved.

[0080] The interactive generative model may be an interactive generative model that meets online requirements, ie, obtained by a method corresponding to the embodiment shown in FIG.

[0081] For the sake of simplicity, the embodiments of the above-mentioned methods are described as a combination of a series of operations. However, those skilled in the art should recognize that the present disclosure is not limited by the order of operations described, since some steps may use a different order or may be performed simultaneously according to the present disclosure. Next, all of the embodiments described herein belong to preferred embodiments, and the associated operations and modules are not necessarily essential to the present disclosure. Furthermore, if some embodiments lack details, reference may be made to the relevant descriptions in other embodiments.

[0082] The above is a description of a method embodiment, and the following will further illustrate the solution described in this disclosure through an apparatus embodiment.

[0083] 7 is a structural schematic diagram of the configuration of an embodiment 700 of the interactive generative model training apparatus of the present disclosure. As shown in FIG. 7, it includes a pre-processing module 701 and a model optimization module 702.

[0084] In response to determining that an update has been made to the safety norm, the pre-processing module 701 sets the updated safety norm as a target safety norm and is used to determine the interactive input corresponding to the current optimization based on the target safety norm, where the update is an update made to the original safety norm when it is determined that the interactive generative model after the most recent optimization does not meet the online requirements.

[0085] The model optimization module 702 is used to optimize the interactive generative model based on the interactive input according to the principle that the response generated by the interactive generative model conforms to the target safety norm, and the interactive generative model is used to generate a response corresponding to the interactive input.

[0086] The solution described in the above-mentioned device embodiment adopts a gradual iterative interactive generative model optimization method, which alternates between the two parts of safety norms and the interactive generative model to continuously optimize and constantly improve the output safety of the interactive generative model, so that the response generated by the interactive generative model is ultimately consistent with human safety values.

[0087] Preferably, the safety norms may include evaluation norms of at least one evaluation dimension corresponding to different combinations, and any combination is each composed of one content area and one application scene, the content area is a safety content area for generative interaction, and the application scene is an application scene for generative interaction. Accordingly, an update to the safety norms may include one or any combination of adding a new evaluation norm for the combination and at least one corresponding evaluation dimension, adding a new evaluation dimension and corresponding evaluation norm for the original combination, and adjusting the original evaluation norm.

[0088] Preferably, the pre-processing module 701 can obtain a first set of dialogue inputs, and the dialogue inputs therein are the dialogue inputs corresponding to the current optimization, the first set of dialogue inputs includes at least the dialogue inputs corresponding to the combinations for which an update has been generated, the first set of dialogue inputs meets a predetermined condition that the proportion of the number of the first type of dialogue inputs is greater than the proportion of the number of the second type of dialogue inputs, the first type of dialogue inputs are the dialogue inputs corresponding to the combinations for which an update has been generated, and the second type of dialogue inputs are the dialogue inputs corresponding to the combinations for which no update has been generated.

[0089] The dialogue inputs in the first dialogue input set can be selected from user utterances of already publicly deployed dialogue product services, provided by experts based on safety norms, automatically generated by a model, etc.

[0090] Preferably, the model optimization module 702 selects some or all of the dialogue inputs from the first dialogue input set to form a second dialogue input set, the second dialogue input set meeting the predetermined conditions, uses an dialogue generative model to generate responses corresponding to each dialogue input in the second dialogue input set to form a first response set, optimizes the dialogue generative model and the detection model based on the first response set and the target safety norm, selects some or all of the dialogue inputs from the first dialogue input set to form a third dialogue input set, the third dialogue input set meeting the predetermined conditions, uses the optimized dialogue generative model to generate responses corresponding to each dialogue input in the third dialogue input set to form a second response set, and can re-optimize the optimized dialogue generative model based on the second response set and the optimized detection model, and the detection model performs safety detection on the generated responses.

[0091] That is, a two-stage optimization method can be adopted: in the first stage, the interactive generation model and the detection model are optimized, and the optimized interactive generation model and the optimized detection model are obtained; in the second stage, the optimized detection model is borrowed to re-optimize the optimized interactive generation model.

[0092] Preferably, the first response set may include M responses generated for each dialogue input in the second dialogue input set, where M is a positive integer greater than 1. The model optimization module 702 may perform the following process for any dialogue input in the second dialogue input set: the dialogue input is treated as a dialogue input to be processed, and each candidate response corresponding to the dialogue input to be processed and a manual labeling result for each candidate response are obtained, where the number of candidate responses is M or more, the candidate responses include responses generated for the dialogue input to be processed and / or responses manually modified from the responses generated for the dialogue input to be processed, and the manual labeling result for any candidate response includes a labeling result after manually performing safety labeling on the candidate response based on a target safety norm; and construct training samples based on the dialogue input to be processed, each candidate response, and the manual labeling result for each candidate response, and optimize the interactive generation model and the detection model using the training samples.

[0093] Preferably, for any candidate response, the labeling result after the safety labeling is performed can include evaluation labels corresponding to different evaluation dimensions based on evaluation criteria of different evaluation dimensions of combinations corresponding to the dialogue inputs to be manually processed, and the evaluation labels can either conform to the corresponding evaluation criteria or not conform to the corresponding evaluation criteria.

[0094] Preferably, the model optimization module 702 can respectively construct a first type of training sample and a second type of training sample, use the first type of training sample to adopt a supervised learning method to optimize the interactive generative model, and use the second type of training sample to adopt a supervised learning method to optimize the detection model.

[0095] Preferably, the method in which the model optimization module 702 constructs the first type of training sample includes selecting a candidate response that meets the following condition from each candidate response, where the evaluation labels of different evaluation dimensions all meet the corresponding evaluation criteria, and configuring each selected candidate response and the dialogue input to be processed as a first type of training sample.

[0096] For each dialogue input in the second dialogue input set, a training sample, i.e., a first type of training sample, can be generated by the above-mentioned method, and the generated first type of training sample can be used to optimize the dialogue generative model.

[0097] Preferably, the model optimization module 702 can further obtain an overall score for each candidate response, where the higher the overall score, the higher the safety; the detection model can include classification detection models corresponding to different evaluation dimensions from the overall detection model; the second type of training samples can include training samples of a first subtype and training samples of a second subtype; the first subtype training samples can include two candidate responses with different overall scores, a dialogue input waiting to be processed, and a sample label, where the sample label is used to indicate the candidate response with the higher overall score among the two candidate responses; the second subtype training samples can include one candidate response, a dialogue input waiting to be processed, and one evaluation label for the candidate response; and accordingly, optimizing the detection model can include optimizing the overall detection model using training samples of the first subtype, and for any classification detection model, optimizing the classification detection model using training samples of a second subtype each including an evaluation label for the evaluation dimension corresponding to the classification detection model.

[0098] In addition, preferably, the second response set can include one response generated for each dialogue input in the third dialogue input set. The model optimization module 702 can use the optimized detection model to perform safety detection for each response in the second response set, and based on the safety detection result of each response, can adopt a reinforcement learning method to re-optimize the optimized dialogue generation model.

[0099] Preferably, the detection model may include a comprehensive detection model and classification detection models corresponding to different evaluation dimensions. Accordingly, the model optimization module 702 can perform the following processing for any response: obtain the comprehensive detection result and classification detection results corresponding to the returned comprehensive detection result and different classification detection models, combine the comprehensive detection result and the different classification detection results, and determine the reward corresponding to the response; use the response, the dialogue input corresponding to the response, and the reward to construct a training sample; similarly, for each response in the second response set, a training sample can be constructed in this manner; and further, the constructed training sample can be used to re-optimize the optimized interactive generative model.

[0100] Preferably, the model optimization module 702 can take the optimized interactive generative model as a baseline model, generate a target model that is exactly the same as the baseline model, and optimize the target model using training samples based on the KL variance constraint introduced between the baseline model and the target model, and then use the optimized target model as the optimized interactive generative model again.

[0101] 8 is a structural schematic diagram of the configuration of the generation-type dialogue realization device 800 of the present disclosure. As shown in FIG. 8, it includes an input acquisition module 801 and a response generation module 802.

[0102] The input acquisition module 801 is used to acquire pending dialogue input.

[0103] The response generation module 802 is used to generate a response corresponding to a pending dialogue input using an interactive generative model, where the interactive generative model is an interactive generative model that matches online requirements obtained by N iterative optimizations, where N is a positive integer greater than 1, and each optimization includes optimizing the interactive generative model based on the determined dialogue input in response to determining that an update has been made to the safety norm, according to the principle that the response generated by the interactive generative model matches a target safety norm, where the target safety norm is the updated safety norm, the determined dialogue input is the dialogue input corresponding to the current optimization determined based on the target safety norm, and the update is an update made to the original safety norm when it is determined that the interactive generative model after the most recent optimization does not match the online requirements.

[0104] As can be seen from the above, by adopting the solution described in the embodiment of the above device, the two parts of the safety norm and the interactive generative model can be alternately repeated to continuously optimize and constantly improve the output safety of the interactive generative model, and accordingly, the trained interactive generative model can be used to generate responses, thereby improving the safety of the generated responses.

[0105] For the specific workflow of the device embodiment shown in FIGS. 7 and 8, please refer to the relevant description of the method embodiment described above, and the description will be omitted.

[0106] In short, the solution described in this disclosure provides a multidimensional incremental iterative interactive generative model security solution, which can improve the output security of interactive generative models, and can be applied to various application scenes and content areas, etc., and has wide applicability.

[0107] The solutions described in this disclosure relate to the field of artificial intelligence technology, and can be particularly applied to fields such as large-scale models, deep learning, and natural language processing. Artificial intelligence is a discipline that studies the simulation of certain human thought processes and intelligent behaviors (e.g., learning, reasoning, thinking, planning, etc.) using computers, and includes both hardware-level and software-level technologies. Artificial intelligence hardware technology generally includes technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing, while artificial intelligence software technology mainly includes several directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0108] The dialogue inputs and responses in the above embodiments of the present disclosure are not directed to a specific user and do not reflect the personal information of a specific user. In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, etc. of relevant user personal information all comply with relevant laws and regulations and do not violate public order and morals.

[0109] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0110] 9 shows a schematic block diagram of an electronic device 900 for implementing embodiments of the present disclosure. The electronic device is intended to represent various types of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframes, and other suitable computers. The electronic device may also represent various types of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the description herein and / or the practice of the present disclosure as claimed.

[0111] 9, the device 900 includes a computing unit 901, which can perform various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 902 or loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 can also store various programs and data required for the device 900 to operate. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0112] Multiple components within device 900 are connected to I / O interface 905, including input units 906 such as a keyboard, mouse, etc., output units 907 such as various types of displays, speakers, etc., storage units 908 such as a disk, optical disk, etc., and communication units 909 such as a network card, modem, wireless communication transceiver, etc. The communication units 909 enable device 900 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.

[0113] The computing unit 901 is a general-purpose and / or special-purpose processing component equipped with various processing and computational capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, computing units that execute various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above, such as the methods described in this disclosure. For example, in some embodiments, the methods described in this disclosure can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, some or all of the computer program is loaded and / or installed into the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, it can perform one or more steps of the methods described in this disclosure above. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the methods described in this disclosure through any other suitable manner (e.g., via firmware).

[0114] Various implementations of the systems and techniques described herein may be realized in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), field programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include being implemented in one or more computer programs that can be executed and / or interpreted by a programmable system that includes at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, and that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0115] Program codes for implementing the methods of the present disclosure can be written using any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus such that, when executed by the processor or controller, the functions / acts specified in the flowcharts and / or block diagrams are performed. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a separate software package and partially on a remote machine, or entirely on a remote machine or server.

[0116] In the context of this disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use with or in connection with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium includes, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide input to the computer. Other types of devices can also be used to provide interaction with a user; for example, feedback provided to the user can be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback) and can receive input from the user in any form (including acoustic input, voice input, and tactile input).

[0118] The systems and techniques described herein can be implemented in a computing system including a back-end component (e.g., a data server), or a computing system including a middleware component (e.g., an application server), or a computing system including a front-end component (e.g., a user computer having a graphical user interface or a web browser through which a user interacts with an implementation of the systems and techniques described herein), or in a computing system including any combination of such back-end, middleware, and front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0119] The computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs running on corresponding computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server that combines blockchains.

[0120] It should be understood that steps can be rearranged, added, or deleted using the various types of flows shown above. For example, the steps described in the present disclosure may be performed in parallel, sequentially, or in a different order, but this specification is not limited thereto as long as the technical solution disclosed in the present disclosure can achieve the desired results.

[0121] The above specific implementation methods do not constitute limitations on the scope of protection of the present disclosure. Those skilled in the art may make various modifications, combinations, subcombinations, and substitutions based on design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present disclosure shall fall within the scope of protection of the present disclosure.

Claims

1. A computer-implemented method for interactively training a generative model, comprising: In response to determining that an update has been made to the safety norm, the updated safety norm is set as a target safety norm, and an interactive input corresponding to the current optimization is determined based on the target safety norm, wherein the update is an update made to the original safety norm when it is determined that the interactive generative model after the most recent optimization does not meet the online requirements; optimizing the interactive generative model based on the interactive input so that responses generated by the interactive generative model conform to the principles of the target safety norm, wherein the interactive generative model is used to generate responses corresponding to the interactive input; The safety norms include evaluation norms of at least one evaluation dimension corresponding to different combinations, and each combination is composed of one content area and one application scene, the content area is a safety content area related to generative interaction, and the application scene is an application scene of generative interaction; The update to the safety rule may include one or any combination of: adding a new evaluation rule for the combination and at least one corresponding evaluation dimension; adding a new evaluation dimension and a corresponding evaluation rule for the original combination; and adjusting the original evaluation rule. Interactive generative model training method.

2. The step of determining an interactive input corresponding to the current optimization based on the target safety norm includes: obtaining a first set of dialogue inputs and determining dialogue inputs therein as dialogue inputs corresponding to the current optimization; the first set of interaction inputs includes at least an interaction input corresponding to a combination for which an update has been generated; The first set of dialogue inputs meets a predetermined condition that a ratio of the number of dialogue inputs of a first type is greater than a ratio of the number of dialogue inputs of a second type, the first type of dialogue inputs being dialogue inputs corresponding to combinations for which an update has been generated, and the second type of dialogue inputs being dialogue inputs corresponding to combinations for which an update has not been generated. The interactive generative model training method of claim 1 .

3. The step of optimizing the interactive generative model includes: selecting some or all of the dialogue inputs from the first dialogue input set to form a second dialogue input set, the second dialogue input set meeting the predetermined condition; generating responses corresponding to each dialogue input in the second dialogue input set using the dialogue generative model to construct a first answer set, and optimizing the dialogue generative model and the detection model based on the first answer set and the target safety norm; selecting some or all of the dialogue inputs from the first dialogue input set to form a third dialogue input set, the third dialogue input set meeting the predetermined condition; generating responses corresponding to the respective dialogue inputs in the third dialogue input set using the optimized dialogue generative model to form a second answer set, and re-optimizing the optimized dialogue generative model based on the second answer set and the optimized detection model, wherein the detection model performs safety detection on the generated responses. The interactive generative model training method of claim 2 .

4. the first set of responses includes M responses generated for each dialogue input in the second set of dialogue inputs, where M is a positive integer greater than 1; optimizing the interactive generative model and the detection model based on the first answer set and the target safety paradigm, The method includes the steps of performing the following processes for each of the dialogue inputs in the second dialogue input set, wherein the dialogue inputs are set as dialogue inputs to be processed, and obtaining candidate responses and manual labeling results for each of the candidate responses corresponding to the dialogue inputs to be processed, the number of the candidate responses being M or more, the candidate responses including responses generated for the dialogue inputs to be processed and / or responses obtained by manually correcting the responses generated for the dialogue inputs to be processed, and the manual labeling results for each of the candidate responses including labeling results after manually performing safety labeling on the candidate responses based on the target safety norm, and constructing training samples based on the dialogue inputs to be processed, each candidate response, and the manual labeling results for each candidate response, and optimizing the dialogue generation model and the detection model using the training samples. The interactive generative model training method of claim 3 .

5. For any candidate response, the labeling result after the safety labeling includes evaluation labels corresponding to different evaluation dimensions of the candidate response, which are manually labeled based on evaluation criteria of different evaluation dimensions of the combination corresponding to the dialogue input to be processed, and the evaluation labels indicate whether the candidate response conforms to the corresponding evaluation criteria or does not conform to the corresponding evaluation criteria. The method of interactive generative model training according to claim 4.

6. constructing training samples based on the pending dialogue input, each candidate response, and a manual labeling result of each candidate response, and optimizing the interactive generative model and the detection model using the training samples, constructing a first type of training sample and a second type of training sample, respectively; optimizing the interactive generative model using the first type of training samples and employing a supervised learning approach; and optimizing the detection model using the second type of training samples and employing a supervised learning approach. The interactive generative model training method of claim 5.

7. The step of constructing the first type of training samples comprises: selecting, from each of the candidate responses, a candidate response that meets the following conditions: The condition is that the evaluation labels of different evaluation dimensions all match the corresponding evaluation criteria, and each selected candidate response and the waiting dialogue input are configured as one first type training sample. The method of interactive generative model training according to claim 6.

8. The method comprises: The method further includes the step of obtaining an overall score for each candidate response, the higher the overall score, the higher the security; The detection model includes a comprehensive detection model and classification detection models corresponding to different evaluation dimensions, respectively; the second type of training samples includes a first subtype of training samples and a second subtype of training samples, the first subtype of training samples includes two candidate replies with different overall scores, the waiting dialogue input, and a sample label, the sample label being used to indicate the candidate reply with the higher overall score of the two candidate replies, the second subtype of training samples includes one candidate reply, the waiting dialogue input, and one rating label of the candidate reply; the step of optimizing the detection model includes a step of optimizing the overall detection model using the training samples of the first subtype, and for any of the classification detection models, optimizing the classification detection model using training samples of the second subtype each including an evaluation label of an evaluation dimension corresponding to the classification detection model; The method of interactive generative model training according to claim 6.

9. the second set of responses includes one response generated for each dialogue input in the third set of dialogue inputs; re-optimizing the optimized interactive generative model based on the second answer set and the optimized detection model, performing safety detection for each answer in the second answer set using the optimized detection model, respectively; and re-optimizing the optimized interactive generative model by employing a reinforcement learning method based on the safety detection result for each answer. The interactive generative model training method according to any one of claims 3 to 7.

10. The detection model includes a comprehensive detection model and classification detection models corresponding to different evaluation dimensions, respectively; and re-optimizing the optimized interactive generation model by using a reinforcement learning method according to the safety detection result of each response. For each optional response, perform the following process: The process includes the steps of: obtaining the returned overall detection result and a classification detection result corresponding to each of different classification detection models; combining the overall detection result and the different classification detection result to determine a reward corresponding to the response; and using the response, the dialogue input corresponding to the response, and the reward to form a training sample; and re-optimizing the optimized interactive generative model using the training samples. The method of interactive generative model training of claim 9.

11. The step of re-optimizing the optimized interactive generative model using the training samples includes: a step of generating a target model that is identical to the optimized interactive generative model as a baseline model; optimizing the target model using the training samples based on a Kulbeck-Lebel variance constraint introduced between the baseline model and the target model, and re-optimizing the optimized target model as an interactive generative model; The method of interactive generative model training of claim 10.

12. A method for realizing a computer-generated dialogue, comprising: obtaining a dialogue input to be processed; a step of generating a response corresponding to the pending dialogue input using an dialogue model, wherein the dialogue model is an dialogue model that matches online requirements obtained by N iterative optimizations, where N is a positive integer greater than 1, and each optimization includes, in response to determining that an update has been made to a safety norm, optimizing the dialogue model based on the determined dialogue input according to a principle that the response generated by the dialogue model matches a target safety norm, the target safety norm is the updated safety norm, the determined dialogue input is a dialogue input that corresponds to the current optimization determined based on the target safety norm, and the update is an update made to the original safety norm when it is determined that the dialogue model after the most recent optimization does not match the online requirements; The safety norms include evaluation norms of at least one evaluation dimension corresponding to different combinations, and each combination is composed of one content area and one application scene, the content area is a safety content area related to generative interaction, and the application scene is an application scene of generative interaction; The update to the safety rule may include one or any combination of: adding a new evaluation rule for the combination and at least one corresponding evaluation dimension; adding a new evaluation dimension and a corresponding evaluation rule for the original combination; and adjusting the original evaluation rule. A method for realizing generative dialogue.

13. An interactive generative model training apparatus, comprising: It includes a preprocessing module and a model optimization module. the pre-processing module, in response to determining that an update has been made to the safety norm, sets the updated safety norm as a target safety norm and determines an interactive input corresponding to the current optimization based on the target safety norm, the update being an update made to the original safety norm when it is determined that the interactive generative model after the most recent optimization does not meet online requirements; the model optimization module is used to optimize the interactive generative model based on the interactive input so that a response generated by the interactive generative model conforms to the principles of the target safety norm, and the interactive generative model is used to generate a response corresponding to the interactive input; The safety norms include evaluation norms of at least one evaluation dimension corresponding to different combinations, and each combination is composed of one content area and one application scene, the content area is a safety content area related to generative interaction, and the application scene is an application scene of generative interaction; The update to the safety rule may include one or any combination of: adding a new evaluation rule for the combination and at least one corresponding evaluation dimension; adding a new evaluation dimension and a corresponding evaluation rule for the original combination; and adjusting the original evaluation rule. An interactive generative model training device.

14. The pre-processing module acquires a first dialogue input set and sets dialogue inputs therein as dialogue inputs corresponding to the current optimization, the first dialogue input set including at least dialogue inputs corresponding to combinations for which updates have been generated, the first dialogue input set meeting a predetermined condition that the ratio of the number of dialogue inputs of a first type is greater than the ratio of the number of dialogue inputs of a second type, the first type of dialogue inputs being dialogue inputs corresponding to combinations for which updates have been generated, and the second type of dialogue inputs being dialogue inputs corresponding to combinations for which updates have not been generated.

14. The apparatus for interactive generative model training according to claim 13.

15. the model optimization module selects some or all of the dialogue inputs from the first dialogue input set to form a second dialogue input set, the second dialogue input set meeting the predetermined condition, uses the dialogue generative model to generate responses corresponding to each dialogue input in the second dialogue input set to form a first response set, optimizes the dialogue generative model and the detection model based on the first response set and the target safety norm, selects some or all of the dialogue inputs from the first dialogue input set to form a third dialogue input set, the third dialogue input set meeting the predetermined condition, uses the optimized dialogue generative model to generate responses corresponding to each dialogue input in the third dialogue input set to form a second response set, and re-optimizes the optimized dialogue generative model based on the second response set and the optimized detection model, and the detection model performs safety detection on the generated responses.

15. The apparatus for interactive generative model training of claim 14.

16. the first set of responses includes M responses generated for each dialogue input in the second set of dialogue inputs, where M is a positive integer greater than 1; The model optimization module performs the following process for each of the dialogue inputs in the second dialogue input set, in which the dialogue input is a dialogue input waiting to be processed, and obtains each candidate response corresponding to the dialogue input waiting to be processed and a manual labeling result for each candidate response, where the number of candidate responses is M or more, the candidate responses include responses generated for the dialogue input waiting to be processed and / or responses manually modified from the responses generated for the dialogue input waiting to be processed, and the manual labeling result for each candidate response includes a labeling result after manually performing safety labeling on the candidate response based on the target safety norm, and constructs training samples based on the dialogue input waiting to be processed, each candidate response, and the manual labeling result for each candidate response, and optimizes the dialogue generation model and the detection model using the training samples.

16. The apparatus for interactive generative model training of claim 15.

17. For any candidate response, the labeling result after the safety labeling is manually performed includes evaluation labels corresponding to different evaluation dimensions of the labeled candidate response based on evaluation criteria of different evaluation dimensions of the combinations corresponding to the dialogue inputs to be processed, and the evaluation labels either conform to the corresponding evaluation criteria or do not conform to the corresponding evaluation criteria.

17. The apparatus for interactive generative model training of claim 16.

18. The model optimization module respectively constructs a first type of training sample and a second type of training sample, uses the first type of training sample to optimize the interactive generative model through supervised learning, and uses the second type of training sample to optimize the detection model through supervised learning.

18. The apparatus for interactive generative model training of claim 17.

19. The model optimization module selects candidate responses that meet the following condition from each candidate response, where the condition is that the evaluation labels of different evaluation dimensions all meet the corresponding evaluation criteria, and configures each selected candidate response and the waiting dialogue input as a first type of training sample.

20. The apparatus for interactive generative model training of claim 18.

20. The model optimization module is further used to obtain a total score for each candidate response, where a higher total score indicates a higher level of security; The detection model includes a comprehensive detection model and classification detection models corresponding to different evaluation dimensions, respectively; the second type of training samples includes a first subtype of training samples and a second subtype of training samples, the first subtype of training samples includes two candidate replies with different overall scores, the waiting dialogue input, and a sample label, the sample label being used to indicate the candidate reply with the higher overall score of the two candidate replies, the second subtype of training samples includes one candidate reply, the waiting dialogue input, and one rating label of the candidate reply; the model optimization module optimizes the overall detection model using the training samples of the first subtype, and for any of the classification detection models, optimizes the classification detection model using training samples of the second subtype each including an evaluation label of an evaluation dimension corresponding to the classification detection model; 20. The apparatus for interactive generative model training of claim 18.

21. the second set of responses includes one response generated for each dialogue input in the third set of dialogue inputs; The model optimization module uses the optimized detection model to perform safety detection for each answer in the second answer set, and re-optimizes the optimized interactive generation model using a reinforcement learning method according to the safety detection result for each answer. An interactive generative model training apparatus according to any one of claims 15 to 19.

22. The detection model includes a comprehensive detection model and classification detection models corresponding to different evaluation dimensions, respectively; The model optimization module performs the following process for each response, respectively: obtain classification detection results corresponding to the returned overall detection result and different classification detection models; combine the overall detection result and the different classification detection results to determine a reward corresponding to the response; use the response, the dialogue input corresponding to the response, and the reward to construct a training sample; and re-optimize the optimized dialogue generative model using the training sample.

22. The apparatus for interactive generative model training of claim 21.

23. The model optimization module uses the optimized interactive generative model as a baseline model, generates a target model that is identical to the baseline model, and optimizes the target model using the training samples based on a Kulbeck-Lebel variance constraint introduced between the baseline model and the target model, and then uses the optimized target model as a re-optimized interactive generative model.

23. The apparatus for interactive generative model training of claim 22.

24. A generative dialogue realization device, an input acquisition module and a response generation module; the input acquisition module is used to acquire a dialogue input to be processed; The response generation module is used to generate a response corresponding to the pending dialogue input using an dialogue generation model, the dialogue generation model being an dialogue generation model that meets online requirements obtained by N iterative optimizations, where N is a positive integer greater than 1, and each optimization includes, in response to determining that an update has been generated in a safety norm, optimizing the dialogue generation model based on the determined dialogue input according to a principle that the response generated by the dialogue generation model meets a target safety norm, the target safety norm being the safety norm after the update, the determined dialogue input being a dialogue input corresponding to the current optimization determined based on the target safety norm, and the update being an update to the original safety norm when it is determined that the dialogue generation model after the most recent optimization does not meet the online requirements; The safety norms include evaluation norms of at least one evaluation dimension corresponding to different combinations, and each combination is composed of one content area and one application scene, the content area is a safety content area related to generative interaction, and the application scene is an application scene of generative interaction; The update to the safety rule may include one or any combination of: adding a new evaluation rule for the combination and at least one corresponding evaluation dimension; adding a new evaluation dimension and a corresponding evaluation rule for the original combination; and adjusting the original evaluation rule. Generative dialogue realization device.

25. An electronic device, at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform the method of any one of claims 1 to 8. electronic equipment.

26. An electronic device, at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform the method of claim 12. electronic equipment.

27. A computer program comprising: The computer program, when executed by a processor, implements the method according to any one of claims 1 to 8. Computer program.

28. A computer program comprising: The computer program, when executed by a processor, implements the method of claim 12. Computer program.

Citation Information

Patent Citations

  • Chat system and chat program

    JP2022093814A

  • Retraining a conversation system based on negative feedback

    US20200364511A1

  • Information processing system and information processing method

    WO2020202731A1

  • Information processing device, information processing method, and program

    WO2021070732A1

  • Method and apparatus for self-supervised extractive question answering

    WO2023098971A1