Question and answer information processing method and device, model training method and device, electronic equipment and medium

By generating training samples from user feedback and adjusting the parameters of the conversational model, the problem of high expert annotation costs and inconsistent user experience improvement was solved, thus improving both model performance and user experience.

CN119168093BActive Publication Date: 2026-02-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411320303.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2026-02-13
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

In existing technologies, relying on annotation by experts in related fields is costly and inefficient, and it does not fully utilize user behavior signals, resulting in a mismatch between model performance and the direction of user experience improvement.

Method used

Training samples are generated by acquiring user feedback information, initial response information is generated using a conversational model, and model parameters are adjusted according to user preference levels to generate response information that is closer to the user's true preferences.

Benefits of technology

It reduces the cost of expert annotation, improves model performance and user experience, and ensures that the model's performance is consistent with user preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119168093B_ABST
    Figure CN119168093B_ABST
Patent Text Reader

Abstract

The present disclosure provides a question and answer information processing method, relates to the technical field of artificial intelligence, and particularly relates to the technical fields of deep learning, large models, intelligent question and answer, and the like. The specific implementation scheme is as follows: generating at least one initial answer information according to question information provided by an object; obtaining at least one feedback information corresponding to the at least one initial answer information, wherein the feedback information is used to indicate a preference degree of the object to the initial answer information; and generating a training sample according to the question information, the at least one initial answer information, and the at least one feedback information. The present disclosure also provides a training method and device of a conversational model, an electronic device, and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of deep learning, large model, intelligent question answering and the like. More specifically, the present disclosure provides a question and answer information processing method, a model training method, an apparatus, an electronic device and a storage medium. BACKGROUND

[0002] With the development of artificial intelligence technology, the application scenarios of conversational models are increasing. SUMMARY

[0003] The present disclosure provides a question and answer information processing method, a training method of a conversational model, an apparatus, a device and a storage medium.

[0004] According to an aspect of the present disclosure, a question and answer information processing method is provided, which includes: generating at least one initial answer information according to question information provided by an object; obtaining at least one feedback information corresponding to the at least one initial answer information, wherein the feedback information is used to indicate a preference degree of the object to the initial answer information; and generating a training sample according to the question information, the at least one initial answer information and the at least one feedback information.

[0005] According to another aspect of the present disclosure, a training method of a conversational model is provided, which includes: adjusting parameters of the conversational model according to at least one training sample, so that the conversational model generates adjusted answer information according to question information in the training sample, the adjusted answer information is close to answer information with a high preference degree in the training sample and far away from answer information with a low preference degree in the training sample, wherein the training sample is generated by: generating at least one initial answer information according to question information provided by an object; obtaining at least one feedback information corresponding to the at least one initial answer information, wherein the feedback information is used to indicate a preference degree of the object to the initial answer information; and generating a training sample according to the question information, the at least one initial answer information and the at least one feedback information.

[0006] According to another aspect of the present disclosure, a question and answer information processing apparatus is provided, which includes: a first generation module configured to generate at least one initial answer information according to question information provided by an object; an obtaining module configured to obtain at least one feedback information corresponding to the at least one initial answer information, wherein the feedback information is used to indicate a preference degree of the object to the initial answer information; and a second generation module configured to generate a training sample according to the question information, the at least one initial answer information and the at least one feedback information.

[0007] According to another aspect of the present disclosure, there is provided a device for training a conversational model, the device comprising: an adjusting module configured to adjust parameters of the conversational model according to at least one training sample, such that the conversational model generates adjusted answer information according to question information in the training sample, the adjusted answer information being closer to answer information with a high degree of preference and further away from answer information with a low degree of preference in the training sample, wherein the training sample is generated by: a first generating module configured to generate at least one initial answer information according to question information provided by a subject; an obtaining module configured to obtain at least one feedback information corresponding to the at least one initial answer information, wherein the feedback information is used to indicate a degree of preference of the subject for the initial answer information; and a second generating module configured to generate the training sample according to the question information, the at least one initial answer information and the at least one feedback information.

[0008] According to another aspect of the present disclosure, there is provided an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method according to the present disclosure.

[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform a method according to the present disclosure.

[0010] According to another aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements a method according to the present disclosure.

[0011] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:

[0013] Figure 1 is an exemplary system architecture schematic diagram of a system to which a question and answer information processing method and device according to one embodiment of the present disclosure can be applied;

[0014] Figure 2 is a flowchart of a question and answer information processing method according to one embodiment of the present disclosure;

[0015] Figure 3 is a schematic diagram of the principle of a question and answer information processing method according to one embodiment of the present disclosure;

[0016] Figure 4 is a schematic flowchart of a method for training a dialog model according to another embodiment of the present disclosure;

[0017] Figure 5 is a schematic schematic diagram of a model training method according to an embodiment of the present disclosure;

[0018] Figure 6 is a block diagram of a question answering information processing apparatus according to an embodiment of the present disclosure;

[0019] Figure 7 is a block diagram of a training apparatus of a dialog model according to another embodiment of the present disclosure; and

[0020] Figure 8 is a block diagram of an electronic device to which a question answering information processing method and / or a method for training a dialog model can be applied according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.

[0022] A large model can include a large language model (LLM), an image large model, a video large model, etc. Taking a large language model as an example, a large-scale unsupervised corpus can be used for pre-training, and then a supervised fine-tuning (SFT) training can be performed using a supervised fine-tuning corpus carefully annotated by experts in the relevant field. Next, an alignment algorithm such as Kahneman-Tversky Optimization (KTO), Direct Preference Optimization (DPO), simple Preference Optimization (simPO), Proximal Policy Optimization (PPO), etc. can be used for alignment training. The upper bound of the performance of the model depends on the quantity and quality of the supervised fine-tuning corpus and the effect of the reward model in the alignment training phase.

[0023] In some embodiments, the effect of the model is highly related to the number and quality of the training samples. For most models, the quality of the corpus can be improved by annotating the corpus by experts in the relevant field. However, the cost of annotation is high and the efficiency is low.

[0024] In some embodiments, for a target application scenario, the data mining process can be customized to continuously mine the problems (queries) that the model is difficult to reply in the target application scenario with high quality, thereby improving the annotation efficiency of experts in the relevant field. However, in the mining process, the behavior signals of the user are not fully used, and the improvement direction of the model effect is difficult to be consistent with the improvement direction of the user experience.

[0025] Therefore, in order to fully improve the performance of the model and the user experience, the present disclosure provides a question and answer information processing method, which will be described below.

[0026] Figure 1 is an exemplary system architecture schematic diagram to which the question and answer information processing method and device according to an embodiment of the present disclosure can be applied. It should be noted that, Figure 1 The diagram shown is only an example of a system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0027] As Figure 1 shown, the system architecture 100 according to the embodiment can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is a medium for providing a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired and / or wireless communication links, etc.

[0028] The user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers and desktop computers, etc.

[0029] The server 105 can be a server providing various services, such as a background management server supporting the website browsed by the user using the terminal devices 101, 102, 103 (only as an example). The background management server can analyze and process the received user request and other data, and feed back the processing result (such as a web page, information or data generated or obtained according to the user request, etc.) to the terminal device.

[0030] It should be noted that the question and answer information processing method provided by the embodiments of the present disclosure can be generally executed by the server 105. Correspondingly, the question and answer information processing apparatus provided by the embodiments of the present disclosure can be generally arranged in the server 105. The question and answer information processing method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal device 101, 102, 103 and / or the server 105. Correspondingly, the question and answer information processing apparatus provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal device 101, 102, 103 and / or the server 105.

[0031] It can be understood that the system architecture of the present disclosure is described above, and the method of the present disclosure will be described below.

[0032] Figure 2 is a flowchart of a question and answer information processing method according to an embodiment of the present disclosure.

[0033] As shown in Figure 2 , the method 200 can include operations S210 to S230.

[0034] At operation S210, at least one initial answer information is generated according to question information provided by a subject.

[0035] In the embodiments of the present disclosure, the subject can include a user. The question information can include a query input by the user. For example, the question information can be information in various forms such as text, image, audio, video, etc.

[0036] In the embodiments of the present disclosure, the initial answer information can be generated in various ways. For example, a preset mapping relationship can be used to determine the initial answer information corresponding to the question information. A conversational model can also be used to determine the initial answer information corresponding to the question information. The initial answer information can also be information in various forms such as text, image, audio, video, etc.

[0037] At operation S220, at least one feedback information corresponding to the at least one initial answer information is obtained.

[0038] In the embodiments of the present disclosure, the feedback information can indicate the degree of preference of the subject for the initial answer information. For example, the feedback information can be provided by the user. It can be understood that after the user authorizes, the question information and the feedback information provided by the user can be obtained.

[0039] At operation S230, a training sample is generated according to the question information, the at least one initial answer information and the at least one feedback information.

[0040] In the embodiments of the present disclosure, according to the question information, the at least one initial answer information and the at least one feedback information, the multi-element data can be determined as the training sample.

[0041] By the embodiments of the present disclosure, the feedback information provided by the object for the initial answer information can be obtained, the dependence on the experts in the related field is reduced, and the labeling cost of the experts can be greatly reduced. The feedback information can be obtained, and the more real user preferences can be obtained. The training sample generated based on the question information, the initial answer information and the feedback information can make the generated answer information closer to the real preferences of the user, and the user experience can be greatly improved.

[0042] It can be understood that the method of the present disclosure is described above, and some ways of generating the initial answer information of the present disclosure will be described below.

[0043] In some embodiments, in some implementations of the above operation S210, at least one initial answer information is generated by using a conversational model according to the question information provided by the object.

[0044] For example, the conversational model can be a large model, or a lightweight model applied to a target application scenario.

[0045] For example, according to the question information of the audio type provided by the user, the conversational model can perform audio recognition to determine the semantics of the question information. Next, the conversational model can generate initial answer information of the audio type or initial answer information of a type corresponding to the semantics. For another example, according to the question information of the text type provided by the user, the conversational model can also determine the semantics of the question information to determine the corresponding initial answer information. It can be understood that the way in which the conversational model generates initial answer information based on question information of the image type or the video type is the same as or similar to the way in which the conversational model generates initial answer information based on question information of the audio type, and the present disclosure will not be repeated here. By the embodiments of the present disclosure, the conversational model can be used to more accurately determine the answer information, and the user experience can be improved.

[0046] It can be understood that some ways of generating the initial answer information of the present disclosure are described above, and the feedback information of the present disclosure will be described below.

[0047] In some embodiments, the feedback information can include a preference degree value.

[0048] In the embodiments of the present disclosure, the preference degree value is determined according to one of a plurality of visual controls triggered by the user, and the plurality of visual controls include a first visual control and a second visual control. The preference degree value corresponding to the first visual control is higher than the preference degree value corresponding to the second visual control.

[0049] In some embodiments, in some implementations of the operation S220 and the operation S230, the preference degree value corresponding to the initial answer information is obtained; and the training sample is generated according to the question information, the initial answer information, and the feedback information.

[0050] For example, the user can provide the question information to the dialog model. The dialog model can generate the initial answer information corresponding to the question information. If the user considers that the initial answer information has a high relevance to the question information, the first visual control can be triggered to determine the preference degree value of the initial answer information as the first preference degree value. Thus, a triple data can be generated as a training sample. The triple data can include the question information, the initial answer information, and the first preference degree value. It can be understood that the first visual control can be a “like” control. The first preference degree value may, for example, be 1.

[0051] For another example, the user can provide the question information to the dialog model. The dialog model can generate the initial answer information corresponding to the question information. If the user considers that the initial answer information has a low relevance to the question information, the second visual control can be triggered to determine the preference degree value of the initial answer information as the second preference degree value. Thus, a triple data can be generated as a training sample. The triple data can include the question information, the initial answer information, and the second preference degree value. It can be understood that the second visual control can be a “dislike” control. The second preference degree value may, for example, be 0.

[0052] For another example, the triple data can be represented as <query, response, signal>. The response can be the initial answer information. The signal can be the feedback signal information corresponding to the first preference degree value or the second preference degree value.

[0053] It can be understood that the triple data including the question information, the initial answer information, and the preference degree value can be pointwise preference data.

[0054] It can be understood that the above describes the present disclosure by taking the feedback information as an example of the first preference degree value or the second preference degree value. However, the present disclosure is not limited thereto, and the feedback information can also be a preference degree evaluation value provided by the object. For example, the value range of the preference degree evaluation value can be 0-100, and the higher the preference degree evaluation value, the more the initial answer information conforms to the real preference of the user.

[0055] It can be understood that the above describes the generation manner of the pointwise preference data. Another type of preference data will be described below.

[0056] In some embodiments, in another implementation of the method 200 described above, according to the question information provided by the subject, a plurality of initial answer information can be generated. A plurality of preference degree values corresponding to the plurality of initial answer information are obtained. According to the question information, the plurality of initial answer information and the plurality of preference degree values, a training sample can be generated.

[0057] For example, a user can provide question information to a conversational model. The conversational model can generate a plurality of initial answer information. Alternatively, the user can also provide the same question information to the conversational model multiple times to obtain a plurality of initial answer information. For the plurality of initial answer information, the user can determine a plurality of preference degree evaluation values respectively. The plurality of preference degree evaluation values can indicate the user's preference degree for different initial answer information. The plurality of preference degree evaluation values can be used as feedback signal information. Thus, taking N initial answer information as an example, an N-tuple data can be generated as a training sample. The N-tuple data can include question information, N initial answer information and feedback signal information. For another example, the N-tuple data can be represented as <query, response1, ……, responseN, signal>. response1, ……, responseN can be N initial answer information. signal can be feedback signal information corresponding to N preference degree evaluation values. N can be an integer greater than 1.

[0058] It can be understood that the N-tuple data can be pairwise preference data.

[0059] It can be understood that the feedback information described above can be generated during the conversation between the subject and the conversational model, and can be used as conversation feedback information. In related application scenarios, the number of users of the conversational model is huge. The use behavior (providing feedback) of the user plays an important role in improving the performance of the model. Through the embodiments of the present disclosure, the evaluation results of the user on the answer information generated by the conversational model can be obtained to obtain the feedback signal of the user in the real scene, and the performance of the conversational model can be efficiently improved.

[0060] It can be understood that the above describes the present disclosure by taking the feedback information as a preference degree value. However, the present disclosure is not limited thereto, and the feedback information can also be target answer information provided by the subject, which will be described below.

[0061] In some embodiments, the target application scenario described above can be an artificial intelligence assisted creation scenario. For example, the target application scenario can be a novel assisted writing scenario, an intelligent coding scenario, or an intelligent presentation creation scenario. According to the question information provided by the user, the conversational model can recommend one or more contents. The user can directly adopt the recommended contents without editing the contents, or can not adopt the recommended contents or partially adopt the recommended contents. The contents can be used as initial answer information.

[0062] In the embodiments of the present disclosure, obtaining the at least one feedback information corresponding to the at least one initial answer information includes: obtaining target answer information provided by the object, corresponding to the question information, and having an association index value between the initial answer information lower than a preset association threshold. For example, in the case that the user considers that the recommended content has low quality, the user can not adopt the recommended content and write target content by himself. The association between the recommended content and the target content is low, which can be lower than the preset association threshold. The target content can be used as target answer information corresponding to the question information.

[0063] In the embodiments of the present disclosure, obtaining the at least one feedback information corresponding to the at least one initial answer information includes: obtaining target answer information provided by the object and edited from the initial answer information. For example, in the case that the user considers that the recommended content can be partially adopted, the user can edit the recommended content to obtain target content. The target content can be used as target answer information.

[0064] In the embodiments of the present disclosure, according to the question information, the initial answer information and the target answer information, a triple data can be generated as a training sample. For example, the triple data can be represented as <query, response_before, response_after>. response_before can be the initial answer information. response_after can be the target answer information. The triple data includes two answer information, and can also be used as pair-wise type preference data.

[0065] It can be understood that in the case that the user directly adopts the recommended content without editing the content, it is difficult to generate pair-wise type preference data based on the content. However, the preference degree value of the content can be determined as a first preference degree value to generate point-wise type preference data.

[0066] It can be understood that the feedback information above is the target answer information generated by the object editing, which can be used as edit feedback information. In related application scenarios, the user can modify and polish the result output by the model. Through the embodiments of the present disclosure, the result modified and polished by the user is obtained as the target answer information, very high quality user preference data can be obtained, and the performance of the conversational model can be fully improved.

[0067] It can be understood that the above describes the present disclosure by taking the user as an example. However, the present disclosure is not limited thereto, and the object can also be a large model, which will be described below.

[0068] In the embodiments of the present disclosure, the feedback information can be generated by the large model according to the preset feedback rule and the initial answer information. For example, the conversational model can be an end-side model with a small size. The large model has a large size and good generation capability, and also has evaluation capability. Based on the preset feedback rule, the large model can be used to determine the preference degree value as the feedback information.

[0069] In the embodiments of the present disclosure, the preset feedback rule can be determined according to the user's preference. For example, the preset feedback rule can include the number of words corresponding to the answer information, the relevance to the question information, etc. The specific values of the number of words and the relevance can be determined based on the user's preference.

[0070] For example, the user can provide the question information to the conversational model. The conversational model can generate a plurality of initial answer information. Alternatively, the user can also provide the same question information to the conversational model multiple times to obtain a plurality of initial answer information. For the plurality of initial answer information, based on the preset feedback rule, the large model can determine a plurality of preference degree evaluation values respectively. The plurality of preference degree evaluation values can also indicate the preference degree of the user to different initial answer information. The plurality of preference degree values can be used as feedback signal information. Thus, taking N initial answer information as an example, an N-tuple data can be generated as a training sample. The N-tuple data can include the question information, the N initial answer information and the feedback signal information. For example, the N-tuple data can be represented as <query, response1, ……, responseN, signal’>. response1, ……, responseN can be the N initial answer information. signal’ can be the feedback signal information corresponding to the N preference degree evaluation values generated by the large model.

[0071] It can be understood that the above describes the present disclosure by taking the large model to determine the preference degree value as an example. However, the present disclosure is not limited thereto, and in some embodiments, the large model can also generate the target answer information. The target answer information generated by the large model can also be used as the feedback information.

[0072] It can be understood that the above is described by taking the large model determining the feedback information as an example. However, the present disclosure is not limited thereto, and the large model can also provide the question information to the dialog model.

[0073] It can be understood that the feedback information in the above is determined by the large model and can be used as artificial intelligence feedback (AI feedback) information. Through the embodiments of the present disclosure, the evaluation capability of the large model can be fully utilized, the richness of the feedback information source can be improved, and the performance of the dialog model can be further improved. The following will be described in combination with Figure 3 The various feedback information of the present disclosure will be further described.

[0074] Figure 3 is a schematic diagram of a question and answer information processing method according to an embodiment of the present disclosure.

[0075] As shown in Figure 3 , the model can be pre-trained to obtain a pre-trained model Mpret30. The pre-trained model Mpret30 can be supervised fine-tuned to obtain a fine-tuned model Msft30. The fine-tuned model Msft30 can be used as a dialog model Mchat30. According to the question information provided by the object, the dialog model Mchat30 can generate initial answer information. Based on the initial answer information, the dialog feedback information CF30, the editing feedback information EF30 and the artificial intelligence feedback information AF30 can be obtained. Based on the dialog feedback information CF30, the point-by-point type preference data pointd30 can be obtained. Based on the dialog feedback information CF30, the editing feedback information EF30 and the artificial intelligence feedback information AF30, the pair-by-pair type preference data paird30 can be obtained. The point-by-point type preference data pointd30 and the pair-by-pair type preference data paird30 can be used as training samples.

[0076] It can be understood that the above describes the way of generating training samples, and the following will describe some ways of model training.

[0077] Figure 4 is a schematic flow chart of a training method of a dialog model according to another embodiment of the present disclosure.

[0078] As shown in Figure 4 , the method 400 can include operation S440.

[0079] In operation S440, according to at least one training sample, the parameters of the dialog model are adjusted so that the dialog model generates adjusted answer information according to the question information in the training sample.

[0080] In the embodiments of the present disclosure, the conversational model can be trained for alignment to adjust the parameters of the conversational model.

[0081] In the embodiments of the present disclosure, the adjusted answer information is close to the answer information with high preference degree in the training sample and far away from the answer information with low preference degree in the training sample.

[0082] In the embodiments of the present disclosure, the training sample is generated by the following operations: generating at least one initial answer information according to the question information provided by the object. At least one feedback information corresponding to the at least one initial answer information is obtained. The feedback information can indicate the preference degree of the object to the initial answer information. The training sample is generated according to the question information, the at least one initial answer information and the at least one feedback information. For example, the training sample can be generated by the method 300.

[0083] Through the embodiments of the present disclosure, the feedback information in the training sample can be at least one of the above-mentioned conversational feedback information, the editing feedback information and the artificial intelligence feedback information, which provides different types of feedback signals for model training. These feedback signals are highly related to the preferences of users in real scenarios, which can align the performance of the model with the real preferences of the users, effectively improve the performance upper limit of the model, and make the model evolve towards the direction of the real preferences of the users. In addition, the labeling cost of the sample can be reduced, and the efficiency of obtaining the training sample can be improved.

[0084] It can be understood that the model training method of the present disclosure is described above, and the model training method of the present disclosure will be further described in combination with the above-mentioned various feedback information.

[0085] Figure 5 is a schematic principle diagram of a model training method according to an embodiment of the present disclosure.

[0086] As Figure 5As shown, the model can be pre-trained to obtain a pre-trained model Mpret50. The pre-trained dialog model Mpret50 can be supervised fine-tuned to obtain a fine-tuned model Msft50. The fine-tuned model Msft50 can be used as the dialog model Mchat50. According to the question information provided by the object, the dialog model Mchat50 can generate initial answer information. Based on the initial answer information, the dialog feedback information CF50, the editing feedback information EF50 and the artificial intelligence feedback information AF50 can be obtained. Based on the dialog feedback information CF50, the point-by-point type preference data pointd50 can be obtained. Based on the dialog feedback information CF50, the editing feedback information EF50 and the artificial intelligence feedback information AF50, the pair-by-pair type preference data paird50 can be obtained. The point-by-point type preference data pointd50 and the pair-by-pair type preference data paird50 can be used as training samples.

[0087] In the embodiments of the present disclosure, the feedback information includes a preference degree value. According to the at least one training sample, adjusting the parameters of the dialog model includes: adjusting the parameters of the dialog model according to the question information, the initial answer information and the preference degree value of the training sample. The parameters of the dialog model can be adjusted based on the Kahneman-Tversky Optimization (KTO) method. For example, the above-mentioned point-by-point type preference data pointd50 can include a triple data <query, response, signal>. The triple data can include question information, initial answer information and preference degree value. Based on the Kahneman-Tversky Optimization method, the parameters of the dialog model can be adjusted. According to the question information in the triple, the adjusted dialog model can generate adjusted answer information. The preference degree value of the adjusted answer information may, for example, be greater than the preference degree value of the initial answer information.

[0088] It can be understood that the above describes the present disclosure by taking the point-by-point type preference data as an example. However, the present disclosure is not limited thereto, and the following will be described by taking the pair-by-pair type preference data as an example.

[0089] In the embodiments of the present disclosure, the training sample includes a plurality of initial answer information, and the feedback information includes a preference degree value. According to at least one training sample, adjusting the parameters of the dialog model includes: adjusting the parameters of the dialog model according to the question information of the training sample, the plurality of initial answer information, and the preference degree value of each of the plurality of initial answer information. The parameters of the dialog model can be adjusted based on at least one of direct preference optimization (DPO), simple preference optimization (simPO), and proximal policy optimization (PPO). For example, the pair-wise type of preference data can include the above-mentioned N-tuple data <query, response1, …, responseN, signal>. The N-tuple data can include question information, N initial answer information, and feedback information. The feedback information can correspond to N preference degree values. For example, based on the direct preference optimization mode, the parameters of the dialog model can be adjusted. According to the question information in the N-tuple, the adjusted dialog model can generate adjusted answer information. The preference degree value of the adjusted answer information can be close to the highest value in the N preference degree values, for example.

[0090] It can be understood that the above describes the present disclosure by taking the pair-wise type of preference data including N-tuple data as an example. However, the present disclosure is not limited thereto, and the pair-wise type of preference data can also include triple data, which will be described below.

[0091] In the embodiments of the present disclosure, the feedback information includes target answer information provided by an object. The preference degree of the object to the target answer information is higher than the preference degree of the initial answer information. According to at least one training sample, adjusting the parameters of the dialog model includes: adjusting the parameters of the dialog model according to the question information of the training sample, the initial answer information, and the target answer information. The parameters of the dialog model can be adjusted based on at least one of direct preference optimization, simple preference optimization, and proximal policy optimization. For example, the pair-wise type of preference data can include the above-mentioned triple data <query, response_before, response_after>. The triple data can include question information, initial answer information, and target answer information. For example, based on the simple preference optimization mode, the parameters of the dialog model can be adjusted. According to the question information in the triple, the adjusted dialog model can generate adjusted answer information. The preference degree of the object to the adjusted answer information can be close to the preference degree to the target answer information, for example.

[0092] It can be understood that the above describes the method of the present disclosure, and the device of the present disclosure will be described below.

[0093] Figure 6 Block diagram of a question and answer information processing device according to an embodiment of the present disclosure.

[0094] As Figure 6As shown, the apparatus 600 can include a first generation module 610, an acquisition module 620, and a second generation module 630.

[0095] The first generation module 610 is configured to generate at least one initial answer information according to the question information provided by the object.

[0096] The acquisition module 620 is configured to acquire at least one feedback information corresponding to the at least one initial answer information.

[0097] In the embodiments of the present disclosure, the feedback information is used to indicate the preference degree of the object to the initial answer information.

[0098] The second generation module 630 is configured to generate a training sample according to the question information, the at least one initial answer information, and the at least one feedback information.

[0099] In some embodiments, the first generation module includes a first generation sub-module configured to generate the at least one initial answer information by using a dialog model according to the question information.

[0100] In some embodiments, the object includes at least one of a user and a large model, and the feedback information includes at least one of a preference degree value and target answer information provided by the object.

[0101] In some embodiments, the preference degree of the object to the target answer information is higher than the preference degree of the object to the initial answer information. The acquisition module includes at least one of a first acquisition sub-module configured to acquire the target answer information provided by the object, corresponding to the question information, and having an association index value between the initial answer information lower than a preset association threshold value, and a second acquisition sub-module configured to acquire the target answer information provided by the object and edited from the initial answer information.

[0102] In some embodiments, the feedback information is generated by the large model according to a preset feedback rule and the initial answer information, and the preset feedback rule is determined according to the preference of the user.

[0103] In some embodiments, the preference degree value is determined according to one of a plurality of visual controls triggered by the user, and the plurality of visual controls include a first visual control and a second visual control, and the preference degree value corresponding to the first visual control is higher than the preference degree value corresponding to the second visual control.

[0104] It can be understood that the above describes the question and answer information processing apparatus of the present disclosure, and the model training apparatus of the present disclosure will be described below.

[0105] Figure 7 is a block diagram of a dialog model training apparatus according to another embodiment of the present disclosure.

[0106] As Figure 7As shown, the apparatus 700 can include an adjusting module 740.

[0107] The adjusting module 740 is configured to adjust the parameters of the dialog model according to the at least one training sample, so that the dialog model generates the adjusted answer information according to the question information in the training sample.

[0108] In the embodiments of the present disclosure, the adjusted answer information is close to the answer information with a high degree of preference in the training sample and far away from the answer information with a low degree of preference in the training sample,

[0109] In the embodiments of the present disclosure, the training sample is generated by the following modules: a first generating module configured to generate at least one initial answer information according to question information provided by the object. An obtaining module configured to obtain at least one feedback information corresponding to the at least one initial answer information. The feedback information is used to indicate the degree of preference of the object for the initial answer information. A second generating module configured to generate the training sample according to the question information, the at least one initial answer information and the at least one feedback information. For example, the training sample can be generated by the above-described apparatus 600.

[0110] In some embodiments, the feedback information includes a preference degree value. The adjusting module includes: a first adjusting submodule configured to adjust the parameters of the dialog model according to the question information, the initial answer information and the preference degree value of the training sample.

[0111] In some embodiments, the adjusting module further includes: a first adjusting unit configured to adjust the parameters of the dialog model based on a Kahneman-Tversky optimization method.

[0112] In some embodiments, the training sample includes a plurality of initial answer information, and the feedback information includes a preference degree value. The adjusting module includes: a second adjusting submodule configured to adjust the parameters of the dialog model according to the question information, the plurality of initial answer information and the preference degree value of each of the plurality of initial answer information.

[0113] In some embodiments, the feedback information includes target answer information provided by the object, and the degree of preference of the object for the target answer information is higher than the degree of preference of the initial answer information. The adjusting module includes: a third adjusting submodule configured to adjust the parameters of the dialog model according to the question information, the initial answer information and the target answer information of the training sample.

[0114] In some embodiments, the adjusting module further includes: a second adjusting unit configured to adjust the parameters of the dialog model based on at least one of direct preference optimization, simple preference optimization and proximal policy optimization.

[0115] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0116] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0117] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0118] As shown in Figure 8 The device 800 includes a computing unit 801 that can perform various appropriate actions and processes in accordance with a computer program stored in a Read-Only Memory (ROM) 802 or loaded into a Random Access Memory (RAM) 803 from a storage unit 808. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An Input / Output (I / O) interface 805 is also connected to the bus 804.

[0119] Various components in the device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc., an output unit 807, such as various types of displays, a speaker, etc., a storage unit 808, such as a magnetic disk, an optical disk, etc., and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0120] The computing unit 801 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a Central Processing Unit (CPU), a Graph Processing Unit (GPU), various special-purpose Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs various methods and processes described above, such as the question-answer information processing method and / or the training method of the conversational model. For example, in some embodiments, the question-answer information processing method and / or the training method of the conversational model can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded onto the RAM 803 and executed by the computing unit 801, one or more steps of the question-answer information processing method and / or the training method of the conversational model described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the question-answer information processing method and / or the training method of the conversational model by other any appropriate means, such as by means of firmware.

[0121] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), an Application Specific Standard Parts (ASSP), a System on Chip (SOC), a Complex Programmable Logic Device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0122] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general or special purpose computer, such that the program code, when executed by the processor or controller, causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0123] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a Cathode Ray Tube (CRT) monitor or a Liquid Crystal Display (LCD)) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0125] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0126] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0127] It should be understood that the steps shown above in the various forms of flow can be reordered, added to, or deleted from. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure can be achieved, which is not limited herein.

[0128] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A question-and-answer information processing method, comprising: Based on the question information provided by the object, generate at least one initial answer; Obtain at least one piece of feedback information corresponding to at least one of the initial response information, wherein the feedback information is used to indicate the object's degree of preference for the initial response information, the feedback information including target response information provided by the object and a preference value, wherein the object's degree of preference for the target response information is higher than its degree of preference for the initial response information, and the preference value is determined based on one of a plurality of visual controls triggered by the user; and Generating training samples based on the question information, at least one initial answer, and at least one feedback includes: generating pairwise typed preference data as the training samples, wherein the pairwise typed preference data includes triplet data comprising the question information, the initial answer, and the target answer. The acquisition of at least one feedback information corresponding to at least one of the initial response information includes at least one of the following: Obtain target answer information provided by the object that corresponds to the question information and whose correlation index value with the initial answer information is lower than a preset correlation threshold; Obtain the target answer information provided by the object and obtained by editing the initial answer information.

2. The method according to claim 1, wherein, The step of generating at least one initial answer based on the question information provided by the object includes: Based on the question information, at least one initial response is generated using a conversational model.

3. The method according to claim 1, wherein, The objects include at least one of users and large models.

4. The method according to claim 3, wherein, The feedback information is generated by the large model based on preset feedback rules and the initial response information. The preset feedback rules are determined based on the user's preferences.

5. The method according to claim 3, wherein, The plurality of visual controls include a first visual control and a second visual control, wherein the preference level value corresponding to the first visual control is higher than the preference level value corresponding to the second visual control.

6. A training method for a conversational model, comprising: Based on at least one training sample, the parameters of the conversational model are adjusted such that the conversational model generates adjusted response information based on the question information in the training sample. The adjusted response information is close to the responses with high preference in the training sample and far away from the responses with low preference in the training sample. The training samples are generated through the following operations: Based on the question information provided by the object, generate at least one initial answer; Obtain at least one piece of feedback information corresponding to at least one of the initial response information, wherein the feedback information is used to indicate the object's degree of preference for the initial response information, the feedback information including target response information provided by the object and a preference value, wherein the object's degree of preference for the target response information is higher than its degree of preference for the initial response information, and the preference value is determined based on one of a plurality of visual controls triggered by the user; and Generating the training samples based on the question information, at least one initial answer, and at least one feedback includes: generating pairwise typed preference data as the training samples, wherein the pairwise typed preference data includes triplet data comprising the question information, the initial answer, and the target answer. The acquisition of at least one feedback information corresponding to at least one of the initial response information includes at least one of the following: Obtain target answer information provided by the object that corresponds to the question information and whose correlation index value with the initial answer information is lower than a preset correlation threshold; Obtain the target answer information provided by the object and obtained by editing the initial answer information.

7. The method according to claim 6, wherein, Adjusting the parameters of the conversational model based on at least one of the training samples includes: The parameters of the conversational model are adjusted based on the question information, the initial answer information, and the preference level value of the training samples.

8. The method according to claim 7, wherein, The adjustment of the parameters of the conversational model includes: The parameters of the conversational model are adjusted based on the Kahneman-Tversky optimization method.

9. The method according to claim 6, wherein, The training samples include multiple initial response messages. Adjusting the parameters of the conversational model based on at least one of the training samples includes: The parameters of the conversational model are adjusted based on the question information of the training samples, the multiple initial response information, and the preference degree values ​​of each of the multiple initial response information.

10. The method according to claim 6, wherein, The feedback information includes the target answer information provided by the object, and the object's preference for the target answer information is higher than its preference for the initial answer information. Adjusting the parameters of the conversational model based on at least one of the training samples includes: The parameters of the conversational model are adjusted based on the question information, the initial answer information, and the target answer information of the training samples.

11. The method according to claim 8 or 9, wherein, The adjustment of the parameters of the conversational model includes: The parameters of the conversational model are adjusted based on at least one of direct preference optimization, simple preference optimization, and proximal policy optimization.

12. A question-and-answer information processing device, comprising: The first generation module is used to generate at least one initial answer based on the question information provided by the object. An acquisition module is configured to acquire at least one piece of feedback information corresponding to at least one initial response, wherein the feedback information is used to indicate the object's preference level for the initial response information, the feedback information includes target response information provided by the object and a preference level value, the object's preference level for the target response information is higher than its preference level for the initial response information, and the preference level value is determined based on one of a plurality of visual controls triggered by the user; and The second generation module is configured to generate training samples based on the question information, at least one initial answer information, and at least one feedback information, including: generating pairwise typed preference data as the training samples, wherein the pairwise typed preference data includes triplet data comprising the question information, the initial answer information, and the target answer information. The acquisition module includes at least one of the following: The first acquisition submodule is used to acquire target answer information provided by the object that corresponds to the question information and whose correlation index value with the initial answer information is lower than a preset correlation threshold; The second acquisition submodule is used to acquire the target answer information provided by the object and obtained by editing the initial answer information.

13. The apparatus according to claim 12, wherein, The first generation module includes: The first generation submodule is used to generate at least one initial response based on the question information using a conversational model.

14. The apparatus according to claim 12, wherein, The objects include at least one of users and large models.

15. The apparatus according to claim 14, wherein, The feedback information is generated by the large model based on preset feedback rules and the initial response information. The preset feedback rules are determined based on the user's preferences.

16. The apparatus according to claim 14, wherein, The plurality of visual controls include a first visual control and a second visual control, wherein the preference level value corresponding to the first visual control is higher than the preference level value corresponding to the second visual control.

17. A training device for a conversational model, comprising: An adjustment module is used to adjust the parameters of a conversational model based on at least one training sample, such that the conversational model generates adjusted response information based on question information in the training sample. The adjusted response information is close to responses with high preference in the training sample and far removed from responses with low preference in the training sample. The training samples are generated through the following modules: The first generation module is used to generate at least one initial answer based on the question information provided by the object. An acquisition module is configured to acquire at least one piece of feedback information corresponding to at least one initial response, wherein the feedback information is used to indicate the object's preference level for the initial response information, the feedback information includes target response information provided by the object and a preference level value, the object's preference level for the target response information is higher than its preference level for the initial response information, and the preference level value is determined based on one of a plurality of visual controls triggered by the user; and The second generation module is configured to generate the training samples based on the question information, at least one initial answer, and at least one feedback, including: generating pairwise typed preference data as the training samples, wherein the pairwise typed preference data includes triplet data comprising the question information, the initial answer, and the target answer. The acquisition module includes at least one of the following: The first acquisition submodule is used to acquire target answer information provided by the object that corresponds to the question information and whose correlation index value with the initial answer information is lower than a preset correlation threshold; The second acquisition submodule is used to acquire the target answer information provided by the object and obtained by editing the initial answer information.

18. The apparatus according to claim 17, wherein, The adjustment module includes: The first adjustment submodule is used to adjust the parameters of the conversational model based on the question information, the initial answer information, and the preference level value of the training samples.

19. The apparatus according to claim 18, wherein, The adjustment module also includes: The first adjustment unit is used to adjust the parameters of the conversational model based on the Kahneman-Tversky optimization method.

20. The apparatus according to claim 17, wherein, The training samples include multiple initial response messages. The adjustment module includes: The second adjustment submodule is used to adjust the parameters of the conversational model based on the question information of the training samples, multiple initial response information, and the preference degree values ​​of each of the multiple initial response information.

21. The apparatus according to claim 17, wherein, The feedback information includes the target answer information provided by the object, and the object's preference for the target answer information is higher than its preference for the initial answer information. The adjustment module includes: The third adjustment submodule is used to adjust the parameters of the conversational model based on the question information, the initial answer information, and the target answer information of the training samples.

22. The apparatus according to claim 20 or 21, wherein, The adjustment module also includes: The second adjustment unit is used to adjust the parameters of the conversational model based on at least one of direct preference optimization, simple preference optimization, and proximal policy optimization.

23. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 11.

24. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 11.

25. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Medical language model training method, medical question and answer method and medical dialogue system

    CN117633166A

  • Recommended question obtaining method and device, electronic equipment and medium

    CN118260484A