Active chat robot construction method, device and medium
By building a user-centric chat quality evaluator and user background data set, and fine-tuning the chat robot model with iterative course learning methods, the problem that existing chat robots find it difficult to actively understand user chat preferences is solved, and the effect of improving user chat experience and human-computer interaction efficiency is achieved.
Patent Information
- Application Number
- CN202411848000.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-06
AI Technical Summary
Existing chatbots find it difficult to actively understand users’ chat preferences, which leads to users being uninterested in chat content and lacking a chat experience.
By building a user-centered chat quality evaluator, the chatbot's active perception of user background information and chat preferences is evaluated, and the chatbot model is fine-tuned through user background data sets and iterative course learning methods to improve its active perception of user background information and chat interests.
The chatbot has realized that the chatbot actively pays attention to the user's background information and chat interests, gives answers that meet users' chat preferences, improves the user's dialogue participation and satisfaction, and improves the human-computer interaction experience.
Smart Images

Figure CN119940407A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device and medium for constructing an active chat robot. Background Art
[0002] The open-domain conversation task aims to build a chatbot that can have a fluent and engaging natural language conversation with users on any topic. It has a wide range of application values in the fields of intelligent personal assistants, virtual customer service, and social robots.
[0003] In human-computer dialogue, users are naturally interested in dialogues that are related to their background and in line with their chat interests. This requires the chatbot to actively perceive the user's background information and chat interests during the chat process, and guide the conversation topic to the topic of interest to the user, that is, to conduct active user-centered chats. This is crucial to improving the user chat experience and improving the efficiency of human-computer interaction.
[0004] Existing large model-based methods rely on the powerful natural language understanding ability and huge knowledge reserve of large models to directly generate smooth and diverse dialogue responses based on user input. However, existing large model-based chatbots are usually trained to passively answer users' questions or perform tasks, and it is difficult to actively understand users' chat preferences. As a result, users are not interested in the chat content and lack chat experience. Another type of task-centric active dialogue system generates a topic path based on the preset target topic through a series of topic planning modules, and gradually guides the conversation to the preset target topic according to the path. Although this type of method can actively guide the topic, its preset target topic is usually a specific task, such as recommending a certain product to the user, rather than focusing on the user's own chat interests, it is difficult to effectively improve the user's chat experience. Summary of the invention
[0005] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a method, device, equipment and medium for constructing a user-centered active chat robot for open domain dialogue.
[0006] The first technical solution adopted by the present invention is:
[0007] A method for constructing an active chat robot comprises the following steps:
[0008] Build a user-centric chat quality evaluator to evaluate the chatbot's ability to proactively perceive user background information and chat preferences;
[0009] Build a user background dataset to allow the big model to play the role of user agents with different background identities and have multiple rounds of conversations with the chatbot;
[0010] Conversation corpus collection and iterative curriculum learning are used to generate high-quality conversation corpus, and the iterative curriculum learning method is used to fine-tune the large model corresponding to the chatbot to enhance the model's ability to proactively perceive user background information and chat preferences.
[0011] Furthermore, the user-centric chat quality evaluator works as follows:
[0012] When the chatbot has multiple rounds of conversations with the user, each time the chatbot outputs a response, the response and the conversation content between the chatbot and the user before the response are input into the chat quality evaluator; the chat quality evaluator quantitatively scores the chatbot's conversation based on three aspects: the user's interest level in the chatbot, the contextual relevance of the chatbot's response, and the value of the response.
[0013] Furthermore, the constructing of the user background data set includes:
[0014] From the ISCO-08 standard, multiple occupational categories are randomly selected. Each occupational category is used to construct a user background generation prompt, which is used to generate multiple specific user backgrounds later. In addition to prompting the big model to generate a specific occupational background for each user according to the occupational category, the user background generation prompt also prompts the big model to add the name, hobbies, personality, and educational background information for each user;
[0015] Generate prompts based on user backgrounds, let the big model generate user backgrounds, and each user background generation prompt is used to generate multiple specific user backgrounds. Each user background is a piece of text, including name, career, education, personality, and hobbies; after obtaining a preset number of user backgrounds, form a user background data set; divide the user background data set into training set, validation set, and test set.
[0016] Furthermore, the dialogue corpus collection and iterative curriculum learning are used to generate high-quality dialogue corpus, and use the iterative curriculum learning method to fine-tune the large model corresponding to the chatbot, including:
[0017] A1. Initialize the chatbot;
[0018] A2, initialize the user agent;
[0019] A3. Let the chatbot engage in a conversation with the user agent for a preset number of rounds and provide feedback through a user-centric chat quality evaluator to generate high-quality dialogue corpus;
[0020] A4, filtering the dialogue data generated in step A3;
[0021] A5. Use the filtered conversation data for supervised fine-tuning of the chatbot model;
[0022] A6. Replace the chatbot before fine-tuning with the fine-tuned chatbot, and repeat steps A3-A5 to perform multiple rounds of data collection and iterative course learning;
[0023] A7. After training in steps A1-A6, a chat robot is obtained that can actively perceive the user's background information and chat interests and can give valuable responses.
[0024] Furthermore, the process of each round of dialogue in step A3 is as follows:
[0025] In the first round of conversation, let the chatbot greet the user agent, and the user agent briefly introduces itself;
[0026] In subsequent rounds of conversation, the chatbot and the user agent conduct open-domain conversations of arbitrary content in turn. Each time the chatbot responds, the chat quality evaluator evaluates it based on three indicators: the user's interest level in the chatbot, the contextual relevance of the chatbot's response, and the value of the response.
[0027] After completing all rounds of dialogue, the background information of the user agent is replaced so that it plays a new user with another identity. The dialogue steps are repeated and new dialogues are generated until all user backgrounds in the training set of the user background dataset are used once, generating a high-quality dataset of the chatbot chatting with users with different backgrounds.
[0028] Furthermore, the chat quality evaluator will evaluate the chatbot based on three indicators: the user's interest level in the chatbot, the context relevance of the chatbot's response, and the value of the response, including:
[0029] If the scores of the three indicators are not lower than the preset scores, the robot's response quality is judged to meet the requirements, and the chatbot response is output to the user agent to continue the subsequent conversation;
[0030] If one or more indicators are lower than the preset score, the chat quality evaluator will inform the chat robot of the reason for the low score of the corresponding indicator, so that the chat robot can regenerate the dialogue response; if the scores of the regenerated responses are all not lower than the preset score, it is determined that the response quality meets the requirements and the chat robot response is output to the user agent, otherwise it will continue to be regenerated; if the number of regenerations reaches the preset number, regardless of the quality of the last dialogue response, the chat robot response will be output to the user agent to continue the subsequent dialogue.
[0031] Furthermore, the step A4 comprises:
[0032] For a round in a conversation, if the score of an indicator in the conversation is lower than the preset score, or when regeneration occurs, the number of improved indicators among the three indicators is lower than another preset score, the conversation data will be discarded; otherwise, the conversation data will be retained.
[0033] Furthermore, in step A6, after each round of iteration, the dialogue quality is verified on the verification set until the average score of all dialogue corpora no longer increases during each iteration; wherein the user agent utilizes the background information of the verification set, and after the chatbot gives a response for the first time, it is no longer repeatedly generated, but is directly evaluated and scored by the evaluator.
[0034] The second technical solution adopted by the present invention is:
[0035] An active chat robot construction device, comprising:
[0036] The evaluator building module is used to build a user-centric chat quality evaluator to evaluate the chatbot's ability to actively perceive user background information and chat preferences;
[0037] The dataset construction module is used to construct a user background dataset, which allows the large model to play the role of user agents with different background identities and conduct multiple rounds of conversations with the chatbot;
[0038] The iterative training module is used for dialogue corpus collection and iterative course learning to generate high-quality dialogue corpus and use iterative course learning methods to fine-tune the large model corresponding to the chatbot to enhance the model's ability to proactively perceive user background information and chat preferences.
[0039] The third technical solution adopted by the present invention is:
[0040] An electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement an active chat robot construction method as described above.
[0041] The fourth technical solution adopted by the present invention is:
[0042] A computer-readable storage medium, wherein at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement an active chat robot construction method as described above.
[0043] The fifth technical solution adopted by the present invention is:
[0044] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.
[0045] The beneficial effects of the present invention are as follows: the present invention enables the chat robot to actively pay attention to the user's background information and chat interests, and give answers that meet the user's chat preferences, thereby enhancing the user's conversation participation and satisfaction, and improving the human-computer interaction experience; and solves the problem of existing chat robots' passive replies and lack of active perception of user background information and chat interests. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0047] Figure 1 is a flowchart of the steps of a method for constructing an active chat robot in an embodiment of the present invention;
[0048] Figure 2 is a schematic diagram of an evaluation prompt design in an embodiment of the present invention;
[0049] Figure 3 is a schematic diagram of prompt of a chat robot system in an embodiment of the present invention;
[0050] Figure 4 Schematic diagram of the user agent system prompt in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0052] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0053] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0054] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.
[0055] Terminology explanation:
[0056] Prompt: A natural language text or instruction input to a large language model, which is used to guide the model to produce a specific output. In other words, prompt is a "hint" or "question" provided to the model, and the model generates an answer based on this prompt.
[0057] Example 1
[0058] like Figure 1 As shown, this embodiment provides a method for constructing a user-centric active chat robot for open domain dialogue, comprising the following steps:
[0059] S1. Build a user-centric chat quality evaluator to evaluate the chatbot's ability to actively perceive user background information and chat preferences.
[0060] In some embodiments, in order to evaluate the chatbot's ability to perceive the user's background and chat interests, this embodiment constructs a user-centric chat quality evaluator. Specifically, this embodiment carefully designs an evaluation prompt and uses it in ChatGPT or other large models with equivalent performance to form the evaluator. When the chatbot has multiple rounds of conversations with the user, each time the chatbot outputs a response, this embodiment inputs the response and the conversation content between the chatbot and the user before the response into the user-centric chat quality evaluator. The evaluator quantitatively scores the chatbot's conversation from three aspects: the user's interest level in the chatbot, the context relevance of the chatbot's response, and the value of the response.
[0061] For example, the design of the evaluation prompt is as follows Figure 2 As shown. This embodiment adopts a role-playing method to allow the evaluator to obtain the user's background information and play the role of the user, so that the evaluator can evaluate from the user's perspective. Then, this embodiment sets the evaluation scores of the user's interest level in the chatbot, the background relevance of the chatbot's response, and the value of the response to 1 to 5 points, and describes 1 point, 3 points, and 5 points in detail, respectively, so that the evaluator can more accurately judge how many points the current chatbot's response should get.
[0062] S2. Build a user background dataset to allow the large model to play the role of user agents with different background identities and engage in multiple rounds of conversations with the chatbot.
[0063] In order to train a chatbot that can actively engage in conversations with users of different backgrounds, this embodiment needs to generate a multi-round conversation dataset as training corpus. To this end, a user background dataset containing various user backgrounds is first constructed, and then the large model uses the user background information in the dataset to play the role of users with different identities (called user agents), thereby engaging in multi-round conversations with the chatbot and generating training corpus.
[0064] As an optional implementation, the process of constructing the user background dataset is as follows:
[0065] This embodiment randomly selects 40 major occupational categories from the ISCO-08 standard proposed by the International Labor Organization to generate specific occupations, and each major occupational category can be used to generate 20 specific user occupational backgrounds. In addition to the occupational background, the present invention also adds name, interests, hobbies, personality, and educational background information for each user. The present invention constructs a user background generation prompt, allowing GPT-4 to generate such user backgrounds, each user background is a text containing 50 to 100 words, including name, career, education, personality, and hobbies. A total of 800 user backgrounds are generated in the above manner, 500 of which are randomly divided into training sets, 100 are divided into verification sets, and 200 are divided into test sets.
[0066] S3, dialogue corpus collection and iterative course learning, is used to generate high-quality dialogue corpus, and use the iterative course learning method to fine-tune the large model corresponding to the chatbot to enhance the model's ability to actively perceive user background information and chat preferences.
[0067] After constructing the user-centric chat quality evaluator in step S1 and the user background dataset in step S2, iterative conversation corpus collection and curriculum learning are performed to train the chatbot.
[0068] In some embodiments, step S3 includes the following steps:
[0069] S31. Initialize the chat robot.
[0070] For example, the chatbot is initialized using Qwen1.5-32B-Chat or other large models with equivalent performance, using a role-playing method. Figure 3 Set its system prompt.
[0071] S32: Initialize the user agent.
[0072] For example, the user intelligence is initialized using Qwen1.5-72B-Chat or other large models with equivalent performance. The role-playing method is also used to select the user background in the training set of the user background data set. Figure 4 Set up its system prompt to let the big model play the role of a user with a specific identity background.
[0073] S33. Collection of feedback dialogue materials.
[0074] Let the chatbot and the user agent have a fixed number of rounds (e.g. 5 rounds) of dialogue, and provide feedback through a user-centric chat quality evaluator to generate high-quality dialogue corpus. The specific process of each dialogue is as follows:
[0075] S331. In the first round of dialogue, the chatbot first greets the user agent, and the user agent makes a brief self-introduction.
[0076] S332. In the subsequent rounds of dialogue, the chatbot and the user agent conduct an open domain dialogue of any content in turn. Whenever the chatbot responds, the user-centered chat quality evaluator will evaluate the three indicators of the user's interest level in the chatbot, the background relevance of the chatbot's response, and the value of the response. If the scores of the three indicators are not less than 4 points, it means that the quality of the robot's response meets the requirements, and the chatbot response is output to the user agent to continue the subsequent dialogue. If one or more indicators are less than 4 points, the evaluator will inform the chatbot of the reasons for the low scores of the corresponding indicators and make it regenerate the dialogue response. If the scores of the regenerated responses are all not less than 4 points, it means that the response quality meets the requirements, and the chatbot response is output to the user agent, otherwise it continues to regenerate. If the number of regenerations reaches 3 times, regardless of the quality of the last dialogue response, the chatbot response is output to the user agent to continue the subsequent dialogue.
[0077] After all rounds of dialogue are completed according to steps S331-S332, the background information of the user agent needs to be changed so that it plays a new user with another identity, and steps S331-S332 are repeated to continue generating new dialogues until all user backgrounds in the training set of the user background dataset are used once. In this way, a high-quality dataset of the chatbot chatting with users with different backgrounds can be generated.
[0078] S34. Dialogue data screening.
[0079] Course learning aims to gradually adapt the model to a task by first learning simple samples and then learning difficult samples. In the open domain dialogue task, different users have different communication difficulties, which is reflected in the quality of the final reply when the chatbot interacts with the user agent. Exemplarily, in this step, the dialogue generated in step S33 is screened. For a round in a certain dialogue, if there is an indicator score lower than 4 points in the round of dialogue, or when regeneration occurs, the number of improved indicators among the three indicators is less than 2, then it is considered that the user faced by the chatbot in this dialogue belongs to the type of user with communication difficulties, and the dialogue data is discarded. The dialogue corresponding to this round of dialogue is discarded. Otherwise, the dialogue data is retained.
[0080] S35. Model supervised fine-tuning.
[0081] The filtered data is used for supervised fine-tuning of the chatbot model. For example, the fine-tuning method may adopt a parameter-efficient fine-tuning method such as QLoRA.
[0082] S36. Replace the chatbot before fine-tuning with the chatbot after fine-tuning, and repeat steps S33-S35 to perform multiple rounds of data collection and iterative course learning.
[0083] After each round of iteration, the dialogue quality is verified on the validation set until the average score of all dialogue corpora no longer improves during each iteration. The verification process is similar to step S33, but the user agent uses the background information of the validation set, and after the chatbot gives its first response, it is no longer generated repeatedly, but is directly evaluated and scored by the evaluator.
[0084] S37. After training from step S31 to step S34, a chat robot can be obtained that can actively perceive the user's background information and chat interests and can give valuable replies.
[0085] In summary, existing open domain dialogue systems usually passively answer questions raised by users or perform tasks given by users, or guide topics according to some predetermined purpose. These methods usually do not actively pay attention to the user's own background identity and chat preferences, resulting in users being uninterested in the chat content and it is difficult to effectively improve the user's chat experience. In response to this problem, the present invention aims to enhance the chat robot's active perception of the user's own background and chat interests. Through a series of model fine-tuning methods, the chat robot's perception of users is improved, the user's chat interest is enhanced, and a more attractive and fascinating chat is brought about.
[0086] Example 2
[0087] This embodiment provides an active chat robot construction device, including:
[0088] The evaluator building module is used to build a user-centric chat quality evaluator to evaluate the chatbot's ability to actively perceive user background information and chat preferences;
[0089] The dataset construction module is used to construct a user background dataset, which allows the large model to play the role of user agents with different background identities and conduct multiple rounds of conversations with the chatbot;
[0090] The iterative training module is used for dialogue corpus collection and iterative course learning to generate high-quality dialogue corpus and use iterative course learning methods to fine-tune the large model corresponding to the chatbot to enhance the model's ability to proactively perceive user background information and chat preferences.
[0091] Since the device is an active chat robot construction device of an embodiment of the present invention, and the principle of solving the problem by the device is similar to that of the method, the implementation of the device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0092] Example 3
[0093] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 1 A method for building an active chatbot is shown.
[0094] It is understood that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0095] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor, but implemented separately through a chip.
[0096] Since the electronic device is an electronic device corresponding to the active chat robot construction method of an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0097] Example 4
[0098] The embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 1 A method for building an active chatbot is shown.
[0099] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0100] Since the storage medium is a storage medium corresponding to an active chat robot construction method of an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0101] Example 5
[0102] In some possible implementations, various aspects of the method of the embodiment of the present invention may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the method for building an active chat robot according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing various embodiments may be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0103] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0104] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0105] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made based on the essence of the content of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for constructing an active chat robot, characterized in that: The following steps are involved: Build a user-centric chat quality evaluator to evaluate the chatbot's ability to proactively perceive user background information and chat preferences; Build a user background dataset to allow the big model to play the role of user agents with different background identities and have multiple rounds of conversations with the chatbot; Conversation corpus collection and iterative curriculum learning are used to generate high-quality conversation corpus, and the iterative curriculum learning method is used to fine-tune the large model corresponding to the chatbot to enhance the model's ability to proactively perceive user background information and chat preferences.
2. The method for constructing an active chat robot according to claim 1, characterized in that: The user-centric chat quality evaluator works as follows: When the chatbot has multiple rounds of conversations with the user, each time the chatbot outputs a response, the response and the conversation content between the chatbot and the user before the response are input into the chat quality evaluator; the chat quality evaluator quantitatively scores the chatbot's conversation based on three aspects: the user's interest level in the chatbot, the contextual relevance of the chatbot's response, and the value of the response.
3. The method for constructing an active chat robot according to claim 1, characterized in that: The constructing of the user background data set includes: From the ISCO-08 standard, multiple occupational categories are randomly selected. Each occupational category is used to construct a user background generation prompt, which is used to subsequently generate multiple specific user backgrounds. In addition to prompting the big model to generate a specific occupational background for each user according to the occupational category, the user background generation prompt also prompts the big model to add the name, hobbies, personality, and educational background information for each user. Generate prompts based on user backgrounds, and let the big model generate user backgrounds. Each user background generation prompt is used to generate multiple specific user backgrounds; each user background is a piece of text, including name, career, education, personality, and hobbies; after obtaining a preset number of user backgrounds, form a user background data set; divide the user background data set into training set, validation set, and test set.
4. The method for constructing an active chat robot according to claim 1, characterized in that: The conversation corpus collection and iterative curriculum learning are used to generate high-quality conversation corpus and fine-tune the large model corresponding to the chatbot using the iterative curriculum learning method, including: A1. Initialize the chatbot; A2, initialize the user agent; A3. Let the chatbot engage in a conversation with the user agent for a preset number of rounds and provide feedback through a user-centric chat quality evaluator to generate high-quality dialogue corpus; A4, filtering the dialogue data generated in step A3; A5. Use the filtered conversation data for supervised fine-tuning of the chatbot model; A6. Replace the chatbot before fine-tuning with the fine-tuned chatbot, and repeat steps A3-A5 to perform multiple rounds of data collection and iterative course learning; A7. After training in steps A1-A6, a chat robot is obtained that can actively perceive the user's background information and chat interests and can give valuable responses.
5. The method for constructing an active chat robot according to claim 4, characterized in that: The process of each round of dialogue in step A3 is as follows: In the first round of conversation, let the chatbot greet the user agent and the user agent introduce itself; In subsequent rounds of conversation, the chatbot and the user agent conduct open-domain conversations of arbitrary content in turn. Each time the chatbot responds, the chat quality evaluator evaluates it based on three indicators: the user's interest level in the chatbot, the contextual relevance of the chatbot's response, and the value of the response. After completing all rounds of dialogue, the background information of the user agent is replaced so that it plays a new user with another identity. The dialogue steps are repeated and new dialogues are generated until all user backgrounds in the training set of the user background dataset are used once, generating a high-quality dataset of the chatbot chatting with users with different backgrounds.
6. The method for constructing an active chat robot according to claim 5, characterized in that: The chat quality evaluator evaluates the chatbot based on three indicators: the user's interest level in the chatbot, the contextual relevance of the chatbot's response, and the value of the response, including: If the scores of the three indicators are not lower than the preset scores, the robot's response quality is judged to meet the requirements, and the chatbot response is output to the user agent to continue the subsequent conversation; If one or more indicators are lower than the preset score, the chat quality evaluator will inform the chat robot of the reason for the low score of the corresponding indicator, so that the chat robot can regenerate the dialogue response; if the scores of the regenerated responses are all not lower than the preset score, it is determined that the response quality meets the requirements and the chat robot response is output to the user agent, otherwise it will continue to be regenerated; if the number of regenerations reaches the preset number, regardless of the quality of the last dialogue response, the chat robot response will be output to the user agent to continue the subsequent dialogue.
7. The method for constructing an active chat robot according to claim 4, characterized in that: The step A4 comprises: For a round in a conversation, if the score of a certain indicator in the conversation is lower than the preset score, or when regeneration occurs, the number of improved indicators among the three indicators is lower than another preset score, then the conversation data will be discarded; Otherwise, the conversation data will be retained.
8. The method for constructing an active chat robot according to claim 4, characterized in that: In step A6, after each round of iteration, the dialogue quality is verified on the validation set until the average score of all dialogue corpora no longer increases during each iteration; wherein, the user agent uses the background information of the validation set, and after the chatbot gives a response for the first time, it no longer generates a response repeatedly, but is directly evaluated and scored by the evaluator.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.