Dialogue information reply processing method and device, equipment, storage medium and product
By retrieval processing and category tag generation during the dialogue information generation process, the low answer accuracy caused by player input non-rights and non-class questions is solved, and the accuracy of reply information and conversation experience are improved.
Patent Information
- Application Number
- CN202510009177.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
During the conversation between the player and the Big Language Model (LLM), the player may enter dialogue information that does not belong to the right and wrong questions, resulting in low accuracy of the generated answers and affecting the game experience.
By obtaining the conversation information input by the target object for the question, searching process is performed to obtain a matching reference conversation example, initial reply information is generated and target reply information is generated based on the category tag.
The accuracy of the reply information generated based on the conversation information is improved, the conversation experience is improved, and the generated reply information is more in line with business requirements.
Smart Images

Figure CN119938839A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, equipment, storage medium and product for processing reply of dialogue information. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, large language models (LLMs) have been cleverly applied to answering questions due to their powerful natural language processing capabilities. For example, in yes-no guessing games, the core is to ask yes-no questions and get yes or no answers. Players gradually reveal the answer through multiple conversations with LLMs. For example, the Turtle Soup game is a representative of yes-no guessing games. Players can ask LLM yes-no questions based on a given suspenseful story description and infer the truth of the story based on the answers generated by LLM.
[0003] However, during the conversation between the player and the LLM, the player may enter dialogue information that is not a yes or no question in an attempt to obtain more clues about the plot from the answer generated by the LLM. In this case, the LLM may still answer the player's input as a yes or no question, making the generated answer less accurate, which may affect the player's gaming experience to a certain extent. Summary of the invention
[0004] The embodiments of the present application provide a method, apparatus, device, storage medium and product for processing reply to conversation information, which are conducive to improving the accuracy of reply information generated based on conversation information and improving the conversation experience to a certain extent.
[0005] In a first aspect, an embodiment of the present application provides a method for processing a reply to a conversation message, the method comprising:
[0006] Obtaining the dialogue information input by the target object in response to the question;
[0007] Performing a search process based on the dialogue information to obtain a reference dialogue example of the dialogue information, wherein text features of the reference dialogue example match text features of the dialogue information;
[0008] Based on the dialogue information, the question and the reference dialogue example, generating initial response information corresponding to the dialogue information, wherein the initial response information includes a category label of the dialogue information;
[0009] Based on the category label, target reply information corresponding to the dialogue information is generated.
[0010] In a second aspect, an embodiment of the present application provides a device for processing a reply to a conversation message, the device comprising:
[0011] An acquisition unit, used to acquire dialogue information input by the target object in response to the question;
[0012] a processing unit, configured to perform a search process based on the dialogue information to obtain a reference dialogue example of the dialogue information, wherein a text feature of the reference dialogue example matches a text feature of the dialogue information;
[0013] A generating unit, configured to generate initial response information corresponding to the dialogue information based on the dialogue information, the question and the reference dialogue example, wherein the initial response information includes a category label of the dialogue information;
[0014] The generating unit is further configured to generate target reply information corresponding to the dialogue information based on the category label.
[0015] In a possible implementation, the processing unit is configured to perform retrieval processing based on the dialogue information to obtain a reference dialogue example of the dialogue information, specifically to:
[0016] Obtaining the associated conversation information of the conversation information;
[0017] extracting text features of the conversation information and text features of the associated conversation information;
[0018] Calculating the similarity between the text features of the dialogue information and the text features of each dialogue example in the dialogue example library, and the similarity between the text features of the associated dialogue information and the text features of each dialogue example;
[0019] A set number of dialogue examples are selected from the dialogue example library in descending order of similarity as the reference dialogue examples.
[0020] In a possible implementation, the processing unit is used to obtain the associated dialog information of the dialog information, and is specifically used to perform at least one of the following steps:
[0021] Acquire a historical conversation record of the target object before the conversation information, and determine the historical conversation record as the associated conversation information;
[0022] The conversation information is input into a text information expansion model, and expanded conversation information generated by the text information expansion model based on the conversation information is obtained, and the expanded conversation information is determined as the associated conversation information.
[0023] In a possible implementation manner, the generating unit is used to generate target reply information corresponding to the dialogue information based on the category label, specifically for:
[0024] Inputting the category label into a stylized rewriting model, obtaining first reply information of a specified text style generated by the stylized rewriting model based on the category label, and determining the first reply information as the target reply information;
[0025] or,
[0026] The category label and the supplementary content included in the initial reply information are input into the stylized rewriting model, and the second reply information of the specified text style generated by the stylized rewriting model based on the category label and the supplementary content is obtained, and the second reply information is determined as the target reply information.
[0027] In a possible implementation, the dialogue information is text information input by the target object in the process of answering the question; the category label includes a control instruction; and the device further includes:
[0028] An exit unit is used to exit the answering process of the question if the control instruction is used to instruct to exit the answering process.
[0029] The generating unit is used for generating a new question if the control instruction is used to instruct to change the question.
[0030] In a possible implementation, the generating unit is used to generate initial reply information corresponding to the dialogue information based on the dialogue information, the question and the reference dialogue example, specifically to:
[0031] Obtaining the historical conversation record of the target object before the conversation information;
[0032] Filling the dialogue information, the question, the reference dialogue example and the historical dialogue record into a set prompt word template to obtain a target prompt word;
[0033] The target prompt word is input into a large language model, and reply information generated by the large language model based on the answer strategy and reply constraint information associated with the prompt word template is obtained to obtain the initial reply information.
[0034] In a possible implementation manner, the generating unit is used to generate target reply information corresponding to the dialogue information based on the category label, specifically for:
[0035] If the category tag indicates that the dialogue information hits key information in the answer to the question, then obtaining the hit key information of the target object for the question;
[0036] If it is determined based on the hit key information that the target object has hit all the key information in the answer, then the category label is updated to a specified category label;
[0037] Based on the specified category label, the target reply information is generated.
[0038] In a possible implementation manner, the device further includes:
[0039] an adding unit, configured to add a reply strategy of the dialogue information to the reply constraint information to obtain an updated prompt word template if it is detected that the dialogue turn of the dialogue information is greater than a set turn threshold, wherein the reply strategy is used to instruct the large language model to output prompt information of an answer to the question;
[0040] The generating unit is used to fill the dialogue information, the question, the reference dialogue example and the historical dialogue record into a set prompt word template to obtain a target prompt word, specifically for:
[0041] The dialogue information, the question, the reference dialogue example and the historical dialogue record are filled into the updated prompt word template to obtain the target prompt word.
[0042] In a possible implementation, the acquisition unit is further used to acquire training questions and multi-round dialogue constraint information, where the multi-round dialogue constraint information is used to constrain the answering process of the training questions;
[0043] A filling unit, used to fill the training question and the multi-round dialogue constraint information into the prompt word template to obtain a training dialogue prompt word;
[0044] The acquisition unit is further used to input the training dialogue prompt words into the dialogue generation model, and acquire the dialogue record generated by the dialogue generation model based on the training dialogue prompt words;
[0045] A training unit is used to train the large language model to be trained based on the conversation record to obtain the large language model.
[0046] In a possible implementation, the conversation record includes multiple rounds of conversation information and reply information of each round of conversation information, wherein the historical conversation records of different rounds of conversation information are different; the training unit is used to train the large language model to be trained based on the conversation record to obtain the large language model, specifically for:
[0047] Performing retrieval processing based on the dialogue information of each round and the historical dialogue records of the dialogue information of each round to obtain training reference dialogue examples of the dialogue information of each round;
[0048] Filling the dialogue information of each round, the historical dialogue record of each round, the training reference dialogue example of each round, and the training question into the prompt word template to obtain the training prompt word;
[0049] Inputting the training prompt words into the large language model to be trained, and obtaining the training response information of each round of dialogue information output by the large language model to be trained;
[0050] Based on the difference between the training reply information of each round of dialogue information and the reply information of each round of dialogue information, the large language model to be trained is trained to obtain the large language model.
[0051] In a possible implementation manner, the device further includes:
[0052] A decomposition unit, configured to decompose the answer to the training question into at least one key information, and generate supplementary explanation content and question examples for each key information;
[0053] The generating unit is further configured to generate supplementary explanation content for the training question based on the training question and the answer;
[0054] A combining unit, configured to combine the training question, the at least one key information, the supplementary explanation content of the training question, the supplementary explanation content of each key information and the question example into a target training question;
[0055] The filling unit is used to fill the training question and the multi-round dialogue constraint information into the prompt word template to obtain the training dialogue prompt word, which is specifically used to:
[0056] The target training question and the multi-round dialogue constraint information are filled into the prompt word template to obtain the training dialogue prompt words.
[0057] In a possible implementation, the acquisition unit is further used to acquire training questions and designated category dialogue information for the training questions;
[0058] The filling unit is further used to fill the training question and the designated category dialogue information into the prompt word template to obtain the training dialogue prompt word;
[0059] The acquisition unit is further configured to input the training dialogue prompt words into a dialogue generation model, and acquire reply information of the designated category dialogue information generated by the dialogue generation model based on the training dialogue prompt words;
[0060] The training unit is further used to train the large language model to be trained based on the designated category dialogue information and the reply information of the designated category dialogue information to obtain the large language model.
[0061] In a possible implementation, the number of the designated category dialogue information is multiple; the training unit is used to train the large language model to be trained based on the designated category dialogue information and the reply information of the designated category dialogue information to obtain the large language model, specifically for:
[0062] Performing retrieval processing based on partial dialogue information selected from multiple designated categories of dialogue information to obtain training reference dialogue examples for the partial dialogue information;
[0063] Filling the partial dialogue information, the training reference dialogue example of the partial dialogue information, and the training question into the prompt word template to obtain the training prompt word;
[0064] Inputting the training prompt words into the large language model to be trained, and obtaining training reply information of the part of the dialogue information output by the large language model to be trained;
[0065] Based on the difference between the reply information of the part of the dialogue information and the training reply information of the part of the dialogue information in the reply information of the designated category of dialogue information, training the large language model to be trained to obtain the large language model;
[0066] The reference dialogue example and the training reference dialogue example are at least one dialogue example in the dialogue examples consisting of the designated dialogue information other than the partial dialogue information and the reply information of the designated dialogue information in the designated category dialogue information.
[0067] In a possible implementation, the acquisition unit is further used to acquire training questions, training prompt dialogue information, and reply prompt information of the training prompt dialogue information, wherein the training prompt dialogue is used to acquire prompt information of answers to the training questions;
[0068] The processing unit is further used to perform retrieval processing based on the training prompt dialogue information to obtain a training reference dialogue example of the training dialogue prompt information;
[0069] The filling unit is further used to fill the training dialogue prompt information, the training reference dialogue example of the training dialogue prompt information and the training question into the prompt word template to obtain the training prompt word;
[0070] The acquisition unit is further used for inputting the training prompt word into the large language model to be trained, and acquiring the training reply prompt information of the training dialogue prompt information output by the large language model to be trained;
[0071] The training unit is further used to train the large language model to be trained based on the difference between the reply prompt information of the training dialogue prompt information and the training reply prompt information of the training dialogue prompt information to obtain the large language model.
[0072] In a third aspect, an embodiment of the present application provides an electronic device, which includes one or more processors; and a memory for storing one or more computer programs, so that when the one or more computer programs are executed by the one or more processors, the electronic device implements the method for replying to conversation information of the first aspect.
[0073] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is executed on a computer, the computer executes the method for replying to the conversation information of the first aspect.
[0074] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, they implement the method for replying to dialogue information as in the first aspect.
[0075] In the technical solutions provided by some embodiments of the present application, by obtaining the dialogue information input by the target object in response to the question, and performing retrieval processing based on the dialogue information, a reference dialogue example of the dialogue information is obtained, and the text features of the reference dialogue example match the text features of the dialogue information. Then, based on the dialogue information, the question and the reference dialogue example, the initial reply information corresponding to the dialogue information is generated, and the initial reply information includes the category label of the dialogue information, and then based on the category label, the target reply information corresponding to the dialogue information is generated. It can be seen that by retrieving the dialogue information before generating the reply information, the reply information is assisted in generating the reply information based on the retrieved reference dialogue example, so that the generated reply information is more consistent with the reply required by the business, which is conducive to improving the accuracy of the generated reply information. In addition, by determining the category label of the dialogue information and further generating the reply information based on the category label, it is possible to reply and process dialogue information of different categories, which can improve the dialogue experience to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0077] Figure 1 It is a schematic diagram of the architecture of a system for processing a reply to a conversation message provided in an embodiment of the present application;
[0078] Figure 2 It is a schematic diagram of the architecture of another system for processing a reply to a conversation message provided in an embodiment of the present application;
[0079] Figure 3 It is a flowchart of a method for replying to a conversation message provided in an embodiment of the present application;
[0080] Figure 4 is a schematic diagram of a dialogue interface provided in an embodiment of the present application;
[0081] Figure 5 This is a timing diagram of a retrieval process for dialog information provided by an embodiment of the present application;
[0082] Figure 6 It is a timing diagram of raw data preparation provided by an embodiment of the present application;
[0083] Figure 7 It is a timing diagram of LLM training provided by an embodiment of the present application;
[0084] Figure 8 It is a timing diagram of an LLM-based reasoning phase provided in an embodiment of the present application;
[0085] Fig. 9 A schematic diagram of the structure of a dialog message reply processing device provided in an embodiment of the present application;
[0086] Fig.10 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0087] It should be noted in advance that, in order to enable those skilled in the art to better understand the technical solutions proposed in the embodiments of the present application, the embodiments of the present application will be combined with one or more drawings to clearly and completely describe the implementation of the technical solutions proposed in the embodiments of the present application. In addition, the various drawings shown in the embodiments of the present application are only exemplary illustrations, and for example, the execution order of the various steps in the drawings can be adaptively adjusted according to the actual application scenario. In addition, in the embodiments of the present application, the block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or these functional entities can be implemented in one or more hardware modules or integrated circuits, or these functional entities can be implemented in different networks and / or processor devices and / or microcontroller devices.
[0088] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0089] It should be noted that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0090] With the rapid development of AI technology, LLM has been cleverly applied to answering questions. For example, questions can be true or false puzzles, and the answers to questions can be applied to true or false guessing games. In a true or false guessing game with LLM, users can have multiple conversations with LLM to gradually reveal the answer. For example, users can ask LLM true or false questions, and LLM will answer the user's questions with yes or no based on the puzzle of the guessing game, and then the user can infer the answer based on the response generated by LLM.
[0091] However, during the conversation between the user and the LLM, the user may enter dialogue information that is not a yes or no question in an attempt to obtain more clues about the answer from the answer generated by the LLM. In this case, the LLM may still answer the user's input as a yes or no question, such as outputting a yes or no reply. In this case, the LLM may not be able to accurately understand the requirements of the dialogue information, thereby generating reply information that does not meet the business requirements, and may even cause an error in which the reply information does not meet the business requirements. Even if the user adjusts the input dialogue information, or the user fine-tunes the input instructions including the dialogue information input into the LLM, the quality of the generation cannot be guaranteed, resulting in a low accuracy of the generated reply information, which may affect the player's gaming experience to a certain extent.
[0092] Based on this, the embodiment of the present application provides a reply processing scheme for dialogue information, which is to retrieve the dialogue information with reference to the dialogue example before generating the reply information based on the dialogue information, and by retrieving the dialogue examples corresponding to different reply strategies, the reply information can be generated based on the dialogue example and the dialogue information, which is conducive to generating the reply information that is highly consistent with the reply required by the business, and improving the accuracy of the generated reply information. In addition, by first determining the category label of the dialogue information, and then generating the final reply information according to the category label, compared with directly obtaining the reply information, the category of the dialogue information can be identified to reply based on the corresponding reply strategy, which is also conducive to improving the accuracy of the answer and can improve the dialogue experience to a certain extent.
[0093] In order to better understand the solutions of the embodiments of the present application, the relevant terms and concepts that may be involved in the embodiments of the present application are first introduced below.
[0094] 1. Language Model (LM): LM is a model used to model the probability distribution of natural language. It is a model based on machine learning or deep learning technology. It is trained by analyzing text sequences to predict the probability distribution of subsequent words or characters. It can not only understand the grammatical structure and expression patterns in natural language, but also generate logically coherent and semantically accurate natural language text.
[0095] 2. Pretrained Language Model (PTM): PTM refers to a neural network model that is pre-trained on a large-scale dataset, such as the Bidirectional Encoder Representations from Transformers (BERT), Generative Pre-trained Transformer (GPT), etc. PTM aims to deeply understand and capture the contextual representation of language by performing unsupervised learning on a large number of text datasets. PTM is able to capture (extract) the deep features (complex features) of natural language and provide rich semantic information for subsequent applications. When applied to specific natural language processing tasks, the task performance can be significantly improved by fine-tuning it.
[0096] 3. LLM: LLM refers to a language model with a very large number of parameters, usually containing billions to hundreds of billions of parameters, such as GPT-4, ChatGPT, Pathways Language Model (PaLM), etc. The training of these models relies on massive data sets and powerful computing resources. Thanks to its huge scale, LLM usually has excellent generalization ability and can achieve outstanding performance in a variety of natural language processing tasks through few-shot learning or zero-shot learning, covering text classification, language translation, question-answering systems and other fields. In addition, LLM can also generate high-quality natural language text, which is suitable for a variety of text generation tasks such as article writing, dialogue interaction, and poetry creation.
[0097] 4. True or False Guessing Games: True or False Guessing Games are intellectual games designed to train players' logical thinking and judgment. These games usually involve judging "yes" or "no" to a series of statements or questions from players, that is, judging whether they are correct or wrong. Players need to gradually approach the answer based on these feedbacks, combined with their own knowledge or logical reasoning. This type of game can be in various forms, for example, it can be in text form, or combined with graphics or multimedia elements to enhance interactivity and fun.
[0098] Among them, the Turtle Soup game is a representative of true or false guessing games. The questioner gives an incomplete and suspenseful story, called the "noodle soup", and at the same time keeps a corresponding complete story, called the "soup base". As the answerer, the player can get clues by asking various possible questions, and the questioner can usually only answer "yes" or "no", that is, he can only answer the player's yes or no questions. The answerer (player) needs to obtain clues through a series of yes and no questions, use logic to deduce the whole picture of the whole story, and finally reveal the "soup base".
[0099] Among them, the guessing game using LLM can be called AI guessing game. In the AI guessing game, the player can interact with the LLM through dialogue. The player can input dialogue information, and the LLM can reply to the player's dialogue information to advance the progress of the guessing game.
[0100] 5. AI role-playing: AI role-playing refers to letting AI interact with users by playing a specific role, interpreting stories or providing emotional value. That is, using AI technology to simulate the behavior and dialogue of a specific role, which can not only provide an immersive interactive experience, but also improve the user experience through learning and adaptation. This technology can be applied to a variety of scenarios, including online games, virtual assistants, customer service, etc.
[0101] 6. Prompt: It can also be called prompt words, etc. It is a key tool for interacting with AI models, especially in LLM applications. It is a piece of text (discrete prompt) or a numerical vector (continuous prompt) that is used to stimulate LLM to generate the expected output. Prompt words in text form are usually carefully designed by humans to clearly indicate the tasks that the model needs to perform or the type of content expected to be produced. By providing specific guidance, prompts can not only accelerate the learning process of LLM for new tasks, but also accurately guide the content and style of the model output to ensure that it meets the needs of specific application scenarios.
[0102] 7. Text Embedding Model: refers to a machine learning model that converts text (such as words, phrases, sentences or paragraphs) into numerical vectors. These vectors can capture the semantic information in the text, so that semantically similar texts have similar vector representations in the embedding space. The core advantage of the text embedding model is that it can map high-dimensional text information to a low-dimensional embedding space while retaining the key features and semantic information of the original data. By reducing the dimension, not only the computational efficiency of the model is improved, but also the accuracy of the model prediction is enhanced.
[0103] Among them, the working principle of the text embedding model is to transform each word or phrase into a real number vector by learning the distributed representation of text data. The positions of these vectors in the embedding space reflect their semantic relationships in the original text. For example, semantically similar words are mapped to similar positions in the embedding space. This method enables the text embedding model to capture the complex patterns and deep semantic relationships of text data, thereby achieving more intelligent and accurate text processing. In addition, the numerical vectors generated by the text embedding model can be easily stored and retrieved, which has important practical value for processing large-scale text data. Embedding models are widely used in various natural language processing tasks, such as information retrieval, text classification, sentiment analysis, etc., and have played a key role.
[0104] 8. Vector Search Library: Vector Search Library is a well-designed software tool built for efficient storage and retrieval of vector data. In the field of data science and machine learning, various data types such as images, text, and audio can be converted into numerical vectors in high-dimensional space. The retrieval of these vectors, especially when finding the most similar elements to a specific item, is highly dependent on vector search technology.
[0105] In order to speed up the retrieval process, the vector retrieval system integrates a variety of efficient algorithms, such as K-dimensional tree (KD tree), approximate nearest neighbor (ANN) algorithm, etc. These algorithms can quickly locate the element (item) closest to the target query vector in a huge data set. Even in the face of large-scale data sets, the vector retrieval system can maintain efficient query performance, which is crucial for applications that require real-time response and tasks that process large-scale data.
[0106] The design advantage of vector search libraries is that they can optimize query speed while maintaining high search accuracy. This ability makes them play a key role in recommendation systems, computer vision, natural language processing and other fields. For example, vector search libraries such as Facebook AI Similarity Search (FAISS) and ChromaDB provide strong support in these applications, ensuring the ability to quickly obtain the most relevant results from massive data.
[0107] 9. Retrieval Augmented Generation (RAG): RAG is a technology that uses information from private or proprietary data sources to assist in text generation. RAG retrieves relevant information from authoritative, predetermined knowledge sources, and can better control the generated text output without retraining the model to significantly improve the quality and accuracy of the generated text, while improving the relevance of the search experience. In the process, users can also gain a deeper understanding of how LLM generates responses.
[0108] RAG's core advantage lies in its ability to leverage a wide range of information sources, including not only the latest information on the Internet (even information that was not used when training the LLM), but also proprietary business background information, or confidential internal documents of the company. By integrating these additional information resources, RAG is able to support generative AI systems to use external information sources to generate more accurate and contextual answers, which is of great value in tasks such as answering questions and content generation. In other words, RAG implements efficient search and retrieval methods to better capture user intent and provide highly relevant results, enabling generative AI systems to dynamically access the latest external information sources to generate responses that are more in line with current situations and needs.
[0109] Based on the above description, please refer to Figure 1 , Figure 1 is a schematic diagram of the architecture of a system for processing a dialogue message reply provided by an embodiment of the present application, such as Figure 1 As shown, the conversation information reply processing system includes a user device 101 and a conversation information reply device 102. The user device 101 and the conversation information reply device 102 can be directly or indirectly connected in a wired or wireless manner. It should be noted that, Figure 1 The number and form of devices shown are for example only and do not constitute a limitation on the embodiments of the present application. In some embodiments, there may be multiple user devices 101. In some embodiments, there may be multiple conversation information reply devices 102 that establish a connection with the user device 101.
[0110] Among them, the user device 101 can be a device used by the user, which is an electronic device that can be used to install and run a platform or software for answering questions, such as a guessing game platform or software. The user as a player can input dialogue information for asking questions (such as game puzzles) based on the installed platform or software. Specifically, the user device 101 can include an input device, such as a touch display, a keyboard, a microphone, etc., and the user can input dialogue information for asking questions based on the input device. In some embodiments, the user device 101 can also include an output device, such as a display, a speaker, etc., which can be used to output the dialogue information and the reply information corresponding to the dialogue information, such as the target reply information.
[0111] The conversation information reply device 102 may be an electronic device that provides a guessing game platform (software), which may be deployed with an LLM that replies to conversation information, or may establish a communication connection with other electronic devices that are deployed with an LLM that replies to conversation information, so as to obtain the function of the LLM to reply based on the conversation information. Taking the application of the answer to the question in the AI guessing game as an example, Figure 1 The system composed of the user device 101 and the dialogue message reply device 102 shown can also be called an AI guessing game dialogue system. The user can conduct an AI guessing game dialogue based on the guessing game platform running in the user device 101 and the LLM deployed in other electronic devices deployed in the dialogue message reply information 102 or establishing a communication connection with it.
[0112] The reply processing of the dialogue information provided in the present application can be performed by the dialogue information reply device 102, the user device 101 can be a terminal device, and the dialogue information reply device 102 can be a terminal device or a server. Among them, the terminal device includes but is not limited to: smart phones (such as Android phones, IOS phones, etc.), tablet computers, portable personal computers, mobile Internet devices (Mobile Internet Devices, MID), intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, wearable devices, etc., and the embodiments of the present application do not limit this. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms, and the embodiments of the present application do not limit this.
[0113] The general process of the reply processing method of the dialogue information provided by this application is as follows:
[0114] The dialogue information reply device 102 can obtain the target object, such as Figure 1 The user shown in the figure inputs the dialogue information for asking questions through the guessing game platform running in the user device 101. Then, the dialogue information replying device 102 can perform retrieval processing based on the dialogue information input by the user to obtain a reference dialogue example of the dialogue information, wherein the text features of the retrieved reference dialogue example match the text features of the dialogue information. Afterwards, the dialogue information replying device 102 can generate initial reply information corresponding to the dialogue information based on the dialogue information, the question asked and the reference dialogue example, which can be the initial reply information generated based on the dialogue information, the question asked and the reference dialogue example using the natural language processing capability of LLM, and the initial reply information can include the category label of the dialogue information. Finally, the dialogue information replying device 102 can generate the target reply information corresponding to the dialogue information based on the category label. After generating the target reply information, the dialogue information replying device 102 can send the target reply information to the user device 101, so that the user can make inferences based on the target reply information to answer the question asked.
[0115] In some embodiments, the conversation information reply device 102 can also obtain training data for training LLM, such as conversation records, in which each round of conversation information and corresponding reply information in multiple rounds of conversation information can be used as training data. For example, it can include conversation information of a specified category and its reply information. For example, it can include training prompt conversation information and its corresponding reply prompt information.
[0116] In some embodiments, the conversation information replying device 102 may also train the LLM based on the acquired training data. The training process may also be referred to as a process of fine-tuning the LLM.
[0117] thus, Figure 1 A solution for a specific rule-based dialogue based on LLM and user interaction to complete the AI guessing game dialogue is proposed. LLM can be used to conduct dialogue based on specific rule-based gameplay. Based on the dialogue information input by the player, LLM makes a response that meets the rule requirements, and generates a final response based on the response output by LLM until the player reaches the ending specified by the rules during the game, such as success or failure.
[0118] Please also read Figure 2 , Figure 2 FIG. 1 is a schematic diagram of the architecture of another system for processing a reply to a conversation message provided in an embodiment of the present application. Figure 2 As shown, the conversation information reply processing system includes a player simulation device 201 and a conversation information reply device 202. The player simulation device 201 and the conversation information reply device 202 can be directly or indirectly connected in a wired or wireless manner. It should be noted that, Figure 2 The number and form of the devices shown are only for example and do not constitute a limitation on the embodiments of the present application. In some embodiments, the number of the player simulation device 201 and the conversation information reply device 202 can be multiple.
[0119] Among them, the player simulation device 201 can be an electronic device that generates dialogue information for asking questions for the user, such as the player simulation device 201 can be an electronic device that installs and runs the guessing game platform or software. The player simulation device 201 can input the generated dialogue information based on the installed platform or software. In the player simulation device 201, a machine learning model for generating dialogue information for asking questions can be deployed, such as an LLM for generating dialogue information, and a communication connection can also be established with other electronic devices that are deployed with a machine learning model (such as an LLM) for generating dialogue information to obtain the dialogue information generated by other electronic devices and input it into the guessing game platform or software.
[0120] In some embodiments, since the dialogue information may be a dialogue of a player in a round during the guessing game, the player simulation device 201 has the ability to generate dialogue information of the next dialogue round based on the reply information generated by the dialogue information reply device 202 to the current dialogue round. For example, the player model deployed by the player simulation device 201 can generate dialogue information of the next dialogue round based on the dialogue record of the current dialogue round or based on the historical dialogue record.
[0121] In some embodiments, the player simulation device 201 may include an input device, such as a touch display screen, a keyboard, a microphone, etc. If the player simulation device 201 is deployed with a machine learning model for generating dialogue information, the user may input a prompt based on the input device, so that the machine learning model in the player simulation device 201 generates dialogue information based on the user's instructions.
[0122] In some embodiments, the player simulation device 201 may also include an output device, such as a display screen, a speaker, etc., which can be used to output the dialogue information and reply information corresponding to the dialogue information, such as target reply information.
[0123] The detailed description of the dialogue information reply device 202 can be found in Figure 1 The detailed description of the dialogue information replying device 102 is omitted here.
[0124] The reply processing of the dialogue information provided in the present application can be executed by the dialogue information reply device 202. The player simulation device 201 and the dialogue information reply device 202 can both be terminal devices or servers. The terminal devices can include but are not limited to: smart phones (such as Android phones, IOS phones, etc.), tablet computers, portable personal computers, MIDs, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, wearable devices, etc., which are not limited to the embodiments of the present application; the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms, which are not limited to the embodiments of the present application.
[0125] The general process of the reply processing method of the dialogue information provided by this application is as follows:
[0126] The dialogue information reply device 202 can obtain the target object, such as Figure 2 The player simulation device 201 shown in the figure inputs the dialogue information for asking questions. Furthermore, the dialogue information reply device 202 can perform retrieval processing based on the text features of the dialogue information generated by the player simulation device 201 to obtain a reference dialogue example of the dialogue information, and the text features of the reference dialogue example match the text features of the dialogue information. Afterwards, the dialogue information reply device 202 can generate initial reply information corresponding to the dialogue information based on the dialogue information, the question asked and the reference dialogue example, which can be the initial reply information generated by LLM based on the dialogue information, the question asked and the reference dialogue example, and the initial reply information can include the category label of the dialogue information. Finally, the dialogue information reply device 202 can generate the target reply information corresponding to the dialogue information based on the category label.
[0127] At this point, this round of conversation ends, and the conversation information reply device 202 can send the target reply information to the player simulation device 201, and the player simulation device 201 can start the next round of conversation, that is, the player simulation device 201 can generate new conversation information and send it to the conversation information reply device 202. For example, the player simulation device 201 can send the new conversation information to the conversation information reply device 202 through the installed guessing game platform or guessing game software.
[0128] thus, Figure 2A solution for completing AI guessing game dialogue based on LLM and electronic devices interacting is given, which can be used to test answers to questions (such as AI guessing game dialogue) and can also be used to generate training dialogue data for fine-tuning LLM. It is understandable that testing AI guessing game dialogue can also be applied to testing AI guessing games with new functions, such as simulating players playing AI guessing games with new game mechanisms. By simulating target objects (such as players), decisions can be made quickly based on data, shortening the product development cycle to evaluate the performance of products in different scenarios, which is conducive to improving development efficiency and reducing development costs.
[0129] based on Figure 1 and Figure 2 In the system shown, the conversation information reply device retrieves the conversation information before generating the reply information, and assists in generating the reply information based on the retrieved reference conversation examples, so that the generated reply information is more consistent with the reply required by the business, which is conducive to improving the accuracy of the generated reply information. In addition, by determining the category label of the conversation information and further generating the reply information based on the category label, it is possible to reply and process conversation information of different categories, which can also improve the accuracy of the reply information to a certain extent.
[0130] In one implementation, the above-mentioned dialogue information, reference dialogue examples, text features of reference dialogue examples, text features of dialogue information, initial reply information corresponding to the dialogue information, category labels of dialogue information, and target reply information of dialogue information can all be stored in the blockchain, which can prevent these information from being tampered with. Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. It is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block.
[0131] It is understandable that the embodiments of the present application are described as follows Figure 1 and Figure 2 The reply processing system for conversation information shown is intended to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. A person of ordinary skill in the art can appreciate that with the evolution of the system architecture and the emergence of new business scenarios, the technical solution provided by the embodiment of the present application is equally applicable to similar technical problems.
[0132] Based on the above-mentioned conversation information reply processing system, the embodiment of the present application provides a conversation information reply processing method. The conversation information reply processing method described in the embodiment of the present application can be executed by an electronic device, which can be Figure 1 The dialogue information replying device 102 in the dialogue information replying processing system shown in FIG. 1 may also be Figure 2 The dialog information replying device 202 in the dialog information replying processing system shown in FIG. Figure 3 , Figure 3 1 is a flow chart of a method for processing a reply to a conversation message provided in an embodiment of the present application. The method for processing a reply to a conversation message includes the following steps S301 to S304:
[0133] S301: Obtaining dialogue information input by the target object in response to a question.
[0134] In an embodiment of the present application, the target object may refer to the user involved in the answer to the question, the question may refer to the puzzle in the guessing game, the answer to the question refers to the answer to the guessing game, and the target object may be understood as the player of the guessing game. The guessing game may refer to an AI guessing game, that is, a guessing game using LLM, specifically, it may refer to a right or wrong guessing game using LLM, and the game process may also be referred to as the answer process for the question. The target object may also be understood as a user participating in the AI guessing game, and the user may complete the game task or personal goal in the guessing game, and is the experiencer and user of the game. The target object may also refer to a specific electronic device, which may simulate the behavior logic of the player in the guessing game to complete the game task or set goal, and may be applied to the test scenario of the AI guessing game, thereby improving the quality of the game product. The electronic device may also simulate different game events (such as reply messages) triggered by the player in the guessing game to collect conversation records for subsequent analysis and processing, such as iterative training of the LLM of the guessing game as training data.
[0135] Among them, the question may refer to a question that is raised, such as a puzzle in a guessing game, which requires the target object to constantly obtain clues to gradually infer and uncover the answer. For the convenience of description, the embodiment of the present application takes the turtle soup game as an example of a true or false guessing game. The question may refer to the riddle of the turtle soup game, that is, the "soup noodles" of the turtle soup game. For example, "Late at night, Xiao Li suddenly slapped himself hard in his sleep. He sat up and smelled the smell of burning gradually in the air. He fell asleep peacefully. What happened?" The question may correspond to its answer, and the answer is the "soup base" of the turtle soup game. For example, for the above question, the answer may be "Xiao Li was awakened by mosquitoes in his sleep, and he got up and lit mosquito coils to repel mosquitoes."
[0136] The dialogue information refers to the dialogue conceived by the target object in order to obtain clues for the answer to the question. The dialogue information can correspond to a variety of different situations:
[0137] 1) The dialogue information may include questions in response to the question, such as yes or no questions. In the above-mentioned turtle soup game, for example, the dialogue information is "Have you lit the mosquito coil?".
[0138] 2) The dialogue information may include topics unrelated to the question being asked or meaningless content, that is, the dialogue information is a dialogue used by the target object for chatting, for example, the dialogue information is "How is the weather today?".
[0139] 3) The dialogue information may include a dialogue in which the target object infers the answer to the question to confirm whether the answer is correct, for example, the dialogue information is "Xiao Li lights the mosquito coil".
[0140] 4) The dialogue information may include other types of questions that the target object tries to get more clues. The questions asked by the target object are related to the game, but no part of them can be judged as right or wrong, that is, the dialogue information can be an open-ended question, for example, the dialogue information is "Why did you slap yourself?"
[0141] In some embodiments, when the dialogue information includes a question, the question may be too brief or too broad, without specifying a specific person or a specific thing. For example, the dialogue information is "delicious", but it does not refer to what is delicious. In this case, the dialogue information can be understood as dialogue information that needs clarification.
[0142] 5) The dialogue information may also be a control instruction for the target object to answer the question, such as exiting the answering process of the question, that is, exiting the guessing game, or changing the question, that is, changing the game puzzle.
[0143] 6) The dialogue information may also be a dialogue in which the target object (such as a player) thinks, or the target object (such as an electronic device) detects, that the answering process is slow, or hopes to obtain prompt information about the answer to the question.
[0144] 7) The conversation information may also be the conversation that the target object wants to review and ask questions.
[0145] 8) The dialogue information may also be a dialogue in which the target object questions the reply information to the historical dialogue information in the historical dialogue.
[0146] Specifically, the guessing game platform may provide a user interface for the target object to input dialogue information, and the target object may input dialogue information for the current question through the user interface. Figure 4 , Figure 4 is a schematic diagram of a dialogue interface provided in an embodiment of the present application, such as Figure 4As shown, the guessing game interface 401 includes an input bar 4011, which is used to obtain the dialogue information (such as natural language text information) input by the target object. In response to the input data in the input bar 4011 being submitted, the historical dialogue information 4013 submitted by the target object before entering the dialogue information is displayed in the content display area 4012 of the guessing game interface; in addition, the reply information 4014 of the historical dialogue information 4013 can also be displayed in the content display area 4012. Thus, based on the dialogue information input by the target object in the guessing game interface, the dialogue information input by the target object in response to the question can be obtained.
[0147] In some embodiments, the target object (such as an electronic device) can establish a communication connection with the guessing game platform, and the target object can send the conversation information to the electronic device corresponding to the guessing game platform through the communication connection, such as Figure 1 and Figure 2 The dialogue information replying device shown is used to obtain the dialogue information input by the target object in response to the question.
[0148] In some embodiments, the target object (such as an electronic device) can also transmit dialogue information through an interface provided by the guessing game platform, such as an application programming interface (API) for obtaining dialogue information, that is, based on a predefined interface, the dialogue information input by the target object in response to the question is obtained. The dialogue information input by the target object in response to the question can also be obtained in other ways, such as forwarding through other electronic devices, etc., which is not limited in this application.
[0149] After the conversation information is acquired, corresponding reply information may be generated based on the conversation information. In the process of generating the corresponding reply information, the conversation information may be first retrieved to generate the reply information based on the retrieved conversation example.
[0150] S302: Perform search processing based on the above conversation information to obtain a reference conversation example of the above conversation information.
[0151] In the embodiment of the present application, the reference dialogue example refers to a dialogue example similar to the dialogue information, which may include a dialogue information example and a reply information example of the dialogue information example. Since the dialogue information example is mainly a question, the dialogue information example and the reply information example may also be referred to as a question-answer pair. Among them, the text features of the retrieved reference dialogue example need to match the text features of the dialogue information. The text features of the reference dialogue example may refer to the numerical vector of the reference dialogue example, that is, the vector representation of the reference dialogue example. The text features of the dialogue information are the vector representation (numerical vector) of the dialogue information.
[0152] In a possible implementation, the process of performing retrieval processing based on the dialogue information to obtain reference dialogue examples of the dialogue information may specifically include first extracting text features of the dialogue information, and then calculating the similarity between the text features of the dialogue information and the text features of each dialogue example in the dialogue example library. After that, a set number of dialogue examples may be selected from the dialogue example library in descending order of similarity as the reference dialogue examples. The dialogue example library may be a vector retrieval library, which may also be called a vector database, a vector library, etc. The vector retrieval library stores the text features of each dialogue example and each dialogue example (text content), and the text content and text features of the dialogue example are stored correspondingly. Each dialogue example may cover as many response strategies and scenarios as possible, that is, the question-answer pairs corresponding to different response strategies are input into the vector retrieval library, so that reference dialogue examples matching the dialogue information can be retrieved based on the dialogue information.
[0153] Among them, extracting text features of conversation information may refer to converting conversation information from text form into vector representation, which may be implemented by inputting into a machine learning model for extracting text features, for example, a text embedding model. The text features of each conversation example refer to the vector representation of each conversation example, and calculating the similarity between the text features of the conversation information and the text features of each conversation example in the conversation example library may be calculating the similarity between the vector representation of the conversation example and the vector representation of each conversation example, such as cosine similarity. Thus, each conversation information of the target object may be searched to obtain the most similar conversation example, which may include typical sentences. The set number may be one or more, and this application does not limit this.
[0154] In some embodiments, the present application provides an electronic device of an AI guessing game platform, such as Figure 1 and Figure 2 The dialogue information reply device in the can build a general RAG process (pipeline), which can also be called a RAG framework, which can interact with the vector retrieval library and LLM, and retrieve at least one dialogue example from the vector retrieval library according to the dialogue information input by the target object, and the retrieved dialogue examples include reply information examples. For example, if the dialogue information input by the target object is "The weather is good today", dialogue examples related to "chatting" may be retrieved from the vector retrieval library, and then the dialogue examples are input into the LLM for processing.
[0155] Calculating the similarity between the text features of the dialogue information and the text features of each dialogue example in the dialogue example library may mean that in the process of generating the reply information for the dialogue information, a similarity matching algorithm, such as the cosine similarity between two numerical vectors, may be used to retrieve reference dialogue examples (question-answer pairs) related to the dialogue information in the dialogue example library. Thus, by retrieving the reference dialogue examples for generating the final reply information, it can be ensured that the retrieved dialogue examples are highly relevant to the dialogue information in terms of content, thereby facilitating the generation of more accurate reply information without losing diversity.
[0156] Please also read Figure 5 , Figure 5 is a timing diagram of a retrieval process for dialog information provided by an embodiment of the present application, such as Figure 5 As shown in the figure, after obtaining the dialogue information input by the target object in response to the question, the dialogue information can be input into the text embedding model, and the text features of the dialogue information extracted by the text embedding model can be obtained, that is, the vector representation of the dialogue information. Then, the similarity between the vector representation of the dialogue information and the vector representation of each dialogue example in the dialogue example library can be calculated, which can be the cosine similarity. After that, a set number of dialogue examples can be selected from the dialogue example library in the order of similarity (cosine similarity) from high to low. Figure 5 Take the setting number as K as an example, where K is a positive integer. Top 1-top K refers to the first to Kth ones with the highest similarity. Therefore, the text content of the dialogue examples in these K dialogue examples can be used as reference dialogue examples for dialogue information.
[0157] In some embodiments, the conversation examples may be stored in different conversation example libraries (vector retrieval libraries) in a dispersed manner, and in the process of calculating the similarity, each conversation example in each conversation example library may be calculated separately.
[0158] In some embodiments, the dialogue example library may also be associated with a question (such as a game puzzle). One question may correspond to one dialogue example library. In the process of calculating the similarity, the dialogue example library corresponding to the question may be determined first to search in the dialogue example library, i.e., to calculate the similarity.
[0159] In another possible implementation, the retrieval method can be optimized so that the dialogue example library is not limited to simple vector retrieval. For example, the dialogue information can be expanded, which can also be called query expansion, and then the retrieval can be based on the dialogue information and the expanded dialogue information. Specifically, the associated dialogue information of the dialogue information can be obtained first, and then the text features of the dialogue information and the text features of the associated dialogue information can be extracted. After that, the similarity between the text features of the dialogue information and the text features of each dialogue example in the dialogue example library, and the similarity between the text features of the associated dialogue information and the text features of each dialogue example can be calculated. Finally, in the order of similarity from high to low, a set number of dialogue examples are selected from the dialogue example library as reference dialogue examples.
[0160] Among them, the associated conversation information may refer to extended conversation information. On the basis of searching based on conversation information, the associated conversation information may be searched for, and then refined sorting may be performed, that is, all similarities (such as cosine similarities) may be sorted from high to low, and a set number of conversation examples may be selected as reference conversation examples.
[0161] In one implementation, obtaining the associated dialogue information of the dialogue information may be to obtain the historical dialogue record of the target object before the dialogue information, and determine the historical dialogue record as the associated dialogue information. The historical dialogue record may include one or more rounds of dialogue before the dialogue information, such as 2-3 rounds of dialogue before the dialogue information. The retrieval process of the dialogue information may be called a one-way recall, and the retrieved reference dialogue example may be called a recall result. That is, if one-way recall is insufficient, it may be considered to perform several more recalls, i.e., multi-way recalls, retrieve the historical dialogue information as the associated dialogue information, and merge all the recall results for screening, which may be fine sorting, so as to obtain the reference dialogue example.
[0162] In another implementation, obtaining the associated dialogue information of the dialogue information may be inputting the dialogue information into a text information expansion model, obtaining the expanded dialogue information generated by the text information expansion model based on the dialogue information, and determining the expanded dialogue information as the associated dialogue information. The text information expansion model may be a text generation model, such as an LLM, and a prompt word for indicating expansion of the dialogue information may be input into the text information expansion model, so that the text information expansion model expands the dialogue information into one or more text information with the same semantics to obtain the expanded dialogue information.
[0163] In another implementation, the expanded dialogue information and the historical dialogue record can be determined as the associated dialogue information. Thus, after retrieving the reference dialogue example, the dialogue information can be replied to based on the reference dialogue example, the question, and the dialogue information to generate reply information corresponding to the dialogue information.
[0164] In the embodiment of the present application, by retrieving the dialogue information with reference to the dialogue example based on the dimension of text features, the most similar dialogue example retrieved can be used as a reference, and with the assistance of the dialogue example, reply information that better meets the business needs can be generated, which is conducive to improving the accuracy of the generated reply information. In addition, in the scenario where LLM is applied and dialogue examples are retrieved, the output of LLM can be better controlled with the help of dialogue examples based on retrieval, so that the content output by LLM is more accurate, thereby reducing the amount of data used to train LLM and improving the training efficiency of LLM.
[0165] S303: Based on the dialogue information, the question and the reference dialogue example, generate initial reply information corresponding to the dialogue information.
[0166] In an embodiment of the present application, the initial reply information corresponding to the dialogue information is not the final reply information of the dialogue information, but an intermediate result for generating the final reply information. The initial reply information may also be referred to as a basic reply, an original basic reply, a reply strategy, etc. Taking the application of answers to questions in AI guessing games as an example, since the current reply rules for AI guessing games are too simple, such as only outputting "yes" or "no" or "irrelevant" replies for any question, the diversity and fun of the reply information are low, which greatly restricts the gameplay of the answering process of asking questions (such as game puzzles), resulting in a low degree of participation of the target object. Therefore, a preliminary reply information can be generated based on the dialogue information, reference dialogue examples, and questions, and then the final reply information can be generated based on the initial reply information.
[0167] The initial reply information includes the category label of the dialogue information, which can be used to indicate the category of the dialogue information, such as [small talk], [open question], etc., and can also be used to indicate the answer to a yes-no question, such as [yes] or [no]. The dialogue label is associated with the reply strategy of the dialogue information of this category. The initial reply information can also include supplementary content, such as supplementary content of the category label. For example, the initial reply information is [no], which means he did not miss the bus because he overslept.
[0168] Specifically, based on the dialogue information, questions, and reference dialogue examples, the process of generating the initial response information corresponding to the dialogue information may be to first obtain the historical dialogue records of the target object before the dialogue information, and then fill the dialogue information, questions, reference dialogue examples, and historical dialogue records into the set prompt word template to obtain the target prompt word. After that, the target prompt word can be input into the LLM, and the response information generated by the LLM based on the answer strategy and response constraint information associated with the prompt word template is obtained to obtain the initial response information.
[0169] In an embodiment of the present application, the historical conversation record of the target object before the dialogue information may include all rounds of conversations before the dialogue information, specifically including the historical dialogue information of each round and the reply information of the historical dialogue information. The set prompt word template is a prompt word (prompt) for input into the LLM, which can be used to clearly indicate the task that the LLM needs to perform or the type of content expected to be generated, and may include a preset text content template. Furthermore, the target prompt word including dialogue information, questions, reference dialogue examples and historical dialogue records can be input into the LLM, so that the LLM generates the initial reply information corresponding to the dialogue information based on the target prompt word.
[0170] That is to say, the dialogue information of the target object is input into the RAG process, and the preset prompt word template is filled in based on the reference dialogue examples retrieved from the vector retrieval library, and then the filled target prompt words are input into the LLM to generate the initial reply information.
[0171] It should be noted that the LLM here refers to the LLM used in the AI guessing game to reply to the dialogue information of the target object, which can be obtained based on training. Since the target prompt word contains the reference dialogue example obtained based on RAG retrieval, the target prompt word improves the quality of the constructed prompt word compared to directly using the dialogue information and the instruction to reply to the dialogue information as the prompt word. Therefore, after the high-quality target prompt word obtained by filling is input into the LLM, the LLM can determine the response strategy corresponding to the dialogue information from the reference dialogue example based on the dialogue example in the target prompt word, so as to generate initial response information that is highly consistent with business requirements.
[0172] The answer strategy associated with the prompt word includes the rules of the answer process of the question, such as the rules of a yes-no guessing game, which may specifically include indication information for indicating the role played by the LLM. The reply constraint information associated with the prompt word may be information constraining the reply information output by the LLM, and may include specific reply requirements, etc.
[0173] In some embodiments, the prompt word template may include an answer strategy and may also include reply constraint information. The prompt word template may also include content for indicating the answer strategy and reply constraint information, which is not limited in the present application.
[0174] In a possible implementation, taking the application of answering questions to an AI guessing game as an example, the prompt word template may include but is not limited to the following key elements:
[0175] 1) Answering strategy: This may include the rules for answering questions, such as the rules and introduction of an AI guessing game, which may include helping LLM understand the role and specific tasks they need to play.
[0176] 2) Reply constraint information: Specifically, it may include a series of specific reply requirements and execution standards. These requirements cover reply strategies in different situations to guide LLM to reply according to the established reply strategy and comply with the AI guessing game mechanism in the embodiment of this application, so as to ensure the logic and correctness of the output initial reply information. In addition, the reply constraint information may also include standards and criteria that need to be followed during the execution of the task, such as keeping the output initial reply information with labels or supplementary content (such as necessary explanations of labels), etc. These standards can be used as a basis for judging the quality of LLM output.
[0177] 3) Questions: Questions can be included in the prompt word template in the form of slots that need to be filled in, such as the turtle soup story {story_scenario_solution}, {..} is the question. The specific question can be filled in the slot.
[0178] 4) Reference dialogue examples can also be included in the prompt word template in the form of slots that need to be filled in. In some embodiments, the reference dialogue examples can include dialogue examples and reply information of different dialogue types, which can be filled in the slots corresponding to different types.
[0179] Therefore, based on the key elements included in the prompt word template, it can be ensured that LLM can accurately grasp the core objectives of the task (AI guessing game) to adapt to business needs.
[0180] For example, a complete prompt word template can be specifically shown in Table 1, which uses the turtle soup game in the true or false guessing game as an example:
[0181]
[0182]
[0183]
[0184]
[0185] Table 1
[0186] In Table 1, "your settings" and "Turtle Soup game rules" can be the key elements of the answer strategy, the content in the "reply strategy" can be the key elements of the reply constraint information, the "user's current question" at the end of Table 1 corresponds to the dialogue information, the "conversation history" corresponds to the historical dialogue record, and the "Turtle Soup Story" corresponds to the question, that is, the puzzle. In the prompt word template shown in Table 1, the slots such as {eg_clarify}, {eg_open_question}, and {eg_chat} in the "other processing situations" correspond to the reference dialogue examples.
[0187] Thus, by combining the retrieved reference dialogue examples, historical dialogue records, specific questions and prompt word templates, a rigorous and reasonable prompt word, namely the target prompt word, is formed. This target prompt word will be input into the trained LLM, and the LLM will generate initial response information that meets the business requirements and the target object dialogue information based on these target prompt words. Since the initial response information is the content used to generate the final response information, the process of generating the initial response information can be called the process of generating the original basic response in combination with RAG.
[0188] It is understandable that Table 1 is an example of a complete prompt word template. In an actual conversation, only part of the prompt word content can be selected as the prompt word template, or only part of the content can be obtained to fill in the corresponding slot and the content of the empty slot can be deleted.
[0189] Exemplarily, the dialogue message may be “Why are you crying?”, and the initial response message generated by LLM may be, for example, “[Other - Open Question] Please ask a specific yes or no question, such as “Is Xiao Li crying because he is so happy to get first place?” Only then can I answer you. As another example, the dialogue message may be “Unhappy”, and the initial response message generated by LLM may be, for example, “[Other - Clarification] Your statement is not clear enough. Are you asking Xiao Li whether she is crying because she is unhappy, or are you describing a fact? Please ask again in the form of a specific yes or no question.”
[0190] It can be seen that through the reply constraint information associated with the prompt word template, such as the reply strategy shown in Table 1, the reply strategy can be expanded, such as adding the recognition of topics such as mid-chat and other non-question questions and the corresponding reply strategy. And add specific tags, such as adding specific tags for types such as open questions (i.e., questions that cannot be answered with yes or no), questions that need to be clarified (i.e., dialogue information that is unclear or includes vague inquiries), etc. The specific tag can refer to the category tag of the dialogue information, which can be understood as the reply strategy for the dialogue information. LLM outputs initial reply information including specific tags. For example, if the dialogue information input by the target object is a developmental question, LLM outputs specific tags (reply strategies) related to open questions, so as to further generate the final reply information based on the specific tags, which solves the problem of only getting a monotonous yes or no reply for any dialogue information in the current AI guessing game to a certain extent, thereby improving the quality of the LLM generated reply and ensuring that the LLM can produce high-quality replies that meet business requirements.
[0191] In a possible implementation, in the current process of answering questions, such as the AI guessing game, any dialogue information input by the target object will only output a yes or no answer, resulting in insufficient guidance for the target object. The target object may not be able to smoothly advance the answering process (such as the game process), resulting in a low degree of participation by the target object. In order to allow the target object to better participate in the process of answering questions, a progressive prompting and guidance method can be used to help the target object reason better.
[0192] Specifically, before the dialogue information, questions, reference dialogue examples and historical dialogue records are filled into the set prompt word template, it can be detected whether the dialogue turn of the current dialogue information is greater than the set turn threshold. If it is detected that the dialogue turn of the dialogue information is greater than the set turn threshold, the reply strategy of the dialogue information can be added to the reply constraint information to obtain an updated prompt word template. Thus, the dialogue information, questions, reference dialogue examples and historical dialogue records can be filled into the updated prompt word template to obtain the target prompt word.
[0193] Among them, the dialogue turn refers to the turn in which the target object conducts a dialogue with respect to the question, and setting the turn threshold may be a trigger threshold pre-set for modifying or adding a response strategy in the response constraint information associated with the prompt word template. The response strategy is used to instruct the LLM to output the prompt information for the answer to the question, and the answer to the question can be understood as the answer to the puzzle, such as the "soup base" in the turtle soup game. For example, in the response constraint information "Current Response Strategy (Priority Followed)" associated with the prompt word template as shown in Table 1, a response strategy for instructing the LLM to output the prompt information for the answer to the question can be added.
[0194] That is to say, when the target object reaches a predetermined number of rounds in the process of answering questions, LLM will receive a special instruction: give a prompt, giving the target object prompt information about the answer. In this way, whether to give a prompt is completely controllable, and after reaching a predetermined number of rounds, it can be determined to give a prompt. Optionally, specific prompt information can also be added to the reply constraint information, that is, the content of the specific prompt can be controlled. While optimizing the game mechanism, it also improves the participation of the target object and the interactive experience of answering questions.
[0195] In some embodiments, if it is detected that the dialogue round of the dialogue information is greater than another set round threshold, if the number of rounds increases further, other reply strategies for the dialogue information can be added to the reply constraint information. The reply strategy is also a special instruction that the LLM will receive. For example, when the target expression requires prompt information, the key information of the answer is directly given to help the target object reasoning. It can be understood that the above is an example of a reply strategy for dialogue information, and this application does not limit this, and it can be determined based on actual task requirements.
[0196] It should be noted that the above-mentioned prompt information about the answer is only triggered by modifying the prompt word template in a specific round, and the target object can express in the dialogue information that he does not know what to ask, has no idea, the question is too difficult, or asks for clues during the dialogue process to obtain prompts about the answer. This is not limited to the dialogue round of the dialogue information. Therefore, through reasonable guidance, when the user's answer progress is slow, or complains and wants prompts, feedback can be given immediately to guide the target object to reason in the right direction (answer the question).
[0197] In some embodiments, in order to better utilize LLM to reply to the dialogue information input by the target object, in the preparation stage of asking questions, a high-quality question data set can be constructed, which can include each question in multiple questions and the supplementary explanation content of each question. In this way, it can be ensured that LLM can accurately identify the user's intention and understand the dialogue information of the target object. The question data set is also the basis for LLM to make reasonable initial reply information.
[0198] In the embodiment of the present application, an example of disassembling and supplementing a question into data that meets the requirements is used for explanation. The question data set includes disassembling and supplementing multiple questions into data that meets the requirements. Specifically, the question and the answer corresponding to the question can be obtained first, and then the answer to the question can be disassembled into at least one key information, and supplementary explanation content and question examples can be generated for each key information. Afterwards, the supplementary explanation content of the question is generated based on the question and the answer to the question. Finally, the question, at least one key information, the supplementary explanation content of the question, the supplementary explanation content of each key information and the question example can be combined into the target puzzle information. The target training question is to disassemble and supplement multiple questions into data that meets the requirements.
[0199] Among them, the key information obtained by disassembling the answer can also be called a key point, which refers to a part of the answer, which is the information or element that best summarizes or represents a part of the main theme of the answer. The answer can be disassembled into at least one key information, and each key information is not repeated. The supplementary explanation content of the key information can be the explanation information for each key point to facilitate the understanding of the target object and LLM. The question example refers to the question example that may be included in the dialogue information input by the target object for the key point. The supplementary explanation content of the training question can be understood as the explanation content for the question.
[0200] Exemplarily, taking the question as a puzzle and the answer to the question as the answer, obtaining the question and the answer corresponding to the question can be manually collected and processed, such as manually collecting a series of AI guessing game puzzles, or can be generated based on LLM, and this application does not limit this. In the process of generating the key information of the answer, it can be based on a set prompt word (prompt) using a machine learning model for text information processing, such as LLM, by indicating the decomposition into at least one key information in the prompt word, and generating supplementary explanation content and question examples for each key information, so as to decompose the answer into several key information through LLM, and generate corresponding explanations for each key information.
[0201] Furthermore, the explanation of the game puzzle can be generated based on the set prompt words using the machine learning model for text information processing, such as the relationship between the soup base and the soup noodles in the turtle soup game, to assist in providing LLM understanding of the AI guessing game. Finally, the game puzzle, at least one key information, the explanation of the puzzle, the explanation of each key point and the question example can be combined into a target question.
[0202] In a possible implementation, after the target question is obtained, it may be manually revised again to ensure the correctness, logic and readability of the target question after preprocessing (manual revision).
[0203] Take the "noodle soup" of the Turtle Soup game as an example: the question is "Late at night, Xiao Li suddenly slapped himself hard in his sleep. He sat up and smelled the burning smell in the air. He fell asleep peacefully. What happened?" The answer is "Xiao Li was awakened by mosquitoes in his sleep, so he got up and lit a mosquito coil to repel the mosquitoes." The key information that can be deconstructed can be "The burning smell is the mosquito coil." The supplementary explanation of this key information can be "Xiao Li confirmed that the mosquito coil has been lit and the mosquito repellent effect has been achieved." The question example of this key information can be "Did Xiao Li light the mosquito coil?" The supplementary explanation of the game puzzle can be "Xiao Li was awakened by mosquitoes in the middle of the night and felt very irritated, so he slapped himself hard to repel the mosquitoes. After slapping, Xiao Li lit the mosquito coil. Xiao Li smelled the burning smell in the air, confirmed that the mosquito coil had been lit and the mosquito repellent effect had been achieved, so he continued to sleep peacefully."
[0204] Therefore, we can obtain questions and combine them with machine learning models such as LLM to break down the answers to the questions and supplement the data that meets the requirements. The process of obtaining the target questions from the question combination can be called the original data preparation stage, and through this stage, we can obtain information-rich questions.
[0205] Please also read Figure 6 , Figure 6 is a timing diagram of raw data preparation provided by an embodiment of the present application, such as Figure 6 As shown, first, the questions and the answers corresponding to the questions can be obtained. The questions and the corresponding answers can also be called scripts. For example, the scripts are manually collected and screened to obtain the questions and the corresponding answers. Furthermore, on the one hand, based on a machine learning model for text information processing, such as LLM, the answer can be assisted in breaking down into at least one key information, and various key supplementary explanation contents and question examples can be generated with the assistance of the machine learning model for text information processing. On the other hand, based on the machine learning model for text information processing, the supplementary explanation content of the question can be assisted in generating based on the questions and answers.
[0206] Afterwards, the question, at least one key information, the supplementary explanation content of the question, the supplementary explanation content of each key information and the question example can be combined into a target question, that is, recombined into a new puzzle structure. Among them, the process of disassembling the answer to the question, generating supplementary explanation content, and reorganizing the structure can be understood as a process of preliminary cleaning of the question in combination with LLM. Subsequently, manual calibration and annotation are performed to obtain the final puzzle, that is, the target question is obtained by manual verification and modification. In this way, all information related to the question to be deduced can be clearly presented to the LLM to ensure that the LLM can understand the question.
[0207] Among them, the LLM applied to the process of answering questions (such as the game process of an AI guessing game) can be obtained by fine-tuning the pre-trained LLM, that is, the LLM input by the target prompt word is a fine-tuned model. It should be noted that in order to enable the fine-tuned LLM to recognize the chat intention of the target object, prompt the target object to clarify unclear dialogue information, and judge whether the dialogue information input by the target object hits the key information, so that the fine-tuned LLM can understand different dialogue information and generate corresponding initial reply information, the training data for fine-tuning the LLM needs to be able to cover more scenarios. In an embodiment of the present application, there are four ways to obtain training data for training LLM. Specifically:
[0208] The first way to obtain training data is to simulate the target object to conduct multiple rounds of dialogue. Specifically, you can first obtain training questions and multi-round dialogue constraint information. Among them, training questions refer to questions used in the training process, such as puzzles used for training in guessing games. In some embodiments, training questions and questions can be the same or different, and this application does not limit this. Multi-round dialogue constraint information refers to information used to constrain the answer process of training questions. Specifically, it can be information that constrains the answer process of training questions, such as information that restricts the game process. For example, it can be limited to normal dialogue during the answer process of training questions, that is, each round of dialogue in the multi-round dialogue is a yes or no question, and each round of dialogue is also a yes or no answer, that is, it is always fun during the game. For another example, it can be limited to occasionally insert topics unrelated to training questions in the process of answering training questions, such as occasionally inserting small talk topics during the game.
[0209] In some embodiments, the training questions may be manually collected and processed, or may be generated based on LLM, which is not limited in this application. The multi-round dialogue constraint information may be set by a technician according to actual business needs.
[0210] Then, the training questions and the multi-round dialogue constraint information are filled into the prompt word template to obtain the training dialogue prompt words. It can be understood that by filling the training questions and the multi-round dialogue constraint information into the specific slots of the prompt word template as shown in Table 1, the training prompt words for generating training data can be obtained. After that, the training dialogue prompt words are input into the dialogue generation model, and the dialogue record generated by the dialogue generation model based on the training dialogue prompt words is obtained. Among them, the dialogue generation model can be a machine learning model with powerful natural language processing capabilities, such as GPT-4. The dialogue record can be a complete dialogue of the target object in different answering processes of the training questions (such as during the game process), and can include multiple rounds of dialogues in the entire answering process (such as a game), as well as the reply information of each round of dialogue.
[0211] Therefore, the conversation information and reply information in the conversation record can be used as training data, and the LLM to be trained can be trained based on the conversation record to obtain the LLM for generating the initial reply information of the conversation information. Specifically, by writing the specified prompt words (i.e., the training conversation prompt words), using a conversation generation model such as GPT-4 based on the cleaned training questions and reply strategies, a simulated conversation that conforms to the pre-assumed answer process (such as the game process) and the game mechanism in the embodiment of the present application (such as interspersed chat topics) can be generated to obtain a conversation record. Furthermore, the conversation record generated by the conversation generation model (such as LLM) can be manually annotated and modified, and the manually corrected and rewritten data can be organized into training data. Using a general LLM for simulation can accelerate the conception of product prototypes, reduce the cost of risk assessment, and verify the feasibility of new technologies and stimulate innovative thinking.
[0212] In some embodiments, there is manual participation in the current training data acquisition scenario. In order to reduce manual participation and improve the efficiency of training data acquisition, existing annotation examples can be used, such as the manually annotated part as an example, combined with corresponding instructions to construct prompt words, and the general LLM can be used to fully automate the annotation modification process.
[0213] In some embodiments, the current conversation record may be generated by a conversation generation model based on a training conversation prompt. Another machine learning model may be introduced to simulate the conversation of the target object, such as a target object model (such as a player model). The conversation may be continued through the player model and the conversation generation model to obtain a large amount of training data.
[0214] The conversation record includes multiple rounds of conversation information and reply information of each round of conversation information, and the historical conversation records of different rounds of conversation information are different. In the process of training the LLM to be trained based on the conversation record, it can be specifically firstly retrieved and processed based on each round of conversation information and the historical conversation record of each round of conversation information to obtain a training reference conversation example of each round of conversation information. The process of retrieving and processing based on each round of conversation information and the historical conversation record of each round of conversation information can refer to the specific implementation method described in step S302, which will not be repeated here.
[0215] Furthermore, each round of dialogue information, the historical dialogue record of each round of dialogue information, the training reference dialogue example of each round of dialogue information, and the training questions can be filled into the prompt word template to obtain the training prompt word. It should be noted that the training prompt word refers to the prompt word corresponding to each round of dialogue information, and each round of dialogue and its reply information in the multiple rounds of dialogue in the dialogue record can be used as training data, that is, based on the multiple rounds of dialogue in the dialogue record, the training prompt words corresponding to the multiple rounds of dialogue can be generated respectively. Afterwards, the training prompt word can be input into the LLM to be trained, and the training reply information of each round of dialogue information output by the LLM to be trained can be obtained.
[0216] Thus, the LLM to be trained can be trained based on the difference between the training reply information of each round of dialogue information and the reply information of each round of dialogue information to obtain the LLM, that is, to provide an LLM such as an AI guessing game. Specifically, loss data can be constructed based on the difference between the training reply information and the reply information of each round of dialogue information, so as to fine-tune the model parameters of the LLM to be trained based on the loss data.
[0217] It should be noted that, whether it is the stage of acquiring training data or the stage of training the LLM to be trained, the training questions filled in the prompt word template can be based on the training questions after LLM cleaning, so that the model can better understand the questions.
[0218] Specifically, the answer to the training question can be broken down into at least one key information, and supplementary explanation content and question examples can be generated for each key information, and then the supplementary explanation content of the training question can be generated based on the training question and the answer, and then the training question, at least one key information, the supplementary explanation content of the training question, the supplementary explanation content of each key information and the question example can be combined into a target training question, which is the cleaned training question. Among them, the cleaning process of the training question can refer to the specific implementation method described in the cleaning process of the question, which will not be repeated here. Therefore, when the training question and the multi-round dialogue constraint information are filled into the prompt word template to obtain the training dialogue prompt word, the training question can be replaced with the cleaned training question, that is, the target training question.
[0219] The second way to obtain training data is to simulate a single-round conversation with the target object. Specifically, you can first obtain training questions and designated category conversation information for the questions. Designated category conversation information can refer to conversation information including open-ended questions, conversation information including fuzzy questions, conversation information including chat content, etc. Designated category conversation information can be understood as special questions for a certain scenario. The designated category conversation information for training questions can be manually constructed for training questions, or it can be an example of manually written designated category conversation information, and it can be obtained by expanding based on a machine learning model, or it can be obtained by other methods, which is not limited in this application. The method for obtaining training questions can refer to the specific description in the first method for obtaining training data, which will not be repeated here.
[0220] Then, the training questions and the designated category dialogue information are filled into the prompt word template to obtain the training dialogue prompt words. That is, after obtaining the training questions and the training questions and the designated category dialogue information including open questions, questions that need clarification, small talk, etc., they can be filled into the corresponding slots in the prompt word template as shown in Table 1 to obtain the training dialogue prompt words for generating the reply information of the designated category dialogue information. After that, the training dialogue prompt words are input into the dialogue generation model, and the reply information of the designated category dialogue information generated by the dialogue generation model based on the training dialogue prompt words is obtained. Thus, the reply information corresponding to the designated category dialogue information can be used as training data, and the LLM to be trained can be trained based on the designated category dialogue information and the reply information of the designated category dialogue information to obtain the LLM for generating the initial reply information of the dialogue information.
[0221] That is to say, in the second method, by constructing special questions involving open questions, small talk, questions that need clarification, etc. (specified category dialogue information), by writing specified prompt words (i.e., training dialogue prompt words), and using dialogue generation models such as GPT-4 to generate rule-compliant reply information, training data is obtained. The training data may include multiple specified category dialogue information and corresponding reply information.
[0222] In a possible implementation, a portion of the multiple designated category dialogue information and corresponding reply information obtained by the second method can be used as training data for training LLM, and the other portion can be stored in the dialogue example library (vector retrieval library) as dialogue examples. For example, after obtaining multiple designated category dialogue information and corresponding reply information, one quarter can be selected for model training, and the remaining three quarters can be stored in the dialogue example library for retrieval processing. This is because the dialogue example library is the data basis of the reasoning stage, and its structure and quality directly affect the effectiveness and scope of application of LLM. Therefore, designated category dialogue information and its reply information covering a variety of situations can be stored in the dialogue example library.
[0223] Specifically, the text features of the dialogue example consisting of the designated dialogue information and the reply information of the designated dialogue information except for part of the dialogue information in the designated category dialogue information can be extracted, such as by inputting the dialogue example consisting of the designated dialogue information and the reply information of the designated dialogue information into the text embedding model, and obtaining the vector representation output by the text embedding model, and storing the vector representation and the text content of the dialogue example in a dialogue example library (vector retrieval library) to build a comprehensive and high-quality dialogue example library. It can be understood that the dialogue example library can be obtained in the process of simulating the dialogue of the target object, and can include the questions of the target object and the ideal replies expected to be obtained. The dialogue example library can be used for retrieval processing in the subsequent construction process of the prompt word to retrieve one or more dialogue examples that are closest to the dialogue information input by the target object.
[0224] Furthermore, the LLM to be trained can be trained based on the designated category dialogue information and the reply information of the designated category dialogue information. Specifically, part of the dialogue information can be selected from multiple designated category dialogue information, and the part of the dialogue information can be retrieved and processed to obtain a training reference dialogue example of the part of the dialogue information. The part of the dialogue information can refer to the data selected from the acquired multiple designated category dialogue information for training the LLM to be trained. The principle of retrieval processing based on part of the dialogue information is the same as the principle of retrieval processing based on dialogue information. Please refer to the specific implementation method described in step S302, which will not be repeated here.
[0225] Then, the partial dialogue information, the training reference dialogue examples of the partial dialogue information, and the training questions are filled into the prompt word template to obtain the training prompt words. Since the partial dialogue information is a single-round dialogue information, each dialogue information in the partial dialogue information does not include the historical dialogue record. When filling the prompt word template, only the partial dialogue information, the training reference dialogue examples of the partial dialogue information, and the training questions can be filled into the corresponding slots in the prompt word template. After obtaining the training prompt words, the training prompt words can be input into the LLM to be trained to obtain the training reply information of the partial dialogue information output by the LLM to be trained. Thus, the LLM to be trained can be trained based on the difference between the reply information of the partial dialogue information in the reply information of the specified category of dialogue information and the training reply information of the partial dialogue information to obtain the LLM. Specifically, the loss data can be constructed based on the difference between the reply information of the partial dialogue information and the training reply information, so as to fine-tune the model parameters of the LLM to be trained based on the loss data.
[0226] After obtaining multiple conversation information of specified categories and corresponding reply information, selected part of the conversation information can be used for training. Therefore, the reference conversation examples and training reference conversation examples retrieved and processed in the training process and the reasoning process (i.e., the process in which LLM generates replies based on the conversation information of the target object) are at least one conversation example among the conversation examples consisting of the specified conversation information other than part of the conversation information and the reply information of the specified conversation information in the conversation information of the specified category.
[0227] Please also read Figure 7 , Figure 7 is a timing diagram of LLM training provided by an embodiment of the present application, such as Figure 7 As shown, the training questions can be obtained first, and then cleaned to obtain the target training questions. Furthermore, on the one hand, the target object's answering process in the question can be simulated based on the target training question to generate a dialogue record of the simulated dialogue, which can include multiple rounds of dialogue and reply information of each round of dialogue. Each round of dialogue and the corresponding reply information can be combined into training data to train the LLM to be trained. On the other hand, the specified category dialogue information and the reply information of the specified category dialogue information can be constructed based on the target training question, and a part of it can be selected as training data for training the LLM to be trained, and the other part can be stored as dialogue examples in the dialogue example library.
[0228] Among them, the process of training the LLM to be trained can be to retrieve the dialogue information in the training data, such as the dialogue information of the specified type and each round of dialogue information in the dialogue record, in the dialogue example library, and then fill in the prompt word template based on the retrieved training reference dialogue example to obtain the training prompt word. It can be understood that since each round of dialogue information in the dialogue record is associated with the historical dialogue record, when filling in the prompt word template, the historical dialogue record of each round of dialogue information needs to be filled into the prompt word template. Then, the training prompt word can be input into the LLM to be trained to obtain the reply information generated by the LLM to be trained. Therefore, based on the difference between the reply information generated by the LLM to be trained and the reply information in the training data, the loss data can be constructed to train the LLM to be trained.
[0229] In some implementations, the data organization format of the dialogue example library may also be optimized, such that the dialogue example library may include not only dialogue information and response information (such as question-answer pairs), but may also include dialogue information and corresponding response information of the target object in the process of answering questions, such as taking all dialogue information and each round of dialogue information of the entire answering process (such as the game process) as dialogue examples to add more real examples.
[0230] The third way to obtain training data is to directly construct training data suitable for prompting answers to questions. Specifically, you can obtain training questions, training prompt dialogue information, and reply prompt information for the training prompt dialogue information. Among them, the training prompt dialogue information is used to obtain prompt information for answers to training questions. In other words, the training prompt dialogue information can be used to ask for prompts about answers, and reply prompt information refers to reply information including prompt information for answers to training questions. For example, the training prompt dialogue information is "I still can't think of it", and the reply prompt information is "
Hint
[0231] Among them, the method of obtaining the training prompt dialogue information and the reply prompt information can be manually constructed, such as manually written, or it can be obtained by manually writing examples of training prompt dialogue information and reply prompt information and expanding based on the machine learning model, or it can be obtained in other ways, and this application is not limited to this. The method of obtaining training questions can refer to the specific description in the first method of obtaining training data, which will not be repeated here. The training prompt dialogue information and reply prompt information can be understood as a specific type of question and answer pair, which can be used to train the model to have the ability to follow special instructions, that is, the model must reply with a reply with a certain label in a certain reply, such as the above-mentioned [prompt] label. The above-mentioned training prompt information can also be understood as a large number of inquiries with special instructions, and the reply prompt information can be understood as reply information that follows special instructions, so that training data can be formed.
[0232] Further, the LLM to be trained can be trained based on the training prompt dialogue information and the reply prompt information. Specifically, the training prompt dialogue information can be first retrieved and processed to obtain a training reference dialogue example of the training dialogue prompt information. The principle of retrieving and processing based on the training prompt dialogue information is the same as the principle of retrieving and processing based on the dialogue information. Please refer to the specific implementation method described in step S302, which will not be repeated here. Then, the training dialogue prompt information, the training reference dialogue example of the training dialogue prompt information, and the training question are filled into the prompt word template to obtain the training prompt word, that is, each part of the content is specifically filled into the corresponding slot in the prompt word template to obtain the training prompt word. After that, the training prompt word is input into the large language model to be trained, and the training reply prompt information of the training dialogue prompt information output by the large language model to be trained is obtained. Based on the difference between the reply prompt information of the training dialogue prompt information and the training reply prompt information of the training dialogue prompt information, the LLM to be trained is trained to obtain the LLM. That is, based on the difference between the training reply prompt information generated by the LLM to be trained and the reply prompt information of the training dialogue prompt information, the loss data is constructed to fine-tune the model parameters of the LLM to be trained based on the loss data.
[0233] The fourth way to obtain training data is to obtain real dialogue information and construct corresponding reply information. For example, after the AI guessing game is deployed online, the dialogue information entered by the target object in the process of answering questions can be collected to obtain a large amount of real dialogue information from the target object (such as the user) in the real interaction process with the target object. Then it can be used as dialogue information in training data, and the corresponding reply information can be constructed, such as manually constructing the corresponding reply information, or manually writing examples of reply information for some real dialogue information, and generating reply information for other real dialogue information based on the machine learning model. Thus, the real dialogue information and the corresponding reply information can be used as training data to train the LLM to be trained.
[0234] It can be seen that the above description includes the process of training data construction and model training. The former is to use a dialogue generation model (such as LLM) to simulate dialogues according to different response strategies and scenario requirements, and to generate response information that meets business requirements based on the constructed dialogue information, such as response prompt information and response information for dialogue information of a specified category. The latter is to construct training data by combining the constructed dialogue example library and the dialogue simulated by the dialogue generation model. This data for training the model based on training data with annotations (reply information) can be called supervised fine-tuning (SFT) data, which can be used to implicitly distill the response strategy into the LLM to be trained, so that the trained LLM is better at answering questions, such as AI guessing game tasks. In the training process, a series of dialogue information with different intentions and a combination of data retrieved based on the dialogue information are used for fine-tuning. In the fine-tuning process, the model learns how to adjust its response strategy based on the given dialogue information.
[0235] In some embodiments, after each training iteration cycle, the reply information output by the trained LLM based on the test dialogue information can be evaluated by the user or the machine learning model used to evaluate the reply information, such as manually or using electronic devices to simulate the answering process of the test questions, to measure the quality of model generation, which can be used to adjust the strategy for further training of the LLM to optimize performance and prevent the model from over-adapting to the training data and losing generalization ability. For example, whether the content of the reply information is too radical, if so, the reply information in the training data can be adjusted, such as reconstructing the reply information.
[0236] In some embodiments, since the reply strategy in the prompt word template indicates that corresponding category labels are generated for different dialogue information, after the end of each training iteration cycle, the correctness of the labels in the reply information output by the trained LLM based on the test dialogue information can be evaluated for the reply information output by the training LLM during the training process. For example, if the accuracy of the reply information labeled "prompt" is low, the number of samples of training prompt dialogue information and reply prompt information of training prompt information can be increased in the training data to further train the performance of the optimization model.
[0237] In some embodiments, the stage of the LLM that fine-tunes the model parameters based on the dialogue information with reply information can be called the SFT stage, and the above training process is the SFT stage. The alignment stage can be further performed to improve the model's capabilities, that is, to improve the model's capabilities through reinforcement learning. For example, a reward function can be constructed to calculate rewards or penalties for the reply information output by the trained LLM based on the answer process of the test question, and then each round of dialogue in the answering process is used as a state, and the reply information of each round of dialogue is used as an action, and the loss data in reinforcement learning is constructed to further adjust the model parameters of the LLM.
[0238] In the embodiment of the present application, by generating initial reply information including a category label, which is the conversation intent identified by LLM for the conversation information, reply information corresponding to the conversation intent can be further generated for the category label, which is conducive to improving the accuracy of the reply information. In addition, the category label can be used for manual quality inspection of LLM. Quality inspection based on category labels is more convenient and improves the efficiency of quality inspection. If the initial reply information is subsequently used as training data for training, the category label can also improve the annotation efficiency of the training data.
[0239] S304: Based on the above category labels, generate target reply information corresponding to the above dialogue information.
[0240] In the embodiment of the present application, the category label is included in the initial reply information, which can refer to the category of the dialogue information or the answer to a yes or no question, and is associated with the reply strategy of the corresponding category. The target reply information is the final reply information of the dialogue information, which is the reply information output to the target object.
[0241] Since the LLM currently used in AI guessing games is obtained by constructing corresponding SFT data to fine-tune the pre-trained LLM, that is, fine-tuning the model parameters of the pre-trained LLM based on the constructed labeled data so that the fine-tuned LLM can adapt to the tasks of AI guessing games. This method of directly constructing simple SFT data may cause the generated reply information to be relatively rigid and inflexible, and cannot be migrated to replies of different styles and personalities, making the reply information less scalable.
[0242] Based on this, in an embodiment of the present application, the specific process of generating target reply information corresponding to the dialogue information based on the category label can be to input the category label into the stylized rewriting model, and obtain the first reply information of the specified text style generated by the stylized rewriting model based on the category label, and determine the first reply information as the target reply information. Among them, the stylized rewriting model can be a machine learning model for converting the initial reply information into a reply information of a specific style (such as a cute style, etc.) to adapt to different style replies, which is also conducive to improving the dialogue experience of the target object. Among them, the specified text style can be the text style of the content output by the style rewriting model.
[0243] If the initial reply information generated by LLM includes not only labels but also supplementary content, such as "[No] He didn't miss the bus because he slept too much", and the supplementary content is "He didn't miss the bus because he slept too much", then the category label and the supplementary content included in the initial reply information can be input into the stylized rewriting model, and the second reply information of the specified text style generated by the stylized rewriting model based on the category label and the supplementary content can be obtained, and the second reply information is determined as the target reply information. Taking the specified text style as cute style as an example, the target reply information can be "No, he missed the bus for other reasons~".
[0244] Thus, the basic reply with reply strategy (initial reply information) is decoupled from the stylized reply, and the unstyled initial reply information is first generated, and then the stylized rewriting model is further used to generate a reply with a specific text style (such as cute style), that is, the stylized reply is generated by decoupling the reply strategy model (such as the LLM obtained by the above training) and the style rewriting model. Thus, by decoupling the two, the reply information to the dialogue information input by the target object is no longer a single unstyled statement, but contains a reply with a specific text style that is not limited to a specific character style.
[0245] In some embodiments, in order to generate a reply message including a specific text style and reduce costs as much as possible, the training data used for the LLM to be trained can be stylized and rewritten to conform to the specific style as much as possible without losing semantic information, such as the specified text style mentioned above, so as to combine the original reply message and the rewritten reply information into training data for training the stylized rewriting model to be trained.
[0246] The specific rewriting process can be to manually rewrite the reply information of the training data, or to manually rewrite part of the reply information in the training data into a reply with a specific text style as an example, and rewrite the reply information in other training data based on the machine learning model. This part of the data can then be handed over to manual calibration to determine whether the rewritten reply information is a lossless rewrite of the original reply information. Thus, the stylized rewriting model to be trained can be trained based on the training data used to train the stylized rewriting model to be trained to obtain a stylized rewriting model that can convert the initial reply information into reply information with a specific text style.
[0247] In a possible implementation, the method of directly constructing simple SFT data may also result in poor control of the target object over the process of answering questions (such as the game process). For example, after a player starts a game, he can only play from the beginning to the end, and the player cannot flexibly control the game process such as exiting, changing game puzzles, etc. Therefore, in an embodiment of the present application, the target object can control the process of answering questions based on dialogue information. The dialogue information is text information input by the target object in the process of answering questions. After being recognized by LLM, the category label in the initial reply information obtained includes control instructions, such as control instructions used to indicate exiting the answering process, and can also be used to indicate changing questions, such as the initial reply information is [Exit], etc.
[0248] Thus, after the category label is input into the stylized rewriting model and the reply information (i.e., the target reply information) of the specified text style generated by the stylized rewriting model based on the category label is obtained, the category label can be parsed to perform the corresponding action based on the result of the analysis, that is, the control operation corresponding to the control instruction is executed. If it is determined that the control instruction included in the category label is used to indicate the exit from the answering process, the answering process of the question is exited. If it is determined that the control instruction included in the category label is used to indicate the replacement of the question, a new question can be generated. Among them, the new question can be obtained from the question library, or it can be a question regenerated using a machine learning model, such as LLM, and this application does not limit this. Thus, the target object has the freedom to exit the answering process or change the question at any time, and can either carry out the normal answering process or chat, and can also choose to enter, exit or change the puzzle at any time. The target object has a high degree of freedom, which improves the interactive experience of the target object.
[0249] In order to further improve the dialogue experience of the target object and make it more interactive, it is possible to determine whether the dialogue information hits the key information in the answer to the question in turn to provide more feedback to the target object. If the dialogue information entered by the target object hits the key information, the initial reply information may include a label of [hit the key point], and indicate the specific key information hit in the supplementary content. Thus, the target reply information generated based on the category label and supplementary content can include a staged success prompt for the target object, so that the target object has a better experience.
[0250] Furthermore, in order to determine whether the dialogue information input by the target object hits the answer to the question (such as the answer to a puzzle), on the one hand, LLM can be used to directly determine whether the answer is revealed based on the dialogue information. If LLM determines that the answer is revealed based on the dialogue information, it can output a category label including [All Hits], indicating that the category of the dialogue information is a hit answer.
[0251] On the other hand, when LLM determines whether the key information is revealed based on the dialogue information, it can indirectly determine whether the answer is revealed by determining whether each key information is hit. If the category label included in the initial reply information output by LLM indicates that the dialogue information hits the key information in the answer to the question, the key information hit by the target object for the question is obtained. Among them, the hit key information can be recorded in the prompt word template filled in the dialogue information, such as the slot "{hit_points}" corresponding to "the key point hit" in Table 1. The prompt word template can be understood as the prompt word template associated with the dialogue information of the current round. For example, at the end of each round of dialogue, if the category label in the initial reply information is [hit key point], the prompt word template can be updated, such as adding the supplementary information in the initial reply information (the specific content of the key information) to the slot "{hit_points}" in Table 1 to obtain the prompt word template for the next round.
[0252] Furthermore, if it is determined based on the hit key information that the target object has hit all the key information in the answer, the category label is updated to the specified category label. Among them, the specified category label can be used to indicate that the target object has hit all the key information of the answer, such as [all hits], so that the target reply information can be generated based on the specified category label. Specifically, the specified category label ([all hits]) can be input into the stylized rewriting model, and the reply information generated by the stylized rewriting model based on the specified category label can be obtained to obtain the target reply information.
[0253] Please also read Figure 8 , Figure 8 is a timing diagram of an LLM-based reasoning phase provided in an embodiment of the present application, such as Figure 8As shown, first, the dialogue information input by the target object in response to the question can be obtained. On the one hand, the dialogue information can be used to fill in the prompt word template. On the other hand, it can be retrieved in the dialogue example library to obtain a reference dialogue example. The reference dialogue example and the dialogue information can be filled in the prompt word template. The filled prompt word template is used to obtain the target prompt word. Then, the target prompt word can be input into the LLM to obtain the initial response information generated by the LLM, and the initial response information can include the category label of the dialogue information.
[0254] Furthermore, from the dimension of the answering process of the question, the initial reply information generated by the LLM can be used for stylized rewriting in the normal answering process, such as inputting the category label into the stylized rewriting model, and obtaining the reply information output by the stylized rewriting model to obtain the target reply information. The category label in the initial reply information generated by the LLM may include a control instruction. When the category label includes a control instruction, the answering process of the question can be controlled, such as exiting the answering process or changing the question. It can be understood that changing the question means restarting the answering process for the new question, that is, exiting the answering process of the current question and starting the answering process of the new question.
[0255] Figure 8 The process shown can be called a model reasoning process based on RAG. The above-mentioned retrieval processing based on dialogue information, based on the retrieval results and the filling prompt word template, and the target prompt word obtained after filling is input into the LLM to generate the initial reply information can be the process of the built RAG. In this process, the dialogue example library (vector example library) and LLM can be interacted respectively. By integrating LLM and RAG technology, the generalization ability of the model in a variety of scenarios is enhanced, and it is conducive to improving the accuracy and consistency of the generated target reply information. For the process of answering questions, such as the game process of AI guessing games, the interactive experience of the target object is improved to a certain extent. The target object can control the answering process and optimize the mechanism of the answering process, such as the mechanism of AI guessing games. The target object can quickly obtain the reply information without having to wait too long for the reply time to reduce the experience.
[0256] It is understandable that, during the training phase in the embodiments of the present application, the capabilities of LLM itself, such as natural language capabilities, are utilized to minimize the cost of manually constructing data, enhance the interactive experience, and optimize resource utilization and computing efficiency. In addition, by designing clear prompt word templates and formulating different reply strategies in the prompt word templates, the interpretability and controllability of the model are improved. By evaluating the reply information output by the trained LLM based on the test dialogue information, solid data support is provided for the continuous learning and iterative optimization of the model, thereby ensuring the continuous improvement of the model performance, which is conducive to meeting specific business needs and enhancing the efficiency and adaptability of the model in handling complex dialogue tasks.
[0257] In the embodiment of the present application, by first generating initial reply information including a category tag, and then generating target reply information based on the category tag, reply information corresponding to the category can be output for the category tag, rather than a simple yes or no, which is conducive to generating more accurate reply information and improving the conversation experience. In addition, further processing can be performed based on the category tag to generate different reply information, such as reply information with different text styles, which can improve the diversity of the generated reply information to a certain extent.
[0258] Therefore, the embodiments of the present application can solve the problem of poor question-asking experience and the problem of unstable quality of reply information generated during a conversation and inaccurate generated reply information through automated training data construction, diversified reply strategies, optimized game mechanisms, decoupling of style rewriting from basic reply information, and a flexible RAG mechanism.
[0259] In the technical solutions provided by some embodiments of the present application, by obtaining the dialogue information input by the target object for the question, and performing retrieval processing based on the dialogue information, a reference dialogue example of the dialogue information is obtained, and the text features of the reference dialogue example match the text features of the dialogue information. Then, based on the dialogue information, the question and the reference dialogue example, the initial reply information corresponding to the dialogue information is generated, and the initial reply information includes the category label of the dialogue information, and then based on the category label, the target reply information corresponding to the dialogue information is generated. It can be seen that by retrieving the dialogue information before generating the reply information, the reply information is assisted in generating the reply information based on the retrieved reference dialogue example, so that the generated reply information is more consistent with the reply required by the business, which is conducive to improving the accuracy of the generated reply information. In addition, by determining the category label of the dialogue information and further generating the reply information based on the category label, it is possible to reply and process dialogue information of different categories, which can improve the dialogue experience to a certain extent.
[0260] The method of the embodiment of the present application is described in detail above. In order to facilitate better implementation of the above scheme of the embodiment of the present application, the device of the embodiment of the present application is provided below accordingly.
[0261] See also Fig. 9 , Fig. 9 A schematic diagram of a dialog message reply processing device provided in an embodiment of the present application. Fig. 9 The dialog information reply processing device shown may be mounted in a computer device, which may specifically be a server. Fig. 9 The dialogue information reply processing device shown can be used to execute the above Figure 3 Some or all of the functions of the described method embodiments. Fig. 9 The dialogue information reply processing device 90 includes:
[0262] The acquisition unit 901 is used to acquire the dialogue information input by the target object in response to the question;
[0263] A processing unit 902 is configured to perform a search process based on the dialogue information to obtain a reference dialogue example of the dialogue information, wherein a text feature of the reference dialogue example matches a text feature of the dialogue information;
[0264] A generating unit 903, configured to generate initial reply information corresponding to the dialogue information based on the dialogue information, the question and the reference dialogue example, wherein the initial reply information includes a category label of the dialogue information;
[0265] The generating unit 903 is further configured to generate target reply information corresponding to the dialog information based on the category label.
[0266] In a possible implementation, the processing unit 902 is configured to perform a search process based on the conversation information to obtain a reference conversation example of the conversation information, specifically to:
[0267] Obtaining the associated conversation information of the conversation information;
[0268] extracting text features of the conversation information and text features of the associated conversation information;
[0269] Calculating the similarity between the text features of the dialogue information and the text features of each dialogue example in the dialogue example library, and the similarity between the text features of the associated dialogue information and the text features of each dialogue example;
[0270] A set number of dialogue examples are selected from the dialogue example library in descending order of similarity as the reference dialogue examples.
[0271] In a possible implementation, the processing unit 902 is configured to obtain the associated dialog information of the dialog information, and is specifically configured to perform at least one of the following steps:
[0272] Acquire a historical conversation record of the target object before the conversation information, and determine the historical conversation record as the associated conversation information;
[0273] The conversation information is input into a text information expansion model, and expanded conversation information generated by the text information expansion model based on the conversation information is obtained, and the expanded conversation information is determined as the associated conversation information.
[0274] In a possible implementation, the generating unit 903 is configured to generate target reply information corresponding to the dialog information based on the category label, specifically:
[0275] Inputting the category label into a stylized rewriting model, obtaining first reply information of a specified text style generated by the stylized rewriting model based on the category label, and determining the first reply information as the target reply information;
[0276] or,
[0277] The category label and the supplementary content included in the initial reply information are input into the stylized rewriting model, and the second reply information of the specified text style generated by the stylized rewriting model based on the category label and the supplementary content is obtained, and the second reply information is determined as the target reply information.
[0278] In a possible implementation, the dialogue information is text information input by the target object in the process of answering the question; the category label includes a control instruction; the device 90 further includes:
[0279] The exit unit 904 is used to exit the answering process of the question if the control instruction is used to instruct to exit the answering process.
[0280] The generating unit 903 is configured to generate a new question if the control instruction is used to instruct to change the question.
[0281] In a possible implementation, the generating unit 903 is configured to generate initial reply information corresponding to the dialogue information based on the dialogue information, the question and the reference dialogue example, and is specifically configured to:
[0282] Obtaining the historical conversation record of the target object before the conversation information;
[0283] Filling the dialogue information, the question, the reference dialogue example and the historical dialogue record into a set prompt word template to obtain a target prompt word;
[0284] The target prompt word is input into a large language model, and reply information generated by the large language model based on the answer strategy and reply constraint information associated with the prompt word template is obtained to obtain the initial reply information.
[0285] In a possible implementation, the generating unit 903 is configured to generate target reply information corresponding to the dialog information based on the category label, specifically:
[0286] If the category tag indicates that the dialogue information hits key information in the answer to the question, then obtaining the hit key information of the target object for the question;
[0287] If it is determined based on the hit key information that the target object has hit all the key information in the answer, then the category label is updated to a specified category label;
[0288] Based on the specified category label, the target reply information is generated.
[0289] In a possible implementation, the device 90 further includes:
[0290] An adding unit 905 is configured to add a reply strategy of the dialogue information to the reply constraint information to obtain an updated prompt word template if it is detected that the dialogue turn of the dialogue information is greater than a set turn threshold, wherein the reply strategy is used to instruct the large language model to output prompt information of an answer to the question;
[0291] The generating unit 903 is used to fill the dialogue information, the question, the reference dialogue example and the historical dialogue record into a set prompt word template to obtain a target prompt word, which is specifically used to:
[0292] The dialogue information, the question, the reference dialogue example and the historical dialogue record are filled into the updated prompt word template to obtain the target prompt word.
[0293] In a possible implementation, the acquisition unit 901 is further used to acquire training questions and multi-round dialogue constraint information, where the multi-round dialogue constraint information is used to constrain the answering process of the training questions;
[0294] A filling unit 906 is used to fill the training question and the multi-round dialogue constraint information into the prompt word template to obtain a training dialogue prompt word;
[0295] The acquisition unit 901 is further configured to input the training dialogue prompt words into a dialogue generation model, and acquire a dialogue record generated by the dialogue generation model based on the training dialogue prompt words;
[0296] The training unit 907 is used to train the large language model to be trained based on the conversation record to obtain the large language model.
[0297] In a possible implementation, the conversation record includes multiple rounds of conversation information and reply information of each round of conversation information, wherein the historical conversation records of different rounds of conversation information are different; the training unit 907 is used to train the large language model to be trained based on the conversation record to obtain the large language model, specifically for:
[0298] Performing retrieval processing based on the dialogue information of each round and the historical dialogue records of the dialogue information of each round to obtain training reference dialogue examples of the dialogue information of each round;
[0299] Filling the dialogue information of each round, the historical dialogue record of each round, the training reference dialogue example of each round, and the training question into the prompt word template to obtain the training prompt word;
[0300] Inputting the training prompt words into the large language model to be trained, and obtaining the training response information of each round of dialogue information output by the large language model to be trained;
[0301] Based on the difference between the training reply information of each round of dialogue information and the reply information of each round of dialogue information, the large language model to be trained is trained to obtain the large language model.
[0302] In a possible implementation, the device 90 further includes:
[0303] A decomposition unit 908, configured to decompose the answer to the training question into at least one key information, and generate supplementary explanation content and question examples for each key information;
[0304] The generating unit 902 is further configured to generate supplementary explanation content for the training question based on the training question and the answer;
[0305] A combining unit 909, configured to combine the training question, the at least one key information, the supplementary explanation content of the training question, the supplementary explanation content of each key information and the question example into a target training question;
[0306] The filling unit 906 is used to fill the training question and the multi-round dialogue constraint information into the prompt word template to obtain the training dialogue prompt word, which is specifically used for:
[0307] The target training question and the multi-round dialogue constraint information are filled into the prompt word template to obtain the training dialogue prompt words.
[0308] In a possible implementation, the acquisition unit 901 is further configured to acquire training questions and designated category dialogue information for the training questions;
[0309] The filling unit 906 is further used to fill the training question and the designated category dialogue information into the prompt word template to obtain the training dialogue prompt word;
[0310] The acquisition unit 901 is further configured to input the training dialogue prompt words into the dialogue generation model, and acquire the reply information of the designated category dialogue information generated by the dialogue generation model based on the training dialogue prompt words;
[0311] The training unit 907 is further configured to train the large language model to be trained based on the designated category dialogue information and the reply information of the designated category dialogue information to obtain the large language model.
[0312] In a possible implementation, the number of the designated category dialogue information is multiple; the training unit 907 is used to train the large language model to be trained based on the designated category dialogue information and the reply information of the designated category dialogue information to obtain the large language model, specifically for:
[0313] Performing retrieval processing based on partial dialogue information selected from multiple designated categories of dialogue information to obtain training reference dialogue examples for the partial dialogue information;
[0314] Filling the partial dialogue information, the training reference dialogue example of the partial dialogue information, and the training question into the prompt word template to obtain the training prompt word;
[0315] Inputting the training prompt words into the large language model to be trained, and obtaining training reply information of the part of the dialogue information output by the large language model to be trained;
[0316] Based on the difference between the reply information of the part of the dialogue information and the training reply information of the part of the dialogue information in the reply information of the designated category of dialogue information, training the large language model to be trained to obtain the large language model;
[0317] The reference dialogue example and the training reference dialogue example are at least one dialogue example in the dialogue examples consisting of the designated dialogue information other than the partial dialogue information and the reply information of the designated dialogue information in the designated category dialogue information.
[0318] In a possible implementation, the acquisition unit 901 is further used to acquire training questions, training prompt dialogue information, and reply prompt information of the training prompt dialogue information, wherein the training prompt dialogue is used to acquire prompt information of answers to the training questions;
[0319] The processing unit 902 is further configured to perform a search process based on the training prompt dialogue information to obtain a training reference dialogue example of the training dialogue prompt information;
[0320] The filling unit 906 is further used to fill the training dialogue prompt information, the training reference dialogue example of the training dialogue prompt information and the training question into the prompt word template to obtain the training prompt word;
[0321] The acquisition unit 901 is further configured to input the training prompt word into the large language model to be trained, and acquire the training reply prompt information of the training dialogue prompt information output by the large language model to be trained;
[0322] The training unit 907 is further configured to train the large language model to be trained based on the difference between the reply prompt information of the training dialogue prompt information and the training reply prompt information of the training dialogue prompt information to obtain the large language model.
[0323] According to one embodiment of the present application, Figure 3 Some steps involved in the reply processing method of the dialogue information shown can be Fig. 9 The reply processing device of the dialogue information shown in the figure is executed by each unit. For example, Figure 2 The step S301 shown in FIG. Fig. 9 The acquisition unit 901 shown in FIG. 1 is executed, and step S902 can be performed by Fig. 9 The processing unit 902 shown in FIG. 1 is executed, and steps S903 to S904 may be performed by Fig. 9 The generation unit 903 shown executes. Fig. 9The various units in the shown dialog information reply processing device can be separately or completely combined into one or several other units to form, or one (some) of the units can be further divided into two or more functionally smaller units to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be implemented by two or more units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the dialog information reply processing device can also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of two or more units.
[0324] According to another embodiment of the present application, the program can be executed by running a program on a general computing device such as a computer device including a central processing unit (CPU), a random access memory medium (RAM), a read-only memory medium (ROM), and other processing elements and storage elements. Figure 3 A computer program (including program code) for each step involved in the corresponding method shown in Fig. 9 The computer program can be recorded on a computer-readable recording medium, loaded into the computing device through the computer-readable recording medium, and run therein.
[0325] Based on the same inventive concept, the principles and beneficial effects of solving the problem by the reply processing device for conversation information provided in the embodiment of the present application are similar to the principles and beneficial effects of solving the problem by the reply processing method for conversation information in the method embodiment of the present application. Please refer to the principles and beneficial effects of the implementation of the method. For the sake of concise description, they will not be repeated here.
[0326] See also Fig.10 , Fig.10 The schematic diagram of the structure of a computer device provided in an embodiment of the present application, the computer device may be a server. Fig.10As shown, the computer device at least includes a processor 1001, a communication interface 1002 and a memory 1003. The processor 1001, the communication interface 1002 and the memory 1003 can be connected via a bus or other means. The processor 1001 (or central processing unit (CPU)) is the computing core and control core of the computer device, which can parse various instructions in the computer device and process various data of the computer device. For example, the CPU can be used to parse the power on and off instructions issued by the object to the computer device, and control the computer device to perform power on and off operations; for example, the CPU can transmit various interactive data between the internal structures of the computer device, and so on. The communication interface 1002 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), which can be used to send and receive data under the control of the processor 1001; the communication interface 1002 can also be used for the transmission and interaction of data within the computer device. The memory 1003 (Memory) is a memory device in the computer device, which is used to store programs and data. It is understandable that the memory 1003 here may include a built-in memory of the computer device, and of course may also include an extended memory supported by the computer device. The memory 1003 provides a storage space, which stores the operating system of the computer device, including but not limited to: Android system, iOS system, Windows Phone system, etc., which is not limited in this application.
[0327] The embodiment of the present application also provides a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space that stores the processing system of the computer device. In addition, a computer program suitable for being loaded and executed by the processor 1001 is also stored in the storage space. It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory, such as at least one disk storage; optionally, it can also be at least one computer-readable storage medium located away from the aforementioned processor.
[0328] In one embodiment, the processor 1001 performs the following operations by running the computer program in the memory 1003:
[0329] Obtaining the dialogue information input by the target object in response to the question;
[0330] Performing a search process based on the dialogue information to obtain a reference dialogue example of the dialogue information, wherein text features of the reference dialogue example match text features of the dialogue information;
[0331] Based on the dialogue information, the question and the reference dialogue example, generating initial response information corresponding to the dialogue information, wherein the initial response information includes a category label of the dialogue information;
[0332] Based on the category label, target reply information corresponding to the dialogue information is generated.
[0333] In a possible implementation, the processor 1001 executes a retrieval process based on the conversation information by running a computer program in the memory 1003 to obtain a reference conversation example of the conversation information, specifically for:
[0334] Obtaining the associated conversation information of the conversation information;
[0335] extracting text features of the conversation information and text features of the associated conversation information;
[0336] Calculating the similarity between the text features of the dialogue information and the text features of each dialogue example in the dialogue example library, and the similarity between the text features of the associated dialogue information and the text features of each dialogue example;
[0337] A set number of dialogue examples are selected from the dialogue example library in descending order of similarity as the reference dialogue examples.
[0338] In a possible implementation, the processor 1001 executes the step of obtaining the associated conversation information of the conversation information by running a computer program in the memory 1003, including at least one of the following steps:
[0339] Acquire a historical conversation record of the target object before the conversation information, and determine the historical conversation record as the associated conversation information;
[0340] The conversation information is input into a text information expansion model, and expanded conversation information generated by the text information expansion model based on the conversation information is obtained, and the expanded conversation information is determined as the associated conversation information.
[0341] In a possible implementation, the processor 1001 generates target reply information corresponding to the dialogue information based on the category label by running a computer program in the memory 1003, specifically for:
[0342] Inputting the category label into a stylized rewriting model, obtaining first reply information of a specified text style generated by the stylized rewriting model based on the category label, and determining the first reply information as the target reply information;
[0343] or,
[0344] The category label and the supplementary content included in the initial reply information are input into the stylized rewriting model, and the second reply information of the specified text style generated by the stylized rewriting model based on the category label and the supplementary content is obtained, and the second reply information is determined as the target reply information.
[0345] In a possible implementation, the dialogue information is text information input by the target object in the process of answering the question; the category label includes a control instruction; the processor 1001 is further configured to execute, by running the computer program in the memory 1003:
[0346] If the control instruction is used to instruct to exit the answering process, then the answering process of the question is exited.
[0347] If the control instruction is used to instruct to change the question, a new question is generated.
[0348] In a possible implementation, the processor 1001 generates initial reply information corresponding to the dialogue information based on the dialogue information, the question, and the reference dialogue example by running a computer program in the memory 1003, specifically for:
[0349] Obtaining the historical conversation record of the target object before the conversation information;
[0350] Filling the dialogue information, the question, the reference dialogue example and the historical dialogue record into a set prompt word template to obtain a target prompt word;
[0351] The target prompt word is input into a large language model, and reply information generated by the large language model based on the answer strategy and reply constraint information associated with the prompt word template is obtained to obtain the initial reply information.
[0352] In a possible implementation, the processor 1001 generates target reply information corresponding to the dialogue information based on the category label by running a computer program in the memory 1003, specifically for:
[0353] If the category tag indicates that the dialogue information hits key information in the answer to the question, then obtaining the hit key information of the target object for the question;
[0354] If it is determined based on the hit key information that the target object has hit all the key information in the answer, then the category label is updated to a specified category label;
[0355] Based on the specified category label, the target reply information is generated.
[0356] In a possible implementation, the processor 1001 is further configured to execute, by running the computer program in the memory 1003:
[0357] If it is detected that the dialogue turn of the dialogue information is greater than the set turn threshold, a reply strategy of the dialogue information is added to the reply constraint information to obtain an updated prompt word template, wherein the reply strategy is used to instruct the large language model to output prompt information of an answer to the question;
[0358] The step of filling the dialogue information, the question, the reference dialogue example and the historical dialogue record into a set prompt word template to obtain a target prompt word includes:
[0359] The dialogue information, the question, the reference dialogue example and the historical dialogue record are filled into the updated prompt word template to obtain the target prompt word.
[0360] In a possible implementation, the processor 1001 is further configured to execute, by running the computer program in the memory 1003:
[0361] Acquire training questions and multi-round dialogue constraint information, where the multi-round dialogue constraint information is used to constrain the answering process of the training questions;
[0362] Filling the training question and the multi-round dialogue constraint information into the prompt word template to obtain the training dialogue prompt words;
[0363] Inputting the training dialogue prompt words into a dialogue generation model, and obtaining a dialogue record generated by the dialogue generation model based on the training dialogue prompt words;
[0364] The large language model to be trained is trained based on the conversation record to obtain the large language model.
[0365] In a possible implementation, the conversation record includes multiple rounds of conversation information and reply information of each round of conversation information, wherein the historical conversation records of different rounds of conversation information are different; the processor 1001 executes training of the large language model to be trained based on the conversation record by running the computer program in the memory 1003 to obtain the large language model, specifically for:
[0366] Performing retrieval processing based on the dialogue information of each round and the historical dialogue records of the dialogue information of each round to obtain training reference dialogue examples of the dialogue information of each round;
[0367] Filling the dialogue information of each round, the historical dialogue record of each round, the training reference dialogue example of each round, and the training question into the prompt word template to obtain the training prompt word;
[0368] Inputting the training prompt words into the large language model to be trained, and obtaining the training response information of each round of dialogue information output by the large language model to be trained;
[0369] Based on the difference between the training reply information of each round of dialogue information and the reply information of each round of dialogue information, the large language model to be trained is trained to obtain the large language model.
[0370] In a possible implementation, the processor 1001 is further configured to execute, by running the computer program in the memory 1003:
[0371] Decomposing the answer to the training question into at least one key information, and generating supplementary explanation content and question examples for each key information;
[0372] Generating supplementary explanation content for the training question based on the training question and the answer;
[0373] Combining the training question, the at least one key information, the supplementary explanation content of the training question, the supplementary explanation content of each key information and the question example into a target training question;
[0374] The processor 1001 executes the computer program in the memory 1003 to fill the training question and the multi-round dialogue constraint information into the prompt word template to obtain the training dialogue prompt word, which is specifically used for:
[0375] The target training question and the multi-round dialogue constraint information are filled into the prompt word template to obtain the training dialogue prompt words.
[0376] In a possible implementation, the processor 1001 is further configured to execute, by running the computer program in the memory 1003:
[0377] Obtaining training questions and designated category dialogue information for the training questions;
[0378] Filling the training question and the designated category dialogue information into the prompt word template to obtain the training dialogue prompt word;
[0379] Inputting the training dialogue prompt words into a dialogue generation model, and obtaining reply information of the designated category dialogue information generated by the dialogue generation model based on the training dialogue prompt words;
[0380] The large language model to be trained is trained based on the designated category dialogue information and the reply information of the designated category dialogue information to obtain the large language model.
[0381] In a possible implementation, the number of the designated category dialogue information is multiple; the processor 1001 executes the training of the large language model to be trained based on the designated category dialogue information and the reply information of the designated category dialogue information by running the computer program in the memory 1003 to obtain the large language model, which is specifically used for:
[0382] Performing retrieval processing based on partial dialogue information selected from multiple designated categories of dialogue information to obtain training reference dialogue examples for the partial dialogue information;
[0383] Filling the partial dialogue information, the training reference dialogue example of the partial dialogue information, and the training question into the prompt word template to obtain the training prompt word;
[0384] Inputting the training prompt words into the large language model to be trained, and obtaining training reply information of the part of the dialogue information output by the large language model to be trained;
[0385] Based on the difference between the reply information of the part of the dialogue information and the training reply information of the part of the dialogue information in the reply information of the designated category of dialogue information, training the large language model to be trained to obtain the large language model;
[0386] The reference dialogue example and the training reference dialogue example are at least one dialogue example in the dialogue examples consisting of the designated dialogue information other than the partial dialogue information and the reply information of the designated dialogue information in the designated category dialogue information.
[0387] In a possible implementation, the processor 1001 is further configured to execute, by running the computer program in the memory 1003:
[0388] Acquire training questions, training prompt dialogue information, and reply prompt information of the training prompt dialogue information, wherein the training prompt dialogue is used to acquire prompt information of answers to the training questions;
[0389] Performing a search process based on the training prompt dialogue information to obtain a training reference dialogue example of the training dialogue prompt information;
[0390] Filling the training dialogue prompt information, the training reference dialogue example of the training dialogue prompt information, and the training question into the prompt word template to obtain the training prompt word;
[0391] Inputting the training prompt words into the large language model to be trained, and obtaining training reply prompt information of the training dialogue prompt information output by the large language model to be trained;
[0392] Based on the difference between the reply prompt information of the training dialogue prompt information and the training reply prompt information of the training dialogue prompt information, the large language model to be trained is trained to obtain the large language model.
[0393] Based on the same inventive concept, the principles and beneficial effects of solving the problem by the computer device provided in the embodiment of the present application are similar to the principles and beneficial effects of solving the problem by the method for replying to the conversation information in the method embodiment of the present application. Please refer to the principles and beneficial effects of the implementation of the method. For the sake of concise description, they will not be repeated here.
[0394] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is suitable for being loaded by a processor and executing the method for replying to the conversation information of the above method embodiment.
[0395] The embodiment of the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the above-mentioned method for processing the reply of the dialogue information.
[0396] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.
[0397] The modules in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.
[0398] In the embodiments of the present application, the "module" or "unit" involved refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, at least one processor (or memory) can be used to implement at least one module or unit. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0399] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the readable storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0400] What is disclosed above is only a preferred embodiment of the present application, and it certainly cannot be used to limit the scope of rights of the present application. Ordinary technicians in this field can understand that all or part of the processes of implementing the above embodiment and equivalent changes made according to the claims of the present application are still within the scope covered by the application.
Claims
1. A method for processing a reply to a conversation message, characterized in that: include: Obtaining the dialogue information input by the target object in response to the question; Performing a search process based on the dialogue information to obtain a reference dialogue example of the dialogue information, wherein text features of the reference dialogue example match text features of the dialogue information; Based on the dialogue information, the question and the reference dialogue example, generating initial response information corresponding to the dialogue information, wherein the initial response information includes a category label of the dialogue information; Based on the category label, target reply information corresponding to the dialogue information is generated.
2. The method according to claim 1, characterized in that The performing a search process based on the conversation information to obtain a reference conversation example of the conversation information includes: Obtaining the associated conversation information of the conversation information; extracting text features of the conversation information and text features of the associated conversation information; Calculating the similarity between the text features of the dialogue information and the text features of each dialogue example in the dialogue example library, and the similarity between the text features of the associated dialogue information and the text features of each dialogue example; A set number of dialogue examples are selected from the dialogue example library in descending order of similarity as the reference dialogue examples.
3. The method according to claim 2, characterized in that The step of obtaining the associated conversation information of the conversation information comprises at least one of the following steps: Acquire a historical conversation record of the target object before the conversation information, and determine the historical conversation record as the associated conversation information; The conversation information is input into a text information expansion model, and expanded conversation information generated by the text information expansion model based on the conversation information is obtained, and the expanded conversation information is determined as the associated conversation information.
4. The method according to claim 1, characterized in that: The generating target reply information corresponding to the dialogue information based on the category label includes: Inputting the category label into a stylized rewriting model, obtaining first reply information of a specified text style generated by the stylized rewriting model based on the category label, and determining the first reply information as the target reply information; or, The category label and the supplementary content included in the initial reply information are input into the stylized rewriting model, and the second reply information of the specified text style generated by the stylized rewriting model based on the category label and the supplementary content is obtained, and the second reply information is determined as the target reply information.
5. The method according to claim 1, characterized in that The dialogue information is text information input by the target object in the process of answering the question; the category label includes a control instruction; after generating the target reply information corresponding to the dialogue information based on the category label, the method further includes: If the control instruction is used to instruct to exit the answering process, then the answering process of the question is exited. If the control instruction is used to instruct to change the question, a new question is generated.
6. The method according to claim 1, characterized in that The generating initial reply information corresponding to the dialogue information based on the dialogue information, the question and the reference dialogue example includes: Obtaining the historical conversation record of the target object before the conversation information; Filling the dialogue information, the question, the reference dialogue example and the historical dialogue record into a set prompt word template to obtain a target prompt word; The target prompt word is input into a large language model, and reply information generated by the large language model based on the answer strategy and reply constraint information associated with the prompt word template is obtained to obtain the initial reply information.
7. The method according to claim 6, characterized in that The generating target reply information corresponding to the dialogue information based on the category label includes: If the category tag indicates that the dialogue information hits key information in the answer to the question, then obtaining the hit key information of the target object for the question; If it is determined based on the hit key information that the target object has hit all the key information in the answer, then the category label is updated to a specified category label; Based on the specified category label, the target reply information is generated.
8. The method according to claim 6, characterized in that The method further comprises: If it is detected that the dialogue turn of the dialogue information is greater than the set turn threshold, a reply strategy of the dialogue information is added to the reply constraint information to obtain an updated prompt word template, wherein the reply strategy is used to instruct the large language model to output prompt information of an answer to the question; The step of filling the dialogue information, the question, the reference dialogue example and the historical dialogue record into a set prompt word template to obtain a target prompt word includes: The dialogue information, the question, the reference dialogue example and the historical dialogue record are filled into the updated prompt word template to obtain the target prompt word.
9. The method according to claim 6, characterized in that The method further comprises: Acquire training questions and multi-round dialogue constraint information, where the multi-round dialogue constraint information is used to constrain the answering process of the training questions; Filling the training question and the multi-round dialogue constraint information into the prompt word template to obtain the training dialogue prompt words; Inputting the training dialogue prompt words into a dialogue generation model, and obtaining a dialogue record generated by the dialogue generation model based on the training dialogue prompt words; The large language model to be trained is trained based on the conversation record to obtain the large language model.
10. The method according to claim 9, characterized in that The conversation record includes multiple rounds of conversation information and reply information of each round of conversation information, wherein the historical conversation records of different rounds of conversation information are different; and training the large language model to be trained based on the conversation record to obtain the large language model includes: Performing retrieval processing based on the dialogue information of each round and the historical dialogue records of the dialogue information of each round to obtain training reference dialogue examples of the dialogue information of each round; Filling the dialogue information of each round, the historical dialogue record of each round, the training reference dialogue example of each round, and the training question into the prompt word template to obtain the training prompt word; Inputting the training prompt words into the large language model to be trained, and obtaining the training response information of each round of dialogue information output by the large language model to be trained; Based on the difference between the training reply information of each round of dialogue information and the reply information of each round of dialogue information, the large language model to be trained is trained to obtain the large language model.
11. The method according to claim 9, characterized in that The method further comprises: Decomposing the answer to the training question into at least one key information, and generating supplementary explanation content and question examples for each key information; Generating supplementary explanation content for the training question based on the training question and the answer; Combining the training question, the at least one key information, the supplementary explanation content of the training question, the supplementary explanation content of each key information and the question example into a target training question; The step of filling the training question and the multi-round dialogue constraint information into the prompt word template to obtain the training dialogue prompt words includes: The target training question and the multi-round dialogue constraint information are filled into the prompt word template to obtain the training dialogue prompt words.
12. The method according to claim 6, characterized in that The method further comprises: Obtaining training questions and designated category dialogue information for the training questions; Filling the training question and the designated category dialogue information into the prompt word template to obtain the training dialogue prompt word; Inputting the training dialogue prompt words into a dialogue generation model, and obtaining reply information of the designated category dialogue information generated by the dialogue generation model based on the training dialogue prompt words; The large language model to be trained is trained based on the designated category dialogue information and the reply information of the designated category dialogue information to obtain the large language model.
13. The method according to claim 12, characterized in that The number of the designated category dialogue information is multiple; the training of the large language model to be trained based on the designated category dialogue information and the reply information of the designated category dialogue information to obtain the large language model includes: Performing retrieval processing based on partial dialogue information selected from multiple designated categories of dialogue information to obtain training reference dialogue examples for the partial dialogue information; Filling the partial dialogue information, the training reference dialogue example of the partial dialogue information, and the training question into the prompt word template to obtain the training prompt word; Inputting the training prompt words into the large language model to be trained, and obtaining training reply information of the part of the dialogue information output by the large language model to be trained; Based on the difference between the reply information of the part of the dialogue information and the training reply information of the part of the dialogue information in the reply information of the designated category of dialogue information, training the large language model to be trained to obtain the large language model; The reference dialogue example and the training reference dialogue example are at least one dialogue example in the dialogue examples consisting of the designated dialogue information other than the partial dialogue information and the reply information of the designated dialogue information in the designated category dialogue information.
14. The method according to claim 6, characterized in that The method further comprises: Acquire training questions, training prompt dialogue information, and reply prompt information of the training prompt dialogue information, wherein the training prompt dialogue is used to acquire prompt information of answers to the training questions; Performing a search process based on the training prompt dialogue information to obtain a training reference dialogue example of the training dialogue prompt information; Filling the training dialogue prompt information, the training reference dialogue example of the training dialogue prompt information, and the training question into the prompt word template to obtain the training prompt word; Inputting the training prompt words into the large language model to be trained, and obtaining training reply prompt information of the training dialogue prompt information output by the large language model to be trained; Based on the difference between the reply prompt information of the training dialogue prompt information and the training reply prompt information of the training dialogue prompt information, the large language model to be trained is trained to obtain the large language model.
15. A device for processing a reply to a conversation message, characterized in that: include: An acquisition unit, used to acquire dialogue information input by the target object in response to the question; a processing unit, configured to perform a search process based on the dialogue information to obtain a reference dialogue example of the dialogue information, wherein a text feature of the reference dialogue example matches a text feature of the dialogue information; A generating unit, configured to generate initial response information corresponding to the dialogue information based on the dialogue information, the question and the reference dialogue example, wherein the initial response information includes a category label of the dialogue information; The generating unit is further configured to generate target reply information corresponding to the dialogue information based on the category label.
16. An electronic device, characterized in that: include: one or more processors; The memory is used to store one or more computer programs, and when the one or more computer programs are executed by the one or more processors, the electronic device implements the method for replying to the dialogue information described in any one of claims 1-14.
17. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for processing the reply of the dialogue information according to any one of claims 1 to 14 is implemented.
18. A computer program product, characterized in that The computer program product includes a computer program, which is stored in a computer-readable storage medium. A processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the method for replying to a conversation message according to any one of claims 1 to 14.