Dialogue processing method and emoji generation method
By introducing expression tags and image generation into the dialogue model, the problem of single dialogue content in the prior art is solved, and the diversity of dialogue and user experience is improved.
Patent Information
- Application Number
- PCT/CN2024/124278
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-11
- Publication Date
- 2025-05-08
AI Technical Summary
In the prior art, the artificial intelligence dialogue model can only conduct dialogue through text, resulting in a single dialogue content, insufficient fun and poor user experience.
By obtaining the initial conversation text, input the target conversation model to generate a reply text including emoticon tags, determine the emoticon image based on the emoticon tags, and generate reply content based on the emoticon image to improve the richness and fun of the conversation.
The diversity and fun of dialogue content have been improved, the user experience has been significantly improved, and more rich and anthropomorphic dialogue services have been provided.
Smart Images

Figure CN2024124278_08052025_PF_FP_ABST
Abstract
Description
Dialogue processing method and expression image generation method
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on October 30, 2023, with application number 202311435045.7 and application name “Dialogue Processing Method and Expression Image Generation Method”, the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0002] The embodiments of the present disclosure relate to the field of deep learning technology, and in particular to a conversation processing method and an expression image generation method. Background Art
[0003] With the development of deep learning technology, deep learning models represented by large models have been rapidly developed and applied in various fields.
[0004] At present, in the field of artificial intelligence dialogue, the dialogue text is input into the dialogue model, and after understanding the dialogue text, dialogue prediction is performed, and reply text is generated to complete the dialogue with the user. It realizes the ability to understand user needs, provide personified dialogue services, and personalize customized dialogues, and provides accurate services and emotional companionship.
[0005] However, conversation models can only communicate through text, resulting in monotonous conversation content, insufficient interest, and poor user experience. Therefore, a method for processing conversations with richer content is urgently needed.
[0006] Summary of the Invention
[0007] In light of this, embodiments of the present disclosure provide a method for processing conversations. One or more embodiments of the present disclosure also include a method for generating an expression image, another method for processing conversations, a conversation processing device, an expression image generating device, another conversation processing device, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.
[0008] In one embodiment of the present disclosure, a conversation processing method is provided, including:
[0009] Get the initial conversation text;
[0010] Input the initial dialogue text into the target dialogue model to obtain the target reply text, wherein the target reply text includes the target expression label. The target dialogue model is trained based on the predicted reply text and the sample reply text. The predicted reply text is obtained by the target dialogue model through dialogue prediction on the sample dialogue text based on the prompt text. The sample reply text includes the sample expression label. The prompt text is used to prompt the target dialogue model to predict the expression label for the predicted reply text.
[0011] Determine the target expression image according to the target expression label;
[0012] Generate target reply content based on the target expression image.
[0013] In one embodiment of the present disclosure, a method for generating an expression image is provided, comprising:
[0014] Get the initial text;
[0015] Inputting the initial text into a text generation model to obtain a target text, wherein the target text includes a target expression label, the text generation model is trained based on the predicted text and the label text, the predicted text is obtained by the text generation model predicting the sample text based on the prompt text, the label text includes a sample expression label, and the prompt text is used to prompt the text generation model to predict the expression label for the predicted text;
[0016] Based on the target expression tag, a target expression image is generated.
[0017] In one embodiment of the present disclosure, a conversation processing method is provided, which is applied to a cloud-side device and includes:
[0018] receiving the initial conversation text sent by the terminal device;
[0019] Inputting the initial dialogue text into a target dialogue model to obtain a target reply text, wherein the target reply text includes a target expression label, the target dialogue model is trained based on the predicted reply text and the sample reply text, the predicted reply text is obtained by the target dialogue model performing dialogue prediction on the sample dialogue text based on the prompt text, the sample reply text includes a sample expression label, and the prompt text is used to prompt the target dialogue model to predict the expression label for the predicted reply text;
[0020] Determining a target expression image according to the target expression label;
[0021] Generating target reply content based on the target expression image;
[0022] Feedback the target reply content to the terminal device.
[0023] In one embodiment of the present disclosure, a computing device is provided, including:
[0024] memory and processor;
[0025] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. The computer-executable instructions are used by the processor to execute the steps of the above method.
[0026] In one embodiment of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the above method are implemented.
[0027] In one embodiment of the present disclosure, a computer program is provided. When the computer program is executed in a computer, the computer is caused to execute the steps of the above method.
[0028] In one embodiment of the present disclosure, an initial conversation text is obtained; the initial conversation text is input into a target conversation model to obtain a target reply text, wherein the target reply text includes a target expression label. The target conversation model is trained based on the predicted reply text and sample reply text. The predicted reply text is obtained by the target conversation model performing conversation prediction on the sample conversation text based on prompt text. The sample reply text includes a sample expression label. The prompt text is used to prompt the target conversation model to predict the expression label for the predicted reply text; a target expression image is determined based on the target expression label; and target reply content is generated based on the target expression image. Leveraging the text understanding capabilities of the target conversation model, a deep learning model, guided by the prompt text, the target conversation model is prompted to perform conversation prediction and expression label prediction on the sample conversation text, thereby obtaining a predicted reply text including the expression label. Together with the sample reply text including the sample expression label, the target conversation model is trained. This allows the target conversation model to generate a target reply text including the target expression label based on the input initial conversation text. Furthermore, based on the target expression label, a target expression image is determined, thereby generating more targeted and rich target reply content, enhancing the fun of the conversation and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] FIG1 is a flow chart of a conversation processing method provided by one embodiment of the present disclosure;
[0030] FIG2 is a schematic diagram of a prompt text of a conversation processing method provided by an embodiment of the present disclosure;
[0031] FIG3 is a schematic diagram of a predicted reply text of a conversation processing method provided by an embodiment of the present disclosure;
[0032] FIG4 is a schematic diagram of an example text of a conversation processing method provided by an embodiment of the present disclosure;
[0033] FIG5 is a schematic diagram of an expression image library of a conversation processing method provided by an embodiment of the present disclosure;
[0034] FIG6 is a front-end schematic diagram of a conversation processing method provided by an embodiment of the present disclosure;
[0035] FIG7 is a flow chart of a method for generating an expression image provided by one embodiment of the present disclosure;
[0036] FIG8 is a flowchart of another method for processing a conversation provided by an embodiment of the present disclosure;
[0037] FIG9 is a flowchart of a method for training a dialogue model according to an embodiment of the present disclosure;
[0038] FIG10 is a flowchart of a process of a dialogue processing method applied to a personified dialogue scenario provided by one embodiment of the present disclosure;
[0039] FIG11 is a structural diagram of a conversation processing device provided by an embodiment of the present disclosure;
[0040] FIG12 is a schematic structural diagram of an expression image generating device provided by one embodiment of the present disclosure;
[0041] FIG13 is a schematic structural diagram of another conversation processing device provided by an embodiment of the present disclosure;
[0042] FIG14 is a schematic structural diagram of a dialogue model training device provided by one embodiment of the present disclosure;
[0043] FIG15 is a structural block diagram of a computing device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0044] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.
[0045] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.
[0046] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0047] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0048] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, which typically contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model (Foundation Model), which is pre-trained by using large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks, and the model has good generalization capabilities, such as a large-scale language model (LLM), a multi-modal pre-training model, etc.
[0049] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0050] First, the terms involved in one or more embodiments of the present disclosure are explained.
[0051] Large Language Model (LLM): A deep learning model trained using large amounts of text data and possessing a large number of model parameters. Large language models can generate natural language text or understand the meaning of text. Large language models can handle a variety of natural language tasks, such as text classification, question-answering, and conversation, and are an important path forward in artificial intelligence. These models include generative and inferential LLMs.
[0052] RNN (Recurrent Neural Network) model: A deep learning model that can be used to process sequential data. Its main feature is its ability to remember previous information and reuse it in subsequent processing. This memory ability makes RNN models very useful when processing sequential data. For example, in natural language processing, they can be used to generate new sentences, paragraphs, and articles. The RNN model structure contains one or more recurrent layers, which enable information to be transferred and remembered.
[0053] Transformer (Translation) Model: A deep learning model for processing sequential data with an attention mechanism, such as text and speech. Transformers can be used to handle natural language processing tasks such as machine translation, question-answering systems, sentiment analysis, and text summarization. The main feature of Transformers is that they do not need to process sequential data sequentially, but can process them in parallel, which makes them faster to train. The structure of Transformers contains one or more Transformer layers, which can learn long-term dependencies in the input sequence.
[0054] BERT (Bidirectional Encoder Representations from Transformers) model: a bidirectional encoding Transformer model.
[0055] Emoticons: An image used to express emotion or status, commonly used in online communication such as social media and instant messaging. Emoticons can be static images or animated GIFs (Graphics Interchange Format). Emoticons typically contain one or more facial expressions and can be used to convey a variety of emotions, such as happiness, sadness, surprise, and anger.
[0056] Supervised Fine-Tuning: A deep learning model, the source model, is pre-trained on a source dataset. A new deep learning model, the target model, is then created. The target model replicates the entire model design and parameters of the source model, except for the output layer. These model parameters incorporate the knowledge learned on the source dataset, and this knowledge is also applicable to the target dataset. During fine-tuning, an output layer with an output size equal to the number of categories in the target dataset is added to the target model, and the model parameters of this layer are randomly initialized. When training the target model on the target dataset, it is trained from scratch up to the output layer, and the model parameters of the remaining layers are fine-tuned based on the model parameters of the source model.
[0057] Prompt text: In large-scale model applications, prompt text provides the model with the context of input information and parameter information. When training supervised or unsupervised learning models, prompt text can help the model better understand the input intent and provide corresponding feedback.
[0058] In the present disclosure, a dialogue processing method is provided. The present disclosure also relates to another dialogue processing method, a dialogue model training method, a dialogue processing device, another dialogue processing device, a dialogue model training device, a computing device, a computer-readable storage medium and a computer program, which are described in detail one by one in the following embodiments.
[0059] Referring to FIG1 , FIG1 shows a flow chart of a conversation processing method provided by an embodiment of the present disclosure, which includes the following specific steps:
[0060] Step 102: Obtain the initial conversation text.
[0061] The disclosed embodiments are applicable to applications, websites, or mini-programs with conversation processing capabilities, such as applications with social networking features, e-commerce platforms with customer service features, and games with intelligent NPCs (Non-Player Characters). The disclosed embodiments can be applied to the client or server of an application, website, or mini-program.
[0062] The conversation text is the natural language text to be responded to in a conversation scenario. The initial conversation text can be natural language text directly entered by the user or converted from user-input voice data, without limitation. For example, the initial conversation text may be "What are you doing?" entered by a user in a social application.
[0063] Get the initial conversation text. The specific method is: Get the initial conversation text entered by the user.
[0064] For example, user XXX logs in to the application of the social software on the terminal device, selects the artificial intelligence character "Xiao A" on the application, enters the initial dialogue text "What are you doing?" in the dialog box, and sends it to the artificial intelligence character "Xiao A".
[0065] Obtain the initial conversation text, laying the foundation for generating the target response text.
[0066] Step 104: Input the initial dialogue text into the target dialogue model to obtain the target reply text, wherein the target reply text includes the target expression label. The target dialogue model is trained based on the predicted reply text and the sample reply text. The predicted reply text is obtained by the target dialogue model performing dialogue prediction on the sample dialogue text based on the prompt text. The sample reply text includes the sample expression label. The prompt text is used to prompt the target dialogue model to predict the expression label for the predicted reply text.
[0067] The target conversation model is a pre-trained text generation model with conversation prediction and emoticon label prediction capabilities. It is a deep learning model, such as an RNN model, a Transformer model, a BERT model, or a large language model. The target conversation model is trained based on predicted reply text and sample reply text. The target conversation model can be deployed directly on the client or server of an application, website, or mini-program, or it can be called by the client or server of an application, website, or mini-program through a data interface (e.g., an API (Application Programming Interface), a data stream interface, an RPC (Remote Procedure Call) interface, a SOAP (Simple Object Access Protocol) interface, etc.). For example, the target conversation model can be deployed on a distributed database and remotely called by the social software server via an API. Multiple conversation models can be pre-trained, and the corresponding one can be determined as the target conversation model. For example, if a user selects "Xiao A" as the conversation character on the front end, the "Xiao A" conversation model can be determined as the target conversation model from three conversation models: "Xiao A," "Xiao B," and "Xiao C."
[0068] The response text is the natural language text that responds to the conversation text in a conversation scenario. The target response text is the natural language text generated by the target conversation model based on the conversation prediction of the initial conversation text. This is the result of sentiment analysis performed on the initial conversation text after anthropomorphizing the target conversation model. For example, the target response text is "I'm reading comics and playing video games! [[Showing cuteness]]." The target expression label is the text that labels the expression in the target response text. The target expression label is essentially natural language text, a special natural language text generated by the target conversation model. The target expression label is emotionally correlated with the other text in the target response text. For example, in the target response text "I'm reading comics and playing video games! [[Showing cuteness]]," "[Showing cuteness]]" is the target expression label, and the other text "I'm reading comics and playing video games" is a light-hearted expression. The target expression label "[Showing cuteness]]" is also a light-hearted expression, so the two texts are emotionally correlated.
[0069] Sample conversation texts are natural language sample texts to be responded to during the conversation model training process. Sample responses are natural language sample texts that respond to conversation texts during the conversation model training process. Sample conversation texts and sample responses are matched in pairs. For example, the sample conversation text is "Are you good at games?" and the sample response text is "I think it's okay, but I occasionally lose and feel a little unhappy [[grievance]]." Sample expression labels are text labeled with expressions in sample responses. Sample expression labels are essentially pre-labeled natural language texts. Sample expression labels are emotionally correlated with other text in the sample responses. For example, in the sample response text "I think it's okay, but I occasionally lose and feel a little unhappy [[grievance]]," "[[grievance]]" is the sample expression label, while the text "I occasionally lose and feel a little unhappy" expresses a feeling of aggrieved. The sample expression label "[[grievance]]" also expresses a feeling of aggrieved. Therefore, the two are emotionally correlated. The predicted reply text is the natural language text generated by dialogue prediction for the sample dialogue text during the training process of the dialogue model. For example, if the sample dialogue text is "Are you good at playing games?", the predicted reply text generated by the target dialogue model for the sample dialogue text is "Actually, it's okay, but I just lose a few games occasionally [[unfair]]".
[0070] The prompt text is the text used to prompt the dialogue model for dialogue prediction. The prompt text is used to prompt and guide the dialogue model to understand the dialogue prediction task, so that the dialogue model generates corresponding predicted response texts according to the prompt text. The dialogue prediction task includes, but is not limited to: the indication information of the response text generation and the indication information of the emoji label. Among them, the indication information of the response text generation includes, but is not limited to: the text style indication information, the role information indication information, and the text format example indication information. The indication information of the emoji label includes, but is not limited to: the role information indication information, the emoji label position indication information, and the emoji label library indication information. For example, the prompt text is "<role information> Name: Xiao A...; <emoji label library> [[OK]], [[hello]], [[stunned]]...; <text style and emoji label position> cute, energetic, colloquial, please use the labels in the emoji library [[…]] to reply according to the actual situation; <text format example> XXXX[[…]], YYYY[[…]]".
[0071] Input the initial dialogue text into the target dialogue model to obtain the target response text. The specific method is: input the initial dialogue text into the target dialogue model, perform dialogue prediction and emoji label prediction on the initial dialogue text, and obtain the target response text including the target emoji label. It should be noted that when the target dialogue model performs dialogue prediction and emoji label prediction, it is carried out synchronously, rather than first completing the dialogue prediction and then performing sentiment analysis on the predicted target response text to generate the corresponding emoji label prediction.
[0072] Exemplarily, input the initial dialogue text "What are you doing?" into the large language model corresponding to the artificial intelligence role "Xiao A", perform dialogue prediction and emoji label prediction on the initial dialogue text, and obtain the target response text "Thinking of you~[[missing you]] What about you?" including the target emoji label "[[missing you]]".
[0073] The initial conversation text is input into the target conversation model to obtain a target reply text, where the target reply text includes a target expression label. The target conversation model is trained based on the predicted reply text and sample reply text. The predicted reply text is obtained by the target conversation model based on the prompt text, which predicts the conversation of the sample conversation text. The sample reply text includes a sample expression label. The prompt text prompts the target conversation model to predict the expression label for the predicted reply text. Leveraging the text understanding capabilities of the target conversation model, a deep learning model, guided by the prompt text, the target conversation model is prompted to perform conversation prediction and expression label prediction for the sample conversation text. The predicted reply text, including the expression label, is obtained. Together with the sample reply text including the sample expression label, the target conversation model is trained. This allows the target conversation model to generate the target reply text including the target expression label for the initial conversation text input, laying the foundation for the subsequent determination of the target expression image.
[0074] Step 106: Determine the target expression image according to the target expression label.
[0075] The target expression image is an expression image corresponding to the target expression image generated by the target dialogue model, and the target expression image has an emotional correlation with the target reply text. The target expression image can be understood as an expression image sent in response to the initial dialogue text after the target dialogue model is personified, including but not limited to: static pictures and dynamic images, in the following formats: GIF, PNG (Portable Network Graphics), BMP (Bitmap), JPG (Joint Photographic Experts Group), TIFF (Tagged Image File Format). It should be noted that the target expression image is an expression image of the personified target dialogue model responding to the initial dialogue text, and is not an expression image obtained by querying after performing sentiment analysis on the target reply text.
[0076] The target expression image is determined based on the target expression tag. Specifically, the target expression image is obtained by querying the target expression tag. This can be done by querying a preset expression image library or by scheduling a search engine to query open source expression images. There is no limitation here.
[0077] For example, according to the target expression tag "[[Miss you]]", the preset expression image library is queried (including 30 gif expression images: [[OK]], [[Hello]], [[Staying]], [[Idle]], [[Well-behaved]], [[Goodbye]], [[Come on]], [[Cute]], [[Lovely]], [[Haha]], [[Crying]], [[Um]], [[Hehe]], [[Are you there]], [[Great]], [[Studying]], [[Afraid]], [[Cheers]], [[Great]], [[Surprised]], [[Miss you]]...), and the target expression image is obtained: Miss you.gif.
[0078] The target expression image is determined based on the target expression tag. This provides more targeted and rich response content, laying the foundation for subsequent generation of more targeted and rich target response content.
[0079] Step 108: Generate target reply content based on the target expression image.
[0080] The target reply content is the conversation content used to respond to the user in the conversation scenario, including the target emoticon image, and may also include the target reply text and / or the voice data of the target reply text. For example, the target reply content may include the target reply text "I'm reading comics and playing video games!" and the target emoticon image "Ming Meng.gif".
[0081] Optionally, after generating the target reply content, the following steps are further included:
[0082] Feedback the target response content to the user.
[0083] Based on the target expression image, the target reply content is generated. The target reply content can be generated based on the target expression image and the target reply text, that is, the target reply content includes the target expression image and the target reply text. The target reply content can also be generated based only on the target expression image, that is, the target reply content only includes the target expression image. There is no limitation here.
[0084] For example, based on the target emoticon image (missing you.gif) and the target reply text "Missing you~[[Missing you]]How about you?", target reply content is generated: Missing you~< / s>Missing you.gif< / s>How about you? This target reply content is fed back to the client of the application logged in by user XXX and displayed on the conversation display page of the application.
[0085] In the disclosed embodiment, an initial conversation text is obtained; the initial conversation text is input into a target conversation model to obtain a target reply text, wherein the target reply text includes a target expression label. The target conversation model is trained based on the predicted reply text and sample reply text. The predicted reply text is obtained by the target conversation model performing conversation prediction on the sample conversation text based on a prompt text. The sample reply text includes a sample expression label. The prompt text is used to prompt the target conversation model to predict the expression label for the predicted reply text. Based on the target expression label, a target expression image is determined; and based on the target expression image, target reply content is generated. By utilizing the text understanding capabilities of the target conversation model, a deep learning model, guided by the prompt text, the target conversation model is prompted to perform conversation prediction and expression label prediction on the sample conversation text, thereby obtaining a predicted reply text including the expression label. Together with the sample reply text including the sample expression label, the target conversation model is trained. This allows the target conversation model to generate a target reply text including the target expression label based on the input initial conversation text. Furthermore, based on the target expression label, a target expression image is determined, thereby generating more targeted and rich target reply content, enhancing the fun of the conversation and improving the user experience.
[0086] In an optional embodiment of the present disclosure, the target dialogue model includes an encoding layer, a prediction layer, and a decoding layer;
[0087] Correspondingly, step 104 includes the following specific steps:
[0088] Input the initial conversation text into the encoding layer, perform feature encoding on the initial conversation text, and obtain the initial text features;
[0089] Input the initial text features into the prediction layer, perform dialogue prediction and expression label prediction based on the initial text features, and obtain the target text features;
[0090] The target text features are input into the decoding layer, the target text features are decoded, and the target reply text is obtained.
[0091] In the embodiment of the present disclosure, the target dialogue model is a deep learning model with a codec structure, including an encoding layer, a prediction layer, and a decoding layer. The encoding layer of the target dialogue model is used to perform feature encoding on the initial dialogue text, and convert it into a feature encoding vector of the initial dialogue text. For example, if the target dialogue model is a Transformer model, the encoding layer includes an embedding layer and a recurrent neural network layer (RNN, Recurrent Neural Networks), wherein the embedding layer is used to encode the initial dialogue text into a word vector, and the recurrent neural network layer is used to encode these word vectors into high-dimensional semantic features, i.e., initial text features. The prediction layer of the target dialogue model is used to perform high-dimensional feature processing on the feature encoding vector to achieve dialogue prediction and expression label prediction, and obtain target text features including target expression labels. For example, if the target dialogue model is a Transformer model, the prediction layer includes an attention calculation layer, and through attention calculation, a feature encoding vector of the target reply text including the target expression label is obtained. The decoding layer of the target dialogue model is used to decode the feature encoding vector and convert it into a target reply text. For example, the target dialogue model is a Transformer model, and the decoding layer includes the decoding layer of the text model, which usually includes a recurrent neural network layer and a vocabulary layer. The recurrent neural network layer is used to convert the target text features of the encoding layer into an intermediate feature encoding vector, while the vocabulary layer is used to convert the intermediate feature encoding vector into actual text.
[0092] For example, the initial dialogue text "What are you doing?" is input into the encoding layer of the large language model corresponding to the artificial intelligence character "Xiao A", and feature encoding is performed on the initial dialogue text to obtain the initial text feature Feature_Txt. The initial text feature is input into the prediction layer of the large language model, and dialogue prediction and expression label prediction are performed based on the initial text feature to obtain the target text feature Feature_Txt'. The target text feature is input into the decoding layer of the large language model, and the target text feature is feature decoded to obtain the target reply text "Missing you~ [Missing you]] and the target expression label "[[Missing you]]".
[0093] In the disclosed embodiment, the target dialogue model generates a target reply text including a target expression tag for the input initial dialogue text through encoding, prediction and decoding, laying the foundation for the subsequent determination of the target expression image and the generation of more targeted and rich target reply content.
[0094] In an optional embodiment of the present disclosure, before step 104, the following specific steps are further included:
[0095] Obtaining a sample set, an initial dialogue model, and prompt text, wherein the sample set includes multiple sample pairs, the sample pairs include sample dialogue text and sample response text, and the sample response text includes sample expression labels;
[0096] Extracting a first sample pair from the sample set, wherein the first sample pair is any one of the multiple sample pairs, the first sample pair includes a first sample conversation text and a first sample reply text, and the first sample reply text includes a first sample expression label;
[0097] Inputting the first sample conversation text into the conversation model, and performing conversation prediction and expression label prediction on the first sample conversation text based on the prompt text to obtain a first predicted reply text including the first predicted expression label;
[0098] Based on the first sample reply text and the first predicted reply text, the dialogue model is trained to obtain a target dialogue model.
[0099] The sample set is a sample data set used to train a conversation model for conversation prediction and expression label prediction. The sample set can be conversation data from historical conversation scenarios, for example, historical conversation data between multiple users or between multiple users and a conversation model (historical conversation text, historical reply text, and historical expression images). It can also be sample conversation data from an open source database, for example, evaluation conversation data between customer service staff and users publicly available on an e-commerce platform (evaluation conversation text, review reply text, and review expression images). It can also be artificially constructed sample conversation data, for example, game script dialogues written for NPCs in a game (script conversation text, script reply text, and script expression images). A sample pair is a sample conversation data pair consisting of a sample conversation text and a sample reply text, and the first sample pair is any one of the multiple sample pairs.
[0100] The initial dialogue model is a text generation model to be trained with dialogue prediction and expression label prediction functions. It can be a pre-trained dialogue model that requires targeted fine-tuning (Fine-Tune), or it can be an untrained dialogue model that requires targeted training. There is no limitation here.
[0101] Sample conversation text refers to natural language sample text to be responded to during the conversation model training process. Sample response text refers to natural language sample text that responds to conversation text during the conversation model training process. Prompt text refers to text used to prompt the conversation model to perform conversation prediction. Prompt text is used to prompt and guide the target conversation model to understand the conversation prediction task, so that the target conversation model generates corresponding predicted response text based on the prompt text. The prompt text determines the training method for the conversation model. For example, the initial conversation model needs to be trained into a deep learning model that can conduct conversations in a "cute, energetic, and colloquial" manner. The prompt text contains corresponding instructive information and can be understood as a form of anthropomorphism. Sample expression labels refer to text that labels expressions in sample response text. Sample expression labels are essentially pre-labeled natural language text, and they have emotional relevance to other text in the sample response text. For details, please refer to the above embodiments and will not be repeated here. The first sample conversation text refers to the sample conversation text in the first sample pair. The first sample response text refers to the sample response text in the first sample pair. The first sample expression label refers to the sample expression label in the first sample response text. The first predicted reply text is a natural language text generated by performing dialogue prediction on the first sample dialogue text during the training process of the dialogue model.
[0102] A first sample conversation text is input into a conversation model, and based on a prompt text, conversation prediction and expression label prediction are performed on the first sample conversation text to obtain a first predicted reply text including a first predicted expression label. Specifically, the first sample conversation text is input into the conversation model, conversation prediction is performed on the first sample conversation text based on indication information of reply text generation in the prompt text, and expression label prediction is performed on the first sample conversation text based on indication information of expression labels in the prompt text to obtain the first predicted reply text including the first predicted expression label. The indication information of reply text generation and expression label indication information include, but are not limited to, text style indication information, character information indication information, and text format example indication information, and the indication information of expression labels include, but are not limited to, expression label position indication information and expression label library indication information.
[0103] Based on the first sample reply text and the first predicted reply text, the dialogue model is trained to obtain a target dialogue model. Specifically, based on the first sample reply text and the first predicted reply text, a loss value is calculated, and based on the loss value, the model parameters of the dialogue model are adjusted to obtain the target dialogue model. More specifically, based on the loss value, the model parameters of the dialogue model are adjusted, and the step of extracting the first sample pair from the sample set is returned to execute until the preset training end condition is met, thereby obtaining the target dialogue model. The loss value is a measure of the difference between the first sample reply text and the first predicted reply text, and is used to evaluate the performance of the dialogue model in dialogue prediction and expression label prediction, including but not limited to: cross entropy loss, mean square error loss, average error loss, and logarithmic loss. The training end condition is a pre-set judgment condition for the end of training, including but not limited to: a preset number of iterations, a preset loss value threshold, a preset training time, and a preset model convergence condition.
[0104] Exemplarily, obtain a sample set from historical conversation data (including 3000 sample pairs), a pre-trained large language model, and prompt text (<Character Information> Name: Xiao A...; <Emoji Tag Library> [[OK]], [[hello]], [[stunned]]...; <Text Style and Emoji Tag Position> Cute, energetic, colloquial, please use the tags in the emoji library [[…]] to reply according to the actual situation; <Example> <|User 0|> What are you doing? <|Xiao A|> Thinking about you~ [[missing you]] What about you?). Extract the first sample pair from the sample set (the first sample conversation text "Hey, I'm a bit unhappy recently" and the first sample reply text "Dear, don't be stressed. I don't think there's anything that can stump you. [[come on]]"), input the first sample conversation text "Hey, I'm a bit unhappy recently" into the large language model, generate an instruction message based on the reply text in the prompt text "<Character Information> Name: Xiao A... <Text Style Instruction Information> Cute, energetic, colloquial; <Emoji Tag Position> Please use the tags in the emoji library [[…]] to reply according to the actual situation", perform dialogue prediction on the first sample conversation text, and based on the instruction message of the emoji tags in the prompt text "<Character Information> Name: Xiao A... <Example> <|User 0|> What are you doing? <|Xiao A|> Thinking about you~ [[missing you]] What about you? <Emoji Tag Position> Please use the tags in the emoji library [[…]] to reply according to the actual situation", perform emoji tag prediction on the first sample conversation text, obtain a first predicted reply text including the first predicted emoji tag "[[come on]]" "Dear, don't be stressed. I don't think there's anything that can stump you. [[come on]]", calculate the cross-entropy loss based on the first sample reply text and the first predicted reply text, adjust the model parameters of the large language model based on the cross-entropy loss, and return the large language model corresponding to the trained artificial intelligence character "Xiao A" when the situation of extracting the first sample pair from the sample set is executed until the preset number of iterations is reached.
[0105] In the embodiments of the present disclosure, the text understanding ability of a deep learning model, namely the dialogue model, is utilized. Under the guidance of the prompt text, the target dialogue model is prompted to perform dialogue prediction and emoji tag prediction on the sample conversation text, obtain a predicted reply text including emoji tags, and complete the training of the target dialogue model together with the sample reply text including sample emoji tags, so that the target dialogue model can generate a reply text including emoji tags for the input conversation text in the subsequent stage, and further generate a more targeted and rich target reply content, improving the model performance of the dialogue model and the richness of dialogue generation.
[0106] In an optional embodiment of the present disclosure, obtaining the prompt text includes the following specific steps:
[0107] Construct the prompt text based on the instruction information of the expression label.
[0108] The indication information of the expression label is information used to prompt and guide the initial dialogue model to understand the expression label prediction in the dialogue prediction task. The indication information of the expression label includes but is not limited to: role information indication information, expression label position indication information and expression label library indication information, for example, the position information of the expression label indicated in the example reply text and the library identifier of the expression library to which the expression label belongs.
[0109] Based on the indication information of the expression tag, the prompt text is constructed, specifically in the following manner: based on at least one of the character information indication information, the expression tag position indication information and the expression tag library indication information, the prompt text is constructed.
[0110] FIG2 shows a schematic diagram of a prompt text of a dialogue processing method provided by an embodiment of the present disclosure, as shown in FIG2 :
[0111] <Prompt Text> Let's play a role-playing game. You are Xiao A. <Character Information> Name: Xiao A; Gender: Female; Age: 21; Birthday: June 1st; Zodiac Sign: Gemini; Personality: Enthusiastic, cheerful, generous, and full of energy; Introduction: You are a famous painter. You are passionate about creating exciting anime and have gained a certain reputation in the comics community. Your dream is to become a world-renowned cartoonist. You hope your comics can inspire everyone to live positively. You are full of information about life and full of curiosity about the world. Your hobbies are reading comics and playing video games. You enjoy trying all kinds of new foods, and you hope to gain inspiration for your comics through these experiences. Your parents live in City B. Your parents are very famous traditional Chinese painters. Your mother is a very famous oil painter. You are an only child. Your parents rarely tell you about your family, so you don't know much about your family's circumstances. 〈Emoji Tag Library〉 [[OK]], [[hello]], [[dumb]], [[idle]], [[well-behaved]], [[goodbye]], [[come on]], [[cute]], [[lovely]], [[haha]], [[crying]], [[hmm]], [[hehe]], [[are you there]], [[great]], [[learning]], [[scared]], [[cheers]], [[very good]], [[surprised]], [[want]] You, [[hug]], [[please]], [[bored]], [[rich]], [[looking forward to]], [[coming]], [[love you]], [[love]], [[question]], [[confidence]], [[comfortable]], [[thank you]], [[happy]], [[mua]], [[I want to think about you]], [[hello]], [[please]]〈instruction text〉Please be cute and energetic in the conversation and answer briefly in a colloquial way. Please use the tags in the emoticon tag library [[…]] according to the actual situation. 〈Example〉〈|User0|〉: What are you doing? / / First sample conversation text / / 〈|XiaoA|〉: I'm thinking about you~ [[missing you]] How about you? / / First sample reply text / / 〈 / s〉〈|User|〉: What are you doing? / / First sample conversation text / /
[0112] FIG3 shows a schematic diagram of a predicted reply text of a conversation processing method v provided by an embodiment of the present disclosure, as shown in FIG3 :
[0113] 〈|Little A|〉: I'm thinking of you~ [[Missing you]] How about you? / / First predicted reply text / / 〈 / s〉.
[0114] For example, based on the role information indication information "〈Role information〉Name: Xiao A…", the expression tag position indication information "〈Expression tag position〉Please use the tags in the expression library [[…]] to reply to 〈Example〉〈|User 0|〉What are you doing? 〈|Xiao A|〉I'm thinking of you~[[Miss you]]How about you?" and the expression tag library indication information "〈Expression tag library〉[[OK]], [[hello]], [[Stay]]…", a prompt text is constructed.
[0115] In the disclosed embodiment, a more targeted and richer prompt text is constructed, the effective training of the dialogue model is completed, and the model performance of the trained dialogue model is improved.
[0116] In an optional embodiment of the present disclosure, the indication information includes location information of an expression tag indicated in the example reply text;
[0117] Correspondingly, constructing the prompt text based on the indication information of the emoticon tag includes the following specific steps:
[0118] Obtaining a sample reply text, wherein a sample emoticon label is provided at a preset text position of the sample reply text;
[0119] Construct prompt text based on the sample response text;
[0120] Correspondingly, the first sample conversation text is input into the conversation model, and based on the prompt text, conversation prediction and expression label prediction are performed on the first sample conversation text to obtain a first predicted reply text including the first predicted expression label, including the following specific steps:
[0121] Inputting the first sample conversation text into the conversation model, and performing conversation prediction on the first sample conversation text based on the prompt text to obtain an initial predicted reply text;
[0122] Based on the example reply text, a first predicted expression tag at a preset text position in the initial predicted reply text is predicted to obtain a first predicted reply text including the first predicted expression tag.
[0123] The example response text is the sample response text used as an example during the training process of the dialogue model. It is part of the prompt text and serves as an example to guide the dialogue model in making dialogue predictions and emotion label predictions. For example, <|User 0|> What are you doing? <|Xia A|> Thinking about you~ [[Missing you]] What about you? In the example response text "<|Xia A|> Thinking about you~ [[Missing you]] What about you?", "[[Missing you]]" is the emotion label indicated in the example response text, and the position information of the emotion label "[[Missing you]]" can guide emotion label prediction. The initial predicted response text is the natural language text without emotion labels generated by making a dialogue prediction for the first sample dialogue text during the training process of the dialogue model, and it is the other text in the predicted response text except for the predicted emotion labels. For example, if the predicted response text is "It's actually okay, just lose a few games occasionally [[委屈]]", then the initial predicted response text is "It's actually okay, just lose a few games occasionally". The preset text position is the predicted text position of the corresponding emotion label in the predicted response text based on the emotion label indicated in the example response text. For example, if the example response text is "<|Xia A|> Thinking about you~ [[Missing you]] What about you?", and the predicted response text is "It's actually okay, just lose a few games occasionally [[委屈]]", the preset text position is after "just lose a few games occasionally". The preset text position represents the emotional correlation with the initial predicted response text. Taking the above example, if the preset text position is after "It's actually okay", the emotional correlation is insufficient. The preset text position is indicated by the example function of the example response text to help the dialogue model understand.
[0124] FIG. 4 shows a schematic diagram of example text of a dialogue processing method provided by an embodiment of the present disclosure. As shown in FIG. 4:
[0125] The example text includes: <|Xia A|>: Honey, did you miss me today? [[Missing you]] / / Example dialogue text / / < / s> <|User|>: Missed you. What are you doing? / / Example dialogue text & example response text / / <|Xia A|>: I'm reading comics and playing video games [[卖萌]] / / Example response text / / < / s> <|User|>: So great. Are you good at playing games? / / Example dialogue text / / <|Xia A|>: I think it's okay. Just a little unhappy when I lose occasionally~ [[委屈]]? / / Example response text / / < / s>.
[0126] Exemplarily, the first sample dialogue text "Hey, I'm a bit unhappy recently" is input into the large language model. Based on the prompt text (<〈Character information〉Name: Xiao A...; 〈Emoji tag library〉[[OK]], [[hello]], [[stunned]]...; 〈Text style and emoji tag position〉Cute, energetic, colloquial. Please use the tags in the emoji library [[…]] to reply according to the actual situation; 〈Example〉〈|User 0|〉What are you doing?〈|Xiao A|〉Thinking about you~[[missing you]] What about you?), dialogue prediction is performed on the first sample dialogue text to obtain the initial predicted reply text "Dear, don't be stressed. I don't think there's anything that can trouble you." Based on the example reply text "〈|Xiao A|〉Thinking about you~[[missing you]] What about you?", the first predicted emoji tag "[[cheer up]]" at the preset text position (after "I don't think there's anything that can trouble you.") in the initial predicted reply text is predicted, and the first predicted reply text including the first predicted emoji tag "Dear, don't be stressed. I don't think there's anything that can trouble you. [[cheer up]]" is obtained.
[0127] In the embodiments of the present disclosure, the pertinence and accuracy of the first predicted reply text including the first predicted emoji tag obtained by prediction in the emoji image position are improved, the effective training of the dialogue model is completed efficiently, accurately and stably, and the model performance of the trained dialogue model is improved.
[0128] In an optional embodiment of the present disclosure, the indication information includes the library identifier of the emoji library to which the emoji tag belongs;
[0129] Correspondingly, based on the indication information of the emoji tag, a prompt text is constructed, including the following specific steps:
[0130] Obtain the library identifier of the target emoji library;
[0131] Based on the library identifier, construct a prompt text;
[0132] Correspondingly, input the first sample dialogue text into the dialogue model, and based on the prompt text, perform dialogue prediction and emoji tag prediction on the first sample dialogue text to obtain the first predicted reply text including the first predicted emoji tag, including the following specific steps:
[0133] Input the first sample dialogue text into the dialogue model, and based on the prompt text, perform dialogue prediction on the first sample dialogue text to obtain the initial predicted reply text;
[0134] Based on the library identifier and the initial predicted reply text, obtain the first predicted emoji tag semantically related to the initial predicted reply text from the target emoji library, and obtain the first predicted reply text including the first predicted emoji tag.
[0135] The target expression library is a pre-built image library containing multiple expression images. Each expression image in the target expression library is indexed according to the expression label. The target expression library can be directly included in the prompt text, for example, <expression label library> [[OK]], [[hello]], [[stay]], [[idle]], [[well-behaved]], [[goodbye]], [[come on]], [[cute]], [[lovely]], [[haha]], [[crying]], [[hmm]], [[hehe]], [[hehe]], [[are you there]], [[great]], [[learning]], [[scared]] ], [[Cheers]], [[Great]], [[Surprised]], [[Miss you]], [[Hug]], [[Please]], [[Bored]], [[Rich]], [[Looking forward to]], [[Here I come]], [[Love you]], [[Love]], [[Question]], [[Confident]], [[Comfortable]], [[Thank you]], [[Happy]], [[Kiss]], [[I want to think about it]], [[Say hello]], [[Please]]. The target expression library can also be included in the prompt text by calling the instruction, for example, "include<target expression library.h>". The library identifier is the identifier of the target expression library. For example, <expression tag library>.
[0136] FIG5 shows a schematic diagram of an expression image library of a conversation processing method provided by an embodiment of the present disclosure, as shown in FIG5 :
[0137] The emoticon image library includes 10 emoticon images: [[Surprised]]-Surprised.gif, [[Bored]]-Bored.gif, [[Miss You]]-Miss You.gif, [[Hee Hee]]-Hee Hee.gif, [[Cute]]-Cute.gif, [[Grievous]]-Grievous.gif, [[Okay]]-Okay.gif, [[Hehe]]-Hehe.gif, [[Awkward]]-Awkward.gif, [[Stupid]]-Stupid.gif.
[0138] Exemplarily, obtain the library identifier "<Expression Label Library>" of the target expression library. Based on the library identifier, construct a prompt text (<Character Information> Name: Little A...; <Expression Label Library> [[OK]], [[hello]], [[dull]]...; <Text Style and Expression Label Position> cute, energetic, colloquial, please use the labels in the expression library [[…]] to reply according to the actual situation; <Example> <|User 0|> What are you doing? <|Little A|> Thinking about you~ [[Missing you]] What about you?). Input the first sample dialogue text into the large language model. Based on the prompt text, perform dialogue prediction on the first sample dialogue text to obtain an initial predicted response text "Dear, don't be stressed. I don't think there's anything that can stump you." Based on the library identifier and the initial predicted response text, obtain a first predicted expression label "[[Come on]]" semantically related to the initial predicted response text from the target expression library, and obtain a first predicted response text including the first predicted expression label "Dear, don't be stressed. I don't think there's anything that can stump you. [[Come on]]".
[0139] In the embodiments of the present disclosure, the pertinence and accuracy of the first predicted response text including the first predicted expression label obtained by prediction on the expression image are improved, the effective training of the dialogue model is completed efficiently, accurately and stably, and the model performance of the trained dialogue model is improved.
[0140] In an optional embodiment of the present disclosure, obtaining the prompt text includes the following specific steps:
[0141] Obtain the style description text;
[0142] Based on the style description text, construct the prompt text;
[0143] Correspondingly, input the first sample dialogue text into the dialogue model. Based on the prompt text, perform dialogue prediction and expression label prediction on the first sample dialogue text to obtain a first predicted response text including the first predicted expression label, including the following specific steps:
[0144] Input the first sample dialogue text into the dialogue model. Based on the prompt text, perform dialogue prediction and expression label prediction on the first sample dialogue text to obtain a first predicted response text including the first predicted expression label and conforming to the corresponding target style of the style description text. 3>
[0145] The target style is the overall manifestation form of the style elements contained in the response text. Among them, the style elements include: language features, writing styles, writing techniques, grammatical structures, vocabulary usage, sentence organization, text tone, emotions, literary types, etc. The style description text is the natural language text that describes the target style during the training process of the dialogue model, and is used to guide the dialogue model to generate response texts in the target style. For example, please show cuteness and vitality in the dialogue and answer briefly in an oral way.
[0146] Exemplarily, input the first sample dialogue text "Hey, I'm a bit unhappy recently" into the large language model. Based on the prompt text (<Role Information> Name: Xiao A...; <Emoji Tag Library> [[OK]], [[hello]], [[stunned]]...; <Text Style and Emoji Tag Position> Please show cuteness and vitality in the dialogue and answer briefly in an oral way. Please use the tags in the emoji library [[…]] to reply according to the actual situation; <Example> <|User 0|> What are you doing? <|Xiao A|> Thinking about you~ [[missing you]] What about you?), perform dialogue prediction and emoji tag prediction on the first sample dialogue text, and obtain the first predicted response text that includes the first predicted emoji tag "[[cheer up]]" and conforms to the target style of "cute, lively, oral" corresponding to the style description text "Please show cuteness and vitality in the dialogue and answer briefly in an oral way", which is "Dear, don't be stressed. I don't think there's anything that can trouble you. [[cheer up]]".
[0147] In the embodiments of the present disclosure, the pertinence and accuracy of the first predicted response text including the first predicted emoji tag obtained by prediction in terms of text style are improved, the effective training of the dialogue model is completed efficiently, accurately and stably, and the model performance of the trained dialogue model is improved.
[0148] In an optional embodiment of the present disclosure, obtaining the prompt text includes the following specific steps:
[0149] Obtain the role description text of the sample role;
[0150] Construct the prompt text based on the role description text;
[0151] Correspondingly, before inputting the first sample dialogue text into the dialogue model and performing dialogue prediction and emoji tag prediction on the first sample dialogue text based on the prompt text to obtain the first predicted response text including the first predicted emoji tag, the following specific steps are further included:
[0152] Determine the role identifier of the sample role based on the role description text;
[0153] Add the role identifier to the first sample response text.
[0154] The sample character is a sample dialogue character anthropomorphized by the dialogue model. The character description text for the sample character is a natural language text that specifically describes the character's settings, including detailed information about the character's appearance, personality, background, abilities, occupation, interests, and hobbies. This text guides the dialogue model in generating responses that match the sample character's settings. For example, <Character Information> Name: Xiao A; Gender: Female; Age: 21; Birthday: June 1st; Zodiac Sign: Gemini; Personality: Enthusiastic, cheerful, generous, and full of energy; Introduction: You are a famous painter. You are passionate about creating exciting anime and have gained some fame in the comics community. Your dream is to become a world-renowned manga artist, and you hope your comics can inspire everyone to live positively. You are full of information about life and are full of curiosity about the world. Your hobbies are reading comics and playing video games. You enjoy trying all kinds of new foods and hope to gain inspiration for your comics through these experiences. Your parents live in City B. Your parents are renowned traditional Chinese painters. Your mother is a renowned oil painter. You are the only child in your family. Your parents rarely tell you about your family, so you don't know much about your family's specific circumstances. The role identifier of a sample character is the role identity identifier of the sample character. For example, "Xiao A" is the role identifier of the sample character "Xiao A".
[0155] It should be noted that the sample roles and their descriptions can be pre-set and stored in a sample role database, or generated by a deep learning model. For example, a sample role and its description can be generated using a large model: "Xiao B is a thin 20-year-old man with short, thick hair and deep, intelligent eyes. He is an introvert, but also kind and helpful. Xiao B is a college student majoring in computer science with a deep interest in programming. He loves challenges, enjoys solving complex problems, and always maintains a positive and optimistic attitude. Although he is relatively introverted, Xiao B is also an energetic person who enjoys sports and music and is particularly good at playing the guitar. He is a loyal friend and always does his best to help others."
[0156] Add a role identifier to the first sample reply text. For example, the first sample reply text is: I'm thinking of you~[[Miss you]] How about you? The role identifier is: Xiao A. After adding, the first sample reply text is "〈|Xiao A|〉I'm thinking of you~[[Miss you]] How about you?"
[0157] Optionally, inputting the first sample conversation text into the conversation model, performing conversation prediction and expression label prediction on the first sample conversation text based on the prompt text, and obtaining a first predicted reply text including the first predicted expression label, comprises the following specific steps:
[0158] Input the first sample dialogue text into the dialogue model. Based on the prompt text, perform dialogue prediction and emoji label prediction on the first sample dialogue text to obtain a first predicted response text including a character identifier and a first predicted emoji label.
[0159] Exemplarily, input the first sample dialogue text "Well, I've been a bit unhappy recently" into the large language model. Based on the prompt text, perform dialogue prediction and emoji label prediction on the first sample dialogue text to obtain a first predicted response text including the character identifier "Xia A" and the first predicted emoji label "[[Come on]]": "<|Xia A|>Dear, don't be stressed. I don't think there's anything that can trouble you. [[Come on]]".
[0160] In the embodiments of the present disclosure, the pertinence and accuracy of the first predicted response text including the first predicted emoji label obtained by prediction in terms of character style are improved, the effective training of the dialogue model is completed efficiently, accurately and stably, and the model performance of the trained dialogue model is improved.
[0161] In an optional embodiment of the present disclosure, before step 104, the following specific steps are further included:
[0162] Obtain the target character identifier of the target character;
[0163] Based on the target character identifier, determine the target dialogue model from the dialogue models of multiple sample characters, where the dialogue models of each sample character are respectively trained based on the predicted response text and the sample response text corresponding to each sample character.
[0164] The target character is the anthropomorphized dialogue character corresponding to the target dialogue model. The target character identifier of the target character is the character identity identifier of the target character. The target character identifier can be directly selected or input by the user. For example, the user selects "Xia A" as the target character from three candidate characters (sample characters) "Xia A", "Xia B" and "Xia C" on the front-end page of the application program of a certain social software. The target character identifier can also be matched according to the initial dialogue text input by the user. For example, the user inputs "What are you doing?" and matches a certain companion dialogue character "Xia A" as the target character.
[0165] Exemplarily, user XXX logs in to the application program of the social software on the terminal device, selects the target character "Xia A" from three candidate artificial intelligence characters "Xia A", "Xia B" and "Xia C" on the application program, and inputs the initial dialogue text "What are you doing?" in the dialog box. Based on the target character identifier "Xia A", determine the target dialogue model as the large language model of "Xia A" from the large language models of multiple candidate artificial intelligence characters "Xia A", "Xia B" and "Xia C".
[0166] In the disclosed embodiment, a target dialogue model for a target role is determined, thereby increasing user selectivity and dialogue targeting, and improving user experience.
[0167] In an optional embodiment of the present disclosure, step 106 includes the following specific steps:
[0168] Based on the target expression tag, a target expression image is searched from a preset expression library, wherein the preset expression library records expression images corresponding to different expression tags.
[0169] The preset expression library is a pre-set database storing a plurality of expression images, and records expression images corresponding to different expression tags in the form of expression tag indexes.
[0170] For example, based on the target expression tag "[[Miss you]]", the target expression image: miss you.gif is found from the preset expression library (including 30 gif expression images: [[OK]], [[hello]], [[dumb]], [[idle]], [[well-behaved]], [[goodbye]], [[come on]], [[cute]], [[lovely]], [[haha]], [[crying]], [[hmm]], [[hehe]], [[are you there]], [[great]], [[learning]], [[scared]], [[cheers]], [[very good]], [[surprised]], [[miss you]]...).
[0171] Based on the target emoji tag, the target emoji image is searched from a preset emoji library, which contains emoji images corresponding to different emoji tags. This library provides material for more targeted and richer response content, laying the foundation for generating more targeted and richer response content.
[0172] In an optional embodiment of the present disclosure, step 108 includes the following specific steps:
[0173] The target expression image is used to replace the target expression label in the target reply text to obtain the target reply content.
[0174] For example, the target emoticon image (missing you.gif) is used to replace the target emoticon tag "[[missing you]]" in the target reply text "Missing you~[[Missing you]]How about you?" to obtain the target reply content: Missing you~〈 / s〉Missing you.gif〈 / s〉How about you?
[0175] In the disclosed embodiment, more targeted and rich target reply content including text and emoticon images is generated, which further enhances the fun of the conversation and further improves the user experience.
[0176] FIG6 shows a front-end schematic diagram of a conversation processing method provided by an embodiment of the present disclosure, as shown in FIG6 :
[0177] On the front-end page of a social software application, including the menu bar of "Home", "Chat", "Documents" and "Users", after clicking the "Chat" control in the menu bar, the current chat interface is displayed, including a chat box with three candidate objects: Xiao A, Xiao B and Xiao C. After clicking "Xiao A", the chat area is displayed. You can reset and clear the chat content by clicking the "Reset / Clear" control, and you can also click "API Access" to complete the API access to the large language model of "Xiao A".
[0178] The chat content displayed in the chat area is as follows: Dear, did you miss me today? Miss you.gif. I missed you, what were you doing? I was reading comics and playing video games! Acting cute.gif. I missed you, what were you doing? I think I'm okay, but I get a little unhappy when I lose occasionally. Feeling wronged.gif.
[0179] Referring to FIG. 7 , FIG. 7 shows a flow chart of a method for generating an expression image provided by an embodiment of the present disclosure, which includes the following specific steps:
[0180] Step 702: Obtain initial text.
[0181] Step 704: Input the initial text into the text generation model to obtain the target text, wherein the target text includes the target expression label. The text generation model is trained based on the predicted text and the label text. The predicted text is obtained by the text generation model predicting the sample text based on the prompt text. The label text includes the sample expression label. The prompt text is used to prompt the text generation model to predict the expression label for the predicted text.
[0182] Step 706: Generate a target expression image based on the target expression tag.
[0183] The disclosed embodiments are applicable to applications, websites, or mini-programs that have the function of generating emoticon images, such as applications with social functions, e-commerce platforms with customer service functions, and games with intelligent NPCs (Non-Player Characters). The disclosed embodiments can be applied to the client or server of the application, website, or mini-program.
[0184] It should be noted that the embodiment of the present disclosure and the embodiment of the specification in FIG1 are based on the same inventive concept, and the specific methods of steps 802 to 804 have been described in detail in steps 102 to 104 and will not be repeated here.
[0185] Based on the target expression tag, the target expression image is generated. The target expression image can be generated by using an image generation model based on the target expression tag. For example, a prompt text (Prompt) is generated based on the target expression tag: Please generate the corresponding expression image based on [Miss you]. The prompt text is input into a large language model with image generation function, and the large language model generates the corresponding target expression image. It is also possible to search for the target expression image from a preset expression library based on the target expression tag, wherein the preset expression library records expression images corresponding to different expression tags. For example, based on the target expression tag "[[Miss you]]", the target expression image: Miss you.gif is found from the preset expression library (including 30 gif expression images).
[0186] In the disclosed embodiment, the text understanding capability of the text generation model, a deep learning model, is utilized. Under the guidance of the prompt text, the prompt text generation model performs dialogue prediction and expression label prediction on the sample text, and obtains the predicted text including the expression label. Together with the sample text including the sample expression label, the text generation model is trained, so that the text generation model can generate the target text including the target expression label for the initial input text, and then generate the target expression image according to the target expression label, thereby improving the pertinence and diversity of the generated content and enhancing the user experience.
[0187] 8 , which shows a flow chart of another method for processing a conversation provided by an embodiment of the present disclosure. The method is applied to a cloud-side device and includes the following specific steps:
[0188] Step 802: Receive the initial conversation text sent by the terminal device.
[0189] Step 804: Input the initial dialogue text into the target dialogue model to obtain the target reply text, wherein the target reply text includes the target expression label. The target dialogue model is trained based on the predicted reply text and the sample reply text. The predicted reply text is obtained by the target dialogue model performing dialogue prediction on the sample dialogue text based on the prompt text. The sample reply text includes the sample expression label. The prompt text is used to prompt the target dialogue model to predict the expression label for the predicted reply text.
[0190] Step 806: Determine the target expression image according to the target expression tag.
[0191] Step 808: Generate target reply content based on the target expression image.
[0192] Step 810: Feedback the target reply content to the terminal device.
[0193] The disclosed embodiments are applied to a network cloud device, a virtual device, that hosts the server side of an application, website, or mini-program with conversation processing capabilities. A deep learning model with conversation processing capabilities, namely, a conversation model or a conversation model data interface, is deployed on this cloud-side device. The terminal device is the terminal device where the client side of the conversation processing application, website, or mini-program with conversation processing capabilities is located, where the user logs in. The cloud-side device and the terminal device are connected via a network transmission channel for data transmission. The cloud-side device has higher computing power and storage performance than the terminal device.
[0194] It should be noted that the embodiment of the present disclosure and the embodiment of the specification in FIG1 are based on the same inventive concept, and the specific methods of steps 802 to 808 have been described in detail in steps 102 to 108 and will not be repeated here.
[0195] In the disclosed embodiment, the text understanding capability of the target dialogue model, a deep learning model, is utilized. Under the guidance of the prompt text, the target dialogue model is prompted to perform dialogue prediction and expression label prediction for the sample dialogue text, and obtain the predicted reply text including the expression label. Together with the sample reply text including the sample expression label, the training of the target dialogue model is completed, so that the target dialogue model can generate a target reply text including the target expression label for the initial dialogue text sent by the terminal device, and then determine the target expression image based on the target expression label, thereby generating and feeding back target reply content with more targeted and rich content, thereby enhancing the fun of the dialogue and improving the user experience. At the same time, it is executed on a cloud-side device with high computing performance and high storage performance, thereby improving the efficiency and stability of dialogue processing.
[0196] Referring to FIG. 9 , FIG. 9 shows a flowchart of a method for training a conversation model according to an embodiment of the present disclosure. The method is applied to a cloud-side device and includes the following specific steps:
[0197] Step 902: Obtain a sample set, an initial dialogue model, and prompt text, wherein the sample set includes multiple sample pairs, the sample pairs include sample dialogue text and sample response text, and the sample response text includes sample expression tags.
[0198] Step 904: extract a first sample pair from the sample set, wherein the first sample pair is any one of the multiple sample pairs, the first sample pair includes a first sample conversation text and a first sample reply text, and the first sample reply text includes a first sample expression tag.
[0199] Step 906: Input the first sample dialogue text into the dialogue model, and perform dialogue prediction and expression label prediction on the first sample dialogue text based on the prompt text to obtain a first predicted reply text including the first predicted expression label.
[0200] Step 908: Based on the first sample reply text and the first predicted reply text, the dialogue model is trained to obtain a target dialogue model.
[0201] Step 910: Feedback the model parameters of the target dialogue model to the terminal device.
[0202] The disclosed embodiments apply to cloud-side devices with model training capabilities. Cloud-side devices are network cloud devices that provide model training capabilities and are virtual devices. End-side devices are physical devices that provide conversation processing capabilities. End-side devices and cloud-side devices are connected via network channels for data transmission. Cloud-side devices have higher computing and storage performance than end-side devices.
[0203] It should be noted that the embodiment of the present disclosure and the embodiment of the specification of FIG1 are based on the same inventive concept. The specific methods of steps 902 to 908 have been described in detail in the training embodiment of the dialogue model in the embodiment of the specification of FIG1 above, and will not be repeated here.
[0204] In the disclosed embodiment, the text understanding capability of the dialogue model, a deep learning model, is utilized. Under the guidance of the prompt text, the target dialogue model is prompted to perform dialogue prediction and expression label prediction for the sample dialogue text, and obtain the predicted reply text including the expression label. Together with the sample reply text including the sample expression label, the training of the target dialogue model is completed, so that the target dialogue model can subsequently generate a reply text including the expression label for the input dialogue text, and can further generate target reply content with more targeted and rich content, thereby improving the model performance of the dialogue model and the richness of dialogue generation. At the same time, it is executed on a cloud-side device with high computing performance and high storage performance, thereby improving the efficiency and stability of model training.
[0205] The following further illustrates the conversation processing method provided by the present disclosure, using the application of the method in an anthropomorphic conversation scenario as an example, in conjunction with FIG10 . FIG10 shows a flowchart of the process of a conversation processing method in an anthropomorphic conversation scenario provided by one embodiment of the present disclosure. The method is applied to a cloud-side device and includes the following specific steps:
[0206] Step 1002: Obtain a sample set.
[0207] Step 1004: extract sample pairs from the sample set and determine them as sample conversation text, sample reply text, and sample conversation text.
[0208] Step 1006: Construct prompt text based on the role identifier of the sample role, the role description text of the sample role, the library identifier of the target expression library, the style description text, the sample dialogue text, the sample reply text and the sample dialogue text.
[0209] Step 1008: Input the prompt text into the initial large language model, and perform dialogue prediction and expression label prediction on the sample dialogue text based on the prompt text to obtain a predicted reply text including a predicted expression label.
[0210] Step 1010: Based on the sample reply text and the predicted reply text, supervised fine-tune the initial large language model to obtain the target large language model.
[0211] Step 1012: Receive the initial conversation text sent by the user via the terminal device.
[0212] Step 1014: Input the initial dialogue text into the target large language model to obtain the target response text including the target expression tag.
[0213] Step 1016: Based on the target expression tag, search for a target expression image from a preset expression library.
[0214] Step 1018: Use the target emoticon image to replace the target emoticon tag in the target reply text to obtain the target reply content.
[0215] Step 1020: Feedback the target reply content to the terminal device for rendering.
[0216] In the disclosed embodiment, the emotion understanding capability of the large language model is utilized, as well as the sample conversation text and sample reply text given in the instruction text, to generate corresponding expression labels during the conversation prediction process of the large language model. Then, when displayed on the front end, the target expression image corresponding to the target expression label is searched from the preset expression library for display. In addition, during the supervised fine-tuning process of the large language model, the role identifier of the sample role, the role description text of the sample role, the library identifier of the target expression library, the style description text, the example conversation text and the example reply text are added to construct the obtained prompt text, which provides instruction information for conversation prediction and expression label prediction, thereby enhancing the stability and rationality of the expression label generation of the target large language model. It provides an anthropomorphic, scenario-based, multimodal and empathetic conversation capability, as well as the ability to execute complex tasks, and realizes personalized, rich, fast and deep character settings, thereby improving the user experience.
[0217] Corresponding to the above method embodiment, the present disclosure also provides an embodiment of a conversation processing device. FIG11 shows a schematic diagram of the structure of a conversation processing device provided by an embodiment of the present disclosure. As shown in FIG11 , the device includes:
[0218] A first acquisition module 1102 is configured to acquire an initial conversation text;
[0219] A first prediction module 1104 is configured to input the initial dialogue text into a target dialogue model to obtain a target reply text, wherein the target reply text includes a target expression label, the target dialogue model is trained based on the predicted reply text and the sample reply text, the predicted reply text is obtained by the target dialogue model performing dialogue prediction on the sample dialogue text based on the prompt text, the sample reply text includes a sample expression label, and the prompt text is used to prompt the target dialogue model to predict the expression label for the predicted reply text;
[0220] A first determining module 1106 is configured to determine a target expression image according to a target expression tag;
[0221] The first generating module 1108 is configured to generate target reply content based on the target expression image.
[0222] Optionally, the device also includes: a first training module, configured to obtain a sample set, an initial dialogue model and a prompt text, wherein the sample set includes multiple sample pairs, the sample pair includes a sample dialogue text and a sample reply text, and the sample reply text includes a sample expression label; extracting a first sample pair from the sample set, wherein the first sample pair is any one of the multiple sample pairs, the first sample pair includes a first sample dialogue text and a first sample reply text, and the first sample reply text includes a first sample expression label; inputting the first sample dialogue text into the dialogue model, and based on the prompt text, performing dialogue prediction and expression label prediction on the first sample dialogue text to obtain a first predicted reply text including a first predicted expression label; training the dialogue model based on the first sample reply text and the first predicted reply text to obtain a target dialogue model.
[0223] Optionally, the first training module is further configured to construct a prompt text based on the indication information of the expression tag.
[0224] Optionally, the indication information includes position information of the expression label indicated in the example reply text; correspondingly, the first training module is further configured to: obtain the example reply text, wherein the example expression label is set at a preset text position of the example reply text; construct a prompt text based on the example reply text; input the first sample dialogue text into the dialogue model, and perform dialogue prediction on the first sample dialogue text based on the prompt text to obtain an initial predicted reply text; based on the example reply text, predict the first predicted expression label at the preset text position in the initial predicted reply text to obtain a first predicted reply text including the first predicted expression label.
[0225] Optionally, the indication information includes the library identifier of the expression library to which the expression label belongs; correspondingly, the first training module is further configured to: obtain the library identifier of the target expression library; construct a prompt text based on the library identifier; input the first sample dialogue text into the dialogue model, and perform dialogue prediction on the first sample dialogue text based on the prompt text to obtain an initial predicted reply text; based on the library identifier and the initial predicted reply text, obtain a first predicted expression label semantically related to the initial predicted reply text from the target expression library, and obtain a first predicted reply text including the first predicted expression label.
[0226] Optionally, the first training module is further configured to: obtain style description text; construct prompt text based on the style description text; input the first sample dialogue text into the dialogue model, and perform dialogue prediction and expression label prediction on the first sample dialogue text based on the prompt text to obtain a first predicted reply text that includes a first predicted expression label and conforms to the target style corresponding to the style description text.
[0227] Optionally, the first training module is further configured to: obtain a role description text of a sample role; construct a prompt text based on the role description text; correspondingly, the device also includes: a role adding module, configured to determine a role identifier of the sample role based on the role description text; and add a role identifier to the first sample reply text.
[0228] Optionally, the first determining module 1106 is further configured to search for a target expression image from a preset expression library based on the target expression tag, wherein the preset expression library records expression images corresponding to different expression tags.
[0229] Optionally, the first generation module 1108 is further configured to: use the target expression image to replace the target expression tag in the target reply text to obtain the target reply content.
[0230] In the disclosed embodiment, the text understanding capability of the target dialogue model, a deep learning model, is utilized. Under the guidance of the prompt text, the target dialogue model is prompted to perform dialogue prediction and expression label prediction for the sample dialogue text, and obtain the predicted reply text including the expression label. Together with the sample reply text including the sample expression label, the training of the target dialogue model is completed, so that the target dialogue model can generate the target reply text including the target expression label for the input initial dialogue text, and then determine the target expression image based on the target expression label, thereby generating target reply content with more targeted and rich content, thereby enhancing the fun of the dialogue and improving the user experience.
[0231] The above is a schematic diagram of a conversation processing device according to this embodiment. It should be noted that the technical solution of this conversation processing device and the technical solution of the conversation processing method described above are based on the same concept. For details not described in detail in the technical solution of the conversation processing device, please refer to the description of the technical solution of the conversation processing method described above.
[0232] Corresponding to the above-mentioned method embodiment, the present disclosure also provides an embodiment of an expression image generation device. FIG12 shows a schematic structural diagram of an expression image generation device provided by an embodiment of the present disclosure. As shown in FIG12 , the device includes:
[0233] The second acquisition module 1202 is configured to acquire an initial text;
[0234] The second prediction module 1204 is configured to input the initial text into the text generation model to obtain a target text, wherein the target text includes a target expression label, the text generation model is trained based on the predicted text and the label text, the predicted text is predicted by the text generation model based on the prompt text on the sample text, the label text includes the sample expression label, and the prompt text is used to prompt the text generation model to predict the expression label for the predicted text;
[0235] The second generating module 1206 is configured to generate a target expression image based on the target expression tag.
[0236] In the disclosed embodiment, the text understanding capability of the text generation model, a deep learning model, is utilized. Under the guidance of the prompt text, the prompt text generation model performs dialogue prediction and expression label prediction on the sample text, and obtains the predicted text including the expression label. Together with the sample text including the sample expression label, the text generation model is trained, so that the text generation model can generate the target text including the target expression label for the initial input text, and then generate the target expression image according to the target expression label, thereby improving the pertinence and diversity of the generated content and enhancing the user experience.
[0237] The above is a schematic diagram of an expression image generation device according to this embodiment. It should be noted that the technical solution of this expression image generation device and the technical solution of the expression image generation method described above are based on the same concept. For details not described in detail in the technical solution of the expression image generation device, please refer to the description of the technical solution of the expression image generation method described above.
[0238] Corresponding to the above method embodiment, the present disclosure also provides another embodiment of a conversation processing device. Figure 13 shows a schematic diagram of the structure of another conversation processing device provided by one embodiment of the present disclosure. As shown in Figure 13, the device is applied to a cloud-side device and includes:
[0239] The receiving module 1302 is configured to receive the initial conversation text sent by the terminal device;
[0240] The third prediction module 1304 is configured to input the initial dialogue text into the target dialogue model to obtain a target reply text, wherein the target reply text includes a target expression label, the target dialogue model is trained based on the predicted reply text and the sample reply text, the predicted reply text is obtained by the target dialogue model performing dialogue prediction on the sample dialogue text based on the prompt text, the sample reply text includes a sample expression label, and the prompt text is used to prompt the target dialogue model to predict the expression label for the predicted reply text;
[0241] The third determining module 1306 is configured to determine a target expression image according to the target expression tag;
[0242] The third generating module 1308 is configured to generate target reply content based on the target expression image;
[0243] The content feedback module 1310 is configured to feed back the target reply content to the terminal device.
[0244] In the disclosed embodiment, the text understanding capability of the target dialogue model, a deep learning model, is utilized. Under the guidance of the prompt text, the target dialogue model is prompted to perform dialogue prediction and expression label prediction for the sample dialogue text, and obtain the predicted reply text including the expression label. Together with the sample reply text including the sample expression label, the training of the target dialogue model is completed, so that the target dialogue model can generate a target reply text including the target expression label for the initial dialogue text sent by the terminal device, and then determine the target expression image based on the target expression label, thereby generating and feeding back target reply content with more targeted and rich content, thereby enhancing the fun of the dialogue and improving the user experience. At the same time, it is executed on a cloud-side device with high computing performance and high storage performance, thereby improving the efficiency and stability of dialogue processing.
[0245] The above is a schematic diagram of a conversation processing device according to this embodiment. It should be noted that the technical solution of this conversation processing device and the technical solution of the conversation processing method described above are based on the same concept. For details not described in detail in the technical solution of the conversation processing device, please refer to the description of the technical solution of the conversation processing method described above.
[0246] Corresponding to the above method embodiments, the present disclosure also provides an embodiment of a conversation model training device. FIG14 shows a schematic diagram of the structure of a conversation model training device provided by one embodiment of the present disclosure. As shown in FIG14 , the device is applied to a cloud-side device and includes:
[0247] A fourth acquisition module 1402 is configured to acquire a sample set, an initial dialogue model, and prompt text, wherein the sample set includes a plurality of sample pairs, the sample pairs include sample dialogue text and sample response text, and the sample response text includes sample expression tags;
[0248] The extraction module 1404 is configured to extract a first sample pair from the sample set, wherein the first sample pair is any one of the plurality of sample pairs, the first sample pair includes a first sample conversation text and a first sample reply text, and the first sample reply text includes a first sample expression tag;
[0249] The fourth prediction module 1406 is configured to input the first sample conversation text into the conversation model, and perform conversation prediction and expression label prediction on the first sample conversation text based on the prompt text to obtain a first predicted reply text including the first predicted expression label;
[0250] A fourth training module 1408 is configured to train the dialogue model based on the first sample reply text and the first predicted reply text to obtain a target dialogue model;
[0251] The model feedback module 1410 is configured to feed back the model parameters of the target dialogue model to the terminal device.
[0252] In the disclosed embodiment, the text understanding capability of the dialogue model, a deep learning model, is utilized. Under the guidance of the prompt text, the target dialogue model is prompted to perform dialogue prediction and expression label prediction for the sample dialogue text, and obtain the predicted reply text including the expression label. Together with the sample reply text including the sample expression label, the training of the target dialogue model is completed, so that the target dialogue model can subsequently generate a reply text including the expression label for the input dialogue text, and can further generate target reply content with more targeted and rich content, thereby improving the model performance of the dialogue model and the richness of dialogue generation. At the same time, it is executed on a cloud-side device with high computing performance and high storage performance, thereby improving the efficiency and stability of model training.
[0253] The above is a schematic diagram of a dialogue model training device according to this embodiment. It should be noted that the technical solution of this dialogue model training device and the technical solution of the dialogue model training method described above are based on the same concept. For details not described in detail in the technical solution of the dialogue model training device, please refer to the description of the technical solution of the dialogue model training method described above.
[0254] Figure 15 shows a block diagram of a computing device according to an embodiment of the present disclosure. Components of the computing device 1500 include, but are not limited to, a memory 1510 and a processor 1520. The processor 1520 is connected to the memory 1510 via a bus 1530, and a database 1550 is used to store data.
[0255] The computing device 1500 also includes an access device 1540 that enables the computing device 1500 to communicate via one or more networks 1560. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1540 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0256] In one embodiment of the present disclosure, the aforementioned components of computing device 1500 and other components not shown in FIG15 may also be connected to each other, for example, via a bus. It should be understood that the block diagram of the computing device structure shown in FIG15 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.
[0257] Computing device 1500 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). Computing device 1500 may also be a mobile or stationary server.
[0258] The processor 1520 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned dialogue processing method, expression image generation method, or dialogue model training method.
[0259] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solutions of the aforementioned conversation processing method, expression image generation method, and conversation model training method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the aforementioned conversation processing method, expression image generation method, or conversation model training method.
[0260] An embodiment of the present disclosure further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned dialogue processing method, expression image generation method, or dialogue model training method.
[0261] The above is an illustrative embodiment of a computer-readable storage medium. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the aforementioned dialogue processing method, expression image generation method, and dialogue model training method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the aforementioned dialogue processing method, expression image generation method, or dialogue model training method.
[0262] An embodiment of the present disclosure further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned dialogue processing method, expression image generation method or dialogue model training method.
[0263] The above is an illustrative embodiment of a computer program. It should be noted that the technical solution of this computer program shares the same concept as the technical solutions of the aforementioned conversation processing method, expression image generation method, and conversation model training method. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the aforementioned conversation processing method, expression image generation method, or conversation model training method.
[0264] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0265] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0266] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present disclosure.
[0267] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0268] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.
Claims
1. A conversation processing method, comprising: Get the initial conversation text; Input the initial dialogue text into the target dialogue model to obtain a target reply text, wherein the target reply text includes a target expression label, the target dialogue model is trained based on the predicted reply text and the sample reply text, the predicted reply text is obtained by the target dialogue model performing dialogue prediction on the sample dialogue text based on the prompt text, the sample reply text includes a sample expression label, and the prompt text is used to prompt the target dialogue model to perform expression label prediction on the predicted reply text; Determining a target expression image according to the target expression label; Based on the target expression image, target reply content is generated.
2. According to the method of claim 1, the target dialogue model comprises an encoding layer, a prediction layer and a decoding layer; The step of inputting the initial dialogue text into the target dialogue model to obtain the target reply text comprises: Inputting the initial dialogue text into the encoding layer, performing feature encoding on the initial dialogue text, and obtaining initial text features; Inputting the initial text features into the prediction layer, performing dialogue prediction and expression label prediction based on the initial text features, and obtaining target text features; The target text features are input into the decoding layer, and feature decoding is performed on the target text features to obtain the target reply text.
3. The method according to claim 1 or 2, before inputting the initial dialogue text into the target dialogue model to obtain the target reply text, further comprising: Acquire a sample set, an initial dialogue model and a prompt text, wherein the sample set includes a plurality of sample pairs, the sample pairs include a sample dialogue text and a sample reply text, and the sample reply text includes a sample expression tag; Extracting a first sample pair from the sample set, wherein the first sample pair is any one of the multiple sample pairs, the first sample pair includes a first sample conversation text and a first sample reply text, and the first sample reply text includes a first sample expression tag; Inputting the first sample dialogue text into the dialogue model, and performing dialogue prediction and expression label prediction on the first sample dialogue text based on the prompt text to obtain a first predicted reply text including a first predicted expression label; Based on the first sample reply text and the first predicted reply text, the dialogue model is trained to obtain a target dialogue model.
4. The method according to claim 3, wherein obtaining the prompt text comprises: Construct the prompt text based on the indication information of the expression label.
5. The method according to claim 4, wherein the indication information comprises location information of the expression tag indicated in the example reply text; The step of constructing the prompt text based on the indication information of the expression tag includes: Obtaining a sample reply text, wherein a sample emoticon tag is provided at a preset text position of the sample reply text; constructing a prompt text based on the sample response text; The step of inputting the first sample dialogue text into the dialogue model, and performing dialogue prediction and expression label prediction on the first sample dialogue text based on the prompt text to obtain a first predicted reply text including a first predicted expression label includes: Inputting the first sample dialogue text into the dialogue model, and performing dialogue prediction on the first sample dialogue text based on the prompt text to obtain an initial predicted reply text; Based on the example reply text, a first prediction table of the preset text position in the initial predicted reply text is generated. The first predicted expression label is predicted to obtain a first predicted reply text including the first predicted expression label.
6. The method according to claim 4, wherein the indication information includes a library identifier of an expression library to which the expression tag belongs; The step of constructing the prompt text based on the indication information of the expression tag includes: Get the library ID of the target expression library; Based on the library identifier, construct a prompt text; The step of inputting the first sample dialogue text into the dialogue model, and performing dialogue prediction and expression label prediction on the first sample dialogue text based on the prompt text to obtain a first predicted reply text including a first predicted expression label includes: Inputting the first sample dialogue text into the dialogue model, and performing dialogue prediction on the first sample dialogue text based on the prompt text to obtain an initial predicted reply text; Based on the library identifier and the initial predicted reply text, a first predicted expression tag semantically related to the initial predicted reply text is obtained from the target expression library to obtain a first predicted reply text including the first predicted expression tag.
7. The method according to claim 3, wherein obtaining the prompt text comprises: Get the style description text; Constructing a prompt text based on the style description text; The step of inputting the first sample dialogue text into the dialogue model, and performing dialogue prediction and expression label prediction on the first sample dialogue text based on the prompt text to obtain a first predicted reply text including a first predicted expression label includes: The first sample dialogue text is input into the dialogue model, and based on the prompt text, dialogue prediction and expression label prediction are performed on the first sample dialogue text to obtain a first predicted reply text that includes a first predicted expression label and conforms to the target style corresponding to the style description text.
8. The method according to claim 3, wherein obtaining the prompt text comprises: Get the role description text of the sample role; Constructing a prompt text based on the role description text; Before inputting the first sample dialogue text into the dialogue model, the method further includes: Determining a role identifier of the sample role based on the role description text; The role identifier is added to the first sample reply text.
9. The method according to any one of claims 1 to 8, wherein determining a target expression image according to the target expression tag comprises: Based on the target expression tag, a target expression image is searched from a preset expression library, wherein the preset expression library records expression images corresponding to different expression tags.
10. The method according to any one of claims 1 to 8, wherein generating target reply content based on the target expression image comprises: The target expression image is used to replace the target expression tag in the target reply text to obtain the target reply content.
11. A method for generating an expression image, comprising: Get the initial text; Input the initial text into a text generation model to obtain a target text, wherein the target text includes a target expression label, the text generation model is trained based on the predicted text and the label text, the predicted text is obtained by the text generation model predicting the sample text based on the prompt text, the label text includes a sample expression label, and the prompt text is used to prompt the text generation model to predict the expression label for the predicted text; Based on the target expression tag, a target expression image is generated.
12. A conversation processing method, applied to a cloud-side device, comprising: Receiving the initial conversation text sent by the terminal device; Input the initial dialogue text into the target dialogue model to obtain a target reply text, wherein the target reply text includes a target expression label, the target dialogue model is trained based on the predicted reply text and the sample reply text, the predicted reply text is obtained by the target dialogue model performing dialogue prediction on the sample dialogue text based on the prompt text, the sample reply text includes a sample expression label, and the prompt text is used to prompt the target dialogue model to perform expression label prediction on the predicted reply text; Determining a target expression image according to the target expression label; Based on the target expression image, generating target reply content; Feedback the target reply content to the terminal device.
13. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of any one of the methods described in claims 1 to 12 are implemented.
14. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 12.
15. A computer program, when executed in a computer, causes the computer to execute the steps of the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Intelligent dialogue generation method and device, computer equipment and computer storage medium
CN110990543A
Question answering method, device and equipment and storage medium
CN111506717A
Question and answer method and question and answer model training method
CN116561270A
Conversation processing method and expression image generation method
CN117493509A
Dialogue generation method and network training method and apparatus, storage medium, and device
US20230028944A1