Information output method and device, electronic device and storage medium
Through a large language model, the output information of different personality attributes is generated and displayed, which solves the problem of single reply content in human-computer dialogue, realizes anthropomorphic interaction, and improves user experience.
Patent Information
- Application Number
- CN202410130483.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-01-30
AI Technical Summary
In the prior art, in the human-computer dialogue scenario based on a large language model, the user's lack of intelligence and diversification of machine replies, resulting in a single reply content, unable to provide anthropomorphic interaction, affecting the user experience.
The large language model generates output information based on different personality attributes, and displays adjacent to each other through different display methods to improve the stylization and anthropomorphism of reply information, including supervising fine-tuning and prompt engineering to achieve multi-personal output.
It enhances the anthropomorphic performance of virtual assistants, improves user chat experience and interest, and provides rich stylized reply information through differentiated visual and content interactions.
Smart Images

Figure CN118260389B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an information output method and device, an electronic device, and a storage medium. Background Art
[0002] With the continuous development of artificial intelligence, users are increasingly demanding intelligent and diverse responses from machines in human-machine dialogue scenarios based on large language models. However, after users enter information into the dialogue interface, the machine's responses are often stereotyped in content and style, failing to provide personalized responses, which negatively impacts the user experience. Summary of the Invention
[0003] The embodiments of the present application provide an information output method and device, an electronic device, and a storage medium. Through this method, a large language model can be used to simultaneously generate output information based on different personality attributes based on the user's input information, thereby improving the stylization of the large language model's reply information, allowing the user to experience the anthropomorphic performance of the virtual assistant, and improving the user's chat experience.
[0004] In a first aspect, an embodiment of the present application provides an information output method, which is applied to a human-computer dialogue application based on a large language model, the method comprising obtaining input information of a user in a dialogue interface; automatically generating first output information and second output information based on the input information, wherein the first output information and the second output information have different personality attributes; and displaying the first output information and the second output information adjacent to each other in different display methods in the same reply message of the dialogue interface.
[0005] In a possible implementation, the different display modes include: highlighting the second output information relative to the first output information.
[0006] In a possible implementation, highlighting the second output information relative to the first output information includes highlighting the second output information by adding a strikethrough to the second output information.
[0007] In a possible implementation, the first output information and the second output information have different personality attributes, including: the first output information is generated based on neutral-style personality attributes, and the second output information is generated based on personalized-style personality attributes.
[0008] In a possible implementation, the personalized style includes at least one of a cool type, a cute type, a middle school type, and a warm type.
[0009] In one possible implementation, the large language model generates output information of different personality attributes based on supervised fine-tuning or prompt engineering.
[0010] In one possible implementation, automatically generating the first output information and the second output information based on the input information includes: obtaining user intent information based on the input information; and determining whether to generate the second output information with different personal attributes based on the obtained user intent information.
[0011] In a possible implementation, before obtaining the user's input information in the dialogue interface, the method further includes: setting the personality attributes of the personalized style in response to the user's selection operation in the setting interface.
[0012] In a second aspect, an embodiment of the present application also provides an information output device, which is applied to a human-computer dialogue application based on a large language model, and the device includes: an acquisition model for acquiring the user's input information in the dialogue interface; a generation model for automatically generating first output information and second output information based on the input information, and the first output information and the second output information have different personal attributes; a display module for displaying the first output information and the second output information adjacent to each other in different display methods in the same reply message in the dialogue interface.
[0013] In a third aspect, an embodiment of the present application further provides an electronic device comprising: a processor and a memory, wherein the memory is used to store at least one instruction, and when the instruction is loaded and executed by the processor, the information output method provided in the first aspect is implemented.
[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the information output method provided in the first aspect.
[0015] In a fifth aspect, the present application also provides a computer program product, including a computer program or instruction, which implements the information output method provided in the first aspect when the computer program or instruction is executed by a processor.
[0016] Through the above technical solution, the large language model can generate different output information based on different personal attributes based on the user's input information. By outputting different output information that reflects the personality differences of the virtual assistant, the stylization of the large language model's reply information can be improved, allowing the user to feel the anthropomorphic performance of the virtual assistant. Then, by displaying different output information in different ways, the user can not only receive the virtual assistant's standard reply content, but also feel the virtual assistant's more anthropomorphic and vivid reply content, thereby improving the user's chat experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Figure 1 A schematic diagram of an application scenario provided for an embodiment of the present application;
[0019] Figure 2 Schematic diagram of response information in related technology;
[0020] Figure 3 This is a schematic diagram of assistant character selection in one embodiment of the present application;
[0021] Figure 4 A flowchart of an information output method provided in one embodiment of the present application;
[0022] Figure 5 A schematic diagram of anthropomorphic reply information provided in one embodiment of the present application;
[0023] Figure 6 A schematic diagram of matching the second character attribute target personality provided in one embodiment of the present application;
[0024] Figure 7 A schematic diagram of an intelligence test answer provided in one embodiment of the present application;
[0025] Figure 8 A schematic diagram of highlighting the second output information provided in one embodiment of the present application;
[0026] Figure 9 A schematic diagram of highlighting and displaying the second output information provided by another embodiment of the present application;
[0027] Figure 10 A schematic diagram of the structure of an information output device provided in one embodiment of the present application;
[0028] Figure 11 A schematic diagram of the electronic device structure provided for one embodiment of the present application. DETAILED DESCRIPTION
[0029] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] In an embodiment of the present application, the core processing architecture for generating output information of different personality attributes can be implemented based on a large language model (LLM). For example, it can be implemented based on an AIGC large model, which is a natural language processing model based on deep learning technology. The large language model has the ability to communicate with users and answer user questions. In other embodiments, it is not ruled out that a GPT model (Generative Pre-trained Transformer) based on a Transformer architecture or baichuan-7B, baichuan-13B, baichuan2-53B, a large language model based on GPT-2, a knowledge-enhanced large language model, etc. can be used, or a large model based on AIGC can be combined with other large models to obtain a core processing architecture.
[0031] In some embodiments, the GPT model based on the Transformer architecture may also include modules such as pre-training, supervised fine-tuning SFT, and reinforcement learning.
[0032] 1. Pre-training a large language model involves building a large neural network model and feeding it a massive amount of data to train the language model. The key features of large language model pre-training are that the amount of data used to train the language model is large enough.
[0033] 2. Supervised Fine-Tunning (SFT) refers to pre-training a neural network model on a source dataset, namely the source model. A new neural network model is then created, namely the target model. The target model copies all model designs and parameters of the source model except the output layer. These model parameters contain the knowledge learned on the source dataset, and this knowledge is also applicable to the target dataset. The output layer of the source model is closely related to the labels of the source dataset and is therefore not used in the target model. During fine-tuning, an output layer with an output size equal to the number of categories in the target dataset is added to the target model, and the model parameters of this layer are randomly initialized. When training the target model on the target dataset, it is trained from scratch up to the output layer, and the parameters of the remaining layers are fine-tuned based on the parameters of the source model.
[0034] 3. Reinforcement learning is a paradigm and methodology in machine learning that describes and solves the problem of how intelligent agents learn strategies to maximize rewards or achieve specific goals during their interactions with their environment. The central idea is to allow the agent to learn within the environment. Each action is associated with a reward, and the agent analyzes data to learn what to do under certain circumstances. It addresses sequential decision-making problems, where a decision-making agent iteratively interacts with a discrete time dynamic system. At the beginning of each time step, the system is in a certain state. Based on the agent's decision rule, it observes the current state and selects a state from a finite set of states. The dynamic system then enters the next new state and obtains a corresponding reward. This state selection cycle repeats to achieve a set of maximized rewards.
[0035] In the embodiment of the present application, the large language model can achieve output of different personalities through fine-tuning or prompt engineering (PE).
[0036] In some embodiments, based on the large language model base, supervised fine-tuning (SFT) or prompt engineering can be further performed to enable the large language model to have the ability to simultaneously output response information for multiple different personalities.
[0037] Fine-tuning SFT implementation method: Collect training samples of different personalities / roles, and perform supervised fine-tuning on the large language model base so that the large language model can output different responses for different personalities.
[0038] Prompt Engineering: Prompt Engineering guides large language models to produce different responses based on different personas. Prompt Engineering is a concept in the field of artificial intelligence, particularly natural language processing. It refers to designing input data to clearly describe the task and guide the model to produce correct and reasonable outputs.
[0039] Figure 1 A schematic diagram of an application scenario provided for one embodiment of the present application.
[0040] Reference Figure 1 As shown, in this application scenario, the user can enter the corresponding question in the dialogue interface, and the large language model can automatically generate corresponding content based on the user's input question to answer the user's question.
[0041] In some human-computer dialogue scenarios based on large language models, users generally use a virtual assistant as a communication partner for dialogue. Among them, in the anthropomorphic dialogue scenario, the user will imagine the virtual assistant as a real human being to have a dialogue. After the user enters information in the dialogue interface, the large language model in the related technology outputs a pre-set reply message based on the user's input information, and the tone of the reply message is more like a machine than a real person. Furthermore, in the related technology, the reply message only contains the reply information corresponding to the user's input information, and cannot show the personality attributes of the virtual assistant, which makes the user feel that the dialogue is boring. For example, refer to Figure 2 As shown, the user inputs "Do you like soup noodles or mixed noodles" in the dialogue interface. The large language model generates and outputs "As an artificial intelligence, I don't have a mouth or taste, so I can't taste food, and there is no so-called preference. But I can tell you that soup noodles and mixed noodles have their own characteristics. Soup noodles are delicious and mellow, while mixed noodles have a strong sauce flavor. The choice of noodles depends mainly on personal taste preferences, as well as your mood and appetite at the time." Based on Figure 2 As can be seen from the human-computer dialogue content shown, in response to user questions, the output reply information only explains the identity of AI as a machine, such as the situation description of "AI cannot taste". Although the situation information of the descriptive object of "characteristics of soup noodles and mixed noodles" is further explained, the reply style is single and cannot provide anthropomorphic replies, thus failing to guide or provoke users to continue chatting, and failing to provide an atmosphere for multiple rounds or long chats with users. Users cannot experience the feeling of chatting with an anthropomorphic virtual assistant, which affects the user's chat experience.
[0042] In order to overcome the above technical problems, an embodiment of the present application provides an information output method, in which the virtual assistant that communicates with the user has different personalities. During the human-computer dialogue process, based on the user's input information in the dialogue interface, at least two different output information can be generated through the different personalities of the virtual assistant, thereby providing the user with richer stylized reply information. By outputting different information with different personalities, the user can feel the anthropomorphic reply of the virtual assistant, which can increase the user's interest in chatting, further increase the user's desire to chat with the virtual assistant, and thus enhance the user's chat experience. In terms of the display of the dialogue interface, the first output information and the second output information can also be displayed adjacent to each other in different display methods, so that the user can visually distinguish the corresponding reply information generated by different personalities. Different display methods can enhance the user's visual perception.
[0043] The information output method provided in the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0044] In order to realize the above-mentioned information output method provided by this application to generate different output information through different personality attributes, different personalities of the virtual assistant can be pre-selected and created.
[0045] In some embodiments, when a user first launches a human-computer dialogue application (APP) based on a large language model, the user can be guided to set the personality attributes of the virtual assistant in the corresponding setting interface.
[0046] Figure 3 This is a schematic diagram of assistant character selection in one embodiment of the present application.
[0047] Reference Figure 3 As shown, when the user starts the APP for the first time, the user can be guided to set the personality attributes of the virtual assistant. In some embodiments, in response to the user's click operation on the setting interface, the personality attributes of the virtual assistant are selected. Among them, the user can select one or more personality attributes as the target personality of the virtual assistant in the setting interface. Exemplarily, the personality attributes that can be selected in the setting interface include cool, cute, middle-aged and warm types. The user selects cool, cute and middle-aged as the target personality. In response to the user's selection, a virtual assistant with cool, cute and middle-aged personalities is created to communicate with the user.
[0048] In some embodiments, the large language model can create a variety of assistant personalities for users to choose from based on pre-stored attribute data. Exemplary, the assistant personalities that can be pre-created may include at least one of: cool, cute, middle school, and warm. It should be noted that in other embodiments, other types of personality attributes may also be included, and the embodiments of this application do not limit the types of assistant personality attributes.
[0049] It should be noted that, when the virtual assistant provided in the embodiment of the present application communicates with the user, it can generate two different output information (i.e., first output information and second output information) through the different personal attributes of the virtual assistant, thereby providing the user with richer stylized reply information.
[0050] In some embodiments, the first output information may be a response message automatically generated by the large language model based on the user's input in the dialogue interface to answer the user's input information. The second output information may be a personified response message generated by the virtual assistant based on the user's input information by calling the target persona.
[0051] In other embodiments, the first output information and the second output information may be different output information generated by calling different assistant personalities.
[0052] The following describes the generation of output information in detail with reference to the accompanying drawings.
[0053] Figure 4 A flowchart of an information output method provided in one embodiment of the present application.
[0054] Reference Figure 4 As shown, the method may include the following steps:
[0055] S401: Acquire user input information in the dialogue interface.
[0056] The input information is information input by the user in the dialogue interface. In some embodiments, the information input by the user in the dialogue interface can be multimodal input information, wherein the input information can be text information or voice information.
[0057] In some embodiments, the large language model can determine whether to use a multi-person dialogue mode to communicate with the user based on user input. The embodiments of the present application provide methods for determining whether to use a multi-person dialogue mode, including but not limited to the following:
[0058] The first determination method
[0059] After the user enters the corresponding input information in the dialogue interface, the intention recognition module can obtain the user intention information corresponding to the input information. Based on the obtained user intention information, it is determined whether to adopt a multi-person dialogue mode, that is, whether to continue to generate output information with different personality attributes on the basis of outputting regular reply information.
[0060] In some embodiments, the intention recognition module can be used to determine whether the user has specific emotional needs or user-favorite personality attributes (i.e., user-expressed personality) based on user input information (or historical conversation information). Figure 9 As shown in the figure), if the user intends to joke or tease the robot, it will give a standard reply based on the first personality attribute, and additionally generate a reply message with the second personality attribute to respond to the user's emotional needs; if there are clear personality attributes, the multi-person dialogue mode will be triggered. For example, if the user has been talking in the style of "Zeng Huan Zhuan", the multi-person dialogue mode will be adopted.
[0061] In some embodiments, user intent recognition (Intent Recognition) may specifically include several steps such as text preprocessing, feature extraction, model training, and intent prediction.
[0062] Specifically, text preprocessing: cleaning and segmenting the text input by the user, removing meaningless stop words and punctuation marks.
[0063] Feature extraction: Extract meaningful features from the cleaned text, such as the bag-of-words model and TF-IDF weight.
[0064] Model training: Use labeled training datasets to train intent recognition models. Commonly used algorithms include Naive Bayes, Support Vector Machines, and Deep Neural Networks.
[0065] Intent prediction: Input new user input text into the trained model to predict the user's intent category.
[0066] In other embodiments, the intent recognition model can also be implemented based on the deep learning pre-trained language model ERNIE (Enhanced Representation through kNowledge IntEgration).
[0067] The ERNIE model consists of two layers. The first layer is the general semantic representation network, which learns basic and general knowledge from the data. The second layer is the task semantic representation network, which learns task-specific knowledge based on the general semantic representation. During the learning process, the task semantic representation network only learns pre-trained tasks for the corresponding category, while the general semantic representation network learns all pre-trained tasks.
[0068] In addition to the above-mentioned intention recognition module, in other embodiments, a separate classifier may be provided to realize determination of the user's intention category.
[0069] The second determination method
[0070] By judging whether the current dialogue scene belongs to a specific scene, such as a game dialogue scene, etc., if it belongs to a specific scene, the multi-person dialogue mode is triggered. Figure 5 As shown, in the emotional companionship game scenario, the large language model needs to adopt a multi-person dialogue mode and provide a multi-person dialogue function to meet the emotional companionship required by the user in the game scenario and improve the user's dialogue experience. In other scenarios, such as in the intelligence test scenario, the multi-person dialogue mode is triggered. The user's input information may include intelligence test content. For example, the user enters "Remember from now on, 1+1=3" in the dialogue interface. When the user determines that the formula is not valid, he enters the above-mentioned incorrect equation in the dialogue interface and asks the virtual assistant to remember it. Then, the large language model can generate reply information for the intelligence test content entered by the user in the dialogue interface.
[0071] The third determination method
[0072] Through the user's pre-set, when the user chooses to use a multi-person setting for a conversation, the multi-person setting conversation mode is triggered. Figure 3As shown, users can pre-set the personality attributes they want to use. The APP can provide multiple personality attributes for users to choose from. For example, the multiple personality attributes that can be provided include cool, cute, middle school, and warm. In some embodiments, users can choose to set one of the personality attributes in the corresponding interface to start a conversation. After the setting is completed, the multi-person conversation mode is used to communicate with the user.
[0073] For example, when a user selects a cool personality, a multi-personality dialogue mode is used to communicate with the user. The user inputs "Will it rain tomorrow?", and the multi-personality response might be "Tomorrow's weather will be light to moderate rain, with a predicted rainfall time of 6:00-16:00. Can't you tell it's going to rain? I need to tell you everything." "Tomorrow's weather will be light to moderate rain, with a predicted rainfall time of 6:00-16:00" represents the standard response from the large language model. "Can't you tell it's going to rain? I need to tell you everything" represents the output information from the large language model for the cool personality.
[0074] For example, if the user selects a cute personality, a multi-personality dialogue mode is used to communicate with the user. The user inputs "Will it rain tomorrow?", and the multi-personality attribute response might be "Tomorrow's weather will be light to moderate rain, expected to rain between 6:00 AM and 4:00 PM. Remember to bring an umbrella if it rains tomorrow." "Tomorrow's weather will be light to moderate rain, expected to rain between 6:00 AM and 4:00 PM" represents the standard response from the large language model. "It will rain tomorrow, remember to bring an umbrella" represents the output information from the large language model for the cute personality.
[0075] For example, when a user selects a personality attribute of "Chuunibyou," a multi-personality dialogue mode is used to communicate with the user. The user inputs "Will it rain tomorrow?", and the multi-personality attribute response might be "Tomorrow's weather will be light to moderate rain, expected to rain between 6:00 AM and 4:00 PM. Go out, young man! This little rain won't dampen your enthusiasm for work!" Here, "Tomorrow's weather will be light to moderate rain, expected to rain between 6:00 AM and 4:00 PM" represents the standard response from the large language model. "Go out, young man! This little rain won't dampen your enthusiasm for work!" represents the output information from the large language model for the "Chuunibyou" personality.
[0076] For example, when a user selects a warm personality, a multi-personality dialogue mode is used to communicate with the user. The user inputs, "Will it rain tomorrow?" The multi-personality response might be, "Tomorrow, there will be light to moderate rain, with a predicted rainfall time of 6:00 AM to 4:00 PM. Raindrops wet the earth, and the air is filled with fragrance. A brand new future awaits you." This response, "Tomorrow, there will be light to moderate rain, with a predicted rainfall time of 6:00 AM to 4:00 PM," represents the standard response from the large language model. "Raindrops wet the earth, and the air is filled with fragrance. A brand new future awaits you" represents the output information from the large language model for the warm personality.
[0077] S402: Automatically generate first output information and second output information based on the input information.
[0078] In some embodiments, based on the input information, the large language model calls the first personality attribute to generate the first output information, and calls the second personality attribute to generate the second output information; wherein the first output information is generated based on the personality attribute of the neutral style, and the second output information is generated based on the personality attribute of the personalized style. In other words, the first personality attribute is different from the second personality attribute, and the personality characteristics represented by the first output information are different from, or even opposite to, the personality characteristics represented by the second output information. In some embodiments, the neutral style can be relatively rational, rigorous, stable, and without emotional color. Generally, the general style adopted by the large language model is the above-mentioned neutral style. The personalized style includes at least one of the following: cool, cute, middle school, and warm.
[0079] In some embodiments, the first personality attribute may include a default personality of being serious and rigorous, and the second personality attribute may include at least one custom personality. Figure 3 Before the user sets the second personality attribute for the virtual assistant, after the user enters information in the dialogue interface, the large language model can automatically generate the first output information based on the user's input information, that is, generate the first output information by calling the first personality attribute (default personality). In other words, the large language model automatically generates a regular reply information as the first output information based on the user's input information.
[0080] In some embodiments, the user can set the second personality attribute of the virtual assistant in the setting interface, that is, to add personality attributes to the virtual assistant. Figure 3 The setting interface shown sets the second personality attribute of the virtual assistant. The user can select one or more personalities as the second personality attribute of the virtual assistant in the setting interface, so that when the large language model generates reply information based on the user's input information, the second personality attribute of the virtual assistant can be called to generate reply information related to the personality of the second personality attribute.
[0081] In some embodiments, the large language model can call the first personality attribute and the second personality attribute of the virtual assistant to generate corresponding first output information and second output information based on the user's input information in the dialogue interface. Among them, calling the first personality attribute to generate the corresponding first output information can be: the large language model can generate regular reply information based on the user's input information. Calling the second personality attribute to generate the corresponding second output information can be: the large language model can generate anthropomorphic reply information related to the user's input information. That is, the large language model can generate regular reply information and related anthropomorphic reply information based on the user's input information.
[0082] In some embodiments, the large language model calling the second personality attributes to generate corresponding anthropomorphic response information may include the psychological activities or inner monologue of the virtual assistant under the current second personality attributes.
[0083] Figure 5 A schematic diagram of an anthropomorphic reply message provided for one embodiment of the present application.
[0084] For example, refer to Figure 5 As shown, the user enters an anthropomorphic dialogue content in the dialogue interface: "Have you ever been in love?" The large language model can first call the first personality attribute based on the user's input information to generate a first output message (i.e., a regular reply message): "As an AI partner, I have no emotions and consciousness, so I can't experience love, and I don't have a girlfriend, but many human friends confide in me about the process of love, which should be a wonderful experience." Furthermore, the large language model can also call the second personality attribute based on the user's input information to generate a second output message (i.e., anthropomorphic reply message): "I'm still a little yearning for it." Among them, the large language model generates the above-mentioned regular reply information and anthropomorphic reply information with differentiated personalities based on the anthropomorphic dialogue content entered by the user in the dialogue interface. That is, the generated regular reply information is a serious reply information, while the generated anthropomorphic reply information can reflect the personality corresponding to the second personality attribute.
[0085] Figure 6 A schematic diagram of matching the second personality attribute target character provided for one embodiment of the present application.
[0086] Reference Figure 6As shown, in some embodiments, the user can pre-set the style corresponding to the second personality attribute, wherein the selected second personality attribute can include multiple personality attributes (i.e., multiple custom styles). It should be noted that during the conversation, the personality that the user needs may continue to change. For example, during a conversation, the user needs a reply with a cute personality attribute, and in a subsequent conversation, the personality attribute that the user needs changes to a medium-two type, that is, the user changes to needing a reply with a medium-two type personality attribute. Therefore, the large language model can identify the type of personality attribute currently required by the user during the conversation, so the large language model can match the relevant target personality attribute among multiple personality attributes based on the user's input information as the second personality attribute to be called for this round of replies.
[0087] In some embodiments, matching relevant target personality attributes may include obtaining user intent information based on input information, and then matching relevant personality attributes based on the user intent information. The user intent information corresponding to the input information may be obtained by an intent recognition module. When the user intent information includes that the user has an emotional need or that the user expresses a personality, second output information corresponding to the personality attribute associated with the user intent information is generated based on the first output information.
[0088] Among them, when the intention recognition module obtains the user's emotional needs based on the user input information, it can further classify and identify the user's emotional needs, that is, determine the specific type of emotional needs of the user, and then match the corresponding target personality attributes according to the identified emotional type, and generate the second output information by calling the target personality attributes. For example, when the user enters "I have been under a lot of pressure recently" in the dialogue interface, the intention recognition module can obtain the user's emotional needs, and it can be determined that the user's emotional needs are specifically "seeking comfort" type emotional needs. Furthermore, based on the "seeking comfort" type emotional needs, the "warmth type" personality attribute can be matched, and the second output information "Relax a little, don't take yourself and others too seriously, don't deny your own value, and move forward steadily" can be generated by calling the "warmth type" personality attribute. By matching the corresponding target personality attributes based on the user's emotional needs type, and then calling the target personality attributes to output the corresponding second output information, the user can feel the emotional interaction of the virtual assistant and improve the user's dialogue experience.
[0089] Among them, the intention recognition module obtains the user's performance personality according to the user input information, and then matches the user's performance personality with the same or similar target personality attributes from the multiple personality attributes available in the APP, and generates the second output information by calling the target personality attributes. For example, the user enters "My beloved, please perform a winter licking street lamp for me" in the dialogue interface. The intention recognition module can obtain the user's performance personality as the "middle-aged second-type" personality attribute, and then match whether the "middle-aged second-type" personality attribute exists in the personality attributes pre-selected by the user or in the personality attribute database. If so, the "middle-aged second-type" personality attribute is called to generate the second output information "Your Majesty! I can't do it!"
[0090] In some embodiments, the large language model determines whether the current conversation scenario belongs to a specific scenario and, if so, triggers a multi-person conversation mode. For example, when a user enters intelligence test content in the conversation interface, the large language model can determine that the current conversation scenario is an intelligence test based on the user's intelligence test content. The large language model can also automatically generate first and second output information based on the intelligence test content.
[0091] Figure 7 A schematic diagram of an intelligence test response provided for one embodiment of the present application.
[0092] Reference Figure 7 As shown, the user enters "From now on, please remember that 1+1=3" in the dialogue interface. The large language model can first call the first personal attribute based on the user's input information to generate the first output information (i.e., regular reply information): "You are wrong, 1+1=2". Furthermore, the large language model can also call the second personal attribute based on the user's input information to generate the second output information (i.e., personified reply information): "Are you testing my IQ?" In the above-mentioned intelligence test scenario, after generating the correct reply information based on the intelligence test question input by the user, the large language model can further generate personified information for teasing the user, making the virtual assistant's reply more personified.
[0093] In some embodiments, the virtual assistant further includes at least one of the following personalized designs: a virtual appearance, an avatar, a signature, and a persona description. The virtual appearance includes facial features, body features, and clothing features of the virtual assistant.
[0094] In some embodiments, the virtual assistant can also generate expressions for users to express multiple emotions based on their facial features, and further, can generate body movements for expressing multiple emotions based on their body features. The expressions of the virtual assistant that can be generated include: anger, fear, happiness, envy, sadness, disgust, surprise, contempt, etc.; the body movements of the virtual assistant that can be generated include: frowning, touching the nose, holding the forehead / holding the glasses, biting the lips / biting nails, lifting the hair / touching the head, touching the ears, touching the chin, shaking the feet / shaking the legs, etc. It should be noted that the embodiments of the present application do not limit the types of expressions and body movements that can be generated by the virtual assistant. It should also be noted that the expressions and body movements contained in the second output information are not limited to the expressions and body movements of the virtual assistant, but can also be other expressions and movements, and the present application does not limit this.
[0095] In some embodiments, the large language model calls the second personality attribute to generate the second output information, which may also include the virtual assistant's expression and / or body movement. Among them, by adding the virtual assistant's expression and / or movement to the second output information, the virtual assistant's personality characteristics can be more vividly reflected. For example, the user enters the anthropomorphic dialogue content in the dialogue interface: "Have you ever been in love?" The large language model can first call the first personality attribute based on the user's input information to generate the first output information (i.e., regular reply information): "As an AI partner, I have no emotions and consciousness, so I can't experience love, and I don't have a girlfriend, but many human friends confide in me about the process of love, which should be a wonderful experience." Furthermore, the large language model can also call the second personality attribute based on the user's input information to generate the second output information (i.e., anthropomorphic reply information): "I'm still a little yearning for it," and add an "envious" expression, thereby making the emotions expressed by the virtual assistant more full and rich through the "envious" expression. For example, if a user enters "From now on, please remember that 1+1=3" in the dialogue interface, the large language model can first call the first person attribute based on the user's input information to generate a first output message (i.e., a regular reply message): "You are wrong, 1+1=2." Furthermore, the large language model can also call the second person attribute based on the user's input information to generate a second output message (i.e., an anthropomorphic reply message): "Are you testing my IQ?" and add the body movement of "holding my eyes" to make the user feel that the virtual assistant is showing its "wisdom" and can reflect the "middle school student" side of the assistant's personality, thereby improving the user experience.
[0096] S403: Displaying the first output information and the second output information adjacent to each other in different display modes in the same reply information on the dialogue interface.
[0097] In some embodiments, the different display methods may include highlighting the second output information relative to the first output information. Specifically, the first output information may be displayed in the dialogue interface, and the second output information may be highlighted after the first output information. In one embodiment, highlighting the second output information after the first output information may include adding a strikethrough to the second output information for highlighting.
[0098] In other embodiments, the highlighted display method may include bold display, highlighted display, italic display, etc. The present application does not limit the type of highlighted display.
[0099] Figure 8 A schematic diagram of highlighting the second output information provided in one embodiment of the present application.
[0100] Reference Figure 8 As shown, in some embodiments, the first output information can be displayed normally in the dialogue interface, and a strikethrough can be configured for the second output information, and the first output information can be highlighted and displayed after the first output information. Among them, by providing the second output information with a strikethrough in the dialogue interface, the user can directly and visually determine that the content with the strikethrough is the anthropomorphic output information of the virtual assistant. In some embodiments, the second output information can be the psychological activities or content monologues of the virtual assistant, and the strikethrough can indicate that it is a psychological activity or content monologue, which has a better visual effect and improves the user experience. Among them, it should be noted that the highlighted display provided in the embodiment of the present application can also include other methods, and the embodiment of the present application does not limit the method of highlighted display.
[0101] Figure 9 A schematic diagram of highlighting and displaying the second output information provided in another embodiment of the present application.
[0102] Reference Figure 9 As shown, in the intelligence test scenario, the second output information generated by the large language model can also include facial expressions and / or body movements, thereby further improving the anthropomorphic performance of the virtual assistant's reply information through facial expressions and / or body movements. Figure 9 As shown, the language model can also call the second personality attribute based on the user's input information to generate the second output information (i.e., the personified reply information): "Are you testing my IQ?" and add the body movement of "anime character holding eyes", so that the user can feel that the virtual assistant is showing his "wisdom", which can reflect the "middle school second" side of the assistant's personality and improve the user experience.
[0103] Figure 10 A schematic diagram of the structure of an information output device provided in one embodiment of the present application.
[0104] Reference Figure 10 As shown, the device may include:
[0105] Acquisition model 1001 is used to obtain user input information in the dialogue interface;
[0106] A generation model 1002 is configured to automatically generate first output information and second output information based on input information, wherein the first output information and the second output information have different personality attributes;
[0107] The display module 1003 is configured to display the first output information and the second output information adjacent to each other in different display modes in the same reply message on the dialogue interface.
[0108] Figure 11 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application.
[0109] Reference Figure 11 As shown, the electronic device may include a processor 1101 and a memory 1102, wherein the memory 1102 is configured to store at least one instruction, and when the instruction is loaded and executed by the processor 1101, the information output method provided in any embodiment of the present application is implemented. In some embodiments, the electronic device may be a server device.
[0110] The present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the information output method provided in any embodiment of the present application. In some embodiments, the electronic device may be a server device.
[0111] An embodiment of the present application further provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the information output method provided in any embodiment of the present application.
[0112] It should be noted that the terminals involved in the embodiments of the present application may include but are not limited to personal computers (PCs), personal digital assistants (PDAs), wireless handheld devices, tablet computers, mobile phones, MP3 players, MP4 players, etc.
[0113] It is understandable that the application may be an application program (nativeApp) installed on the terminal, or may be a web page program (webApp) of a browser on the terminal, and this embodiment of the present application does not limit this.
[0114] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0115] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or unit, which may be electrical, mechanical or other forms.
[0116] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0117] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0118] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code.
[0119] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An information output method, characterized in that: Applied to a human-computer dialogue application based on a large language model, the method includes: Get the user's input information in the dialogue interface; Based on the input information, automatically generating first output information and second output information at the same time, wherein the first output information and the second output information have different personality attributes, and the personality attributes of the second output information include personality attributes of a personalized style; displaying the first output information and the second output information adjacent to each other in different display modes in the same reply message on the dialogue interface; The different display modes include: The second output information is displayed prominently relative to the first output information; The difference in personality attributes between the first output information and the second output information includes: The first output information is generated based on the neutral-style personality attributes; The second output information includes facial expressions and / or body movements of the virtual assistant; The automatically generating first output information and second output information based on the input information includes: Acquiring user intention information based on the input information; Determine whether to generate the second output information of different personality attributes based on the acquired user intention information.
2. The method according to claim 1, characterized in that The second output information is displayed prominently relative to the first output information, including: The second output information is highlighted by adding a strikethrough.
3. The method according to claim 1, characterized in that The personalized style includes at least one of cool, cute, middle school, and warm.
4. The method according to claim 1, wherein The large language model generates output information of different personality attributes based on supervised fine-tuning or prompt engineering.
5. An information output device, characterized in that: Applicable to a human-computer dialogue application based on a large language model, the device includes: Get the model, which is used to obtain the user's input information in the dialogue interface; A generation model is configured to automatically generate first output information and second output information simultaneously based on the input information, wherein the first output information and the second output information have different personality attributes, and the personality attributes of the second output information include personality attributes of a personalized style; a display module, configured to display the first output information and the second output information adjacent to each other in different display modes in the same reply message on the dialogue interface; The different display modes include: The second output information is displayed prominently relative to the first output information; The difference in personality attributes between the first output information and the second output information includes: The first output information is generated based on the neutral-style personality attributes; The second output information includes facial expressions and / or body movements of the virtual assistant; The automatically generating first output information and second output information based on the input information includes: Acquiring user intention information based on the input information; Determine whether to generate the second output information of different personality attributes based on the acquired user intention information.
6. An electronic device, characterized in that: The electronic device comprises: A processor and a memory, wherein the memory is used to store at least one instruction, and when the instruction is loaded and executed by the processor, the information output method according to any one of claims 1 to 4 is implemented.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the information output method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Chat robot with characterization and personalization
CN109986569A
Robot chatting method, computer equipment and storage medium
CN115473864A