Dialogue support system, dialogue support method, and program

The dialogue support system uses large-scale language models to generate persona prompts and control avatars, addressing the limitations of existing systems by enhancing dialogue variability and realism in role-playing scenarios.

JP7893330B2Inactive Publication Date: 2026-07-22TOPPAN HOLDINGS INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TOPPAN HOLDINGS INC
Filing Date
2025-03-11
Publication Date
2026-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing dialogue systems struggle to dynamically adjust dialogue content based on the attributes of the dialogue partner, scene, and situation, limiting the variability and realism of role-playing scenarios.

Method used

A dialogue support system utilizing large-scale language models to generate persona generation and specification prompts, controlling avatars to interact with users based on persona information and unique conversation details, enabling dynamic adjustment of dialogue attributes.

Benefits of technology

Enhances the variability of dialogue partners and improves the realism of role-playing by accurately reflecting the attributes of the dialogue partner, facilitating effective training scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007893330000001
    Figure 0007893330000001
  • Figure 0007893330000002
    Figure 0007893330000002
  • Figure 0007893330000003
    Figure 0007893330000003
Patent Text Reader

Abstract

To easily increase variations of attributes of a dialogue partner and to support dialogue in accordance with the attributes of the dialogue partner.SOLUTION: A dialogue support system of one aspect of the present disclosure includes: a receiving unit for receiving persona information indicating a persona that characterizes a dialogue partner of a user; a persona generation unit for generating a persona generation prompt indicating features of the persona by using a large-scale language model on the basis of the persona information received by the receiving unit; and a dialogue control unit for controlling an avatar corresponding to the persona on the basis of the persona generation prompt generated by the persona generation unit and controlling dialogue between the avatar and the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a dialogue support system, a dialogue support method, and a program.

Background Art

[0002] Conventionally, as a technology for supporting dialogue with a user, for example, the technology described in Patent Document 1 is known. Patent Document 1 describes a training system for supporting the education of professionals who require conversation skills, in which communication AI is implemented. The training system includes a processor configured to execute acquiring the utterance of a trainee, analyzing the content of the acquired utterance of the trainee, creating the next utterance content directed to the trainee based on the analyzed utterance content, and synthesizing the voice representing the utterance content.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Although various dialogue partners, scenes, and situations desired by the user are assumed depending on the user's use, the training system described in Patent Document 1 has a problem that it is difficult to change the dialogue content according to the attributes of the dialogue partner, scene, and situation. In particular, it is difficult to increase the variations for existing role-playing or to perform role-playing with reality.

[0005] In view of such circumstances, the present disclosure has been made, and an object thereof is to provide a dialogue support system, a dialogue support method, and a program that can easily increase the variations of the attributes of the dialogue partner and support a dialogue according to the attributes of the dialogue partner. [Means for solving the problem]

[0006] This disclosure was made to solve the above-mentioned problems, and one aspect of this disclosure includes a receiving unit that receives persona information indicating a persona that characterizes the user's conversation partner, a unique information acquisition unit that acquires unique information which is information specific to the conversation, and a first large-scale language model that processes the persona information and the unique information received by the receiving unit. Enter it, Characteristics of the persona To instruct the second large-scale language model to generate The dialogue support system comprises: a persona generation unit that generates a persona generation prompt and generates a persona specification prompt that indicates the characteristics of the persona using the second large-scale language model based on the generated persona generation prompt; and a dialogue control unit that generates response text using a third large-scale language model based on the persona specification prompt and the unique information generated by the persona generation unit, controls an avatar corresponding to the persona, and controls the dialogue between the avatar and the user.

[0007] Other aspects of this disclosure include the steps of a server device receiving persona information that characterizes the user's conversation partner, the server device acquiring unique information which is information specific to the conversation, and the server device processing the received persona information and the unique information into a first large-scale language model. Enter it, Characteristics of the persona To instruct the second large-scale language model to generate The dialogue support method includes the steps of: generating a persona generation prompt; generating a persona specification prompt that indicates the characteristics of the persona using the second large-scale language model based on the generated persona generation prompt; and the server device generating response text using a third large-scale language model based on the generated persona specification prompt and the unique information, controlling an avatar corresponding to the persona, and controlling the interaction between the avatar and the user.

[0008] Other aspects of this disclosure include the steps of: receiving persona information indicating a persona that characterizes the user's conversation partner into a computer of a server device; obtaining unique information which is information specific to the conversation; and processing the received persona information and the unique information into a first large-scale language model. Enter it, Characteristics of the persona To instruct the second large-scale language model to generate The program performs the following steps: generating a persona generation prompt, generating a persona specification prompt that indicates the characteristics of the persona using the second large-scale language model based on the generated persona generation prompt, and generating response text using the third large-scale language model based on the generated persona specification prompt and the unique information, controlling an avatar corresponding to the persona, and controlling the interaction between the avatar and the user. [Effects of the Invention]

[0009] According to one aspect of the present invention, it is possible to easily increase the variety of attributes of the dialogue partner for role-playing and to realize role-playing that is appropriate to the attributes of the dialogue partner. [Brief explanation of the drawing]

[0010] [Figure 1] This is a block diagram showing one example configuration of the dialogue support system 1 in the embodiment. [Figure 2] This diagram illustrates the processing overview of the dialogue support system 1 in the embodiment. [Figure 3] A flowchart illustrating an example of the processing procedure of the dialogue support system 1 in the embodiment. [Figure 4] This figure shows an example of customer-defined information, including variable names and strings, in the embodiment. [Figure 5] This figure shows an example of a persona generation prompt in an embodiment. [Figure 6] This figure shows an example of variables created from the persona generation prompt in the embodiment. [Figure 7]This figure shows an example of a persona specification prompt in an embodiment. [Figure 8] This figure shows an example of persona designation information in an embodiment. [Figure 9] This figure shows an example of an evaluation prompt in the embodiment. [Modes for carrying out the invention]

[0011] Hereinafter, a dialogue support system, a dialogue support method, and a program to which the present invention is applied will be described with reference to the drawings.

[0012] Figure 1 is a block diagram showing one example configuration of the dialogue support system 1 in an embodiment. The dialogue support system 1 of this embodiment supports a conversation between a user and an avatar characterized by a specific persona. The dialogue support system 1 generates a persona based on information specified by the user, and controls an avatar corresponding to the generated persona, thereby enabling the user and the avatar to role-play. Role-playing can include, for example, conducting new employee training, sales training, language training, or communication training with the user and a simulated avatar as conversation partners. Furthermore, the role-playing in this embodiment may include interactions between people of different nationalities or origins.

[0013] The dialogue support system 1 includes, for example, a processing server device 100, a generation server device 200, and a user terminal device 300. The processing server device 100, the generation server device 200, and the user terminal device 300 are communicably connected via a network NW such as the Internet. Note that the processing server device 100, the generation server device 200, and the user terminal device 300 may be connected by either wired communication or wireless communication, and may include a general-purpose network such as the Internet and a private network such as Local 5G or WiFi (registered trademark). The processing server device 100, the generation server device 200, and the user terminal device 300 may have a communication interface such as a NIC (Network Interface Card) or a wireless communication module for connecting to the network, and may exchange information with each other.

[0014] The user terminal device 300 is, for example, an information processing device operated by a user who interacts with an avatar. The user terminal device 300 includes, for example, a speaker, a microphone, a display device, an operation unit, a processing unit such as a CPU, and the like.

[0015] The processing server device 100 is a server device that includes a processor that performs processing in response to requests received from the generation server device 200 and the user terminal device 300, and transmits the processing results to the generation server device 200 and the user terminal device 300. The processing server device 100 includes, for example, a customer generation unit 110, a dialogue control unit 120, an operation control unit 130, and a storage unit 140. The customer generation unit 110, the dialogue control unit 120, and the operation control unit 130 are functional units that are realized by information processing circuits that perform various processing by having a CPU (Central Processing Unit) execute a program. Furthermore, some or all of these functional units may be realized by hardware such as an LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), or FPGA (Field-Programmable Gate Array), or they may be realized by the cooperation of software and hardware. The storage unit 140 can be implemented by, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), flash memory, an EEPROM (Electrically Erasable Programmable Read Only Memory), a ROM (Read Only Memory), or a RAM (Random Access Memory), or a hybrid storage device using multiple of these. Part or all of the storage unit 140 may be implemented by an external storage device accessible via various networks. An example of an external storage device is a NAS (Network Attached Storage) device.

[0016] The customer generation unit 110 generates customer information. The customer information is information indicating the customer assumed by the user. The customer corresponds to, for example, an avatar that becomes the user's interlocutor in role-playing. The customer generation unit 110 includes, for example, a reception unit 111 and a customer definition unit 112. The reception unit 111 receives persona information based on the information received from the user terminal device 300. The persona information is customer definition information indicating a persona that characterizes the customer (interlocutor) for the user. The persona may be a virtual human image or a human image based on the information of an actual person. Also, the persona information may be information based on the information of a partially existing person or information based on the information of a person who has already passed away. The customer definition unit 112 generates customer information based on the persona information received by the reception unit 111.

[0017] The dialogue control unit 120 performs a process of controlling the avatar corresponding to the persona based on the persona generation prompt generated by the persona generation unit 211 and controlling the dialogue between the avatar and the user. The dialogue control unit 120 includes, for example, a speech acquisition unit 121, an emotion parameter processing unit 122, a response prompt generation unit 123, a response text conversion unit 124, and a conversation history generation unit 125. The speech acquisition unit 121 acquires speech information indicating the user's speech input from the user terminal device 300 and converts the acquired speech information into text data. The emotion parameter processing unit 122 performs a process of setting and updating emotion parameters. The emotion parameter is a numerical value indicating the emotion of the avatar (customer). The emotion parameter is, for example, information representing each of emotions such as joy, anger, sadness, fun, confidence, confusion, and fear in five levels from 1 to 5. In this embodiment, a configuration related to the emotion of the customer such as the emotion parameter processing unit 122 is described, but it is not limited thereto, and a configuration related to the emotion of the customer may not be provided. The response prompt generation unit 123 generates a response prompt including the text data of the user voice and the emotion parameter, and transmits the generated response prompt to the generation server device 200. The response text conversion unit 124 converts the response text obtained from the generation server device 200 into audio data. The conversation history generation unit 125 generates history information that shows the history of conversations between the user and the avatar.

[0018] The motion control unit 130 performs processing to control the avatar's movements. The motion control unit 130 includes, for example, an avatar generation unit 131, a voice generation unit 132, a voice tone information processing unit 133, a motion processing unit 134, an emote processing unit 135, and a lip-sync processing unit 136. The avatar generation unit 131 generates an avatar. The avatar generation unit 131 creates component information that represents content to display the avatar based on an image showing the customer's appearance, for example.

[0019] The voice generation unit 132 generates voice data to be output to the user. For example, the voice generation unit 132 generates voice data that reproduces the customer's own voice reading aloud. If the persona information includes nationality, place of origin, or region, the voice generation unit 132 may generate voice data in the language corresponding to the nationality or place of origin based on the persona generation prompt generated by the persona generation unit 211.

[0020] The voice color information processing unit 133 processes the audio data based on voice color information corresponding to the emotion parameters. If the persona information includes nationality, place of origin, or region, the voice color information processing unit 133 may process the audio data generated by the audio generation unit 132 based on the nationality or place of origin included in the persona information. The voice color information processing unit 133 may process the audio data to reflect, for example, the tone of voice (e.g., speaking speed), pitch, or voice character. The voice color information processing unit 133 may also process the audio data to reflect dialects and intonations corresponding to differences in nationality or place of origin.

[0021] The elemental data for generating speech may include multiple elemental data corresponding to each of several languages. The speech generation unit 132 selects one of the multiple languages ​​based on the nationality or place of origin included in the persona information and generates speech data in the selected language. As a result, the speech generation unit 132 generates speech data using synthesized speech data corresponding to each of the multiple languages.

[0022] The motion processing unit 134 controls the avatar's motion based on emotion parameters and the content of response text. The avatar's motion can represent, for example, the movement of the entire avatar or the movement of the avatar's hands. The emote processing unit 135 controls the avatar's facial expressions based on emotion parameters and the content of response text. For example, the emote processing unit 135 controls the movements of the avatar's eyes, eyebrows, mouth, etc. The lip-sync processing unit 136 controls the movement of the avatar's lips based on emotion parameters and the content of the response text.

[0023] The memory unit 140 stores, for example, customer information 141, response information 142, voice information 143, and action information 144. The customer information 141 includes, for example, persona information, utterance information, persona generation prompts, and persona specification prompts. The persona generation prompts are detailed information for generating a persona. The persona specification prompts are information indicating the persona that is specified when the user and the avatar actually engage in a dialogue such as role-playing. Response information 142 includes, for example, user voice text and response text, but may also include initial values ​​and current values ​​of emotion parameters. Voice information 143 includes, for example, voice data such as user voice and response voice, and voice tone information, but may also include emotion parameters. Action information 144 includes, for example, emotion parameters, component information, motion information, emote information, and lip-sync information. Motion information is a default value representing the avatar's motion, emote information is a default value representing the avatar's emote, and lip-sync information is a default value representing the avatar's lip-sync.

[0024] The generation server device 200 is, for example, a server device that processes requests received from the processing server device 100 and transmits the processing results. The generation server device 200 comprises, for example, a generation unit 210, a storage unit 220, an evaluation unit 230, and an LLM learning unit 240. The generation unit 210, the evaluation unit 230, and the LLM learning unit 240 are functional units that are implemented by information processing circuits that perform various processes by having a CPU execute a program. The storage unit 220 is implemented by, for example, a recording device such as an HDD or SSD, or a hybrid storage device using multiple such devices, and may also be implemented by an external storage device that can be accessed via various networks such as a NAS device.

[0025] The generation unit 210 includes, for example, a persona generation unit 211, a response text generation unit 212, an emotion parameter generation unit 213, and a unique information acquisition unit 214. The persona generation unit 211 inputs persona information acquired from the processing server device 100 into a first large-scale language model and generates a persona generation prompt based on the output of the first large-scale language model. The persona information may be existing items including the persona's gender, age, personality, place of origin (including within Japan), way of speaking, tone of voice, or dialect. The first large-scale language model may also be at least one of the following: items specified based on user operations, items relating to customer characteristics in a specific industry, or items relating to customer characteristics in a specific generation. The first large-scale language model learns persona information, including existing items such as the persona's gender, age, personality, place of origin (including within Japan), way of speaking, tone of voice, or dialect; items specified based on user actions; items relating to customer characteristics in a specific industry; items relating to customer characteristics in a specific generation; and persona generation prompts as training data. It is configured to output a persona generation prompt when at least one of the following is input: the persona information, including existing items such as the persona's gender, age, personality, place of origin (including within Japan), way of speaking, tone of voice, or dialect; items specified based on user actions; items relating to customer characteristics in a specific industry; or items relating to customer characteristics in a specific generation.

[0026] The persona generation unit 211 may input persona information and information about a specific field into a first large-scale language model and create a persona generation prompt that shows the characteristics of a persona corresponding to the specific field based on the output of the first large-scale language model. Information about a specific field is various information about the field that is discussed in the dialogue. Information about a specific field may be customer characteristic information such as customer issues related to product purchase that are empirically assumed by a specific industry, a specific generation, a specific nationality, or a specific place of origin. Information about a specific field is acquired as unique information by the unique information acquisition unit 214. A specific field may be a field, industry, task, etc., that the user wants to improve.

[0027] Persona information includes nationality or place of origin. The persona generation unit 211 may input existing items, including the persona's nationality or place of origin, as persona information into the persona generation prompt (first large-scale language model), and create a persona generation prompt based on the output of the persona generation prompt.

[0028] The response text generation unit 212 generates response text from the response prompt generated by the response prompt generation unit 123, the conversation history generated by the conversation history generation unit 125, and the unique information acquired by the unique information acquisition unit 214. For example, the response text generation unit 212 inputs the response prompt, the conversation history between the user and the avatar, and the unique information into a second large-scale language model and generates response text based on the second large-scale language model. The response text generation unit 212 may also extract contextual information from the conversation and generate response text based on the contextual information in addition to the response prompt, conversation history, and unique information. The first large-scale language model is, for example, a large-scale language model (LLM) using a neural network. The second large-scale language model may be the same LLM as the first large-scale language model, or they may be different LLMs.

[0029] The emotion parameter generation unit 213 generates or updates emotion parameters according to the content of the generated response text. The unique information acquisition unit 214 acquires unique information, which is information specific to the dialogue, such as role-playing. Unique information is acquired from a storage device that contains, for example, customer characteristic information (not shown), specific field information, specific industry information, specific generational information, and information about a specific region (both domestic and international).

[0030] The memory unit 220 includes, for example, unique information 221 and LLM information 222. Unique information 221 includes customer characteristics information, specific field information, specific industry information, specific generational information, and specific country and region information. Customer characteristics information is information that shows the characteristics of the customer interacting with the user. For example, customer characteristics information includes information such as age, gender, occupation, way of speaking, tone of voice, personality, nationality, and place of origin. Specific field information is information that shows the field of interaction between the user and the customer. Specific industry information is information that shows the industry of interaction between the user and the customer. Specific generational information is information that shows the generation of the customer. LLM information 222 is parameter information for the LLM (First Large-Scale Language Model) for generating persona generation prompts. LLM information 222 may include parameter information for the LLM that generates persona specification prompts based on the persona generation prompts. The LLM for generating persona generation prompts is a machine learning model trained using historical data of persona generation prompts and persona specification prompts as training data, and is configured to output a persona specification prompt when a persona generation prompt is input. LLM information 222 may include parameter information for an LLM (Second Large-Scale Language Model) for generating response text based on a persona-specified prompt. The LLM for generating response text is a machine learning model trained on historical data of persona-specified prompts and response texts as training data, and is configured to output response text when a persona-specified prompt is input. The LLM information 222 may include parameter information for the LLM that generates an evaluation prompt based on the conversation history. The LLM that generates the evaluation prompt is a machine learning model trained on past conversation history and evaluation prompt data as training data, and is configured to output an evaluation prompt when conversation history is input. Note that the LLM for generating the persona generation prompt, the LLM for generating the persona specification prompt, the LLM for generating the response text, and the LLM for generating the evaluation prompt may be a single LLM, or they may be different LLMs.

[0031] The LLM learning unit 240 performs the process of learning an LLM (first large-scale language model) for generating persona generation prompts and an LLM (second language model) for generating response text. The LLM learning unit 240 may also learn an LLM for generating persona specification prompts and an LLM for generating evaluation prompts.

[0032] In this embodiment, the dialogue support system 1, as shown in Figure 1, distributes its functional components (function units) between the processing server device 100 and the generation server device 200. However, it is not limited to this configuration, and the function units may be distributed in other configurations. The function units of the processing server device 100 and the generation server device 200 may be consolidated into a single device. Multiple function units may be combined into a single function unit, or a single function may be distributed across multiple function units.

[0033] Figure 2 is a diagram illustrating the processing overview of the dialogue support system 1 in the embodiment. The reception unit 111 and the customer definition unit 112 generate customer definition information D10 and send it to the generation server device 200. The persona generation unit 211 inputs the customer definition information D10 to the persona generation LLM (P10) and generates a persona generation prompt D12 based on the output of the persona generation LLM (P10). The persona generation LLM (P10) is configured to learn, for example, past data of customer definition information and persona generation prompts as training data, and to output a persona generation prompt D12 when customer definition information D10 is input. The persona generation unit 211 generates a persona specification prompt D14 when the user performs role-playing. The persona specification prompt D14 is output to the response text generation LLM.

[0034] The generation server device 200 acquires unique information D20, such as specific fields, for role-playing, and processes the unique information D20 in the following order: text extraction process P20, chunking process P21, and vectorization process P22, storing the vectors corresponding to the unique information D20 in a vector database (storage unit 220). The processing server device 100 performs speech recognition processing P40 on the user's spoken voice acquired from the user terminal device 300, performs vectorization processing P41 on the text information processed by speech recognition processing P40, and extracts the result of referencing the vector database using the vectors corresponding to the spoken voice as queries from the vector database. The vectorized unique information D20 and the text information processed by speech recognition processing P40 are output to the response text generation LLM (P30) along with the persona specification prompt D14.

[0035] The generation server device 200 inputs the persona specification prompt D14, unique information D20, and text information processed by speech recognition processing P40 into the response text generation LLM (P30), performs speech synthesis processing P31 on the response text output from the response text generation LLM (P30), and performs avatar control processing P32 based on the emotion parameters output from the response text generation LLM (P30), thereby transmitting the avatar content D30 to the user terminal device 300. As a result, the user terminal device 300 can display or output audio using the avatar content D30.

[0036] Figure 3 is a flowchart showing an example of the processing procedure of the dialogue support system 1 in the embodiment. First, the processing server device 100 inputs user information about the user interacting with the avatar (step S100). User information is string data that characterizes the user, such as new employee, sales manager, specific nationality, or place of origin. Next, the processing server device 100 defines the customer the user is expecting (step S102). At this time, the processing server device 100 receives persona information from the user terminal device 300 via the reception unit 111, which indicates a persona that characterizes the user's conversation partner, and stores it as a variable, for example, as shown in Figure 4. Figure 4 shows an example of customer-defined information including variable names and strings in the embodiment. The reception unit 111 determines whether there is further input (step S104), and if input is received, it repeats the process in step S102, and if there is no input, it confirms the customer-defined information. The processing server device 100 transmits the customer-defined information to the generation server device 200.

[0037] The persona generation unit 211 inputs customer-defined information and unique information stored in the storage unit 220 to the persona generation LLM, and generates a persona generation prompt based on the output of the persona generation LLM (step S106). Figure 5 shows an example of a persona generation prompt in an embodiment. The persona generation prompt includes, for example, text data indicating the customer's preconditions, manner of speaking, and personality traits. The customer's preconditions include, for example, items appropriate for role-playing, such as age, gender, occupation, family structure, residential area, nationality, place of origin, and insurance information. If, for example, the place of origin is set to Osaka as a precondition, a persona generation prompt corresponding to the Kansai dialect can be created, and it is also possible to control the tone of voice or dialect according to the residential area. Furthermore, by setting the place of origin to California as a precondition, for example, a persona generation prompt that can reproduce a character strongly influenced by Californian culture can be generated. Customer speech patterns include, for example, first-person and second-person pronouns, verbal tics, interjections, and dialects. Customer personality traits include, for example, neuroticism, extroversion, openness, conscientiousness, and agreeableness. The persona generation unit 211 determines whether there is input of other items from the user terminal device 300 (step S108). If input is received, the process in step S106 is repeated. If there is no input, the persona generation prompt is confirmed. The persona generation unit 211 stores the created persona generation prompt as a variable in the storage unit 220. Figure 6 shows an example of variables created from the persona generation prompt in the embodiment. The variables generated from the persona generation prompt are information predicted based on the output of the persona generation LLM, after inputting customer-defined information into the persona generation LLM.

[0038] The persona generation unit 211 inputs existing items, including at least one of the persona's gender, age, personality, place of origin (including within Japan), way of speaking, tone of voice, and dialect, as persona information (customer definition information) into the persona generation LLM (first large-scale language model). Furthermore, it may input at least one of the following into the persona generation LLM: an item specified based on user operation, an item relating to customer characteristics in a specific industry, and an item relating to customer characteristics in a specific generation. Based on the output of the persona generation LLM, it may create a persona generation prompt. The persona generation unit 211 may also input existing items, including the persona's way of speaking in addition to gender, age, and personality, into the persona generation LLM. By inputting various information into the persona generation LLM, the persona generation unit 211 can realize Retrieval-Augmented Generation (RAG) and create a highly accurate persona generation prompt that meets the user's requirements.

[0039] The persona generation unit 211 may input items related to customer challenges regarding a specific product into a persona generation LLM (first large-scale language model) and create a persona generation prompt based on the output of the persona generation LLM. This persona generation LLM is a machine learning model that, for example, is trained using information on customer challenges regarding a specific product and past data of persona generation prompts as training data, and is configured to output a persona generation prompt when items related to customer challenges regarding a specific product are input. Items related to customer challenges regarding a specific product include information empirically assumed in a specific field or industry, and information such as customer challenges related to product purchase assumed based on market analysis or survey results. Items related to customer challenges regarding a specific product may be added or modified by the user as variables. Alternatively, the persona generation unit 211 may input existing items, including the persona's nationality or place of origin, as persona information into a persona generation LLM (first large-scale language model), and create a persona generation prompt based on the output of the persona generation LLM. This persona generation LLM is configured to learn from past data of existing items, including the persona's nationality or place of origin, and persona generation prompts as training data, and to output a persona generation prompt when an existing item, including the persona's nationality or place of origin, is input.

[0040] Next, the persona generation unit 211 creates a persona specification prompt (step S110). Figure 7 shows an example of a persona specification prompt in this embodiment. The persona generation unit 211 may input variables corresponding to the persona generation prompt into the LLM and create a persona specification prompt based on the output of the LLM. The variables corresponding to the persona generation prompt may be variables selected based on user operations, or they may be variables extracted from the persona generation prompt randomly or according to a predetermined rule. Based on the persona specification prompt, the persona generation unit 211 transmits persona specification information as shown in Figure 8 to the processing server device 100. Figure 8 shows an example of persona specification information in this embodiment.

[0041] Next, the dialogue control unit 120 controls an avatar corresponding to a persona based on the persona designation information and controls the dialogue between the avatar and the user (step S112). In this way, the dialogue control unit 120 performs role-playing through dialogue between the user and the avatar. The speech acquisition unit 121 acquires speech information indicating the user's speech, and the conversation history generation unit 125 stores the conversation history (step S114).

[0042] Next, the processing server device 100 determines whether or not to perform a user evaluation (step S116). If the processing server device 100 does not perform a user evaluation, it repeats the processes in steps S112 and S114. If the processing server device 100 detects a user utterance such as "evaluate the role-playing," it determines to perform a user evaluation and sends the conversation history to the generation server device 200.

[0043] The timing for evaluating the user (role-playing) is not limited to evaluating the entire role-play (such as a business negotiation) at the end of the role-play; evaluation may be performed after each rally (one round trip of conversation) during the role-play. For example, at the start of the role-play, the processing server device 100 allows the user to choose between an overall evaluation (evaluation at the end) or evaluation after each exchange. Overall evaluation is a process that evaluates the entire conversation history, while evaluation after each exchange is a process that evaluates each rally (including sets of responses and answers) during the conversation. If overall evaluation is selected, the processing server device 100 sends the conversation history to the generation server device 200 when it detects a user utterance such as "evaluate the role-playing," and if evaluation after each exchange is selected, it sends the conversation history for one rally to the generation server device 200 when it detects a break in the rally during the conversation. The generation server device 200 stores the results of evaluations performed each time based on the conversation history for one rally in the storage unit 220, and transmits one or more evaluation results to the processing server device 100 when the role-playing is completed. This allows the evaluation results to be presented to the user.

[0044] Furthermore, it is possible to simultaneously evaluate the entire role-play (such as a business negotiation) and evaluate each rally (one round trip of conversation) during the role-play. For example, at the start of the role-play, the processing server device 100 allows the user to choose whether to perform both an overall evaluation and a step-by-step evaluation. When the processing server device 100 detects a break in the rally during the conversation, it sends the conversation history for one rally to the generation server device 200, and further sends the conversation history to the generation server device 200 when it detects the user's utterance, "Evaluate the role-play." The generation server device 200 stores the results of the step-by-step evaluation based on the conversation history for one rally in the storage unit 220. When the role-play ends, the generation server device 200 sends the result of the overall evaluation and the results of one or more step-by-step evaluations to the processing server device 100. This allows the user to see the result of the overall evaluation and the result of the step-by-step evaluation on the same screen.

[0045] The evaluation unit 230 evaluates the user based on the speech information acquired by the speech acquisition unit 121 (step S118). At this time, the evaluation unit 230 generates an evaluation prompt based on the conversation history acquired from the processing server device 100. The evaluation unit 230 may input the conversation history into an LLM that generates evaluation prompts and generate an evaluation prompt based on the output of the LLM. For example, when a user gives a product description in a role-playing scenario, the evaluation unit 230 may input the conversation history and information stored in the product information database into an LLM that generates evaluation prompts and generate an evaluation prompt based on the output of the LLM that generates evaluation prompts. Figure 9 shows an example of an evaluation prompt in an embodiment. The evaluation prompt may include, for example, evaluation items, evaluation criteria, evaluation points, customer understanding evaluation, product knowledge evaluation, and communication evaluation for a conversation for sales activities.

[0046] The reception unit 111 may receive usage purpose information indicating the purpose of use for which the user interacts with the avatar, and the evaluation unit 230 may evaluate the user based on evaluation items set according to the usage purpose. The evaluation unit 230 may, for example, input the usage purpose information into an LLM that generates an evaluation prompt, and have it generate an evaluation prompt that includes an evaluation of the usage purpose.

[0047] The evaluation unit 230 may acquire user speech information and avatar speech information using the speech acquisition unit 121. The evaluation items should include at least one of the following based on the user speech information and avatar speech information: customer understanding, accuracy of product knowledge, accuracy of specialized knowledge, and communication ability. The evaluation items may also be based on individual company evaluation criteria (individual company evaluation). Examples of individual company evaluations include whether a mobile phone store can propose advantageous plans (such as family bundle plans) based on customer information, whether a car dealership can propose three or more types of quotes, and whether an insurance store can schedule the next appointment and close the deal.

[0048] The evaluation unit 230 outputs evaluation information to the processing server device 100 as a result of evaluating the user's conversation based on the evaluation prompt (step S120). As a result, the processing server device 100 transmits the evaluation information to the user terminal device 300, and the user terminal device 300 can present the evaluation to the user.

[0049] As described above, the dialogue support system 1 of this embodiment receives persona information indicating a persona that characterizes the user's dialogue partner, creates a persona generation prompt indicating the persona's characteristics using a first large-scale language model based on the persona information, controls an avatar corresponding to the persona based on the persona generation prompt generated by the persona generation unit, and controls the dialogue between the avatar and the user. With this dialogue support system 1, for example, if persona information is input based on the user's operation, a persona generation prompt can be created by the first large-scale language model, so that the variations in the attributes of the dialogue partner for role-playing can be easily increased and role-playing according to the attributes of the dialogue partner can be realized.

[0050] The functions of the processing server device 100, generation server device 200, and user terminal device 300 in the above-described embodiment may be implemented using a computer. In that case, the functions may be implemented by recording a program for implementing these functions on a computer-readable recording medium, loading the program recorded on this recording medium into a computer system, and executing it. Here, "computer system" includes hardware such as an OS and peripheral devices. Furthermore, "computer-readable recording medium" refers to portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and storage devices such as hard disks built into a computer system. Moreover, "computer-readable recording medium" may also include those that dynamically hold programs for a short period of time, such as communication lines used when transmitting programs via networks such as the Internet or communication lines such as telephone lines, and those that hold programs for a certain period of time, such as volatile memory inside a computer system that acts as a server or client in such a case. Furthermore, the above-mentioned program may be for implementing a part of the functions described above, or it may be a program that can implement the above-mentioned functions in combination with a program already recorded in the computer system, or it may be implemented using a programmable logic device such as an FPGA (Field Programmable Gate Array).

[0051] Although various embodiments and variations have been described, these are merely examples and are not limited to these. For example, one embodiment or variation, or a part of one embodiment or variation, may be combined with one or more other embodiments or variations to realize one aspect of the present invention. [Explanation of symbols]

[0052] 1...Dialogue support system, 100...Processing server device, 110...Customer generation unit, 111...Reception unit, 112...Customer definition unit, 120...Dialogue control unit, 121...Utterance acquisition unit, 122...Emotion parameter processing unit, 123...Response prompt generation unit, 124...Response text conversion unit, 125...Conversation history generation unit, 130...Motion control unit, 131...Avatar generation unit, 132...Voice generation unit, 133...Voice tone information processing unit, 134...Motion processing unit, 135 ...Emote processing unit, 136...Lip sync processing unit, 140...Memory unit, 141...Customer information, 142...Response information, 143...Voice information, 144...Motion information, 200...Generation server device, 210...Generation unit, 211...Persona generation unit, 212...Response text generation unit, 213...Emotion parameter generation unit, 214...Unique information acquisition unit, 220...Memory unit, 222...LLM information, 230...Evaluation unit, 240...LLM learning unit, 300...User terminal device

Claims

1. A reception desk that receives persona information, which describes the persona that characterizes the user's conversation partner, A unique information acquisition unit that acquires unique information which is information specific to the dialogue, A persona generation unit inputs persona information and unique information received by the reception unit into a first large-scale language model to generate a persona generation prompt to instruct a second large-scale language model to generate persona features, and generates a persona specification prompt that indicates the persona features using the second large-scale language model based on the generated persona generation prompt. A dialogue control unit generates response text using a third large-scale language model based on the persona specification prompt and unique information generated by the persona generation unit, controls an avatar corresponding to the persona, and controls the interaction between the avatar and the user. A dialogue support system equipped with the following features.

2. The first large-scale language model learns using past data of the persona information, the unique information, and the persona generation prompt as training data, and outputs the persona generation prompt when the persona information and the unique information are input. The second large-scale language model is a machine learning model trained using past data of the persona generation prompt and the persona specification prompt as training data, and outputs the persona specification prompt when the persona generation prompt is input. The third large-scale language model is a machine learning model trained using past data of the persona specification prompt, the unique information, and the response text as training data, and outputs a response text based on the persona specification prompt and the unique information. The dialogue support system according to claim 1.

3. The aforementioned unique information acquisition unit acquires information relating to a specific field as the unique information, The dialogue support system according to claim 1, wherein the persona generation unit inputs the persona information and information relating to a specific field into the first large-scale language model, creates the persona generation prompt corresponding to the specific field based on the output of the first large-scale language model, and generates the persona specification prompt using the second large-scale language model.

4. The unique information acquisition unit acquires items relating to the characteristics of customers in a specific industry and items relating to the characteristics of customers in a specific generation as the unique information. The persona generation unit inputs existing items, including the persona's gender, age, personality, place of origin (including within Japan), way of speaking, tone of voice, or dialect, as persona information into the first large-scale language model, and further inputs at least one of the following into the first large-scale language model: an item specified based on user operation, an item relating to customer characteristics in a specific industry, or an item relating to customer characteristics in a specific generation, and creates a persona specification prompt based on the output of the second large-scale language model. The first large-scale language model learns persona information including existing items such as the persona's gender, age, personality, place of origin (including within Japan), way of speaking, tone of voice, or dialect, items specified based on user operations, items relating to customer characteristics in a specific industry, items relating to customer characteristics in a specific generation, and a persona generation prompt as training data, and outputs a persona generation prompt when at least one of the following is input: persona information including existing items such as the persona's gender, age, personality, place of origin (including within Japan), way of speaking, tone of voice, or dialect, items specified based on user operations, items relating to customer characteristics in a specific industry, and items relating to customer characteristics in a specific generation. This is the dialogue support system according to claim 1.

5. The aforementioned persona information includes nationality or place of origin, The persona generation unit inputs existing items, including the persona's nationality or place of origin, as persona information into the first large-scale language model, and creates a persona instruction prompt based on the output of the second large-scale language model. The dialogue support system according to claim 1, wherein the first large-scale language model learns existing items including the nationality or place of origin of a persona and past data of persona generation prompts as training data, and outputs a persona generation prompt when an existing item including the nationality or place of origin of a persona is input.

6. A voice generation unit generates voice data in a language corresponding to nationality or place of origin based on the persona generation prompt generated by the persona generation unit, A voice tone information processing unit processes the voice data generated by the voice generation unit based on the nationality or place of origin included in the persona information. The dialogue support system according to claim 5, comprising:

7. The dialogue support system according to claim 6, wherein the voice generation unit selects one of several languages ​​based on the nationality or place of origin included in the persona information and generates voice data in the selected language.

8. The aforementioned unique information acquisition unit acquires items related to customer-perspective issues regarding a specific product as the unique information, The persona generation unit inputs items related to customer issues regarding a specific product into the first large-scale language model, and creates a persona generation prompt based on the output of the first large-scale language model. The first large-scale language model is trained using historical data of customer-related issues with a specific product and persona generation prompts as training data, and outputs a persona generation prompt when items related to customer-related issues with a specific product are input. The dialogue support system according to claim 1.

9. A conversation history generation unit that stores the conversation history between the user and the avatar, The system includes an evaluation unit that evaluates the user based on the conversation history stored by the conversation history generation unit using a fourth large-scale language model, The dialogue support system according to claim 1.

10. The fourth large-scale language model is a machine learning model trained using the conversation history and past evaluation data as training data, and outputs an evaluation when the conversation history is input, as described in claim 9 for the dialogue support system.

11. A conversation history generation unit that stores the conversation history between the user and the avatar, The dialogue support system according to claim 1, further comprising: an evaluation unit that outputs an evaluation of the user based on the conversation history stored by the conversation history generation unit and the unique information, using a fourth large-scale language model.

12. The fourth large-scale language model is a machine learning model trained using the conversation history, the unique information, and the past data of the evaluation as training data, and outputs an evaluation when the conversation history is input, as described in claim 11.

13. The reception unit receives information indicating the purpose of use for which the user interacts with the avatar. The dialogue support system according to claim 9 or 11, wherein the evaluation unit evaluates the user based on evaluation items set according to the purpose of use.

14. It includes a speech acquisition unit that acquires user speech information and avatar speech information, The aforementioned evaluation items include at least one of the following: customer understanding based on user speech information and avatar speech information, accuracy of product knowledge, accuracy of specialized knowledge, and communication ability. The dialogue support system according to claim 13.

15. The server device receives persona information that characterizes the user's conversation partner, The server device acquires unique information which is information specific to the interaction. The server device inputs the received persona information and the unique information into a first large-scale language model to generate a persona generation prompt to instruct a second large-scale language model to generate persona features, and generates a persona specification prompt that indicates the persona features using the second large-scale language model based on the generated persona generation prompt. The server device generates response text using a third large-scale language model based on the generated persona specification prompt and the unique information, controls an avatar corresponding to the persona, and controls the interaction between the avatar and the user. A method of supporting dialogue, including...

16. On the server device's computer, A step of receiving persona information that describes the persona that characterizes the user's conversation partner, The steps include obtaining unique information, which is information specific to the dialogue, The steps include: inputting the received persona information and the unique information into a first large-scale language model to generate a persona generation prompt to instruct a second large-scale language model to generate persona features; and generating a persona specification prompt that indicates the persona features using the second large-scale language model based on the generated persona generation prompt; The steps include generating response text using a third large-scale language model based on the generated persona-specific prompt and the unique information, controlling an avatar corresponding to the persona, and controlling the interaction between the avatar and the user, A program that executes something.