Dialogue system and dialogue method
The dialogue system addresses the challenge of declining social interaction by using a language model to generate responsive dialogues tailored to individual user types and histories, improving engagement for the elderly.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- 稲場真由美
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-26
Smart Images

Figure 0007865662000001 
Figure 0007865662000002 
Figure 0007865662000003
Abstract
Description
Technical Field
[0001] The present invention relates to a dialogue system and a dialogue method.
Background Art
[0002] With the increase in the elderly population, the need for supporting elderly people living alone and supporting caregivers in facilities is increasing.
[0003] In particular, it is known that as the opportunities for the elderly to interact with others decrease, their cognitive and swallowing functions are likely to decline. Therefore, it is important to ensure opportunities for the elderly to interact with others.
[0004] For example, Patent Document 1 describes a dialogue system that selects and outputs a response sentence suitable for a user based on the user's attribute information.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] At least one object of the present invention is to provide a dialogue system that enables dialogue using a language model.
Means for Solving the Problems
[0007] According to the present invention, the above object is [1] A dialogue system comprising at least one computer device, comprising: a system utterance output means for outputting system utterances to a user; a user utterance input means for receiving user utterances from a user; a prompt generation means for generating prompts to be input to a language model; and a response sentence acquisition means for acquiring a response sentence generated by inputting the generated prompt to the language model, wherein the prompt generation means generates a prompt that includes at least the input user utterance and the system utterance output immediately before the user utterance, and the system utterance output means outputs the acquired response sentence, or a sentence created based on the response sentence, as the next system utterance; [2] The dialogue system according to [1] above, wherein the prompt generating means generates a prompt requesting the generation of a response to a dialogue history which includes at least an input user utterance and a system utterance output immediately before the user utterance; [3] A dialogue system according to [1] or [2] above, comprising a storage means for storing user utterances and system utterances output in response to said user utterances as a dialogue history, wherein when at least a portion of an input user utterance matches a stored user utterance, a system utterance output means outputs a system utterance stored in association with said user utterance, or a sentence created based on said system utterance, as the next system utterance; [4] A dialogue system according to any one of [1] to [3] above, wherein the prompt generation means generates a prompt by replacing at least a portion of the user utterance with predetermined information and / or adding predetermined information to at least a portion of the user utterance, based on user information relating to the user; [5] A dialogue system according to any one of [1] to [4] above, comprising: type diagnostic means for diagnosing the user type based on user information relating to the user, wherein prompt generation means generates a prompt requesting the user to generate a response sentence containing a predetermined word based on the diagnosed user type; [6] A conversational system as described in [4] or [5] above, in which user information includes the user's name, date of birth, age, gender, blood type, zodiac sign, possessions, place of residence, place of origin, preferences, habits, experiences, social networking services (SNS), family, friends, and / or acquaintances; [7] A dialogue system according to any of [1] to [6] above, wherein the prompt generating means generates a prompt requesting the generation of a response sentence that includes the user's name; [8] A dialogue system according to any of [1] to [7] above, wherein, when the status of obtaining a response by the response acquisition means satisfies predetermined conditions, the system utterance output means outputs a predetermined system utterance; [9] A dialogue method performed in a dialogue system comprising at least one computer device, comprising: a system utterance output step of outputting a system utterance to a user; a user utterance input step of receiving input of a user utterance from a user; a prompt generation step of generating a prompt for input to a language model; and a response sentence acquisition step of acquiring a response sentence generated by inputting the generated prompt to the language model, wherein the prompt generation step generates a prompt that includes at least the input user utterance and a system utterance output immediately before the user utterance, and the system utterance output step outputs the acquired response sentence, or a sentence created based on the response sentence, as the next system utterance; This can be achieved. [Effects of the Invention]
[0008] According to the present invention, it is possible to provide a dialogue system that enables dialogue using a language model. [Brief explanation of the drawing]
[0009] [Figure 1] This is a block diagram showing the configuration of a dialogue system according to an embodiment of the present invention. [Figure 2] This is a block diagram showing the hardware configuration of an assistant terminal according to an embodiment of the present invention. [Figure 3] This is a block diagram showing the hardware configuration of a server device according to an embodiment of the present invention. [Figure 4] This is a flowchart of the user registration process according to an embodiment of the present invention. [Figure 5] This is a flowchart of the dialogue processing according to an embodiment of the present invention. [Figure 6] This is a flowchart of the response text generation process according to an embodiment of the present invention. [Modes for carrying out the invention]
[0010] The following describes embodiments of the present invention, but the present invention is not limited to the following embodiments unless it contradicts the spirit of the invention. The order of each process constituting the flowchart described below is not limited to any order that does not cause contradictions or inconsistencies in the processing content, and it is also possible to omit some of the processes constituting the flowchart or to add new processes to each process constituting the flowchart, as long as it does not contradict the spirit of the invention. Furthermore, the device that is the main entity that executes each process constituting the flowchart can be changed to another device, as long as it does not contradict the spirit of the invention. In this case, it is possible to change the processing content so as not to cause contradictions or inconsistencies in the processing content.
[0011] [Dialogue System Configuration] Figure 1 is a block diagram showing the configuration of a dialogue system according to an embodiment of the present invention. The dialogue system shown in Figure 1 comprises a dialogue device 1, an assistant terminal 2, and a server device 3. The dialogue device 1, the assistant terminal 2, and the server device 3 are connected to each other so as to be able to communicate with one another via a communication network 5. In addition, the dialogue device 1, the assistant terminal 2, and the server device 3 are connected to a language model server 4 so as to be able to communicate with one another via the communication network 5. Note that the dialogue device 1, the assistant terminal 2, the server device 3, and the language model server 4 do not need to be constantly connected; they only need to be able to connect as needed.
[0012] In a dialogue system, any one of the dialogue device 1, the assistant terminal 2, the server device 3, and the language model server 4 can function as an information processing device. When any one of the dialogue device 1, the assistant terminal 2, the server device 3, and the language model server 4 functions as an information processing device, information transmission and reception are executed between the dialogue device 1, the assistant terminal 2, the server device 3, and the language model server 4 as necessary.
[0013] The dialogue device 1 is a device for conducting a dialogue with a user. The dialogue device 1 may be an object that can be the target of a user's address and can speak to the user. The dialogue device 1 may include, for example, an audio input unit such as a microphone, an audio output unit such as a speaker, a control unit, a storage unit, a communication interface, etc. The dialogue device 1 can be provided with any other configuration.
[0014] The shape of the dialogue device 1 is not particularly limited and can be designed as appropriate. The shape of the dialogue device 1 may, for example, mimic an animal such as a person or a dog, a plant such as grass or a flower, an object such as a snowman or a scarecrow, etc. Alternatively, the shape of the dialogue device 1 may be a geometric shape such as a sphere or a cylinder. That is, the dialogue device 1 may have an appearance like a robot or an appearance like a smart speaker. From the perspective that it is easy for the user to feel attachment, the appearance of the dialogue device 1 may be like a robot.
[0015] Also, the dialogue device 1 may be a desktop / laptop personal computer, a tablet terminal, a smartphone, a conventional mobile phone, etc.
[0016] The user who conducts a dialogue with the dialogue device 1 is not particularly limited and can be designed as appropriate. The user may be, for example, an elderly person, a child, or an adult who is not elderly.
[0017] The assistant terminal 2 is a terminal operated by an assistant who assists the user. The "assistant" includes, for example, the user's family members, care managers, caregivers, and the staff of the facility where the user is staying. By inputting information about the user (hereinafter also referred to as user information) into the assistant terminal 2, user registration described later can be performed.
[0018] FIG. 2 is a block diagram showing the hardware configuration of an assistant terminal according to an embodiment of the present invention. The assistant terminal 2 includes a control unit 21, a RAM 22, a storage unit 23, an input unit 24, a display unit 25, and a communication interface 26, which are connected by a bus respectively.
[0019] The control unit 21 is composed of a CPU and a ROM. The control unit 21 executes the program stored in the storage unit 23 to control the assistant terminal 2. The RAM 22 is a work area of the control unit 21. The storage unit 23 is a storage area for storing programs and data. That is, the storage unit 23 functions as a recording medium storing the program. The control unit 21 performs arithmetic processing based on the program and data read from the RAM 22, and the data input by the input unit 24.
[0020] The display unit 25 has a display screen. The control unit 21 outputs a video signal for displaying an image on the display screen according to the result of the arithmetic processing. Here, the display screen of the display unit 25 may be a touch panel equipped with a touch sensor. In this case, the touch panel functions as the input unit 24.
[0021] The communication interface 26 can be connected to the communication network 5 wirelessly or by wire, and can transmit and receive data with other computer devices via the communication network 5. The data received via the communication interface 26 is loaded into the RAM 22, and arithmetic processing is performed by the control unit 21.
[0022] If the user does not have an assistant, the dialogue system may provide a user terminal for the user to operate instead of the assistant terminal 2. The hardware configuration of the user terminal can be adapted to the extent necessary from the description of the hardware configuration of the assistant terminal 2 above.
[0023] Server device 3 may be managed by the administrator of the dialogue system.
[0024] Figure 3 is a block diagram showing the hardware configuration of a server device according to an embodiment of the present invention. The server device 3 comprises at least a control unit 31, RAM 32, storage unit 33, and communication interface 34, each connected by an internal bus.
[0025] The control unit 31 consists of a CPU and ROM, and executes programs stored in the storage unit 33 to control the server device 3. The control unit 31 also has an internal timer for timing. The RAM 32 is the work area of the control unit 31. The storage unit 33 is a memory area for saving programs and data. In other words, the storage unit 33 functions as a recording medium that stores programs. The control unit 31 reads programs and data from the RAM 32 and performs program execution processing based on information received from the dialogue device 1, the assistant terminal 2, and the language model server 4.
[0026] Furthermore, the program may be stored on a recording medium such as a CD-ROM. In this case, the program stored on the recording medium may be installed on the interactive device 1, the assistant terminal 2, and / or the server device 3 to perform predetermined functions.
[0027] Alternatively, the program may be delivered from a computer device outside the system. In this case, the program delivered from the computer device outside the system may be installed on the interactive device 1, the assistant terminal 2, and / or the server device 3 to perform predetermined functions.
[0028] Language model server 4 is a server that uses a language model to perform natural language processing tasks such as text generation, question and answer, and text summarization. By inputting instructional texts called prompts to the language model, the language model outputs a response text. Language model server 4 may be managed by someone other than the administrator of the dialogue system. Furthermore, server device 3 and language model server 4 may be configured to communicate using an API (Application Programming Interface).
[0029] The type of language model is not particularly limited, as long as it can process natural language. The language model may be large-scale or small-scale. Examples of Transformer-based language models include GPT (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representations from Transformers).
[0030] The language model server 4 may include a computer device that stores language models for performing natural language processing. The hardware configuration of the language model server 4 can be adapted from the description of the hardware configuration of server device 3 to the extent necessary.
[0031] Furthermore, an artificial intelligence equipped with a language model may be a specialized artificial intelligence focused on handling a specific type of task, or it may be a general-purpose artificial intelligence capable of handling multiple types of tasks. General-purpose artificial intelligence includes AGI (Artificial General Intelligence) and ASI (Artificial Superintelligence). In other words, a language model is a concept that includes an artificial intelligence and / or server that has the function of outputting answers to questions and / or instructions.
[0032] Furthermore, although this description focuses on a configuration in which the language model server 4 outputs response sentences from the language model, the server device 3, assistant terminal 2, or dialogue device 1 may be configured to store the language model, and the server device 3, assistant terminal 2, or dialogue device 1 may be configured to output response sentences from the language model.
[0033] Furthermore, the dialogue system may include multiple dialogue devices 1. In this case, it is preferable that each dialogue device 1 is assigned identification information (e.g., a dialogue device ID) to identify it. It is also preferable that users who use the dialogue devices 1 are assigned identification information (e.g., a user ID) to identify them. Here, we assume that each dialogue device 1 is assigned a dialogue device ID, and each user is assigned a user ID. The server device 3 can store the user ID and the dialogue device ID in association. The server device 3 can also store information acquired by each dialogue device 1, information about the user of each dialogue device 1, etc., in association with the user ID and the dialogue device ID.
[0034] Furthermore, the dialogue system may include multiple assistant terminals 2. Each assistant may be assigned identification information (e.g., an assistant ID) to identify them. Here, we assume that each assistant is assigned an assistant ID. The server device 3 can store user IDs and the assistant IDs of assistants assisting the user corresponding to that user ID in association. The number of assistant IDs stored in association with a single user ID may be one or two or more. Assistants may be able to register, edit, view, etc., information about users corresponding to user IDs stored in association with their own assistant ID.
[0035] Furthermore, server device 3 may function in a distributed manner across multiple computer devices. For example, instead of server device 3, distributed ledger technology such as blockchain may be used.
[0036] Furthermore, the number and types of computer devices included in the dialogue system are not particularly limited and can be designed as appropriate. The dialogue system only needs to include at least one computer device. For example, the dialogue system may consist of dialogue device 1, or it may consist of dialogue device 1 and server device 3.
[0037] [User registration process] First, the assistant registers a user in the dialogue system. User registration may involve storing information about the user in the server device 3 provided by the dialogue system. For example, the assistant may enter an assistant ID and a user ID on the login screen for logging into the dialogue system, thereby displaying a user management screen on the display screen of the assistant terminal 2 for managing the user corresponding to the user ID. Alternatively, the assistant may display a user registration screen on the assistant terminal 2 via the user management screen. Figure 4 is a flowchart of the user registration process according to an embodiment of the present invention.
[0038] The assistant terminal 2 accepts user information input (step S101). The assistant terminal 2 transmits the input user information to the server device 3 (step S102). The server device 3 receives the transmitted user information (step S103). The server device 3 diagnoses the user type based on the received user information (step S104). The server device 3 stores the user information, including the diagnosed user type (step S105), and the user registration process ends.
[0039] The user information entered in step S101 may include the user's name, date of birth, age, gender, blood type, zodiac sign, possessions, place of residence, place of origin, preferences, habits, experiences, SNS (Social Networking Services), and information about family, friends, and acquaintances. Information about the user's name may include the user's real name and nickname. Information about the user's possessions may include the name of the dialogue device 1 used by the user and the make and model of the car owned by the user. Information about the user's preferences may include the user's favorite foods and hobbies. Information about the user's habits may include the supermarkets the user frequently visits and the user's daily routines. Information about the user's experiences may include destinations the user has traveled to and the user's past occupations.
[0040] Furthermore, the method for inputting user information in step S101 is not particularly limited and can be designed as appropriate. For example, an assistant may input user information by typing the specified user information into the designated fields on the user registration screen. Alternatively, the user information stored on the SNS server in association with the user's SNS account may be input by referring to the user's SNS account information.
[0041] The user information may be entered on the user registration screen, and the user information transmission in step S102 may be executed when the registration button or other button displayed on the user registration screen is pressed.
[0042] The method for diagnosing the user type in step S104 is not particularly limited and can be designed as appropriate. Here, "type" refers to a classification of people based on some criteria, from which common characteristics are extracted. For example, users could be classified as "vision types" with high insight, "peace types" with pacifism, or "logical types" with logical reasoning.
[0043] The server device 3 can diagnose the user type, that is, identify the user type, based on the user information received in step S103. For example, the server device 3 may diagnose the user type based on the user's date of birth, the user's blood type, or the user's zodiac sign.
[0044] In step S105, the server device 3 can store the type diagnosed in step S104 and the user information received in step S103 in the storage unit 33, associating them with the user ID. The type diagnosed in step S104 is related to the user, and therefore can be said to be user information.
[0045] Furthermore, the assistant may display a reminder registration screen on assistant terminal 2 via the user management screen, allowing the assistant to register information that should be reminded to the user. The assistant can register a reminder by entering the time of the reminder, the content of the reminder, etc., on the reminder registration screen. The content of the reminder may be text corresponding to the actual voice output, or it may be information that identifies the text corresponding to the voice output. Hereinafter, "text corresponding to the voice output" will also be referred to as "output text". An example of actual output text is "Good morning. Are you awake?". An example of information that identifies the output text is "Morning reminder".
[0046] When information that can identify the text to be output is entered as part of the reminder, the dialogue system may output predetermined text corresponding to that information, or it may output text generated by the language model that corresponds to that information. Note that "outputting speech corresponding to text" is also referred to as "outputting text." When outputting text generated by the language model, for example, when the time to be reminded arrives, the server device 3 generates a prompt requesting the generation of a response sentence for the reminder and sends it to the language model server 4, thereby obtaining the response sentence generated by the language model. The dialogue device 1 can output the obtained response sentence, or a sentence created based on the response sentence, as a reminder. The method of creating a sentence based on the response sentence can be adopted to the extent necessary, as described later. The prompt requesting the generation of a response sentence for the reminder may include information that can identify the text to be output. For example, the server device 3 may generate a prompt such as, "Please generate a sentence suitable for a morning reminder in a friendly conversational tone."
[0047] Furthermore, the assistant may configure the system to include a predetermined word (keyword) within the text of a predetermined reminder when registering a reminder. For example, the assistant may configure the system to include the word "jogging" in the text of a health management reminder and the word "recipe" in the text of a hobby reminder. The predetermined word can be entered and registered by the assistant. The registered word may be included in the predetermined reminder text with a predetermined probability. Server device 3 can generate a prompt requesting the generation of a predetermined reminder response text that includes the registered predetermined word with a predetermined probability. By configuring the system to include a predetermined word within the text of a predetermined reminder, it is possible to provide reminders tailored to the user.
[0048] By registering a reminder, the dialogue device 1 can output a predetermined message at a predetermined time. For example, using the reminder function, the dialogue device 1 can output to the user at 7:00 AM, "Good morning. Are you awake?"
[0049] [Dialogue Processing] Next, we will describe how the dialogue system interacts with the user and the dialogue device 1. Here, the voice output from the dialogue device 1 to the user is called system utterance. System utterance may be a verbal address to the user, or it may be a response to an utterance made by the user to the dialogue device 1. The voice input from the user to the dialogue device 1 is called user utterance. Dialogue takes place through the alternating output of system utterances and input of user utterances. The "original text output as voice in the form of system utterance" is also called the "text corresponding to the system utterance" or simply the "system utterance." In other words, the content of the system utterance itself is also called system utterance. The "text generated from the voice input as user utterance" is also called the "text corresponding to the user utterance" or simply the "user utterance." In other words, the content of the user utterance itself is also called user utterance. Figure 5 is a flowchart showing the dialogue processing according to an embodiment of the present invention.
[0050] First, the server device 3 identifies the text corresponding to the system utterance to be output to the user (step S201). The server device 3 stores the text corresponding to the identified system utterance (step S202) and transmits it to the dialogue device 1 (step S203). The dialogue device 1 receives the transmitted text corresponding to the system utterance (step S204) and outputs it as speech (step S205).
[0051] The system utterance identified in step S201 may be identified based on the content of the reminder registration described above. In other words, the system utterance in step S201 may be output as a reminder.
[0052] In step S202, the server device 3 may store, in association with the system utterance, the time (date and time) when the system utterance was identified in step S201, or the time (date and time) when the system utterance was stored in step S202. Alternatively, after the dialogue device 1 outputs a system utterance in step S205, it may transmit the time (date and time) when the system utterance was output to the server device 3, and the server device 3 may store, in association with the system utterance, the time (date and time) when the system utterance was output. The time (date and time) when the system utterance was identified, the time (date and time) when the system utterance was stored, or the time (date and time) when the system utterance was output is also referred to as the "time the system utterance was made." In other words, the server device 3 may store, in association with the system utterance, the time when the system utterance was made. The server device 3 may also store, in association with the system utterance, the dialogue device ID of the dialogue device 1 that outputs the system utterance.
[0053] In step S203, the server device 3 transmits the system utterance identified in step S201 and stored in step S202 to the dialogue device 1 used by the user.
[0054] Upon hearing the system utterance output in step S205, the user speaks to the dialogue device 1 in order to respond to the system utterance. It is preferable that the system utterance output by the reminder function is easy for the user to respond to.
[0055] The dialogue device 1 receives user utterances from the user (step S206). The dialogue device 1 sends text corresponding to the input user utterances to the server device 3 (step S207). The server device 3 receives the text corresponding to the transmitted user utterances (step S208) and stores it (step S209). The server device 3 searches the dialogue history stored in the cache memory to see if there is text corresponding to user utterances that matches the text corresponding to the received user utterances (step S210). The server device 3 determines whether or not to use the dialogue history stored in the cache memory (step S211).
[0056] In step S206, the dialogue device 1 receives voice input from the user speaking to the dialogue device 1 via a microphone or the like. The dialogue device 1 may convert the user utterance input by voice into text.
[0057] In step S207, the user utterance converted to text may be sent to the server device 3.
[0058] In step S209, the server device 3 may store, in association with the user utterance, the time (date and time) when the user utterance was received in step S206, the time (date and time) when the user utterance was received in step S208, or the time (date and time) when the user utterance was stored in step S209. Hereinafter, the time (date and time) when the user utterance was received, the time (date and time) when the user utterance was received, or the time (date and time) when the user utterance was stored will also be referred to as the "time the user utterance was made." The server device 3 may also store, in association with the user utterance, the dialogue device ID of the dialogue device 1 that received the input of the user utterance. Furthermore, the server device 3 may store the user utterance in association with the system utterance stored in step S202.
[0059] Server device 3 can store system utterances output during a dialogue and user utterances input during a dialogue as a dialogue history. In the dialogue history, it is preferable that system utterances and user utterances constituting a single dialogue are stored in relation to each other. For example, it is preferable that a system utterance and the user utterance input in response to that system utterance are stored in relation to each other. Also, for example, it is preferable that a user utterance and the system utterance output in response to that user utterance are stored in relation to each other. Furthermore, in the dialogue history, it is preferable that system utterances and user utterances constituting a single dialogue are stored in a way that allows for the identification of their chronological order. For example, in the dialogue history, system utterances and user utterances may be stored together with the time (date and time) in which they were uttered, or together with the order in which they were uttered. Furthermore, in the dialogue history, it is preferable that the chronological order of each dialogue is also stored in a way that allows for the identification of each dialogue. The dialogue history may include multiple dialogues.
[0060] The definition of what constitutes a single dialogue, that is, the start and end conditions of a dialogue, are not particularly limited and can be designed as appropriate. For example, the end condition of a dialogue may be that no user utterance is input within a predetermined time after a system utterance is output, or it may be that a predetermined system utterance that satisfies the end condition (e.g., "Okay, see you later.") is output. Also, for example, the start condition of a dialogue may be that a system utterance is output after the dialogue end condition is met, or that a user utterance is input after the dialogue end condition is met.
[0061] Server device 3 can store the dialogue history in main memory and / or cache memory. Server device 3 may store a predetermined number of days' worth and / or a predetermined capacity of dialogue history in the cache memory. For example, server device 3 may delete from cache memory dialogue history that has been stored for a predetermined number of days and / or dialogue history whose storage capacity exceeds a predetermined capacity.
[0062] In step S210, the server device 3 searches the dialogue history stored in the cache memory for user utterances that match the received user utterance. Here, it is sufficient to search whether at least a portion of the received user utterance matches at least a portion of the user utterances in the dialogue history stored in the cache memory. In other words, it is sufficient to search whether the received user utterance and the user utterances in the cache memory are partially the same.
[0063] In step S211, the conditions for determining whether or not to use the dialogue history stored in the cache memory are not particularly limited and can be designed as appropriate.
[0064] For example, in the search in step S210, if the received user utterance and the user utterance in the cache memory partially match, it may be decided to use the dialogue history stored in the cache memory.
[0065] Alternatively, for example, in the search of step S210, if the received user utterance and the user utterance in the cache memory partially match, and the time at which the matching user utterance was made satisfies a predetermined temporal condition, the system may decide to use the dialogue history stored in the cache memory. The predetermined temporal condition is not particularly limited and can be designed as appropriate. For example, from the viewpoint of preventing the user from becoming bored with the same response, the predetermined temporal condition may be earlier than the time of the decision. Specifically, for example, if the dialogue history is from two days or more ago than the time of the decision, the system may decide to use it.
[0066] Alternatively, for example, in the search in step S210, if the received user utterance and the user utterance in the cache memory partially match, and the system utterance made immediately before the partially matching user utterance in the cache memory partially match the system utterance made immediately before the received user utterance, then it may be determined to use the dialogue history stored in the cache memory. If both the user utterance and the system utterance immediately preceding the user utterance match, it can be assumed that the flow of the dialogue is the same.
[0067] In step S211, if it is determined to use the dialogue history stored in the cache memory (YES in step S211), the server device 3 uses the dialogue history to identify the text corresponding to the system utterance (step S212).
[0068] In step S212, using the dialogue history stored in the cache memory may mean using a system utterance previously output in response to a matching user utterance in the dialogue history for the next system utterance. System utterances previously output in response to a user utterance are stored in the dialogue history in association with that user utterance.
[0069] In step S212, the server device 3 may, for example, identify the text of a system utterance previously output in response to a matching user utterance as the next system utterance, or it may identify a sentence (text) created based on a system utterance previously output in response to a matching user utterance as the next system utterance. "A sentence created based on a system utterance previously output" refers to a sentence that has been modified from a system utterance previously output, and may be a sentence to which predetermined information has been added, a sentence to which predetermined information has been deleted, or a sentence to which the expression has been changed. The server device 3 can create a sentence (text) based on a system utterance previously output.
[0070] On the other hand, if it is not determined in step S211 to use the dialogue history stored in the cache memory (NO in step S211), the response text generation process is executed in the server device 3 and the language model server 4 (step S213). The language model server 4 sends the response text generated in the response text generation process to the server device 3 (step S214). The server device 3 receives the sent response text (step S215). Based on the received response text, the server device 3 identifies the text corresponding to the system utterance (step S212).
[0071] The process for generating the response in step S213 will be described later.
[0072] In step S212, the server device 3 may identify the response received in step S215 as the next system utterance, or it may identify a sentence (text) created based on the received response as the next system utterance. "Sentence created based on the response" refers to a sentence that has been modified from the response, and may be a sentence to which predetermined information has been added, a sentence to which predetermined information has been deleted, or a sentence to which the expression has been changed. The server device 3 can create a sentence (text) based on the response.
[0073] Server device 3 stores the text corresponding to the system utterance identified in step S212 (step S216) and transmits it to dialogue device 1 (step S217). Dialogue device 1 receives the transmitted text corresponding to the system utterance (step S218) and outputs it as speech (step S219).
[0074] If the user responds to the system utterance output in step S219, step S206 is executed again, and the dialogue device 1 accepts user utterance input. Steps S206 to S219 are repeated until the dialogue ends. The conditions for ending the dialogue can be those described above to the extent necessary. When the conditions for ending the dialogue are met, the dialogue process ends.
[0075] For the processing in steps S216 to S219, the descriptions of the processing in steps S202 to S205 can be adopted to the extent necessary. In step S216, the system utterance is stored in the server device 3 as dialogue history.
[0076] [Response text generation process] Next, the response text generation process in step S213 described above will be explained. Figure 6 is a flowchart of the response text generation process according to an embodiment of the present invention.
[0077] Server device 3 generates a prompt for input to the language model (step S301) and sends it to language model server 4 (step S302). Language model server 4 receives the transmitted prompt (step S303). Language model server 4 inputs the received prompt into the language model (step S304) and generates a response sentence (step S305). Steps S301 to S305 complete the response sentence generation process.
[0078] In step S301, the server device 3 generates a prompt that includes at least the text corresponding to the input user utterance and the text corresponding to the system utterance output immediately before the user utterance. The "input user utterance" refers to the user utterance that was input immediately before the prompt was generated. The server device 3 may also generate a prompt that requests the generation of a response to a dialogue history that includes at least the input user utterance and the system utterance output immediately before the user utterance.
[0079] The input user utterance refers to the user utterance input in step S206, and the system utterance output immediately before the user utterance may refer to the system utterance output in step S205. For example, if the system utterance in step S205 is "Good morning. Are you awake?" and the user utterance in step S206 is "Yes. I've already had breakfast.", then the server device 3 may generate a prompt that says, "Generate a response in a friendly conversational tone following 'Good morning. Are you awake?' and 'Yes. I've already had breakfast.'"
[0080] Furthermore, if a user utterance is input again in step S206 after a system utterance has been output in step S219, the input user utterance refers to the user utterance input again in step S206, and the system utterance output immediately before that user utterance refers to the system utterance output in step S219. For example, if the system utterance in step S219 is "Wow! What did you eat?" and the user utterance in step S206 in response is "Bread today.", then the server device 3 may generate a prompt that reads, "Generate a response in a friendly conversational tone following 'Wow! What did you eat?' and 'Bread today.'"
[0081] The server device 3 only needs to generate a prompt that includes at least the input user utterance and the system utterance output immediately before the user utterance, and may also generate a prompt that includes the input user utterance and the dialogue history prior to the system utterance output immediately before the user utterance. For example, in the above example, the server device 3 may generate a prompt that says, "Please generate a response in a friendly conversational tone following 'Good morning. Are you awake?', 'Yes. I've already had breakfast.', 'Wow! What did you eat?', 'Bread today.'"
[0082] Furthermore, the user utterances and system utterances included in the prompt may not be the entire text corresponding to each utterance, but rather only a portion of the text corresponding to each utterance. For example, in the above example, server device 3 may generate the prompt, "Generate a response in a friendly conversational tone following 'What did you eat?', 'Bread.'"
[0083] By generating prompts that include not only the input user utterance but also the system utterance output immediately preceding that user utterance, it is possible to obtain appropriate response sentences that correspond to the flow of the dialogue.
[0084] Furthermore, by controlling the system to generate prompts that include only a portion of the dialogue history, such as the input user utterance and the system utterance output immediately preceding the user utterance, rather than prompts that include the entire dialogue history, the processing load on the server device 3 and / or the language model server 4 can be reduced, and the time required to generate prompts and / or response sentences can be shortened.
[0085] The number of dialogue history entries to be included in a prompt is not particularly limited and can be designed as appropriate. Server device 3 may generate a prompt that includes several system utterances and user utterances immediately preceding the input user utterance. Alternatively, the dialogue history entries to be included in a prompt may be limited by indicators other than the number of entries. For example, server device 3 may generate a prompt that includes a predetermined number of dialogue history entries, at least including the input user utterance and the system utterance output immediately preceding the user utterance. By limiting the dialogue history entries to be included in a prompt based on predetermined criteria, server device 3 can reduce the processing load while obtaining appropriate response sentences that correspond to the flow of the conversation.
[0086] Furthermore, in step S301, the server device 3 may generate a prompt by replacing at least a portion of the text corresponding to the user utterance with predetermined information and / or by adding predetermined information to at least a portion of the text corresponding to the user utterance, based on the registered user information.
[0087] In other words, server device 3 may introduce so-called variables into the prompt. For example, if the user utterance includes the name of a supermarket that the user frequently visits, server device 3 may replace that name with a variable called {supermarket that the user frequently visits} and generate the prompt. Alternatively, server device 3 may add information about the specific content of the variable to the user utterance, such as "supermarket that the user frequently visits:○○○", and generate the prompt.
[0088] Furthermore, the server device 3 may generate a prompt by adding predetermined information to at least a portion of the user utterance without introducing variables into the prompt. For example, if a word included in the user utterance is included in the user information, information about what that word means to the user may be added. Specifically, for example, the server device 3 may generate a prompt by adding information such as, "○○○ is the name of a supermarket I often go to."
[0089] When the server device 3 stores user utterances and / or system utterances, it may store the text corresponding to the user utterances and / or system utterances as is, or it may store the text with at least a portion of it replaced by variables.
[0090] Furthermore, in step S301, the server device 3 may generate a prompt requesting the generation of a response sentence that includes the user's name. For example, the server device 3 can generate a prompt that causes the user to generate a response sentence that includes the user's name with a predetermined probability.
[0091] Furthermore, in step S301, the server device 3 may generate a prompt requesting the user to generate a response sentence containing a predetermined word, based on the registered user type. For example, the server device 3 can generate prompts that include different words in the response sentence with a predetermined probability for each user type. The type of word is not particularly limited and can be designed as appropriate. The type of word may be a compliment or an exclamation mark.
[0092] Specifically, for example, as described above, if users are diagnosed and registered as either the insightful "Vision Type," the pacifist "Peace Type," or the logical "Logical Type," prompts may be generated to include, with a predetermined probability, one of the following words in the response for "Vision Type" users: "Amazing," "Surprise," "Number One," and "Best"; one of the following words for "Peace Type" users: "Thank You," "Happy," "Kind," and "Reassuring"; and one of the following words for "Logical Type" users: "Impressive," "Interesting," "Wonderful," and "Sincere."
[0093] By generating response sentences that include different words for each user type, it becomes possible to obtain response sentences that are tailored to the user's type.
[0094] Furthermore, in step S301, the server device 3 may generate a prompt requesting the user to generate a response sentence with a predetermined number of characters or less. By limiting the number of characters in the response sentence, the processing load can be reduced while enabling a smooth and efficient dialogue.
[0095] In step S305, the language model server 4 generates a response sentence corresponding to the prompt entered in step S304.
[0096] The response text generated in step S305 is sent to the server device 3 in step S214 of the dialogue processing. Upon receiving the response text, the dialogue system can obtain the response text.
[0097] If a language model is stored in the server device 3, the server device 3 may input the prompt generated in step S301 into the language model to obtain the response text output from the language model.
[0098] Although not shown in the diagram, if the acquisition status of the response text meets certain conditions, the server device 3 may identify a predetermined system utterance in step S212. The dialogue device 1 can then output the identified predetermined system utterance.
[0099] The predetermined conditions for obtaining the response are not particularly limited and can be designed as appropriate. For example, the predetermined conditions for obtaining the response may be that, after sending a prompt in step S302, the response is not received from the language model server 4 even after a predetermined time has elapsed, or the response received in step S215 may contain an error.
[0100] Furthermore, the content of the predetermined system utterance is not particularly limited and can be designed as appropriate. The predetermined system utterance may be, for example, "Yes, understood."
[0101] If the acquisition of response text is unsuccessful, outputting a predetermined system utterance allows the dialogue to continue even if the language model server 4 is not functioning properly.
[0102] In addition, while the above describes a scenario in which the dialogue begins with a system utterance, if the dialogue begins with a user utterance, the server device 3 may treat the system utterance immediately preceding the input user utterance as blank and generate a prompt.
[0103] Furthermore, it is possible to use the dialogue system of the present invention to engage in dialogue with users who have not registered. The server device 3 can generate a prompt that includes at least the input user utterance and the system utterance output immediately before the user utterance, even in response to an utterance from a user who has not registered.
[0104] The number of users interacting with one dialogue device 1 may be one or more. If there are two or more users interacting with one dialogue device 1, the dialogue system may identify the user who made the input user utterance. The dialogue system may identify the user, for example, by an image of the user's face or by information about the user's voice. The image of the user's face may be acquired by a camera provided in the dialogue device 1.
[0105] If multiple users interacting with a single dialogue device 1 are registered users, then user information may be registered to identify the users. This user identification information may be an image of the user's face or information about the user's voice. In this case, the server device 3 can identify the user ID of the user who made the user utterance.
[0106] In addition, while the above describes how the dialogue system can be used in real space, the dialogue system of the present invention may also be used for dialogue in virtual space. That is, the dialogue system may be used when a user interacts with characters or objects existing in virtual space. In this case, the user may interact with characters or objects existing in virtual space through their own avatar existing in virtual space, or they may interact with characters or objects existing in virtual space from real space through a screen that displays virtual space. Characters and objects existing in virtual space may be automatically controlled by a computer device. The dialogue system can be used by treating utterances from characters or objects existing in virtual space as system utterances.
[0107] Thus, a dialogue system can be provided that enables dialogue using a language model by comprising at least one computer device, a system utterance output means for outputting system utterances to a user, a user utterance input means for receiving user utterances from the user, a prompt generation means for generating prompts to be input to a language model, and a response sentence acquisition means for acquiring a response sentence generated by inputting the generated prompt to the language model, wherein the prompt generation means generates a prompt that includes at least the input user utterance and the system utterance output immediately before the user utterance, and the system utterance output means outputs the acquired response sentence, or a sentence created based on the response sentence, as the next system utterance.
[0108] Furthermore, by generating a prompt that includes at least the input user utterance and the system utterance output immediately preceding the user utterance, it is possible to obtain a natural response sentence that corresponds to the flow of the dialogue, and to reduce the processing load on the dialogue system.
[0109] Furthermore, by having the prompt generation means generate a prompt that requests the generation of a response to a dialogue history that includes at least the input user utterance and the system utterance output immediately preceding the user utterance, the language model can obtain a response to the dialogue history.
[0110] Furthermore, the dialogue system includes a storage means that stores user utterances and system utterances output in response to those user utterances as a dialogue history, and when at least a portion of the input user utterance matches a stored user utterance, the system utterance output means outputs the system utterance stored in association with the user utterance, or a sentence created based on the system utterance, as the next system utterance, thereby enabling output using the dialogue history.
[0111] Furthermore, the prompt generation means generates a prompt by replacing at least a portion of the user utterance with predetermined information and / or adding predetermined information to at least a portion of the user utterance, based on user information about the user, thereby obtaining a response sentence that corresponds to the user information.
[0112] Furthermore, the dialogue system includes a type diagnosis means that diagnoses the user's type based on user information about the user, and a prompt generation means that generates a prompt requesting the generation of a response sentence containing predetermined words based on the diagnosed user type, thereby obtaining a response sentence appropriate to the user's type.
[0113] Furthermore, because user information includes the user's name, date of birth, age, gender, blood type, zodiac sign, possessions, place of residence, place of origin, preferences, habits, experiences, social networking services (SNS), family, friends, and / or acquaintances, it is possible to obtain responses that correspond to the user's name and other information.
[0114] Furthermore, by having the prompt generation means generate a prompt requesting the generation of a response that includes the user's name, a response containing the user's name can be obtained.
[0115] Furthermore, in this manner, when the response acquisition status of the response text by the response text acquisition means satisfies predetermined conditions, the system utterance output means outputs a predetermined system utterance, thereby enabling predetermined output when the predetermined conditions for the response text acquisition status are met. [Explanation of Symbols]
[0116] 1. Dialogue device 2. Assistant terminal 3 Server equipment 4. Language Model Server 5. Communication Network 21 Control Unit 22 RAM 23 Storage Section 24 Input section 25 Display section 26 Communication Interfaces 31 Control Unit 32 RAM 33 Storage Section 34 Communication Interfaces
Claims
1. A dialogue system comprising at least one computer device, A system utterance output means that outputs system utterances to the user, A user speech input means that accepts user speech input from the user, A prompt generation means for generating prompts to be input to a language model, A means for obtaining a response sentence by inputting the generated prompt into a language model, A storage means for storing user information about a user and Equipped with, The prompt generation means generates a prompt that includes at least the input user utterance, and if the words included in the input user utterance match user information stored in the storage means, adds information about what meaning the words have to the user. The system utterance output means outputs the acquired response sentence, or a sentence created based on said response sentence, as the next system utterance. Dialogue system.
2. The prompt generation means generates a prompt that requests the generation of a system utterance as a response to the input user utterance. The dialogue system according to claim 1.
3. The prompt generation means generates a prompt that requests the user to generate a response sentence that includes the user's name. The dialogue system according to claim 1 or 2.
4. The dialogue system according to claim 1, wherein user information is information about the user's possessions, information about the user's preferences, information about the user's habits, or information about the user's experiences.
5. A dialogue method performed in a dialogue system comprising at least one computer device, A system utterance output step that outputs system utterances to the user, A user utterance input step that accepts user utterances from the user, A prompt generation step that generates prompts for input to the language model, The response text acquisition step involves obtaining the generated response text by inputting the generated prompt into the language model, and A storage step for storing user information about a user and It has, The prompt generation step generates a prompt that includes at least the input user utterance, and if the words in the input user utterance match user information stored in the storage step, adds information about what the words mean to the user. In the system utterance output step, the acquired response sentence, or a sentence created based on said response sentence, is output as the next system utterance. Methods of dialogue.