Dialogue system, dialogue method, dialogue program and dialogue device
The dialogue system addresses the lack of personalization in conventional systems by extracting user background and behavioral information to deliver natural and personalized interactions, enhancing user intimacy and functional support through device integration.
Patent Information
- Application Number
- JP2024193384
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-11-05
AI Technical Summary
Conventional dialogue systems fail to realize natural conversations based on a user's personal background and status, and lack the ability to provide functions according to the user's intentions.
A dialogue system comprising a dialogue unit, extraction unit, and memory unit that acquires user utterances, extracts background information and behavioral purposes, and registers them in a memory unit to output personalized responses, and further integrates with devices to execute actions based on user inputs.
Enables natural and personalized dialogue, allowing the system to feel like it 'knows' the user, provides tailored device functions, and supports user actions through integrated device interactions.
Smart Images

Figure 0007761307000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a dialogue system, a dialogue method, a dialogue program, and a dialogue device. [Background technology]
[0002] Conversational robots using artificial intelligence (AI) are widely known as conventional technologies for dialogue systems and dialogue methods. These systems utilize speech recognition technology and natural language processing technology to generate appropriate responses based on voice input from the user. Many conventional dialogue systems have been used to advance dialogue according to pre-set scenarios and provide responses to specific questions. These technologies are useful in situations such as customer support and information guidance.
[0003] For example, Patent Document 1 discloses an invention related to a robot that can speak more closely to the user than conventional dialogue systems. This invention is characterized by having an utterance control unit that utters a voice message consisting of a combination of message content and voice characteristics, and further acquiring and analyzing user reaction information and making the next utterance based on the results. In this way, by providing a response according to the reaction, a more natural dialogue with the user is realized. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2022-180232 Summary of the Invention [Problem to be solved by the invention]
[0005] However, the above technologies were unable to realize natural conversations based on the user's personal background, etc. Furthermore, they were unable to check the user's status during natural conversations or provide functions according to the user's intentions.
[0006] In view of the above circumstances, an object of the present invention is to provide a novel technology that realizes more natural dialogue. [Means for solving the problem]
[0007] [1] A dialogue system comprising a dialogue unit, an extraction unit, and a memory unit, wherein the dialogue unit acquires the content of a user's utterance, the extraction unit acquires background information regarding the user's background and a behavioral purpose regarding the user's future actions based on the content of the utterance, links them to the user, and registers them in the memory unit, and the dialogue unit outputs a response to the user based on the background information and behavioral purpose linked to the user who is the speaker.
[0008] This configuration allows the robot to naturally acquire user information from the user's utterances and realize personalized dialogue based on that information. This allows the user to feel as if the robot is "getting to know" them, giving the user a sense of intimacy with the robot and characters that provide the dialogue system.
[0009] [2] The dialogue system described in [1], further comprising a setting unit and a plurality of devices used by the user, wherein the storage unit stores device information for each device linked to the user, the setting unit sets an action specifying a function for supporting the user's behavior based on the content of the utterance, linked to the device capable of executing the action, and the dialogue unit causes the device linked to the action to output information regarding the action.
[0010] With this configuration, it is possible to provide functions according to user comments.
[0011] [3] The dialogue system according to [2], wherein the setting unit sets the action based on the purpose of the action.
[0012] This configuration makes it possible to set actions that match the behavioral goals recognized from the user's utterances, thereby helping the user in natural conversations.
[0013] [4] The dialogue system described in [2] or [3], wherein the device includes an in-vehicle device, the extraction unit obtains the user's destination based on the content of the utterance, the setting unit sets route guidance processing to the destination as the action by linking it to the in-vehicle device, and the in-vehicle device executes the route guidance processing.
[0014] With this configuration, route guidance to the destination can be provided based on a dialogue with the user.
[0015] [5] The dialogue system described in any of [2] to [4], wherein the dialogue unit acquires information identifying the device from which the utterance was input along with the content of the utterance, and outputs a response to the user based on the background information and the device information relating to the device.
[0016] This configuration allows output to be tailored to the device. For example, it is possible to change the content of a message between a device shared by a family and a device used by only one user, or to output messages tailored to the functions of each device.
[0017] [6] A dialogue system described in any of [1] to [5], wherein the extraction unit inputs the utterance content into a large-scale language model and acquires the background information based on the output of the large-scale language model, and the dialogue unit inputs the utterance content and the background information into a large-scale language model and outputs the output of the large-scale language model as the response.
[0018] With this configuration, it is possible to use a large-scale language model to output natural responses that correspond to the content of the statement and the background information.
[0019] [7] The dialogue system described in any of [1] to [6], wherein the memory unit stores a question schedule, the dialogue unit periodically outputs questions to the user at times specified in the question schedule and acquires answers from the user, and the extraction unit registers answer information including the date and time of the answer and the content of the answer in the memory unit, linking it to the user based on the content of the statement that states the answer to the question.
[0020] With this configuration, it is possible to collect information on changes in physical condition and situation, for example, through everyday conversations.
[0021] [8] The dialogue system according to [7], further comprising a linking unit that aggregates the response information and transmits the aggregation result to an external service.
[0022] [9] The dialogue system described in [8], wherein the questions include questions about the user's health condition, the external service provides services in the fields of health, medicine, or nursing care, the collaboration unit receives suggested information based on the aggregation results from the external service, and the dialogue unit outputs the suggested information.
[0023] With this configuration, it is possible to periodically link information to external systems such as a care information system or a system for recording health conditions.
[0024]
[10] A dialogue system according to any one of [1] to [9], further comprising an acquisition unit that acquires sensor information from a sensor related to a device used by the user and identifies the state of the device based on the sensor information, and the dialogue unit outputs a utterance to the user based on the state of the device.
[0025] With this configuration, it is possible to make suggestions and checks regarding the state of the device, and to alert the user.
[0026]
[11] A dialogue method for engaging in dialogue with a user, in which a computer acquires the content of a statement made by the user, acquires background information about the user's background and a behavioral purpose about the user's future actions based on the content of the statement, links the information to the user and registers them in a memory unit, and outputs a response to the user based on the background information and behavioral purpose linked to the user who made the statement.
[0027]
[12] An interactive program that causes a computer to execute the interactive method described in
[11] .
[0028]
[13] An interactive device connected to the interactive system described in
[10] , which receives sensor information from the sensor, transmits it to the acquisition unit, receives the output of the interactive unit, and outputs a statement to the user by voice or display.
[0029]
[14] The dialogue device described in
[14] is used in a vehicle, receives sensor information from a sensor capable of detecting information indicating the status of components of the vehicle, transmits the sensor information to the acquisition unit, receives the output of the dialogue unit, and outputs remarks to the user by voice or display.
[0030]
[15] An interactive device that functions as a device used by a user in the interactive system described in [2], and outputs information about the action by voice or display based on instructions from the interactive unit. [Effects of the Invention]
[0031] According to the present invention, a novel technique can be provided that realizes more natural dialogue. [Brief explanation of the drawings]
[0032] [Figure 1] FIG. 1 is a configuration diagram of an embodiment of the present invention. [Figure 2] Hardware configuration diagram. [Figure 3]FIG. 2 is a functional block diagram of the present embodiment. [Figure 4] 10 shows an example of information stored in a storage unit of the present embodiment. [Figure 5] 3 is a flowchart relating to a dialogue according to the present embodiment. [Figure 6] FIG. 3 is a sequence diagram relating to a dialogue according to the present embodiment. [Figure 7] 10 shows an example of dialogue and background information extraction according to the present embodiment. [Figure 8] 10 shows an example of setting dialogue, behavioral purpose, and action in this embodiment. [Figure 9] 10 shows an example of information stored in a storage unit regarding background information, purpose of action, and action according to this embodiment. [Figure 10] FIG. 10 is a functional block diagram of a modified example. [Figure 11] FIG. 10 is a hardware configuration diagram of an in-vehicle device according to a modified example. [Figure 12] FIG. 10 is a sequence diagram relating to a dialogue according to a modified example. DETAILED DESCRIPTION OF THE INVENTION
[0033] <1. Overview> The present invention relates to a technology for providing natural dialogue, and more particularly to a technology for acquiring personal user information through dialogue and utilizing the information for further dialogue or other services.
[0034] In this embodiment, a system will be described in which, during a dialogue with a user, information about the user's own background, the user's schedule, the purpose of the action, and the like is acquired, and speech and processing are executed based on the acquired information.
[0035] The present invention will now be described in more detail with reference to the accompanying drawings, in which preferred embodiments are shown, but which may be embodied in many different forms and are not limited to the embodiments set forth herein.
[0036] For example, in this embodiment, the configuration, operation, etc. of the dialogue system will be described, but similar effects can be achieved by a device having similar functions, a server or terminal device constituting the system, a method executed by the device, a computer program that causes a computer device to execute the method, etc. The program may be provided as a non-transitory computer-readable recording medium, or may be provided so as to be downloadable from an external server.
[0037] In the following embodiments, the term "unit" may include, for example, a combination of hardware resources implemented by a broadly defined circuit and software information processing that can be specifically realized by these hardware resources. In this embodiment, "information" is represented by, for example, the physical value of a signal value representing voltage or current, the high or low value of a signal value as a binary bit set consisting of 0 or 1, or quantum superposition (so-called quantum bits), and communication and calculation can be performed on a broadly defined circuit.
[0038] A circuit in the broad sense is a circuit realized by appropriately combining a circuit, a processor, a memory, etc. For example, it is a circuit including any of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), etc.
[0039] FIG. 1 is a diagram showing an example of the configuration of a dialogue system according to this embodiment. The dialogue system according to this embodiment is configured by a dialogue server 1 and a plurality of dialogue devices 2 connected to each other so that they can communicate with each other via a network NW. The dialogue server 1 can also access an external generation AI server 3 via the network NW. In this embodiment, the network NW is an IP (Internet Protocol) network, but there are no restrictions on the type of communication protocol, the type of network, etc.
[0040] The dialogue device 2 is any device used by a user and has a mechanism for connecting to a network NW. A user may use multiple dialogue devices 2, each of which communicates with the dialogue server 1 to realize a dialogue with the user. Here, it is preferable that the dialogue system authenticates the user or device using any method, such as password or biometric authentication, thereby ensuring security. Note that authentication may be performed by either the dialogue server 1 or the dialogue device 2.
[0041] The user speaks to the dialogue device 2, and the dialogue device 2 transmits the user's utterance to the dialogue server 1, and receives a response from the dialogue server 1 and outputs it to the user by voice or on-screen display.
[0042] For example, a user may regularly interact with a portable interaction device 2a or an interaction device 2b installed in the living room of the home, and when out and about, interact with an in-car interaction device 2c to obtain directions to a destination, information about the vehicle's status, etc. The content of the interaction in each interaction device 2 is associated with the user in the interaction server 1, and the content of the interaction in one interaction device 2 is taken over to realize linked interactions in other interaction devices 2.
[0043] The dialogue server 1 executes processing based on the user's utterances using the various units described below, determines a response to the user, and transmits it to the dialogue device 2. When analyzing the user's utterances and determining how to process them, the dialogue server 1 accesses the generation AI server 3 via the network NW and uses the large-scale language model 31 provided by the generation AI server 3.
[0044] <2. Hardware configuration> Next, the hardware configuration of the dialogue system according to this embodiment will be described. One or more information processing devices 10 (computer devices) such as a general-purpose server or a personal computer can be used as the dialogue server 1. In this embodiment, the dialogue server 1 is an information processing device 10 in which a computer program (dialogue program) that executes a dialogue method is installed.
[0045] Furthermore, as the dialogue device 2, a terminal device 9 (computer device) having a communication function with the dialogue server 1, such as a personal computer, a smartphone, a tablet terminal, a smart speaker, or a connected car, can be used.
[0046] Fig. 2(a) is a hardware configuration diagram of the information processing device 10. As shown in Fig. 2, the information processing device 10 has a control unit 101, a storage unit 102, and a communication unit 103, which are used to perform the functions of each unit and each process.
[0047] The control unit 101 has a processor such as a CPU that can execute an instruction set, and executes an OS and programs. The storage unit 102 includes a volatile memory such as a RAM capable of storing an instruction set, and a non-volatile recording medium such as an HDD or SSD capable of recording an OS, an interactive program, a DBMS, and the like. The communication unit 103 has an interface for physically connecting to a network, and controls communication with the network NW to input and output information.
[0048] Fig. 2(b) is a hardware configuration diagram of the terminal device 9. As shown in Fig. 2, the terminal device 9 has a control unit 901, a storage unit 902, a communication unit 903, an input unit 904, and an output unit 905, which are used to perform the functions of each unit and each process.
[0049] The control unit 901 has a processor such as a CPU that can execute an instruction set, and executes an OS, application programs, and the like. The storage unit 902 includes a volatile memory such as a RAM capable of storing an instruction set, and a non-volatile recording medium such as an HDD or SSD capable of recording an OS, any application program, and the like. The communication unit 903 has an interface for physically connecting to a network, and controls communication with the network NW to input and output information. The input unit 904 includes an operation input device capable of input processing, such as a touch panel or keyboard, and an audio input device capable of audio input, such as a microphone. The output unit 905 includes a display device capable of display processing, such as a display, and an audio output device, such as a speaker.
[0050] Furthermore, when a connected car is used as the interactive device 2, the information processing device 10 is further connected to the drive system of the automobile or the like. In this case, the control unit 901 is connected to the drive system, including the motor, brakes, and battery, and acquires drive system information acquired by sensors. The information processing device 10 is also connected to an input unit 904, such as a LiDAR (Light Detection and Ranging), a camera, or a GPS (Global Positioning System) receiver, and is configured to be able to acquire information about the surrounding situation.
[0051] <3.Definition> Next, definitions of key terms used in this embodiment will be explained. In the present invention, background information refers to information related to the personal background of each user, and indicates information related to the user's past or present. For example, background information includes any information related to personal background, such as attributes such as age, gender, and occupation, family structure, lifestyle, past experiences, hobbies, and preferences.
[0052] In the following, dialogue refers to communication between the user and the dialogue system via the dialogue device 2. While communication is primarily assumed to be verbal, the dialogue of the present invention also includes the execution of functions in response to user utterances.
[0053] Furthermore, a behavioral intent refers to a goal related to a user's future behavior. In particular, a typical example of a behavioral intent is information that expresses the intention of a user's statement regarding the purpose of the behavior the user intends to carry out. A behavioral intent is, in principle, a short-term request, and in this embodiment, is deleted when a specified time arrives. More specifically, depending on the content, a short period such as until the next goal is extracted, only on the day of setting, or until a deadline is set is assumed. An action refers to a function that can be executed by the interactive device 2. For example, the actions of the present invention include executing a search for a specific word, providing route guidance to a destination, and the like.
[0054] <4. Functional configuration> The details of the processing in this embodiment will be explained below. Fig. 3 is a block diagram showing the functional configuration of the dialogue system of this embodiment. The dialogue server 1 has a dialogue unit 11, an extraction unit 12, a setting unit 13, an acquisition unit 14, and a storage unit 16. This is a specific implementation of software-based information processing using hardware. Note that these components do not need to be implemented by a single computer, but may be implemented by multiple computers working together. Furthermore, some of these components may be provided in the dialogue device 2. For example, by providing the same function as the dialogue unit 11 in the dialogue device 2, even if communication with the network NW is temporarily interrupted, the dialogue device 2 can provide a response.
[0055] The dialogue unit 11 receives the user's utterance from the dialogue device 2, acquires the content of the utterance, and determines a reply to the user and outputs it to the dialogue device 2. The dialogue language can be determined by recognizing the user's prior settings or the language of the utterance from the user. The dialogue unit 11 of this embodiment acquires the content of the user's utterance by receiving the user's utterance as speech and converting it into a character string through speech recognition processing. Note that the speech recognition processing may be performed in an external device. Alternatively, the speech recognition processing may be performed in the dialogue device 2.
[0056] The dialogue unit 11 transmits an instruction sentence including the utterance content and background information linked to the user and registered in advance to the generation AI server 3 to input the content to the large-scale language model 31, and determines a response based on the output from the large-scale language model 31. In addition to the utterance content and background information, the instruction sentence may also include a behavioral purpose linked to the user. Note that the "response" here broadly refers to output to the user and is not limited to linguistic output. For example, the display of a screen proposing an action is also envisioned as a response from the dialogue unit 11. In the following, the terms "utterance", "response", etc. can all be replaced with display on a screen, etc. Details of the processing will be described later.
[0057] Furthermore, the dialogue unit 11 of this embodiment accepts input of images in addition to voice or text utterances. The dialogue unit 11 recognizes still images and videos taken by a camera provided in the dialogue device 2, image files stored in the dialogue device 2, and the like. The recognition results are generally output in text, explaining the content of the image. Note that the image recognition process may be performed in an external device or the dialogue device 2.
[0058] The extraction unit 12 acquires background information about the user's background based on the utterance content acquired by the dialogue unit 11, and links it to the user and registers it in the storage unit 16. The extraction unit 12 also acquires the user's purpose of action based on the utterance content in addition to the background information, and similarly links it to the user and registers it in the storage unit 16. This allows the dialogue unit 11 to determine a response according to the user's background information and purpose of action.
[0059] Based on the content of the utterance, the setting unit 13 associates an action, which specifies a function for supporting the user's behavior, with the interaction device 2 and sets it in the storage unit 16. The setting unit 13 of this embodiment identifies and sets an action and an interaction device 2 corresponding to the action based on the behavioral purpose extracted by the extraction unit 12.
[0060] The acquisition unit 14 acquires sensor information from sensors related to the device used by the user and identifies the state of the device based on the sensor information. For example, assuming that the device used by the user is an automobile, examples of sensor information include sensors that detect the operating status of the engine, motor, battery, brakes, etc. of the automobile used by the user, sensors that detect the state or amount of gasoline or oil, sensors that detect the driving speed or the approach of objects, etc. Such sensor information can be used to identify the state of the device, such as whether maintenance is required, whether there is an abnormality, whether the speed is excessive, etc.
[0061] The linking unit 15 compiles answer information indicating answers to questions posed to the user according to the question schedule, and transmits the compilation result to an external service.
[0062] The memory unit 16 stores the history of the dialogue by the dialogue unit 11, user information, group information indicating the user's group, device information, etc., as well as information such as background information, purpose of action, and action obtained through dialogue with the user.
[0063] 4 is a diagram showing an example of user information, group information, and device information stored in the storage unit 16 of this embodiment. In this way, information about each user is identified by an ID, and a user group is defined in the group information by linking it to the user ID. Groups may be set arbitrarily, but it is envisioned that a group will be set up in which multiple people, such as family members living together, share and use the dialogue device 2. In addition, user settings such as speech length, speed, voice quality, and type of avatar (if an avatar is displayed) are stored in association with the user information.
[0064] The device information also includes information such as the owner ID (user ID), device type, device name, location, and whether the device is being used by anyone other than the owner. The device type specifies the functions that the device can execute. For example, for a smartphone, the available actions include launching a specific app or searching the web, while for a smart speaker, the available actions include controlling IoT-enabled home appliances in the location where the device is installed. Similarly, for an in-car device (connected car), the actions may include controlling the air conditioning or operating the audio system inside the car. Note that available functions may be set individually for each device rather than by type. Furthermore, a configuration may be adopted in which the content of the interaction for each device can be specified by linking it to device information. The information stored in the storage unit 16 is used to determine a response, set an action, and so on.
[0065] The dialogue device 2 has an input unit 21, an output unit 22, and a communication unit . The input unit 21 accepts user input via voice, a touch panel, a keyboard, or the like. The output unit 22 outputs the response of the dialogue system to the user's utterance by voice, screen display, or the like. The communication unit 23 provides communication with the dialogue server 1 through communication via the network NW. The communication unit 23 also communicates with sensors related to devices used by the user to acquire sensor information.
[0066] <5. Interactive processing> Next, the process of the dialogue will be described in detail with reference to the flowchart shown in FIG. First, in step S501, the dialogue unit 11 acquires a user's utterance and generates an instruction sentence for the large-scale language model 31 based on the content of the utterance. Here, before transmitting the instruction sentence, the dialogue unit 11 may analyze or convert the text indicating the content of the utterance. Note that part of the instruction sentence may be given to the large-scale language model 31 in advance. For example, it is assumed that an instruction to reply to the content of the utterance, notes to be taken when replying, etc. may be input in advance.
[0067] More specifically, it is preferable to give instructions such as "You are a therapist who calms the user's mind. Have a friendly conversation with the user as if you were a good friend. Conduct the conversation while referring to the user's past conversation memories, background information, purpose of the action, set actions, etc. If the conversation or topic seems to be coming to a halt, change the topic and try to keep the conversation going for as long as possible. Occasionally look back on past events of the user and move on to the next topic," in advance, and have the user refer to information stored in the memory unit 16, such as the conversation history and background information.
[0068] In addition to the dialogue history and background information, any other information such as general knowledge, health information, current events information, etc. may be stored in the storage unit 16 and used in the dialogue. To refer to the information stored in the storage unit 16, a technology called RAG (Retrieval Augmented Generation) can be used, which compares the contents in vector format and uses the information.
[0069] Next, in step S502, the extraction unit 12 extracts background information and a purpose of action from the content of the user's utterance. In this embodiment, the extraction unit 12 transmits an instruction statement to the generation AI server 3 to extract background information and a purpose of action based on the content of the utterance, and extracts information using the large-scale language model 31. Note that information extraction may be performed in parallel with the dialogue, but the content of the user's utterance and responses by the dialogue system may also be stored in the storage unit 16 as a dialogue history, and information extraction may be performed at regular intervals from the dialogue history in the storage unit 16.
[0070] For example, when extracting a behavioral purpose, the extraction unit 12 sends to the generation AI server 3 an instruction such as, "You are a concierge who determines the user's intention from the dialogue history between the user and the dialogue system. Please extract specific actions the user wants to take from the user's most recent dialogue history. Please make sure to extract not only the user's state such as 'I'm hungry,' but also their intention such as 'I want to go to an Italian restaurant for dinner tonight.'" Similarly, when extracting background information, the background information and important points to be extracted are sent as an instruction. The generation AI server 3 then inputs such an instruction into the large-scale language model 31, obtains the extraction results, and sends them to the dialogue server 1. The extraction process may be performed without using the large-scale language model 31, using a rule-based or statistical method such as extracting pre-registered keywords.
[0071] If background information is extracted from the content of the comment (Y in step S503), the process proceeds to step S504, where the extraction result received by the extraction unit 12 from the generation AI server 3 is registered in the storage unit 16.
[0072] Similarly, if an action purpose is extracted from the content of the statement (Y in step S505), the process proceeds to step S506, where the extraction result received by the extraction unit 12 from the generation AI server 3 is registered in the storage unit 16. Furthermore, when a behavioral purpose is extracted, if an action corresponding to the behavioral purpose exists, then in step S507 the setting unit 13 sets the action in association with the interaction device 2. Specifically, the content of the action is registered in the storage unit 16 in association with the ID of the corresponding interaction device 2.
[0073] As described above, the dialogue system of this embodiment extracts background information and purpose of action from user utterances, associates the results with the user along with the dialogue history, and stores them in the storage unit 16. Then, in subsequent dialogues, utterances are made based on the background information, purpose of action, etc. associated with the user, making it possible to conduct dialogues based on the background, such as each user's experience, and previous dialogues, which gives the user a sense of intimacy and allows the dialogue to continue for a longer period of time. Specific examples of the content of the statement, background information, purpose of the action, and action will be described separately below.
[0074] <6. Communication between devices> Next, the flow of information when a voice dialogue is conducted in the above-mentioned dialogue processing will be explained. Figure 6 is a sequence diagram showing the flow of processing between the dialogue device 2, the dialogue server 1, and the generation AI server 3 in the above-mentioned dialogue processing. Figure 6 shows an example in which some kind of behavioral purpose is extracted from the user's utterance and an action is set.
[0075] First, in step S601, the dialogue device 2 acquires the user's speech. Here, the start of dialogue may be conditional on the user's speech, movement, detection of the user by any method, such as a human sensor, etc. Then, in step S602, the dialogue device 2 transmits the speech data together with its own device ID to the dialogue server 1. Note that the dialogue device 2 may acquire the utterance content by text instead of by voice.
[0076] The dialogue server 1 executes a speech recognition process (S603), and transmits an information extraction instruction including the speech content converted into a character string to the generation AI server 3 (S604). Once the generation AI server 3 extracts information (S605), it returns the results to the dialogue server 1 (S606). Assuming that a behavioral purpose has been extracted, the dialogue server 1 registers the extracted behavioral purpose in the memory unit 16 in the setting unit 13, and sends an instruction statement to the generation AI server 3 specifying an action corresponding to the behavioral purpose (S607). Here, if information such as location information is also obtained in addition to the utterance, it is preferable to store the supplementary information in the memory unit 16 along with the dialogue history.
[0077] The generation AI server 3 refers to the storage unit 16, identifies an action according to the behavioral purpose and the device ID of the interaction device that will execute the action (S608), and returns the action and the device ID to the interaction server 1 (S609). Identifying an action according to the behavioral purpose will be described later.
[0078] In the dialogue server 1, the setting unit 13 associates the action with the dialogue device 2 and sets it in the memory unit 16 (S610), and transmits the action setting to the dialogue device 2 that acquired the user's utterance and to the dialogue device 2 associated with the action.
[0079] Then, the interaction device 2 that acquires the user's utterance outputs a report of the set action to the user, and the interaction device 2 associated as the destination for the action is set to execute the action.
[0080] <7. Dialogue example> Next, specific examples of a user's utterance, extracted background information, set action information, and a response from the dialogue system will be described. The dialogue system generates a response using information such as the user's utterance, date and time, day of the week, location information, and weather information. When the user's background, feelings, intentions, etc. are extracted from the user's utterance, the contents are stored in the storage unit 16 in association with the utterance content, acquisition date and time, the user, the device used, etc.
[0081] FIG. 7 shows an example of acquiring background information from an everyday conversation. On January 5, 2024, the dialogue system (Emo-chan) receives a greeting of "Good morning" from a user (Jack), and responds with a greeting and a reply based on the weather forecast for that day. When the user says, "I'm glad it's sunny," background information indicating the user's feelings of "I'm glad it's sunny" is extracted in association with the date. The system also checks the schedule for that day, and further extracts background information indicating the schedule based on the reply.
[0082] In addition, by asking about past experiences, such as "Do you have any memories of Osaka?", the dialogue system can extract background information indicating the user's past experiences, such as "I lived in Osaka until I went to college," based on the answer.
[0083] 8 is a diagram showing an example of extracting a behavioral purpose from a user's utterance and setting an action. The dialogue system extracts a behavioral purpose from the user's utterance, and then asks questions to extract more specific behavioral purposes from the answers.
[0084] Furthermore, when the purpose of going out, "I want to go buy a Phillips head screwdriver," is extracted, a web search is performed as an executable action on the interactive device 2 with which the conversation is taking place, and a specific destination is suggested, "Shall we go to X Home Center?" When the user specifies the destination, the outing plan is registered. Furthermore, when the user gets out of the car and operates the portable interactive device, the location of the "Phillips screwdriver section" at X Home Center is searched from the Web and displayed on the portable device based on the user's purpose and the characteristics of the portable device.
[0085] In this way, if there is a destination for the outing, and if the user is using the in-car interaction device 2, it is assumed that route guidance to the destination will be performed on the in-car interaction device 2. When an action is set in this way and the in-car interaction device 2 is started up, the set action of "route guidance to X home center" will be executed.
[0086] 9 is a diagram showing an example of background information, purpose of action, and action registered based on the above-mentioned dialogue. As such, the background information and purpose of action each include information such as the recording date and time, user ID, event (content), confidence level, information source, degree of overlap, location where the dialogue that was the basis for recording took place, and device ID.
[0087] The confidence level is information indicating the degree of confidence in the recorded event. Possible information sources include initial settings, conversations, images, and location information. For example, a user's basic profile can be registered by entering it during initial settings, in which case the information source is the "initial settings." In addition to acquiring information through conversations, for example, if a photograph is taken using the conversation device 2, which is a smartphone, the conversation system can register the visit history as background information based on the location information at the time of taking the photograph and the content of the image.
[0088] Additionally, for each action, information such as a deadline, the purpose ID of the purpose corresponding to the action, the device ID of the device that will execute the action, the action content, and the proposed conditions is registered. The deadline is the time limit for executing the action. For example, if an action "provide route guidance to a restaurant" is set in response to a purpose of going to eat dinner that day, inconvenience may occur if the action setting remains after dinner is eaten by another method. Therefore, it is possible to set a deadline for executing the action depending on the content, and automatically delete actions that have exceeded the deadline.
[0089] The proposed condition is a condition for proposing the execution of an action. For example, a web search can be executed without any particular condition, or it can be executed with the user's consent. Actions such as guidance to a destination need to be executed at the timing when the user takes action. Therefore, it is expected that conditions for suggesting the execution of an action, such as starting up an in-vehicle device, will be registered.
[0090] In this way, interactions that have taken place in different interaction devices 2 are linked, and actions to be executed in each interaction device 2 are set, thereby making it possible to effectively support the user's actions.
[0091] <8. How to decide on an action> Next, we will explain an example of a specific procedure for determining an action according to a behavioral purpose. Here, we will assume an example of an action in which the conversation device 2, which is an in-vehicle device (connected car), uses the autonomous driving system to go to the X home center in accordance with the behavioral purpose of "going to the X home center to buy a Phillips screwdriver" shown in Figure 8.
[0092] In this embodiment, the setting unit 13 sets an action based on the purpose of the action. For example, it is assumed that the setting unit 13 acquires information on the action to be set from the large-scale language model 31 by inputting, into the large-scale language model 31, an instruction statement to set an action by referring to information associated with the user in the storage unit 16, together with the content of the purpose of the action ("event" in FIG. 9 ).
[0093] More specifically, the types of executable actions are set in advance in the storage unit 16 for each interaction device 2, and an instruction to determine an action is issued by referring to the settings of the device information linked to the user. For example, the interaction device 2 with device ID D0002 is an in-vehicle device, and executable actions include route guidance by a navigation system and driving to a destination by an automated driving system. Therefore, the large-scale language model 31 determines an action according to the behavioral purpose from among such actions and transfers it to the setting unit 13. Route guidance by a navigation system and driving to a destination by an automated driving system are specific examples of route guidance processing to a destination in this embodiment.
[0094] Furthermore, an action may be determined according to a predetermined rule without using the large-scale language model 31. For example, a method is conceivable in which types of purpose of action are defined in advance and actions corresponding to each type are associated with each purpose of action, thereby setting an action associated with the purpose of action. For example, for a purpose of action such as going to a specific destination, an action for route guidance processing is associated with the purpose of action, thereby making it possible to set an action according to the purpose of action. In this case, the setting unit 13 determines and sets an action according to the predetermined rule.
[0095] Furthermore, if the action cannot be clearly determined from the purpose of the action, the dialogue unit 11 may be controlled to acquire information for determining the action by asking a detailed purpose of the action.
[0096] In this way, an action is determined based on the purpose of the action and the device information registered in the storage unit 16, and set in the storage unit 16 as shown in Fig. 9. In addition to the purpose of the action and the device information, other information stored in the storage unit 16, such as the content of the user's remarks (dialogue history) linked to the purpose of the action and background information, may also be used to determine the action.
[0097] <9. Speech Generation> Next, the generation of a response to a user's utterance will be described in more detail. When the dialogue unit 11 acquires the content of a statement from the user, it transmits an instruction statement to the generation AI server 3 to instruct the generation AI server 3 to generate a reply by referring to the background information and purpose of the action associated with the user. The instruction statement includes the content of the user's statement.
[0098] At this time, it is also preferable to transmit the device ID of the interactive device 2 that acquired the user's utterance, and further generate an instruction sentence to generate a reply based on the device information. For example, if the conversation is on an interactive device 2 that is used by someone other than the owner, it is preferable to set precautions in advance according to the individual or type of interactive device 2, such as avoiding mention of private topics, and transmit the instruction content to the generation AI server 3.
[0099] The dialogue unit 11 acquires the output of the generation AI server 3 and uses the content as a reply to the user. If it is desirable to add an image (still image or video) or sound to the output, it is desirable to acquire this information from web search results or the generation AI server 3, etc., and output it together with the remarks to the user.
[0100] Furthermore, it is preferable that the dialogue unit 11 creates instructions for the generation AI server 3 according to the situation and content of the dialogue. For example, for general dialogue without a particular purpose, it is assumed that the dialogue unit 11 will create instructions to keep the dialogue going by referring to background information and other dialogue history. Furthermore, if the user's remarks include questions or requests, it is assumed that the dialogue unit 11 will create instructions to meet the user's intentions by referring to the background information, dialogue history, and the purpose of the action.
[0101] Furthermore, when generating a utterance to the user, sensor information acquired by the dialogue device 2 can also be used. In this embodiment, the interactive system includes sensors related to the devices used by the user, and sensor information obtained from the sensors is received by the interactive device 2. The acquisition unit 14 then acquires the sensor information transmitted from the interactive device 2 and identifies the state of the device.
[0102] The sensor information may be, for example, information indicating the operating status of any device used in the user's home. Specifically, it is expected that sensors for detecting the operating status of the engine, motor, battery, brakes, etc., in the vehicle parts of a connected car connected to the in-vehicle dialogue device 2, sensors for detecting the state or amount of gasoline or oil, and sensors for detecting the driving speed or the approach of an object, may be used. The dialogue device 2 preferably converts the sensor information into structured text such as XML, transmits it to the dialogue server 1, and stores the sensor information in the storage unit 16 in association with the date, time, user, and device.
[0103] The acquisition unit 14 can, for example, identify the presence or absence of an abnormality as the state of the vehicle (device) based on various operating conditions, and register the details thereof in the storage unit 16. In addition, the running state of the vehicle, such as the running speed or the approach of an object, may be identified as the state of the device and registered in the storage unit 16.
[0104] The dialogue unit 11 generates a utterance based on the state of the device registered in the storage unit 16 in the dialogue device 2 connected to the vehicle. Specifically, for example, it is assumed that the acquisition unit 14 periodically acquires sensor information, and if there is a suspected abnormality in the vehicle, a utterance is generated to inform the user of this and encourage maintenance such as a vehicle inspection. Here, when generating a utterance using the large-scale language model 31, the dialogue unit 11 generates the utterance by sending an instruction sentence including the state of the device based on the sensor information to the large-scale language model 31 and receiving the utterance content as a reply.
[0105] If an action has been set, the dialogue unit 11 generates a utterance suggesting the execution of that action. If the proposal conditions are met and an action associated with the target user and the dialogue device 2 used in the dialogue is registered, a utterance suggesting the action is generated. The above-mentioned instructions may be given to the generation AI server 3 in advance, and the generation AI server 3 may generate the utterance by referring to the action set in the memory unit 16. It is preferable that the dialogue unit 11 checks and corrects the utterance before outputting it so as not to cause any disadvantage to the user in terms of ethics, safety, etc. The check may also be achieved by sending an instruction to the generation AI server 3.
[0106] The dialogue unit 11 also generates utterances based on periodic questions and external linkages, which will be described later. As will be described in detail later, the dialogue unit 11 inputs information related to periodic questions and external linkage information stored in the storage unit 16 into the large-scale language model 31 so as to generate utterances, thereby enabling the generation of appropriate utterances.
[0107] More specifically, the dialogue unit 11 first determines the type, purpose, and action in the device information, and identifies background information related to the content.Then, the dialogue unit 11 includes this information in an instruction sentence and sends the instruction sentence to the generation AI server 3, thereby generating a utterance for the user.
[0108] <10. Regular Questions> The dialogue system of this embodiment asks the user periodic questions (hereinafter referred to as "periodic questions") and externally links the answers to the periodic questions and information obtained during the dialogue. Details of the periodic questions and external linkage will be described below.
[0109] The storage unit 16 stores a question schedule for periodic questioning, linked to the user. The question schedule may be information indicating a rough time for questioning, such as once a week, once a month, once every six months, or a predetermined period after a specific event such as a health check. In this way, the question schedule is not limited to specific dates, but may also specify a rough time.
[0110] The dialogue unit 11 periodically outputs questions at times specified in the question schedule. Specifically, it is assumed that the dialogue unit 11 refers to the question schedule in the storage unit 16 and instructs the large-scale language model 31 to ask a question when the time matches. Note that questions may be registered in advance, and the dialogue unit 11 may output questions without using the large-scale language model 31.
[0111] Whether the time indicated in the question schedule matches can be determined by any method, such as specifying a specific period in the question schedule and outputting a question when the first dialogue occurs within that period. Another method is to specify the start of a period in the question schedule and output a question when the first dialogue occurs after the start of the period. Alternatively, a condition can be specified in advance for the generation AI server 3, such as when an opportunity for dialogue occurs around the date indicated in the question schedule.
[0112] The dialogue unit 11 then acquires the answer to the question and stores the content of the answer in the storage unit 16 in association with the user, the dialogue device 2, and the date and time.
[0113] The questions are expected to mainly concern, for example, the user's health condition, the status of surrounding support, changes in the environment, etc. More specifically, the questions are expected to ask the user about their health condition and mood at multiple stages. In this embodiment, the linking unit 15 compiles the answer information and transmits the compilation results to the external service.
[0114] In addition to the answers to the questions generated by the dialogue unit 11, if information on the user's health condition or the like is obtained in daily dialogue, the extraction unit 12 may extract the information and store it in the storage unit 16, similar to background information and purpose of action. In this case, the linking unit 15 may also aggregate the information obtained through dialogue and link it to an external service.
[0115] Examples of external services include services related to health, medical care, nursing care, etc., and the linking unit 15 transmits response information related to the user's health, etc. to these services and obtains proposal information based on the aggregation results from the external services. More specifically, for example, a scientific nursing information system (LIFE) or an SNS that provides a function to contact family members can be used as an external service. In addition, a reservation system for a hospital or restaurant may be used as an external service, and the linking unit 15 may have a function to make a reservation based on the response information.
[0116] As described above, the dialogue system of this embodiment can acquire personal information of the user during natural dialogue and provide dialogue according to the content of the information. Furthermore, by providing actions according to the user's purpose of action through the dialogue device 2, it is possible to more effectively support the user.
[0117] Furthermore, by periodically asking questions about the user's condition, acquiring and compiling the answers, and linking them to external services, it is possible to provide services such as status checks for users who require care or supervision, and regular health checks, etc. This makes it possible to effectively detect illness early, contact relevant parties, and propose related services.
[0118] <11. Variations> A specific example of processing in another embodiment will be described below with reference to FIGS. Fig. 10 shows the detailed functional configuration of the dialogue server 1 in the modified example. Note that communication between the dialogue server 1 and external services including the dialogue device 2 and the generation AI server 3 is generally configured using TCP / IP, an encryption module, HTTP, etc., but the description of the communication unit is omitted here.
[0119] Moreover, 10120 to 10111, 10121 to 10123, and 10127 described below correspond to the dialogue unit 11 in the above-described embodiment. 10124 corresponds to the extraction unit 12 in the above-described embodiment. 10125 corresponds to the extraction unit 12 and setting unit 13 in the above-described embodiment. 10128 and 10131 to 10133 correspond to the linking unit 15 in the above-described embodiment. 102 corresponds to the storage unit 16 . Each component of the modified example will be described below.
[0120] 10101 is an authentication unit that authenticates users or devices. It authenticates access to this service using a predetermined method such as password or face recognition, ensuring security.
[0121] Reference numeral 10102 denotes a text input unit. When the interactive device 2 transmits text (including structured text such as XML), simple analysis and conversion of the text is performed here. It is assumed that structured text is used, for example, when transmitting sensor information possessed by a terminal, by converting it into structured text on the terminal side and transmitting it.
[0122] Reference numeral 10103 denotes a voice recognition unit. When a voice waveform or compressed voice is sent from the interactive device 2 by streaming or file transfer, voice recognition processing is performed here. The voice recognition itself may be realized by an external service (not shown).
[0123] An image recognition unit 10104 recognizes still images and videos taken by the camera of the interactive device 2, image files stored in the terminal, etc. The recognition result is generally output as text, explaining the content of the image.
[0124] 10106 is a text output unit. Like 10102, it may be structured text such as XML. It is also used when synthesizing text to speech in the dialogue device 2 or when using structured text to control a terminal (such as setting a destination in a car navigation system or controlling an air conditioner).
[0125] Reference numeral 10107 denotes an image generation unit. If adding an image to the output text will make it easier for the user to understand, the image is generated here. For example, from the text "Set XX department store as destination", an image of "XX department store" can be obtained from the Internet, or an image can be generated using an external service including the generation AI server 3, and sent to the dialogue device 2.
[0126] Reference numeral 10108 denotes an audio / video output unit, which also synthesizes text from text and generates video, or obtains audio or video data from an external service and outputs the data to the interactive device 2.
[0127] Reference numeral 10109 denotes a language switching unit, which switches the language of voice and text to various languages such as Japanese, English, and Chinese according to user settings.
[0128] 10110 is a device recognition unit. It recognizes the device currently used by the user. It does this by acquiring the device type and device ID stored in advance in the interactive device 2. It is assumed that the output content and format (text, audio, image, etc.) can be dynamically switched depending on the device type.
[0129] 10111 is an information output filter that checks dialogue output and deletes or modifies it. Dialogue responses generated by 10121 to 10128 are checked to ensure they do not cause any harm to the user in terms of ethics or safety, and inappropriate responses are deleted or replaced with other appropriate expressions. In this case too, a large-scale language model can be utilized. It is also possible to change the strength of the filter using the user's age stored in the personal setting database.
[0130] 10121 is a sequencer unit that controls each module that generates a dialogue. In this system, it is necessary to operate multiple blocks depending on the content of the dialogue, the type of dialogue purpose, the type of device, etc. The sequencer unit operates the necessary blocks according to various situations.
[0131] 10122 is a general dialogue unit. For example, when driving long distances, there are cases where a dialogue without a specific purpose is required, such as to wake up from drowsiness. The general dialogue unit aims to continue a general dialogue for a long time while referring to the background information DB 1022 and the dialogue memory DB 1023.
[0132] 10123 is a situation switching unit. When using a generation AI, it is difficult to have it completely answer all the information with a single instruction. Therefore, it is desirable to select a group of information and instructions that are appropriate for each situation from the previous dialogue content and then have the dialogue proceed. For example, depending on the purpose, such as satisfying hunger, shopping, or seeking health advice, the necessary information is obtained from the purpose-specific DB 1021, and the dialogue is dynamically set to respond to the user's intentions.
[0133] Reference numeral 10124 denotes a background information learning unit. The background information of the user is learned from the dialogue history stored in the dialogue memory DB and stored in the background information DB. The contents of the background information DB are referenced in situation-specific dialogues and general dialogues, and are intended to enable dialogue that is tailored to the user as the dialogue progresses. Background information is information that is associated with an individual and does not involve actions, such as preferences and experiences. Furthermore, background information is mainly information that indicates past or present situations. By weaving background information into dialogue, dialogue related to the user's background information becomes possible, thereby realizing a dialogue device that creates a sense of familiarity. Background information is relatively long-term information, and the background information in this modified example does not have a deletion rule, such as deleting it once the purpose is completed, as is the case with purpose information extracted by the purpose extraction unit.
[0134] 10125 is the intent extraction unit. It is similar to situation switching, but it extracts the current user intent more precisely, such as "I want to eat," "I want to go to a restaurant," or "I want to eat curry." An intent is an intent that the user wants to achieve in the future. It also extracts actions according to the intent and device. For the intent "I want to eat curry," the action for the portable dialogue device 2 is "select a curry restaurant," and for the dialogue device 2 in the car (in-vehicle device), the action is "set automatic driving with the selected curry restaurant as the destination."
[0135] Since the dialogue purpose and actions are temporarily stored in the purpose information DB 1023, if a dialogue is held on a mobile device and the purpose of going to Restaurant X is set, the purpose will be carried over and the destination will be set in the car's navigation system even if the dialogue device 2 is changed when getting into a car. Purposes are basically short-term requests that involve action, and depending on the content, they will be retained in memory for a short period of time, such as until the next purpose is extracted, only on the day it was set, or until a set deadline, at which point the information will be deleted.
[0136] A scheduler unit 10126 schedules periodic actions and is used to operate modules such as periodic queries performed by 10127 and monitoring the status of devices connected to the interactive device 2.
[0137] Reference numeral 10127 denotes a questioning unit that asks questions to the user and the interactive device 2. The user is periodically asked questions about depression or dementia assessment, for example, to understand the user's current state. The interactive device 2 is periodically queried about fuel efficiency, braking status, travel distance, etc., to understand the state of the devices connected to the interactive device 2. The answers to the questions are stored in the background information DB 1022.
[0138] 10128 is a compilation unit that compiles the questions and answers of 10127 stored in the background information DB of 1022 and creates a report. The report can be communicated to the user by voice, a graph can be generated and displayed on the terminal screen, or it can be communicated to a third party via an external service. For example, by periodically asking about depression, it is possible to understand the changes in the mental state of a user living alone in numerical terms.
[0139] Reference numeral 1021 denotes a purpose-specific information DB. It stores restaurant information, health information, time-killing information, information about the interactive device 2 currently in use, and the like. This information is divided into two types: dialogue instructions and dialogue content information. Dialogue instructions include instructions to the generation AI server 3, etc., to realize a dialogue purpose, such as "recommend a restaurant that suits the user's preferences." Dialogue content information includes restaurant information, hospital information, and the like. Of course, in addition to the contents of the DB stored locally, related information may be obtained online by performing a web search at that time. The dialogue content information may also be used in a format called RAG (Retrieval-Augmented Generation), which retrieves information by checking content matches in a vector format.
[0140] Reference numeral 1022 denotes a background information DB. It stores the user's attributes such as age, sex, occupation, etc., obtained by the background information learning unit 10124, as well as family structure, lifestyle habits, past experiences, hobbies, preferences, etc. Some of this information may be input during initial setup.
[0141] 1023 is a purpose information DB. 10125 acquires and stores the user's purpose and the corresponding device and action from the dialogue information. It also deletes the purpose and action upon request from the purpose extraction unit.
[0142] Reference numeral 1024 denotes a dialogue memory DB. It stores dialogues over time. In this case, it stores both utterances from the system and utterances from the user (including information from the dialogue device 2). By processing this dialogue memory DB retroactively, the accuracy of extracting the user's intentions is ensured.
[0143] Reference numeral 1025 denotes a personal setting DB. The personal setting is set by the user and stores therein information such as the length and speed of speech, voice quality, type of avatar, and dialogue content for each device.
[0144] 10131a, 10131b, and 10131c are plug-in modules for processing the settings and protocols for connecting to external services 10141a, 10141b, and 10141c, respectively. These services have different authentication and communication procedures, so by installing the corresponding plug-in, it becomes possible to exchange information with external services.
[0145] An authentication unit for external services 10132. Although authentication differs depending on the external service, basic information is stored in a personal setting DB 1024, and authentication corresponding to each service is performed using this information. An external service interface unit 10133 communicates with external services and performs caching and other processing as needed. 10141a, 10141b, and 10141c are external services. For example, it is assumed that the user's dialogue status can be grasped by sending a monthly dialogue status report to family members in remote locations, that changes in depression status can be registered in an external database by linking with a health-related database, and that restaurant reservations can be made as needed.
[0146] Next, we will explain the hardware configuration of the dialogue device 2, which is an in-vehicle device in particular. Figure 11 shows an example of the implementation of the dialogue device 2 in a mobility device such as an automobile.
[0147] The terminal part of 9 is the same whether it is a car, a smartphone, or a stationary or mobile robot. 903 is a communication unit that communicates with external servers and services. 902 is a non-volatile storage unit that stores the terminal and user status, programs, data, etc. 9011 is memory that is used when the CPU of 9012 performs processing. 9012 is a CPU that performs various calculation processes. 9041 is a touch device that inputs user instructions, and 9042 is a camera that takes images of the user and the scenery. 9051 is a speaker that outputs voice and music. 9052 is a display that outputs images and videos to the user.
[0148] 501 to 509 and 511 to 515 are block diagrams specific to devices with drivetrains such as automobiles. 501 is an ECU (Electronic Control Unit) that processes drivetrain information. ECU 501 is connected to the CPU 9012 via an interface such as a bus or USB. Because the hardware and software of 501 often differ depending on the manufacturer and model, it is desirable to communicate information with the CPU in a text-based, common language. For this reason, it is desirable to convert information about brakes, electrical systems, etc. into structured text such as XML or natural language and then communicate it to the CPU. Of course, it is also possible to convert information specific to the model on the CPU side of 9012, and then communicate the information to the cloud in a common language.
[0149] 502 is a communication path (bus) that transmits information about the drive system. 503 is the driving motor, 504 is the brake, and 505 is the battery, and it is assumed that the ECU can grasp the status by sensing the status of these.
[0150] 506 is a communication path (bus) for the sensor system. 507 is a camera that takes pictures of the inside and outside of the vehicle. 508 is a LiDAR (Light Detection And Ranging). Of course, the LiDAR has a sub-CPU that interprets the information, understands the driving situation, and judges danger, but we have omitted the description of this. Here, we are assuming that the situation of driving over the center line is obtained from the LiDAR information, and that driving caution is spoken in a dialogue based on that information.
[0151] 509 is a GPS receiver. It is used in conjunction with LiDAR for autonomous driving and collision prevention, and is intended to use GPS information to announce driving hazards specific to a location (such as warning of falling rocks).
[0152] Reference numeral 511 is a sub-ECU that controls the infotainment system. Reference numeral 512 is a communication path (bus) for the infotainment system. Reference numeral 513 is a navigation system. It is assumed that when a destination is determined through dialogue, the destination is set in the navigation system using information from the CPU, and waypoints along the way are added.
[0153] 514 is an air conditioning system that communicates the air conditioning status to the CPU and controls the air conditioning temperature setting from the CPU. 515 is the audio system. The CPU plays downloaded music and adjusts the volume.
[0154] The sensing information and control information from 501 to 515 are stored in the background information DB in the 1022 cloud. This allows the user to control the vehicle as desired even if the vehicle model changes. Furthermore, even if the physical operation of the vehicle changes, control via dialogue is possible, allowing the user to control various functions through familiar dialogue, regardless of the physical interface.
[0155] The details of the process in the modified example will be explained below. Figure 12 shows the dialogue process flow using background information and purpose extraction. In the diagram, the dialogue processing, purpose extraction processing, and background information extraction processing are executed by the dialogue server 1. In S1001, when the dialogue device 2 is connected to the dialogue server 1, the dialogue server 1 authenticates the user or device. In S1002, the dialogue begins when the user speaks into the microphone or makes a movement towards the camera. Of course, the dialogue server can also speak to the user by sending a user detection event to the dialogue server using an image or human presence sensor.
[0156] In S1003, text after speech recognition and image recognition is sent to the dialogue server 1. Here, an example is described in which the dialogue device has speech recognition and image recognition functions, but when recognition is performed by the dialogue server 1, speech data and image data are sent. In S1004, dialogue processing is performed. Details of dialogue processing will be described later.
[0157] In S1005, the results of recognizing the user's voice, etc., and the response results created in the dialogue processing are stored along with the time. If location information is available using GPS, etc., the location is also stored. The dialogue data is stored in a dialogue memory DB with 1024 entries. The DB (database) may be in the form of a database such as MySQL (registered trademark), or may be in any on-memory data format other than a database, as long as the data can be accessed and referenced from other processes.
[0158] In S1006, the text, voice, image, and action created in the dialogue processing in S1004 are returned to the dialogue device 2. Actions are descriptions for operating the device's functions, and include setting target values for autonomous driving of a car, moving the body of a dialogue robot, and changing facial expressions.
[0159] In S1007, the response in S1006 is sent to the interactive device 2. In S1008, the response data sent in S1007 is interpreted by the interactive device 2, converted into a format suitable for the user, and transmitted to the user.
[0160] In S1009, the dialogue history stored in the 1024 dialogue memory DB is referenced to extract the user's purpose and next action. In S1010, the purpose and action extracted in S1009 are stored in the purpose information DB 1023. At this time, a deletion deadline may be estimated and set along with the purpose and action. For example, the deletion deadline for a food-related goal such as "I want to eat curry" will be the end of the day or when you have eaten a meal, unless otherwise specified.
[0161] In S1011, the dialogue history stored in the 1024 dialogue memory DB is referenced to extract background information of the user. In S1012, the background information extracted in S1011 is stored in the background information DB 1022. The background information is simply a fact, and there is no corresponding action. In S1013, if the response output created in S1006 is addressed to an external service, an appropriate action is executed for the external service. Actions include searching for an external service, making a restaurant reservation, and registering data in an external database. In S1014, the processes from S1002 to S1013 are repeated to continue the dialogue.
[0162] Next, the interactive processing in S1004 will be described. The data in the purpose information DB 1021 and the personal setting DB 1025 is rewritten at the time of initial setup or system update, and is not changed for each conversation. The information stored in the purpose information DB 1023 and the information stored in the background information DB 1022 is updated for each conversation or for each block of conversation.
[0163] In S1004, the following various pieces of information 1) to 5) are set as preconditions, and a dialogue is generated based on the preconditions. 1) Set the device type. The device type is determined by the device type associated with the pre-authenticated device ID. 2) Refer to the purpose information DB 1023 and set the current purpose. There may be no purpose. If there is an action, set the action as well. For example, "I want to eat lunch." 3) From the background information of 1022, obtain and set information that matches the purpose of 2). We are assuming information such as "I like curry" or "I ate spicy curry yesterday." 4) Information that matches the purpose of 2) is obtained and set from the personal setting database of 1025. Possible information includes: today is your birthday, and you are 52 years old. 5) Information that matches the purpose of 2) is selected from the 1021 purpose-specific information database. There could be web information about nearby restaurants or recipes for making curry.
[0164] In the dialogue processing in S1004, the above information is listed as text, and a response is generated along with that information to achieve the purpose of 2). The information from 1) to 5) is expressed in sentences as prompts (instructions) and sent to the generation AI server 3, which then generates an answer using the large-scale language model 31.
[0165] If there is no purpose, the system uses information excluding the purpose to generate a general dialogue response that follows the flow of the dialogue in the dialogue memory database. In this case, by using background information such as "I spent my student days in Osaka" to create an utterance from the system such as "I'd like to go to Osaka once in a while," it is possible to realize a dialogue that is tailored to the user based on the background information.
[0166] 12 may be executed using the large-scale language model 31. These processes may be executed in parallel in the large-scale language model 31.
[0167] For example, when extracting goals using a conversational generation AI, it is assumed that the following prompt (instruction) will be sent to the generation AI server 3: "You are a kind butler who is attentive to the wishes of your customers. Please extract the goal that the current customer wants to achieve from the conversation with the customer attached below. A goal is something that the customer can achieve in a relatively short period of time, such as 'I want to eat' or 'I want to buy a Phillips screwdriver.' Also, if there is no particular goal and you are having a general conversation, please answer 'None' as the goal." In addition to the above prompts, it is preferable to also transmit a dialogue history, which is assumed to be the most recent dialogue sequence obtained from a dialogue memory DB. [Explanation of symbols]
[0168] 1: Interactive server 11: Dialogue section 12:Extraction part 13: Setting section 14: Acquisition part 15: Collaboration Department 16: Storage section 2: Interactive devices 21: Input section 22: Output section 23: Communications Department 3: Generative AI server 31: Large-scale language model 10: Information processing device 101: Control unit 102: Storage section 103: Communications Department 9: Terminal device 901: Control unit 902: Storage section 903: Communications Department 904: Input section 905: Output section NW: Network 10101: User device authentication unit 10102: Text input section 10103: Speech recognition unit 10104: Image recognition unit 10105: Text input section 10106: Speech synthesis unit 10107: Image generation unit 10108: Audio / video output section 10109: Language switching section 10110: Device recognition unit 10111: Information output filter 10121: Sequencer section 10122: General Conversation Department 10123: Status switching section 10124: Background Information Learning Department 10125: Purpose extraction part 10126: Scheduler section 10127: Questions 10128: Aggregation section 10131: External service plugin 10132: Service authentication section 10133: External service I / F section 102: Storage section 1021: Purpose-specific information database 1022: Background information DB 1023: Purpose information DB 1024: Dialogue memory DB 1025: Personal settings DB
Claims
1. A dialogue system comprising a dialogue unit, an extraction unit, and a storage unit, The dialogue unit acquires the content of a user's statement, the extraction unit acquires background information relating to the user's background and a behavioral purpose relating to the user's future behavior based on the content of the utterance, and associates the background information, the behavioral purpose, and its deadline with the user and registers them in the storage unit; The action goal is deleted if the action goal is completed before the deadline and if the deadline has passed; The dialogue unit determines a response to the user based on the background information associated with the user who is the speaker and the valid purpose of the action.
2. the storage unit includes a purpose-specific information DB that stores information that can be provided to users; 2. The dialogue system according to claim 1, wherein the dialogue unit identifies and acquires information to be acquired from the purpose-specific information DB based on the content of the dialogue up to that point and the effective behavioral purpose, and determines a response to the user by inputting a prompt based on the information acquired from the purpose-specific information DB into a large-scale language model.
3. The system further includes a setting unit and a plurality of devices used by the user, the storage unit stores device information for each device in association with the user; the setting unit sets, based on the content of the utterance, an action specifying a function for supporting the user's behavior and a condition for proposing execution of the action, by linking the action to the device capable of executing the action; The dialogue system according to claim 1 , wherein, when the condition is satisfied, the dialogue unit causes the device that associates the action with the condition to output a proposal to execute the action.
4. The dialogue system according to claim 3 , wherein the setting unit sets the action based on the purpose of the action.
5. the device includes an in-vehicle device; The extraction unit acquires a destination of the user based on the content of the comment, the setting unit sets, as the action, a route guidance process to the destination in association with the in-vehicle device; The interactive system according to claim 3 or 4, wherein the in-vehicle device executes the route guidance process.
6. The dialogue system according to claim 3 , wherein the dialogue unit acquires information specifying a device through which the utterance was input along with the content of the utterance, and outputs a response to the user based on the background information and the device information related to the device.
7. the extraction unit inputs the utterance content into a large-scale language model and acquires the background information based on an output of the large-scale language model; The dialogue system according to claim 1 , wherein the dialogue unit inputs the utterance content and background information into a large-scale language model and outputs an output of the large-scale language model as the response.
8. The storage unit stores a question schedule, the dialogue unit periodically outputs questions to the user at times specified in the question schedule and acquires answers from the user; The dialogue system according to claim 1 , wherein the extraction unit registers answer information including a date and time of the answer and a content of the answer in the storage unit, in association with the user, based on the content of the statement that states the answer to the question.
9. Further provided with a linking section, The dialogue system according to claim 8 , wherein the linking unit tally the response information and transmits a result of the tallying to an external service.
10. the questions include questions about the user's health status; The external service provides a service in the fields of health, medicine, or nursing care; the linking unit receives proposal information based on the aggregation result from the external service; The dialogue system according to claim 9 , wherein the dialogue unit outputs the suggested information.
11. further comprising an acquisition unit, the acquisition unit acquires sensor information from a sensor related to a device used by the user and identifies a state of the device based on the sensor information; The dialogue system according to claim 1 , wherein the dialogue unit outputs a utterance to the user based on a state of the device.
12. 1. A method for interacting with a user, comprising: Acquire the content of the user's comments; Based on the content of the statement, background information regarding the user's background and a behavioral purpose regarding the user's future behavior are acquired, and the background information, the behavioral purpose, and its deadline are linked to the user and registered in a storage unit; Delete the action goal if the action goal is completed before the deadline and if the deadline has passed; A dialogue method in which a reply to the user is determined based on the background information linked to the user who is the speaker and the valid purpose of the action.
13. An interaction program that causes a computer to execute the interaction method according to claim 12.
Citation Information
Patent Citations
Speech dialogue method and device
JP2006003743A
Interactive household electrical system, server device, interactive household electrical appliance, method for household electrical system to interact, and program for realizing the same by computer
JP2015184563A
Automatic activation of smart responses based on activation from remote devices
JP2016534616A
Voice interaction control method
JP2018169624A
Voice interactive method, and voice interactive agent server
JP2020173477A