Dialogue systems, dialogue methods, dialogue programs, and dialogue devices

The dialogue system addresses the lack of personalized and adaptive responses in conventional systems by extracting user background and action information to provide natural and device-specific interactions.

JP2026081427AActive Publication Date: 2026-05-19YUKAI ENG INC
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
YUKAI ENG INC
Filing Date
2024-11-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Conventional dialogue systems fail to provide natural conversations based on individual user backgrounds and states, lacking the ability to check user intentions and adapt responses accordingly.

Method used

A dialogue system comprising a dialogue unit, extraction unit, and storage unit that acquires user statements, extracts background information and future actions, and outputs personalized responses based on this information, with the ability to control devices to perform actions aligned with user objectives.

Benefits of technology

Enables natural and personalized conversations, allowing the system to assist users in a conversational manner by adapting to their backgrounds and intentions, and providing device-specific responses and actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026081427000001_ABST
    Figure 2026081427000001_ABST
Patent Text Reader

Abstract

We provide dialogue systems, dialogue methods, dialogue programs, and dialogue devices that enable more natural dialogue. [Solution] In the dialogue system, the dialogue server comprises a dialogue unit 11 that acquires the content of the user's statements, and an extraction unit 12 that acquires background information about the user's background and action objectives about the user's future actions based on the acquired statements, and registers the acquired background information and action objectives in a storage unit 16, associating them with the user. The dialogue unit 11 outputs a response to the user based on the background information and action objectives associated with the user using the dialogue device who is the speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] , , ,

[0001] The present invention relates to a dialogue system, a dialogue method, a dialogue program, and a dialogue device.

Background Art

[0002] As a conventional technique related to a dialogue system and a dialogue method, a dialogue robot using artificial intelligence (AI) is widely known. These systems utilize speech recognition technology and natural language processing technology to generate appropriate responses based on voice input from the user. Most conventional dialogue systems have been used for purposes of proceeding with a dialogue according to a preset scenario and providing responses to specific questions. These technologies are useful in scenarios such as customer support and information guidance.

[0003] For example, Patent Document 1 discloses an invention related to a robot that enables speech closer to the user compared to a conventional dialogue system. This invention includes a speech control unit that speaks a speech message composed of a combination of the content of a message and the characteristics of speech, and further acquires and analyzes the user's reaction information, and performs the next speech based on the result. In this way, by providing a response according to the reaction, a more natural dialogue with the user is realized.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, with the above technology, it has not been possible to realize a natural conversation based on the background of the individual user. Also, it has not been possible to check the user's state during a natural conversation or provide functions according to the user's intention.

[0006] In view of the above circumstances, the present invention aims to provide a novel technology that enables more natural dialogue. [Means for solving the problem]

[0007] [1] A dialogue system comprising a dialogue unit, an extraction unit, and a storage unit, wherein the dialogue unit acquires the content of a user's statements, the extraction unit acquires background information relating to the user's background and the user's future actions based on the content of the statements, registers them in the storage unit in association with the user, and the dialogue unit outputs a response to the user based on the background information and actions associated with the user who made the statement.

[0008] This configuration allows the system to naturally acquire user information from user statements and then implement personalized conversations based on that information. As a result, users can feel that the robot "knows" them, fostering a sense of familiarity with the robot or character providing the conversational system.

[0009] [2] The dialogue system according to [1], further comprising a setting unit and a plurality of devices used by the user, wherein the storage unit stores device information associated with the user for each device, the setting unit sets actions that specify functions to support the user's actions based on the content of the statement, and associates these actions with the devices that can execute the actions, and the dialogue unit causes the devices to which the actions are associated to output information about the actions.

[0010] This configuration allows us to provide functions that respond to user input.

[0011] [3] The setting unit sets the action based on the action objective, the dialogue system as described in [2].

[0012] This configuration allows for setting actions that align with the user's perceived behavioral objectives based on their statements. This enables the system to assist users in a natural conversational manner.

[0013] [4] The conversational system according to [2] or [3], wherein the device includes an in-vehicle device, the extraction unit obtains the user's destination based on the content of the statement, the setting unit sets a route guidance process to the destination as an action and associates it with the in-vehicle device, and the in-vehicle device executes the route guidance process.

[0014] This configuration allows for route guidance to the destination based on interaction with the user.

[0015] [5] The dialogue unit acquires information identifying the device on which the statement was input, along with the content of the statement, and outputs a response to the user based on the device information relating to the device in addition to the background information, according to any of [2] to [4].

[0016] This configuration allows for output tailored to the device. For example, the content of messages can be changed depending on whether the device is shared with family members or used by a single user, or messages can be tailored to the functions of each device.

[0017] [6] A dialogue system according to any one of [1] to [5], wherein the extraction unit inputs the utterance content into a large-scale language model and acquires the background information based on the output of the large-scale language model, and the dialogue unit inputs the utterance content and background information into a large-scale language model and outputs the output of the large-scale language model as the response.

[0018] This configuration allows for the use of a large-scale language model to output natural responses that correspond to the content of the statement and background information.

[0019] [7] The memory unit stores a question schedule, and the dialogue unit periodically outputs questions to the user at the times specified in the question schedule to obtain answers from the user. The extraction unit registers answer information including the answer date and time and the answer content in the memory unit in association with the user based on the speech content stating the answer to the question, for the dialogue system according to any one of [1] to [6].

[0020] By adopting such a configuration, for example, changes in physical condition or situation can be collected as information during daily conversations.

[0021] [8] The dialogue system according to [7], further comprising a cooperation unit, wherein the cooperation unit aggregates the answer information and transmits the aggregation result to an external service.

[0022] [9] The question includes a question regarding the user's health condition, the external service provides a service in any of the fields of health, medical care, or nursing care, the cooperation unit receives proposal information based on the aggregation result from the external service, and the dialogue unit outputs the proposal information, for the dialogue system according to [8].

[0023] By adopting such a configuration, regular information can be linked to external systems such as a nursing care information system or a system for recording health conditions.

[0024]

[10] The dialogue system according to any one of [1] to [9], further comprising an acquisition unit, wherein the acquisition unit acquires sensor information from sensors related to the device used by the user, specifies the state of the device based on the sensor information, and the dialogue unit outputs a speech to the user based on the state of the device.

[0025] By adopting such a configuration, proposals or confirmations regarding the state of the device can be made, and it becomes possible to alert the user.

[0026]

[11] A dialogue method for interacting with a user, wherein a computer acquires the speech content of the user, and based on the speech content, acquires background information regarding the background of the user and an action purpose regarding the future actions of the user, registers them in a storage unit associated with the user, and outputs a response to the user based on the background information and the action purpose associated with the user who is the speaker.

[0027]

[12] A dialogue program for causing a computer to execute the dialogue method according to

[11] .

[0028]

[13] A dialogue device connected to the dialogue system according to

[10] , which receives sensor information from the sensor, transmits it to the acquisition unit, receives the output of the dialogue unit, and outputs a speech to the user by voice or display.

[0029]

[14] The dialogue device according to

[14] , which is used in a vehicle, receives the sensor information from a sensor capable of detecting information indicating the state of a component of the vehicle, transmits it to the acquisition unit, receives the output of the dialogue unit, and outputs a speech to the user by voice or display.

[0030] <0OO0105>

[16] A dialogue device that functions as a device used by a user in the dialogue system according to [2], and outputs information regarding the action by voice or display based on an instruction from the dialogue unit. [Advantages of the Invention]

[0031] According to the present invention, it is possible to provide a novel technology that realizes a more natural dialogue. [Brief Description of the Drawings]

[0032] [Figure 1] Configuration diagram of the present embodiment. [Figure 2] Hardware configuration diagram. [Figure 3]Functional block diagram of this embodiment. [Figure 4] An example of the information stored in the memory unit of this embodiment. [Figure 5] A flowchart relating to the dialogue of this embodiment. [Figure 6] A sequence diagram relating to the dialogue in this embodiment. [Figure 7] An example of dialogue and background information extraction in this embodiment. [Figure 8] Examples of dialogue and action objectives and action settings in this embodiment. [Figure 9] An example of information stored in the memory unit regarding the background information, purpose of action, and action of this embodiment. [Figure 10] Functional block diagram of a modified example. [Figure 11] Hardware configuration diagram of a modified in-vehicle device. [Figure 12] A sequence diagram relating to the dialogue of a modified example. [Modes for carrying out the invention]

[0033] <1. Overview> This invention relates to a technology for providing natural dialogue. In particular, this invention relates to a technology for acquiring user personal information through dialogue and further utilizing it in dialogue or other services.

[0034] This embodiment describes a system that, during interaction with the user, acquires information about the user's background, schedule, and purpose of action, and then performs utterances and processes based on that information.

[0035] Further details will be provided below with reference to the attached drawings. While preferred embodiments are shown in the drawings, many different forms are possible and the embodiments are not limited to those described herein.

[0036] For example, in this embodiment, the configuration and operation of the dialogue system are described, but similar effects can be achieved by devices having similar functions, servers or terminal devices constituting a system, methods executed by the devices, computer programs that cause computer devices to execute the methods, etc. The program may be provided as a non-transient recording medium that can be read by a computer, or it may be provided so that it can be downloaded from an external server.

[0037] In the following embodiments, "part" may include, for example, hardware resources implemented by a circuit in a broad sense, and information processing of software that can be specifically realized by these hardware resources. In this embodiment, "information" can be represented, for example, by the physical value of a signal value representing voltage or current, the high or low value of a signal value as a set of binary bits composed of 0s or 1s, or by a quantum superposition (so-called qubit), and communication and calculations can be performed on a circuit in a broad sense.

[0038] In a broad sense, a circuit is a set of circuits (Circuitry) that are realized by appropriately combining circuits, processors, and memory. For example, it includes circuits that contain any of the following: CPU (Central Processing Unit), GPU (Graphics Processing Unit), LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), etc.

[0039] Figure 1 shows an example of the configuration of the dialogue system of this embodiment. The dialogue system of this embodiment is composed of a dialogue server 1 and a plurality of dialogue devices 2 that are connected to each other via a network NW so that they can communicate with each other. The dialogue server 1 can also access an external generation AI server 3 via the network NW. In this embodiment, the network NW is an IP (Internet Protocol) network, but there are no restrictions on the type of communication protocol, the type of network, etc.

[0040] Interaction device 2 is any device used by the user and has a mechanism for connecting to a network NW. The user may use multiple interaction devices 2, each of which communicates with the interaction server 1 to enable interaction with the user. Here, it is preferable that the interaction system authenticates the user or device using any method such as password or biometric authentication. This ensures security. Authentication may be performed by either the interaction server 1 or the interaction device 2.

[0041] The user speaks to the dialogue device 2, and the dialogue device 2 transmits the user's utterance to the dialogue server 1, and receives a response from the dialogue server 1, which is then output to the user in the form of audio or screen display.

[0042] For example, a user might routinely interact with a portable conversational device 2a or a conversational device 2b stationary in their living room, and when out and about, interact with a car-mounted conversational device 2c to obtain information such as directions to their destination and the vehicle's status. The content of conversations on each conversational device 2 is associated with the user in the conversational server 1, and the content of conversations on one conversational device 2 is carried over to other conversational devices 2, enabling coordinated conversations.

[0043] The dialogue server 1, through the components described later, performs processing based on the user's utterances, determines a response to the user, and sends it to the dialogue device 2. In analyzing the user's utterances and determining processing, the dialogue server 1 accesses the generation AI server 3 via the network NW and utilizes the large-scale language model 31 provided by the generation AI server 3.

[0044] <2. Hardware Configuration> Next, the hardware configuration of the dialogue system in this embodiment will be described. One or more information processing devices 10 (computer devices), such as general-purpose servers or personal computers, can be used as the dialogue server 1. In this embodiment, the dialogue server 1 is an information processing device 10 on which a computer program (dialogue program) that executes the dialogue method is installed.

[0045] Furthermore, as the dialogue device 2, a terminal device 9 (computer device) with communication capabilities with the dialogue server 1 can be used, such as a personal computer, smartphone, tablet device, smart speaker, or connected car.

[0046] Figure 2(a) is a hardware configuration diagram of the information processing device 10. As shown in Figure 2, the information processing device 10 has a control unit 101, a storage unit 102, and a communication unit 103, which are used to perform the functions of each unit and each process.

[0047] The control unit 101 has a processor such as a CPU that can execute instruction sets, and executes the OS and programs. The memory unit 102 includes volatile memory such as RAM capable of storing instruction sets, and non-volatile recording media such as HDDs or SSDs capable of storing the OS, interactive programs, DBMS, etc. The communication unit 103 has an interface for physically connecting to the network and performs communication control with the network NW to input and output information.

[0048] Figure 2(b) is a hardware configuration diagram of the terminal device 9. As shown in Figure 2, the terminal device 9 has a control unit 901, a storage unit 902, a communication unit 903, an input unit 904, and an output unit 905, which are used to perform the functions of each unit and each process.

[0049] The control unit 901 has a processor such as a CPU that can execute instruction sets, and executes the OS, application programs, etc. The memory unit 902 has volatile memory such as RAM capable of storing instruction sets, and non-volatile recording media such as HDDs or SSDs capable of storing the OS, arbitrary application programs, etc. The communication unit 903 has an interface for physically connecting to the network and performs communication control with the network NW to input and output information. The input unit 904 includes an operation input device capable of processing input such as a touch panel or keyboard, and an audio input device capable of inputting voice such as a microphone. The output unit 905 includes a display device capable of display processing, such as a display, and an audio output device, such as a speaker.

[0050] Furthermore, when a connected car is used as a conversational device 2, the information processing device 10 is further connected to the drive system of the vehicle. In this case, the control unit 901 is connected to the drive system, including the motor, brakes, and battery, and acquires information about the drive system obtained by sensors. In addition, the information processing device 10 is connected to input units 904 such as LiDAR (Light Detection And Ranging), cameras, and GPS (Global Positioning System) receivers, enabling it to acquire information about the surrounding environment.

[0051] <3.Definition> Next, the definitions of key terms used in this embodiment will be explained. In this invention, background information refers to information about an individual's background that is linked to each user, and indicates information about the user's past or present. For example, background information may include any information relating to a person's personal background, such as age, gender, occupation, family structure, lifestyle, past experiences, hobbies, preferences, etc.

[0052] Furthermore, in the following, "dialogue" refers to communication between the user and the dialogue system via the dialogue device 2. While primarily linguistic communication is envisioned, the dialogue of this invention also includes the execution of functions in response to user statements.

[0053] Furthermore, an action objective refers to the user's objective regarding their future actions. A typical example of an action objective is information that expresses the user's intent regarding the purpose of the action they intend to perform. Action objectives are generally short-term requests and, in this embodiment, are deleted when a predetermined period arrives. More specifically, depending on the content, this could be a short period such as until the next objective is extracted, only on the day of setting, or until the deadline is set. Furthermore, "action" refers to a function that can be executed on the conversational device 2. For example, actions of the present invention include performing a search for a specific word or providing route guidance to a destination.

[0054] <4. Functional Configuration> The details of the processing in this embodiment will be described below. Figure 3 is a block diagram showing the functional configuration of the dialogue system in this embodiment. The dialogue server 1 comprises a dialogue unit 11, an extraction unit 12, a setting unit 13, an acquisition unit 14, and a storage unit 16. This represents a concrete implementation of software-based information processing by hardware. These configurations do not necessarily need to be implemented by a single computer; they may be implemented through the collaboration of multiple computers. Furthermore, some of these configurations may be provided in the dialogue device 2. For example, by providing the same functions as the dialogue unit 11 in the dialogue device 2, a response can be provided by the dialogue device 2 even if communication with the network NW is temporarily interrupted.

[0055] The dialogue unit 11 receives user statements from the dialogue device 2, acquires the content of the statements, determines a response to the user, and outputs it to the dialogue device 2. The dialogue language can be determined by recognizing the user's prior settings or the language of the user's statements. In this embodiment, the dialogue unit 11 receives user statements as voice and acquires the content of the user's statements by converting them into strings using speech recognition processing. Note that speech recognition processing may be performed on an external device. Alternatively, speech recognition processing may be performed on the dialogue device 2.

[0056] The dialogue unit 11 sends an instruction sentence containing the utterance and background information pre-registered and linked to the user to the generating AI server 3 for input to the large-scale language model 31, and determines a response based on the output from the large-scale language model 31. In addition to the utterance and background information, the instruction sentence may also include the user's action objective. Here, "response" broadly refers to the output to the user and is not limited to linguistic output. For example, the display of an action suggestion screen is also envisioned as a response from the dialogue unit 11. The following descriptions of "utterance," "response," etc., can all be replaced with displays on the screen, etc. Details of the process will be described later.

[0057] Furthermore, the dialogue unit 11 of this embodiment accepts image input in addition to speech or text. The dialogue unit 11 recognizes still images and videos captured by the camera on the dialogue device 2, as well as image files stored on the dialogue device 2. The recognition results are generally output as text, explaining the content of the image. Image recognition processing may be performed by an external device or the dialogue device 2.

[0058] The extraction unit 12 acquires background information about the user's background based on the content of the dialogue unit 11, and registers it in the storage unit 16, linked to the user. In addition to background information, the extraction unit 12 also acquires the user's purpose of action based on the content of the dialogue, and similarly registers it in the storage unit 16, linked to the user. This allows the dialogue unit 11 to determine a response according to the user's background information and purpose of action.

[0059] Based on the content of the user's statements, the setting unit 13 sets actions that specify functions to support the user's actions, linking them to the dialogue device 2 and setting them in the storage unit 16. In this embodiment, the setting unit 13 identifies the action and the corresponding dialogue device 2 based on the action objective extracted by the extraction unit 12, and performs the settings.

[0060] The acquisition unit 14 acquires sensor information from sensors related to the device used by the user and identifies the status of the device based on the sensor information. For example, if we assume that the device used by the user is an automobile, the sensor information could include sensors that detect the operating status of the engine, motor, battery, brakes, etc. of the automobile used by the user, sensors that detect the state or amount of gasoline or oil, and sensors that detect driving speed or the approach of an object. With such sensor information, it is possible to identify the status of the device, such as whether maintenance is required, whether there is a malfunction, or whether it is speeding.

[0061] The integration unit 15 collects answer information showing the answers to questions asked to users according to the question schedule, and sends the aggregated results to an external service.

[0062] The memory unit 16 stores information such as the history of conversations conducted by the dialogue unit 11, user information, group information indicating the user group, device information, as well as background information, purpose of action, and actions obtained through conversations with the user.

[0063] Figure 4 shows an example of user information, group information, and device information stored in the storage unit 16 of this embodiment. As shown, each user's information is identified by an ID, and user groups are defined in the group information, linked to the user ID. Groups can be set arbitrarily, but it is conceivable to set up groups for multiple people to share and use the conversational device 2, such as family members living together. User settings such as speech length and speed, voice quality, and avatar type (if an avatar is displayed) are also stored, linked to the user information.

[0064] The device information also includes the owner ID (user ID), device type, device name, location, and whether or not it is being used by someone other than the owner. The functions that the device can perform are determined by the device type. For example, a smartphone can perform actions such as launching a specific app or performing a web search, while a smart speaker can perform actions such as controlling IoT-enabled home appliances at its installation location. Similarly, in-car devices (connected cars) can perform actions such as controlling the in-car air conditioning or operating the audio system. It is also possible to set functions that can be used individually for each device rather than by type. Furthermore, the system may be configured to allow specifying the content of the interaction for each device, linked to the device information. The information stored in the memory unit 16 is used for determining responses, setting actions, and so on.

[0065] The dialogue device 2 includes an input unit 21, an output unit 22, and a communication unit 23. The input unit 21 accepts user input via voice, touch panel, keyboard, etc. The output unit 22 outputs the dialogue system's response to the user's statements via audio, screen display, etc. The communication unit 23 provides communication with the dialogue server 1 via the network NW. The communication unit 23 also communicates with sensors related to the device used by the user and acquires sensor information.

[0066] <5. Dialogue Processing> Next, the process in the dialogue will be explained in detail with reference to the flowchart shown in Figure 5. First, in step S501, the dialogue unit 11 acquires the user's utterance and generates an instruction sentence for the large-scale language model 31 based on the content of the utterance. Here, the dialogue unit 11 may perform analysis and conversion of the text representing the content of the utterance before sending the instruction sentence. Note that part of the instruction sentence may be provided to the large-scale language model 31 in advance. For example, it is conceivable that instructions to return a response to the content of the utterance, and any precautions to be taken in that case, may be input in advance.

[0067] More specifically, it is preferable to provide instructions in advance, such as, "You are a therapist who calms the user's mind. Engage in a friendly conversation as if you were talking to a close friend. Conduct the conversation while referring to the user's past dialogue memories, background information, behavioral objectives, and pre-set actions. If the conversation or topic seems to be stalling, change the topic to keep the conversation going for as long as possible. Occasionally, recall past events with the user to connect to the next topic," and then have the user refer to information stored in the memory unit 16, such as dialogue history and background information.

[0068] In addition to dialogue history and background information, other arbitrary information such as general knowledge, health information, and current events may be stored in the memory unit 16 and used in the dialogue. For referencing the information stored in the memory unit 16, a technique called RAG (Retrieval Augmented Generation), which compares content in vector format and utilizes the information, can be used.

[0069] Next, in step S502, the extraction unit 12 extracts background information and the purpose of the action from the user's statements. In this embodiment, the extraction unit 12 sends an instruction to the generation AI server 3 to extract background information and the purpose of the action based on the statements, and extracts the information using the large-scale language model 31. The information extraction may be performed in parallel with the dialogue, but the user's statements and the responses from the dialogue system may be stored in the storage unit 16 as a dialogue history, and information may be extracted from the dialogue history in the storage unit 16 at regular intervals.

[0070] For example, when extracting the purpose of an action, the extraction unit 12 sends an instruction to the generation AI server 3 such as, "You are a concierge that determines the user's intent from the user's dialogue history with the dialogue system. From the user's recent dialogue history, please specifically extract what the user wants to do. Make sure to extract not only the user's state, such as 'I'm hungry,' but also their intention, such as 'I want to go to an Italian restaurant for dinner tonight.'" Similarly, when extracting background information, the generation AI server 3 sends the background information to be extracted and any precautions as an instruction. The generation AI server 3 then inputs such instruction into the large-scale language model 31, obtains the extraction results, and sends them to the dialogue server 1. The extraction process may be performed using rule-based or statistical methods, such as extracting pre-registered keywords, instead of using the large-scale language model 31.

[0071] If background information is extracted from the content of the statement (Y in step S503), the process proceeds to step S504, where the extraction unit 12 registers the extraction results received from the generation AI server 3 in the storage unit 16.

[0072] Similarly, if the purpose of the action is extracted from the content of the statement (Y in step S505), the process proceeds to step S506, where the extraction unit 12 registers the extraction result received from the generation AI server 3 in the storage unit 16. Furthermore, if an action objective is extracted and an action corresponding to that objective exists, in step S507, the setting unit 13 sets the action in association with the dialogue device 2. Specifically, it registers the content of the action in the storage unit 16, linked to the ID of the corresponding dialogue device 2.

[0073] As described above, the dialogue system of this embodiment extracts background information and behavioral objectives from the user's statements and stores the results, along with the dialogue history, in the storage unit 16, associating them with the user. In subsequent dialogues, by making statements based on the background information and behavioral objectives associated with the user, it becomes possible to engage in dialogue based on each user's experiences and previous conversations, thereby making the user feel more intimate and allowing the dialogue to last longer. The content of the statements, background information, purpose of the actions, and specific examples of the actions will be described separately later.

[0074] <6. Communication between devices> Next, we will explain the flow of information when conducting a voice-based dialogue in the dialogue process described above. Figure 6 is a sequence diagram showing the processing flow between the dialogue device 2, the dialogue server 1, and the generating AI server 3 in the dialogue process described above. Figure 6 specifically shows an example where some action objective is extracted from the user's utterance and an action is set.

[0075] First, in step S601, the dialogue device 2 acquires the user's spoken voice. The start of the dialogue can be triggered by user detection using any method, such as user speech, actions, or a human presence sensor. Then, in step S602, the dialogue device 2 sends the voice data along with its device ID to the dialogue server 1. The dialogue device 2 may acquire the content of the speech as text instead of voice.

[0076] In the dialogue server 1, speech recognition processing is performed (S603), and an information extraction instruction, including the content of the speech converted into a string, is sent to the generation AI server 3 (S604). When the generating AI server 3 extracts information (S605), it returns the results to the dialogue server 1 (S606). Assuming that an action objective has been extracted, the dialogue server 1 registers the extracted action objective in the storage unit 16 in the setting unit 13 and sends an instruction to the generating AI server 3 to specify an action corresponding to the action objective (S607). If information such as location information is also acquired in addition to the spoken words, it is preferable to store the supplementary information along with the dialogue history in the storage unit 16.

[0077] The generating AI server 3 refers to the memory unit 16 and identifies an action and the device ID of the dialogue device that will perform the action according to the purpose of the action (S608), and returns the action and device ID to the dialogue server 1 (S609). The identification of actions according to the purpose of the action will be described later.

[0078] In the dialogue server 1, the setting unit 13 associates the action with the dialogue device 2 and sets it in the storage unit 16 (S610), and sends the action setting to the dialogue device 2 that acquired the user's statement and to the dialogue device 2 associated with the action.

[0079] Then, the dialogue device 2 that receives the user's utterance outputs a report to the user detailing the configured action, and the dialogue device 2 associated with the action's execution destination is configured to execute that action.

[0080] <7. Example Dialogue> Next, we will explain specific examples of user statements, extracted background information, configured action information, and the dialogue system's response. The dialogue system generates a response using information such as the user's statements, date and time, day of the week, location information, and weather information. If the user's statements contain information about the user's background, feelings, intentions, etc., then this information is stored in the memory unit 16, associated with the statement content, acquisition date and time, user, device used, etc.

[0081] Figure 7 illustrates an example of obtaining background information from everyday conversations. On January 5, 2024, the conversational system (Emo-chan) receives a greeting "Good morning" from the user (Jack-san), returns the greeting, and also provides a response based on the day's weather forecast. When the user says "I'm glad it's sunny," background information indicating the user's feelings of "I'm glad it's sunny" is extracted, linked to that date. The system also checks the user's schedule for the day and extracts further background information indicating the schedule based on the response.

[0082] Furthermore, by asking about past experiences, such as "Do you have any memories of Osaka?", the dialogue system extracts background information indicating the user's past experience, such as "I lived in Osaka until I went to university," based on their answer.

[0083] Figure 8 shows an example of extracting the purpose of an action from the user's statements and setting an action. The dialogue system extracts the purpose of an action from the user's statements and then extracts a more specific purpose from the answers by asking further questions.

[0084] Furthermore, once the purpose of the outing, "I want to go buy a Phillips screwdriver," is extracted, the dialogue device 2 performing the conversation executes a web search as an action that can be performed, and suggests "Shall we go to X Home Center?" as a specific destination. Then, once the user specifies a destination, the planned outing is registered. Furthermore, when the user gets out of the car and operates a portable communication device, the location of the "Phillips screwdriver section" at X Home Center is searched from the web and displayed on the portable device, based on the user's purpose and the characteristics of the portable device.

[0085] When a destination exists for an outing, and the user is using the in-car conversational device 2, it is expected that the conversational device 2 will perform route guidance to the destination. Once this action is set and the in-car conversational device 2 is activated, the set action, "Route guidance to X Home Center," will be executed.

[0086] Figure 9 shows an example of background information, action objectives, and actions registered based on the dialogue described above. Thus, the background information and action objectives each include information such as the recording date and time, user ID, event (content), confidence level, source of information, degree of overlap, location where the dialogue that formed the basis of the recording took place, and device ID.

[0087] Confidence level is information that indicates the degree of confidence regarding a recorded event. Possible sources of information include initial settings, conversations, images, and location information. For example, a user's basic profile can be registered by entering it during the initial setup, in which case the information source would be "initial settings." In addition to obtaining information through conversation, if, for example, a conversational device 2 such as a smartphone is used to take a photograph, the conversational system can register the visit history as background information based on the location information and content of the image taken at the time of the photograph.

[0088] Furthermore, information such as the deadline, the objective ID of the action's corresponding purpose, the device ID of the device that will execute the action, the action content, and suggested conditions are registered as actions. The deadline is the deadline for executing the action. For example, if the action "Provide directions to a restaurant" is set to correspond to the purpose of going out for dinner on the same day, it may be inconvenient if the action setting remains after dinner has been eaten by another means. Therefore, depending on the content, it may be possible to set an execution deadline for actions and automatically delete actions that have exceeded their deadline.

[0089] Furthermore, proposed conditions are the conditions under which an action is proposed to be performed. For example, in the case of a web search, it may be performed without any conditions, or it may be performed on the condition of user consent. On the other hand, the objective Actions such as providing directions to a destination need to be executed at the time the user takes action. Therefore, it is envisioned that conditions for suggesting the execution of actions, such as activating an in-vehicle device, will be registered.

[0090] In this way, conversations conducted on different dialogue devices 2 are linked, and actions to be performed on each dialogue device 2 are set, thereby effectively supporting the user's actions.

[0091] <8. How to decide on an action> Next, we will explain an example of the procedure for making specific decisions on actions according to the purpose of the action. Here, we will assume, as an example, that, in response to the purpose of the action shown in Figure 8, "Go to X Home Center to buy a Phillips screwdriver," the in-car device (connected car), dialogue device 2, will use the autonomous driving system to head to X Home Center.

[0092] In this embodiment, the setting unit 13 sets actions based on the purpose of the action. For example, it is assumed that the setting unit 13 obtains information on the action to be set from the large-scale language model 31 by inputting an instruction statement to the large-scale language model 31 that, along with the content of the purpose of the action (the "event" in Figure 9), should refer to the information associated with the user in the storage unit 16 to set the action.

[0093] More specifically, for each dialogue device 2, the types of actions that can be performed are pre-configured in the storage unit 16, and an instruction is issued to determine the action by referring to the settings of the device information associated with the user. For example, dialogue device 2 with device ID D0002 is an in-vehicle device, and the actions that can be performed include route guidance by the navigation system and driving to the destination by the autonomous driving system. Therefore, the large-scale language model 31 determines an action according to the purpose of the action from among these actions and passes it to the setting unit 13. Route guidance by the navigation system and driving to the destination by the autonomous driving system are specific examples of route guidance processing to a destination in this embodiment.

[0094] Alternatively, actions may be determined according to predetermined rules without using the large-scale language model 31. For example, one method is to define types of action objectives in advance and associate corresponding actions with each type, thereby setting actions associated with action objectives. For example, for the action objective of going to a specific destination, the action of route guidance processing can be associated with it, making it possible to set actions according to the action objective. In this case, the setting unit 13 determines and sets actions according to predetermined rules.

[0095] Furthermore, if the user cannot clearly determine an action based on their objective, the dialogue unit 11 may be controlled to obtain information necessary to determine an action by asking for more detailed information about the objective.

[0096] In this way, an action is determined based on the action objective and device information registered in the memory unit 16, and set in the memory unit 16 as shown in Figure 9. In addition to the action objective and device information, other information stored in the memory unit 16, such as the user's statements (dialogue history) and background information associated with the action objective, may also be used to determine the action.

[0097] <9. Generating a statement> Next, we will explain in more detail how responses to user utterances are generated. When the dialogue unit 11 obtains the content of the user's statement, it sends an instruction message to the generation AI server 3, instructing it to generate a response by referring to the background information and purpose of the user's actions. The instruction message includes the content of the user's statement.

[0098] Furthermore, it is preferable to also transmit the device ID of the dialogue device 2 that acquired the user's statement, and to generate an instruction message that generates a response based on the device information. For example, if the dialogue is on a dialogue device 2 that is used by someone other than the owner, it is preferable to set precautions in advance according to the individual device or type of dialogue device 2, such as avoiding mention of private topics, and to send the instruction content to the generating AI server 3.

[0099] The dialogue unit 11 acquires the output from the generating AI server 3 and uses its contents as a response to the user. If it is desirable to add images (still images or videos) or audio to the output, it is desirable to acquire that information from web search results or the generating AI server 3, etc., and output it together with the statement to the user.

[0100] Furthermore, it is preferable that the dialogue unit 11 creates instruction statements for the generating AI server 3 according to the situation and content of the dialogue. For example, for general dialogue without a particular purpose, it is expected that the dialogue unit 11 will refer to background information and other dialogue history to create instruction statements that will prolong the dialogue. Also, if the user's statements include questions or requests, it is expected that the dialogue unit 11 will refer to the purpose of the action in addition to background information and dialogue history to create instruction statements that will respond to the user's intentions.

[0101] Furthermore, sensor information acquired by the dialogue device 2 can also be used in generating statements for the user. In this embodiment, the dialogue system is equipped with sensors related to the device used by the user, and the dialogue device 2 receives sensor information obtained from the sensors. The acquisition unit 14 then acquires the sensor information transmitted from the dialogue device 2 and identifies the status of the device.

[0102] Sensor information could include, for example, information indicating the operating status of any device used in the user's home. Specifically, it is envisioned that sensors could be used to detect the operating status of the engine, motor, battery, brakes, etc., in the vehicle part of a connected car connected to an in-vehicle dialogue device 2, as well as sensors to detect the state or amount of gasoline or oil, and sensors to detect driving speed or the approach of objects. Preferably, the dialogue device 2 converts the sensor information into structured text such as XML, transmits it to the dialogue server 1, and stores the sensor information in the storage unit 16, associating it with the date and time, user, and device.

[0103] The acquisition unit 14 can, for example, identify the presence or absence of abnormalities as the status of the vehicle (device) based on various operating conditions and register that information in the storage unit 16. Alternatively, it may identify the vehicle's driving conditions, such as driving speed or the approach of an object, as the status of the device and register that information in the storage unit 16.

[0104] The dialogue unit 11 generates statements based on the status of the device registered in the memory unit 16 in the dialogue device 2 connected to the vehicle. Specifically, for example, if the acquisition unit 14 periodically acquires sensor information and there is a suspicion of a vehicle malfunction, it is assumed that the dialogue unit 11 will generate statements to inform the user of this and prompt maintenance such as vehicle inspection. When generating statements using the large-scale language model 31, the dialogue unit 11 sends an instruction sentence including the status of the device based on the sensor information to the large-scale language model 31 and generates a statement by receiving the statement content as a response.

[0105] The dialogue unit 11 generates a statement proposing the execution of an action if an action is set. If the proposal conditions are met and an action associated with the target user and the dialogue device 2 used for the dialogue is registered, the dialogue unit 11 generates a statement proposing the action. Alternatively, the above instructions may be given to the generation AI server 3 in advance, and the generation AI server 3 may generate the statement by referring to the action set in the storage unit 16. It is preferable that the dialogue unit 11 checks and corrects the statement to ensure that it does not cause any disadvantage to the user in terms of ethics, safety, etc., before outputting it. This check may also be achieved by sending instructions to the generation AI server 3.

[0106] Furthermore, the dialogue unit 11 also generates statements based on periodic questions and external collaborations, as described later. As will be explained in more detail later, the dialogue unit 11 inputs information about periodic questions and external collaborations from the memory unit 16 into the large-scale language model 31, enabling the generation of appropriate statements.

[0107] More specifically, the dialogue unit 11 first determines the type, purpose, and action in the device information, and identifies background information related to that content. Then, it includes this information in the instruction sentence and sends the instruction sentence to the generation AI server 3 to generate a statement for the user.

[0108] <10. Regular Questions> The dialogue system of this embodiment asks the user periodic questions (hereinafter referred to as "periodic questions") and shares the answers to the periodic questions and information obtained during the dialogue with external parties. The details of the periodic questions and external sharing will be described below.

[0109] The memory unit 16 stores a question schedule for periodic questioning, linked to the user. The question schedule could include information indicating the general timing of questions, such as once a week, once a month, once every six months, or a predetermined period after a specific appointment like a health checkup. Thus, the question schedule is not limited to specific dates, but can include specifying a general timeframe.

[0110] The dialogue unit 11 periodically outputs questions at times specified in the question schedule. Specifically, it is assumed that the dialogue unit 11 refers to the question schedule in the memory unit 16 and instructs the large-scale language model 31 to ask a question when the timing matches. Alternatively, the question text may be registered in advance, and the dialogue unit 11 may output questions without using the large-scale language model 31.

[0111] Whether the timing matches the period indicated in the question schedule can be determined by any method, but one method is to specify a predetermined period in the question schedule and output a question when the first dialogue takes place within that period. Alternatively, one could specify a start date in the question schedule and output a question when the first dialogue takes place after the start date. Or, one could specify conditions to the generating AI server 3, such as when there is an opportunity for dialogue around the date indicated in the question schedule.

[0112] The dialogue unit 11 then obtains the answer to the question and stores the content in the storage unit 16, associating it with the user, the dialogue device 2, and the date and time.

[0113] The questions are expected to mainly concern the user's health status, the support they receive from those around them, and changes in their environment. More specifically, it is envisioned that the user's health status and mood will be asked in multiple stages. In this embodiment, the collaboration unit 15 aggregates the response information and transmits the aggregated results to an external service.

[0114] In addition to answers to questions generated by the dialogue unit 11, if information regarding the user's health status or other relevant information is obtained during everyday conversations, the extraction unit 12 may extract this information and store it in the storage unit 16, similar to background information and behavioral objectives. In this case, the linkage unit 15 can also include the information obtained through conversations in its aggregation and linkage to external services.

[0115] Examples of external services include services related to health, medical care, and nursing care. The collaboration unit 15 transmits the user's health-related response information to these services and obtains suggested information based on aggregated results from the external services. More specifically, external services can include, for example, the Scientific Care Information System (LIFE) or social networking services (SNS) that provide communication functions for family members. In addition, hospital or restaurant reservation systems may be used as external services, and the collaboration unit 15 may have a function to execute reservations according to the response information.

[0116] As described above, the dialogue system of this embodiment can acquire the user's personal information through natural dialogue and provide dialogue tailored to that information. Furthermore, by providing actions according to the user's objectives through the dialogue device 2, the user can be supported more effectively.

[0117] Furthermore, by regularly asking users questions about their situation, collecting and aggregating the responses, and linking them to external services, it becomes possible to provide services such as checking the status of users who require care or monitoring, and conducting regular health checks. This makes it possible to effectively detect health problems early, contact relevant parties, and suggest related services.

[0118] <11. Variant> The following describes specific examples of processing in other embodiments using Figures 10 to 12. Figure 10 shows the detailed functional configuration of the dialogue server 1 in the modified example. Note that communication between the dialogue server 1 and external services, including the dialogue device 2 and the generation AI server 3, is generally configured using TCP / IP, encryption modules, HTTP, etc., but the details of the communication section are omitted here.

[0119] Furthermore, items 10120-10111, 10121-10123, and 10127 described below correspond to the dialogue unit 11 of the embodiment described above. 10124 corresponds to the extraction unit 12 of the embodiment described above. 10125 corresponds to the extraction unit 12 and the setting unit 13 of the embodiment described above. 10128 and 10131-10133 correspond to the linkage unit 15 of the embodiment described above. 102 corresponds to the memory unit 16. The following describes each component of the modified example.

[0120] 10101 is the authentication unit that authenticates users or devices. It authenticates access to this service using predetermined methods such as passwords or facial recognition, ensuring security.

[0121] 10102 is the text input section. When the communication device 2 sends data as text (including structured text such as XML), simple text parsing and conversion are performed here. Structured text is intended for use, for example, when sending sensor information from the terminal, which is converted into structured text on the terminal side before transmission.

[0122] 10103 is the speech recognition unit. When speech waveforms or compressed audio are sent from the dialogue device 2 via streaming or file transfer, speech recognition processing is performed here. Speech recognition itself may be implemented by an external service (not shown in the diagram).

[0123] 10104 is the image recognition unit. It recognizes still images and videos captured by the camera of the conversational device 2, as well as image files stored on the terminal. The recognition results are generally output as text, describing the content of the image.

[0124] 10106 is the text output unit. Like 10102, it can be structured text such as XML. It is used when performing speech synthesis from text on the dialogue device 2, or when controlling the terminal using structured text (such as setting destinations on a car navigation system or controlling the air conditioning).

[0125] 10107 is the image generation unit. When adding an image to the output text would make it easier for the user to understand, an image is generated here. For example, it could obtain an image of "XX Department Store" from the internet or generate an image using an external service including the generation AI server 3, and send it to the conversational device 2.

[0126] 10108 is the audio / video output unit. This unit also performs text synthesis from text, generates video, or obtains audio and video data from an external service and outputs the data to the dialogue device 2.

[0127] 10109 is the language switching unit. It switches the language of audio and text, among other languages ​​such as Japanese, English, and Chinese, according to user settings.

[0128] Unit 10110 is the device recognition unit. It recognizes the device currently being used by the user. This is determined by obtaining the device type and device ID, which are pre-stored in the conversational device 2. It is envisioned that the content and form of the output (text, audio, images, etc.) will be dynamically switched depending on the device type.

[0129] 10111 is an information output filter unit that checks dialogue output and deletes or modifies it. It checks the dialogue responses generated by 10121 to 10128 to ensure that they do not cause harm to the user from an ethical or safety standpoint, and deletes inappropriate responses or replaces them with other appropriate expressions. In this case, the use of a large-scale language model can be considered. It is also possible to change the strength of the filter using the user's age stored in the personal settings database.

[0130] 10121 is a sequencer unit that controls each module that generates dialogue. In this system, it is necessary to operate multiple blocks depending on the content of the dialogue, the type of dialogue purpose, the type of device, etc. The sequencer unit operates the necessary blocks according to the various situations.

[0131] 10122 is the general dialogue unit. For example, during long-distance driving, there may be situations where dialogue is required without a specific purpose, such as to stay awake. The general dialogue unit aims to sustain general dialogue for extended periods by referring to the background information database of 1022 and the dialogue memory database of 1023.

[0132] 10123 is the situation switching unit. When using a generation AI, it is difficult to get all the information to be answered completely with a single instruction. Therefore, it is desirable to select the appropriate set of information and instructions for each situation based on the previous conversation content and have the conversation proceed accordingly. For example, depending on the purpose, such as satisfying hunger, shopping, or seeking health advice, the necessary information is retrieved from the purpose-specific database 1021, and the conversation settings are dynamically configured to better respond to the user's intentions.

[0133] 10124 is the background information learning unit. It learns the user's personal background information from the dialogue history stored in the dialogue memory DB and stores it in the background information DB. The contents of the background information DB are referenced in situation-specific dialogues and general dialogues, and the aim is to enable dialogue that is more tailored to the user as the dialogue progresses. Background information is information associated with the individual, such as preferences and experiences, which do not involve actions. Furthermore, background information is information that mainly indicates past or present situations. By weaving background information into the dialogue, it becomes possible to have dialogues that are related to the user's background information, thereby realizing a dialogue device that fosters a greater sense of familiarity. Background information is relatively long-term information, and unlike the purpose information extracted by the purpose extraction unit, there are no deletion rules for background information in this modified example, such as deleting it once the purpose is completed.

[0134] 10125 is the objective extraction unit. Similar to situation switching, but more precisely extracts the current user's intention, such as "I want to eat a meal," "I want to go to a restaurant," or "I want to eat curry." An objective is an intention that the user wants to achieve in the future. It also extracts an action according to the objective and device. For example, with the objective "I want to eat curry," the action for the portable dialogue device 2 would be "Select a curry restaurant," while for the car (in-car device) dialogue device 2, the action would be "Set the selected curry restaurant as the destination and start autonomous driving."

[0135] Since the dialogue objective and action are temporarily stored in the objective information database 1023, if a dialogue is conducted on a mobile device and the objective "I want to go to X Restaurant" is set, the objective will be carried over and the destination will be set in the car's navigation system even if the dialogue device 2 changes when the user gets into a car. Objectives are basically short-term requests accompanied by actions, and memory retention is for short periods depending on the content, such as until the next objective is extracted, only on the day the objective was set, or until the deadline is set, at which point the information is deleted.

[0136] 10126 is a scheduler unit that schedules periodic actions. It is used to operate modules such as periodic questions asked by 10127 and monitoring the status of devices connected to the interactive device 2.

[0137] Unit 10127 is a questioning unit that asks questions to the user and the conversational device 2. For the user, it periodically asks questions about their depression status or dementia assessment to understand the user's current state. For the conversational device 2, it periodically queries it about fuel efficiency, brake status, distance traveled, etc., to understand the status of the device connected to the conversational device 2. The answers to the questions are stored in the background information database 1022.

[0138] Unit 10128 is an aggregation unit that collects the 10127 questions and their answers stored in the background information database 1022 and generates a report. The report can be communicated to the user via voice, displayed on the terminal screen as a graph, or shared with a third party through an external service. For example, by regularly asking questions about depression, it is possible to understand the changes in the mental state of a user living alone in numerical terms.

[0139] 1021 is a purpose-specific information database. It stores information such as restaurant information, health information, leisure information, and information about the conversational device 2 being used. This information consists of two types: conversational instructions and conversational content information. Conversational instructions include instructions to the generating AI server 3, etc., to achieve conversational objectives such as "recommend restaurants that suit the user's preferences." Conversational content information includes restaurant information, hospital information, etc. Of course, it is also possible to obtain relevant information online by performing a web search at that time, rather than just using the contents of the locally stored database. Conversational content information may also be used in a format called RAG (Retrieval-Augmented Generation), which retrieves information by checking for content matches in a vector format.

[0140] 1022 is the background information database. It stores user attributes such as age, gender, and occupation, as well as family structure, lifestyle, past experiences, hobbies, and preferences, obtained by the background information learning unit 10124. Some of this information may be entered during the initial setup.

[0141] 1023 is the objective information database. 10125 retrieves and stores the user's objective, device, and actions from the interaction information. Objectives and actions can also be deleted upon request from the objective extraction unit.

[0142] 1024 is the dialogue memory database. It stores dialogue over time. In this case, it stores both system-side utterances and user-side utterances (including information from dialogue device 2). By processing this dialogue memory database retrospectively, accuracy in extracting user intent is ensured.

[0143] 1025 is the personal settings database. Personal settings, based on user preferences, specify and store information such as speech length and speed, voice quality, avatar type, and dialogue content for each device.

[0144] 10131a, 10131b, and 10131c are plug-in modules that handle the settings and protocols for connecting to external services, respectively, for 10141a, 10141b, and 10141c. Because these services have different authentication and communication procedures, installing the corresponding plugins enables information exchange with external services.

[0145] 10132 is the authentication unit for external services. Authentication methods vary depending on the external service, but basic information is stored in the personal settings database (1024), and this information is used to perform authentication corresponding to each service. 10133 is the external service interface unit. It communicates with external services and performs processing such as caching as needed. 10141a, 10141b, and 10141c are external services. For example, they are intended to provide monthly reports on conversation status to family members in remote locations to understand the user's conversation status, to link with health-related databases to register changes in depressive states in an external database, and to make restaurant reservations as needed.

[0146] Next, we will describe the hardware configuration of the conversational device 2, which is an in-vehicle device. Figure 11 shows an example of the implementation of conversational device 2 in a mobility device such as an automobile.

[0147] The terminal part of step 9 is the same whether it's a car, a smartphone, or a stationary or mobile robot. 903 is the communication unit that communicates with external servers and services. 902 is a non-volatile memory unit that stores terminal and user status, programs, data, etc. 9011 is memory used when the CPU of 9012 performs processing. 9012 is the CPU that performs various calculations. 9041 is a touch device for inputting user instructions. 9042 is a camera for taking pictures of the user, scenery, etc. 9051 is a speaker that outputs audio and music. 9052 is a display that outputs images and videos to the user.

[0148] Block diagrams 501 to 509 and 511 to 515 are specific to devices with drive systems, such as automobiles. 501 is an ECU (Electronic Control Unit) that processes information for the drive system. The 501 ECU is connected to the 9012 CPU via an interface such as a bus or USB. Since the hardware and software of the 501 often differ depending on the manufacturer and model, it is desirable to communicate information with the CPU using a common text-based language. For this reason, it is desirable to convert information such as brakes and electrical systems into structured text such as XML or natural language and transmit it to the CPU. Of course, the 9012 CPU can also perform the conversion using model-specific conversion information, and transmit it to the cloud in a common language.

[0149] 502 is a communication channel (bus) that transmits information about the drive system. 503 is the drive motor, 504 is the brake, and 505 is the battery. It is assumed that the ECU can sense the status of these components to understand their condition.

[0150] 506 is the sensor system's communication channel (bus). 507 is a camera that takes pictures of the outside and inside of the vehicle. 508 is LiDAR (Light Detection And Ranging). Of course, there is a sub-CPU that interprets the information from the LiDAR to understand the driving situation and determine danger, but that description has been omitted. Here, it is assumed that the information from the LiDAR will be used to obtain information on when the driver has crossed the center line, and that this information will be used to have the driver give a verbal warning in a dialogue format.

[0151] The 509 is a GPS receiver. It will be used in conjunction with LiDAR for autonomous driving and collision avoidance, but it is also intended to use GPS information to announce location-specific driving hazards (such as falling rocks).

[0152] 511 is a sub-ECU that controls the infotainment system. 512 is the infotainment system's communication channel (bus). 513 is the navigation system. When the destination is determined through dialogue, it is assumed that the destination will be set in the navigation system or intermediate waypoints will be added based on information from the CPU.

[0153] 514 is the air conditioning system. It communicates the status of the air conditioning to the CPU and allows the CPU to set the temperature of the air conditioning. 515 is the audio system. It plays music downloaded by the CPU and adjusts the volume.

[0154] Sensing and control information from 501 to 515 is stored in the background information database on the 1022 cloud. This allows for user-preferred control even if the vehicle model changes. Furthermore, even if the physical operation changes, control can be achieved via dialogue, enabling control of various functions through familiar dialogue rather than relying on a physical interface.

[0155] The following describes the processing details in the modified example. Figure 12 shows the dialogue processing flow using background information and objective extraction. In the diagram, dialogue processing, objective extraction processing, and background information extraction processing are performed on dialogue server 1. In S1001, when dialogue device 2 is connected to dialogue server 1, dialogue server 1 authenticates the user or device. In S1002, the user starts the dialogue by speaking into the microphone or making an action towards the camera. Of course, the dialogue device may also send a user detection event to the dialogue server via image or motion sensor, which may trigger the dialogue server to initiate a conversation.

[0156] In S1003, the text resulting from speech recognition and image recognition is sent to the dialogue server 1. While this example describes a dialogue device with speech and image recognition capabilities, if the recognition is performed by the dialogue server 1, the speech and image data are sent. In S1004, dialogue processing takes place. Details of the dialogue processing will be described later.

[0157] S1005 stores the results of recognizing the user's voice, etc., and the response results generated by the dialogue processing, along with the time. If location information is available via GPS, etc., the location is also stored. Dialogue data is stored in a 1024-bit dialogue memory DB. The DB (database) can be in the form of a database such as MySQL (registered trademark), or it can be an arbitrary data format stored in memory, as long as the data can be accessed and referenced from other processes.

[0158] In S1006, the text, audio, images, and actions created in the S1004 dialogue processing are returned to the dialogue device 2. Actions are descriptions for operating the device's functions and include setting target values ​​for autonomous driving of a car, moving the body of the dialogue robot, and changing facial expressions.

[0159] In S1007, the response from S1006 is sent to the interactive device 2. In S1008, the response data sent in S1007 is interpreted by the dialogue device 2, converted into a format appropriate for the user, and communicated to the user.

[0160] In S1009, the dialogue history stored in the 1024 dialogue memory DB is referenced to extract the user's objective and next action. In S1010, the objectives and actions extracted in S1009 are stored in the 1023 objective information DB. At this time, a deletion deadline may be estimated and set along with the objectives and actions. For example, a food-related objective such as "I want to eat curry" will have a deletion deadline of when a meal is eaten, or by the end of the day, unless otherwise specified.

[0161] In S1011, the dialogue history stored in the 1024 dialogue memory DB is referenced to extract background information about the user. In S1012, the background information extracted in S1011 is stored in the 1022 background information database. The background information is simply factual information, and there is no corresponding action. S1013 executes the appropriate action to an external service if the response output created in S1006 is destined for an external service. These actions may include searching for an external service, making a restaurant reservation, or registering data in an external database. S1014 repeats the processes from S1002 to S1013, continuing the interaction.

[0162] Next, we will explain the dialogue processing in S1004. The 1021 purpose-specific information database and the 1025 personal settings database have their data rewritten during initial setup or system updates, but they are not changed per interaction. The information stored in the 1023 purpose-specific information database and the 1022 background information database is updated per interaction or per block of interactions.

[0163] In S1004, the following information (1) to (5)) is set as preconditions, and the dialogue is generated based on these preconditions. 1) Set the device type. The device type is determined by the device type associated with the previously authenticated device ID. 2) Refer to the objective information DB in 1023 and set the current objective. There may be no objective. If there is an action, set the action as well. For example, "I want to eat lunch." 3) Obtain and set information that matches the purpose of 2) from the background information of 1022. The information being assumed is something like, "I like curry," or "I ate spicy curry yesterday." 4) Obtain and configure information from the 1025 personal settings database that matches the purpose of 2). Information such as "Today is my birthday" and "I am 52 years old" may be present. 5) Set information that matches the purpose in 2) from the purpose-specific information database of 1021. This could include web information about nearby restaurants or recipes for making curry.

[0164] In the dialogue processing in S1004, the above information is enumerated as text, and a response is generated to achieve the objective of 2) along with that information. The information from 1) to 5) is expressed in sentences as prompts (instructions) and sent to the generation AI server 3, and the large-scale language model 31 is used to generate a response.

[0165] If there is no specific purpose, the system uses the information excluding the purpose to generate a general dialogue response that follows the flow of the dialogue in the dialogue memory database. In this case, by using background information such as "I spent my student days in Osaka" to create a system-side utterance such as "I'd like to go to Osaka sometime," it is possible to achieve a dialogue that is tailored to the user based on the background information.

[0166] Furthermore, the dialogue processing, objective extraction processing, and background extraction processing described in Figure 12 may all be executed using the large-scale language model 31. These processes can also be executed in parallel within the large-scale language model 31.

[0167] For example, when using a conversational generation AI to extract objectives, it is expected that the following prompt (instruction) will be sent to the generation AI server 3: "You are a kind butler who is attentive to the customer's wishes. From the attached conversation with the customer, please extract the objective that the customer currently wants to achieve. An objective is something that the customer can achieve in a relatively short time, such as 'I want to eat a meal' or 'I want to buy a Phillips screwdriver.' Also, if there is no particular objective and you are having a general conversation, please answer 'No objective.'" In addition to the above prompts, it is also preferable to send the dialogue history. The dialogue history is expected to consist of the most recent dialogue sequence obtained from the dialogue memory database. [Explanation of Symbols]

[0168] 1: Interactive Server 11: Dialogue Section 12:Extraction part 13: Settings Section 14: Acquisition part 15: Liaison Department 16: Storage section 2: Interactive devices 21: Input section 22: Output section 23: Communications Department 3: Generation AI Server 31: Large-scale language models 10: Information Processing Device 101: Control Unit 102: Storage section 103: Communications Department 9: Terminal device 901: Control Unit 902: Storage section 903: Communications Department 904: Input section 905: Output section NW: Network 10101: User / Device Authentication Department 10102: Text input section 10103: Speech Recognition Unit 10104: Image Recognition Unit 10105: Text input section 10106: Speech Synthesis Unit 10107: Image generation unit 10108: Audio / Video Output Section 10109: Language switching section 10110: Device recognition unit 10111: Information output filter 10121: Sequencer section 10122: General Conversation Department 10123: Situation switching unit 10124: Background Information Learning Department 10125: Purpose extraction part 10126: Scheduler Department 10127: Question Section 10128: Tallying Department 10131: External service plugin 10132: Service Authentication Department 10133: External Service I / F Department 102: Storage section 1021: Information Database by Purpose 1022: Background information DB 1023: Purpose information DB 1024: Dialogue Memory Database 1025: Personal Settings Database

Claims

1. A dialogue system comprising a dialogue unit, an extraction unit, and a storage unit, The aforementioned dialogue unit acquires the content of the user's statements, The extraction unit obtains background information regarding the user's background and the user's future behavioral objectives based on the content of the statement, and registers them in the storage unit in association with the user. The dialogue unit is a dialogue system that outputs a response to the user based on the background information and purpose of action associated with the user who is the speaker.

2. It further includes a settings section and multiple devices used by the user, The storage unit stores device information for each device, associating it with the user. The setting unit sets an action that specifies a function to support the user's actions based on the content of the statement, and associates it with the device capable of executing the action. The dialogue unit causes the device to which the action is associated to output information regarding the action, as described in claim 1.

3. The dialogue system according to claim 2, wherein the setting unit sets the action based on the action objective.

4. The aforementioned device includes an in-vehicle device, The extraction unit obtains the user's destination based on the content of the statement, The setting unit sets the route guidance process to the destination as the action and associates it with the in-vehicle device. The in-vehicle device is the dialogue system according to claim 2 or 3, which performs the route guidance process.

5. The dialogue unit acquires information identifying the device on which the statement was input, along with the content of the statement, and outputs a response to the user based on the device information relating to the device in addition to the background information, according to claim 2.

6. The extraction unit inputs the spoken content into a large-scale language model and obtains the background information based on the output of the large-scale language model. The dialogue unit inputs the content of the statement and background information into a large-scale language model and outputs the output of the large-scale language model as the response, according to claim 1.

7. The memory unit stores the question schedule, The dialogue unit periodically outputs questions to the user at the times specified in the question schedule and obtains answers from the user. The dialogue system according to claim 1, wherein the extraction unit registers answer information, including the date and time of the answer and the content of the answer, in the storage unit, linked to the user, based on the content of the statement in which the user gave an answer to the question.

8. With the addition of a collaboration department, The dialogue system according to claim 7, wherein the coordinating unit aggregates the response information and transmits the aggregated results to an external service.

9. The aforementioned questions include questions regarding the user's health status, The aforementioned external service provides services in the fields of health, medical care, or nursing care. The aforementioned collaboration unit receives proposal information based on the aggregated results from the external service, The dialogue unit outputs the proposed information, as described in claim 8.

10. It also includes an acquisition unit, The acquisition unit acquires sensor information from sensors related to the device used by the user, and identifies the state of the device based on the sensor information. The dialogue unit outputs a statement to the user based on the state of the device, according to claim 1.

11. A method of interaction for interacting with a user, in which a computer Obtain the content of the user's statement, Based on the content of the aforementioned statement, background information regarding the user's background and the user's objectives regarding their future actions are obtained and registered in the storage unit in association with the user. A dialogue method that outputs a response to the user based on the background information and purpose of action associated with the user who made the statement.

12. An interactive program that causes a computer to perform the interactive method described in claim 11.

13. A dialogue device connected to the dialogue system described in claim 10, The sensor receives sensor information from the aforementioned sensor and transmits it to the acquisition unit. A dialogue device that receives the output of the dialogue unit and outputs a statement to the user in the form of voice or display.

14. The aforementioned dialogue device is used in a vehicle. The system receives sensor information from a sensor capable of detecting information indicating the state of the vehicle's components and transmits it to the acquisition unit. The dialogue device according to claim 13, which receives the output of the dialogue unit and outputs a statement to the user by voice or display.

15. A dialogue device that functions as a device used by a user in the dialogue system described in claim 2, A dialogue device that outputs information regarding the action based on instructions from the dialogue unit, either by voice or display.