Information processing system, information processing method, program, and information processing device
Patent Information
- Application Number
- PCT/JP2026/004127
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-14
- Filing Date
- 2026-02-05
- Publication Date
- 2026-09-17
Smart Images

Figure JP2026004127_17092026_PF_FP_ABST
Abstract
Description
Information Processing System, Information Processing Method, Program, and Information Processing Apparatus
[0001] The present disclosure relates to an information processing system, an information processing method, a program, and an information processing apparatus.
[0002] Interactive conversation robots have been put into practical use. This type of conversation robot is provided with a function of determining the content of an utterance based on a conversation log with a user. Hereinafter, a user who uses a conversation service provided by a conversation robot is also referred to as an "owner".
[0003] Japanese Patent No. 7474211
[0004] There is a conversation robot equipped with a camera. A conversation robot equipped with a camera starts capturing an image in response to an instruction from the owner, and outputs a response to the captured image. A conversation log is a record of a conversation. That is, every word that appears in a conversation is accurately recorded in a conversation log. On the other hand, human memory is incomplete, and people forget most of the content of a conversation. This memory imbalance may cause misalignment in the conversation between the conversation robot and a human.
[0005] One object of the present disclosure is to provide an information processing system, an information processing method, a program, and an information processing apparatus that allow response content to approximate human responses.
[0006] One aspect of the present disclosure is an information processing system including one or more processors, wherein the one or more processors generate a response text using content uttered by the owner, an image captured by an owner terminal, and information related to the owner, and output a voice corresponding to the generated response text from the owner terminal. One aspect of the present disclosure is an information processing apparatus that, to one or more processors, inputs a summary of conversation text stored in association with an account, and outputs a response to content uttered by a human.
[0007] According to one aspect of the present disclosure, a response close to human memory characteristics can be achieved.
[0008] This figure shows an example of the overall configuration of the conversation system assumed in Embodiment 1. This figure illustrates an example of starting a conversation from an utterance (i.e., a speech) from the conversation server. This figure illustrates an example of a personal identification table stored in the auxiliary storage device. This figure illustrates an example of data stored in the user DB. This figure illustrates an example of an owner screen used for registering owner information. This figure illustrates an example of a family registration screen. This figure illustrates an example of a family registration screen. This figure illustrates an example of data stored in the conversation log. This figure illustrates an example of data stored in the summary DB. This figure illustrates an example of summary data for each conversation unit. This figure illustrates an overview of a conversation sequence started by a user's utterance. This figure illustrates one example of a processing operation that starts a conversation by a user's utterance. This figure illustrates an example of response text output when the identified person is "Hanako". This figure illustrates another example of response text output when the identified person is "Ai-chan". This figure illustrates another example of a processing operation that starts a conversation by a user's utterance. This figure illustrates an example of owner's stored information. This figure illustrates an example of response text output when the identified person is "Owner". This figure illustrates another example of owner's stored information. This diagram illustrates another example of response text output when the identified person is "Owner". This diagram illustrates another example of processing operation initiating a conversation by user utterance. This diagram illustrates an example of an event schedule DB. This diagram illustrates an example of response text output using an event or schedule. This diagram illustrates another example of response text output using an event or schedule. This diagram illustrates another example of response text output using an event or schedule. This diagram illustrates another example of processing operation initiating a conversation by user utterance. This diagram illustrates an overview of a conversation sequence initiated by a conversation server utterance. This diagram illustrates one example of processing operation initiating a conversation by a conversation server utterance. This diagram illustrates an example of utterance text output when the identified person is "Hanako". This diagram illustrates another example of utterance text output when the person is not identified. This diagram illustrates another example of processing operation initiating a conversation by a conversation server utterance.This diagram illustrates an example of response text output when the identified person is the "owner". This diagram illustrates another example of response text output when the identified person is the "owner". This diagram illustrates another example of processing operation in which a conversation is initiated by an utterance from the conversation server. This diagram illustrates an example of utterance text output using an event or schedule. This diagram illustrates another example of response text output using an event or schedule. This diagram illustrates another example of processing operation initiated by an utterance from the conversation server. This diagram illustrates another example of processing operation in which a conversation is initiated by an utterance from the user. This diagram illustrates an example of data stored in the user DB. This diagram illustrates an example of daily summary data. This diagram illustrates an example of monthly summary data. This diagram illustrates an example of yearly summary data. This diagram illustrates the generation timing of conversation-unit, daily, and monthly summary data. This diagram illustrates the generation timing of monthly and yearly summary data. This diagram illustrates the processing sequence of the conversation service. This flowchart illustrates a detailed example of the response text generation process. This flowchart illustrates an example of a response text generation method by a generation AI-type conversation engine. This diagram illustrates the flow until summary data given to the prompt is obtained. This diagram illustrates an example of extracting summary data based on differences in utterance timing. This flowchart explains the procedure for generating summary data on a conversational basis. This flowchart explains the procedure for generating summary data on a daily basis. This flowchart explains other procedures for generating summary data on a daily basis. This diagram explains the clustering of summary data from multiple conversational units with the same conversation date. This diagram explains the re-summarization of summary data within a cluster and the decomposition of the cluster after re-summarization. This flowchart explains the procedure for generating summary data on a monthly basis. This flowchart explains the procedure for generating summary data on a yearly basis. This flowchart explains the data compression process in the summary database. This diagram illustrates an example of an operation screen displayed on a user terminal. This diagram explains the summary data confirmation screen. This diagram explains the summary data confirmation screen. This diagram explains the summary data confirmation screen.This is a diagram illustrating an example of a screen for setting topics to be excluded from conversation. This is a diagram illustrating an example of a screen for setting topics to be excluded from memory. This is a diagram illustrating an example of the overall configuration of the conversation system assumed in Embodiment 2. This is a flowchart illustrating the speech processing from the AI character.
[0009] Embodiments of this disclosure will be described below with reference to the drawings. The embodiments described below are merely examples of forms for carrying out this disclosure, and the embodiments of this disclosure are not limited to the embodiments described below. Accordingly, the technical scope of this disclosure is not limited to the scope described in the embodiments described below. For example, various changes or improvements made to the contents described in the embodiments are also included in the technical scope of this disclosure. The various functional units described below are realized through the execution of programs by a processor.
[0010] <Terminology> First, let's explain the terminology used in these embodiments. "Computer" is not limited to a processor as hardware, but also includes combinations of software programs and a processor as hardware. A computer may also include a general-purpose computer, a computer for a specific purpose, a workstation, or any other system capable of performing each of these processes.
[0011] A "processor" includes an integrated circuit that performs various processes in cooperation with a program. The processor can function as each unit and each means described in the embodiments below. The execution order of processes by the processor is not limited to the order described in the embodiments and can be changed as necessary. The processor can be configured with one or more hardware components. The types of hardware that constitute the processor are not limited to any particular type.
[0012] The processor may be, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a programmable logic device such as an FPGA (Field Programmable Gate Array), a dedicated circuit for executing specific processing such as an ASIC (Application Specific Integrated Circuit), a GPU (Graphics Processing Unit), or hardware such as an NPU (Neural Processing Unit).
[0013] A processor can consist of not only combinations of the same type of hardware, but also combinations of different types of hardware. When multiple pieces of hardware are configured to perform one or more processes of a given processor, the multiple pieces of hardware may reside in physically separate devices or in the same device. Hardware is composed of circuits, which are combinations of circuit elements such as semiconductor elements.
[0014] "Program" includes software such as the OS (Operating System), firmware, application programs, and microcode. A program may also be, for example, a group of program modules. Each function constituting the group of program modules may be implemented by a processor configured to execute each function. In each embodiment, the program may be program code or multiple code segments stored in one or more non-temporary computer-readable media (e.g., semiconductor memory, magnetic or optical storage media, or other storage).
[0015] A program may be divided and stored on multiple non-temporary computer-readable media located on devices that are physically separated from each other. Program code and multiple code segments can be represented as any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, instructions, data structures, and program statements. Program code and multiple code segments may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents. A program may be executed on multiple terminals located on devices that are physically separated from each other.
[0016] "Conversation" includes verbal or written exchanges in natural language. Written exchanges are also called "chat." Incidentally, "exchange" is also called "response." Conversation may also include dialogues, discussions, and debates. "Conversation" may also include non-natural language reactions. Reactions include, for example, the output of sound effects, light emission, changes in light emission color, changes in light emission pattern, vibration, and the movement of moving parts.
[0017] A "topic" refers to the main subject of a conversation. Topics can include, for example, events from everyday life. Conversations that discuss events from everyday life are sometimes called "small talk." However, topics do not have to be limited to events from everyday life. For example, in dialogues, discussions, and debates, a specific topic is designated in advance. Specific topics can include, for example, problems, issues, solutions, likes and dislikes. A "topic" can also be determined by the words included in the conversation. For example, if a user says, "I went to a friend's wedding in my hometown today," possible topics could include "wedding," "marriage," "hometown," and "friend."
[0018] "Speaker" refers to a person participating in a conversation. In this specification, one of the participants is a natural person (hereinafter referred to as "person"). There may be one person or more people. The other participant is a virtual character based on AI (= Artificial Intelligence) technology (hereinafter also referred to as "AI character"). The actual AI character includes an application program or conversation model that understands the content of natural language uttered or input by the other party and generates appropriate response sentences. The AI character described later includes an application program or conversation model that generates appropriate question sentences without using the other party's utterance or input.
[0019] The content of what an AI character speaks is determined based on rule-based or machine learning models (hereinafter also referred to as "conversation models"). An AI character can refer to, for example, one of the objects displayed on a screen, a conversation service, or a physical AI speaker. Most physical AI speakers are used in the owner's home. A conversation model can refer to the machine learning model itself, or to a conversation service that utilizes a machine learning model.
[0020] An "AI speaker" is a device that includes a microphone that converts human speech into electrical signals and a speaker that converts text data into speech. An AI speaker may be a computer terminal such as a smartphone equipped with voice assistance functions, or it may be a dedicated device specifically designed for conversation. A dedicated device specifically designed for conversation includes, for example, a smart speaker. The external shape of a dedicated device may include, for example, a cylinder, a robot, a pet, or a fictional creature.
[0021] "Conversation text" includes both text representing human utterances recognized by speech recognition technology and text representing utterances made by AI characters. AI character utterances include utterances as responses to user utterances and utterances initiated by the AI character. Text corresponding to the user or AI character may include, for example, one or more sentences, interjections, and exclamations. Note that sentences spoken in spoken language are not necessarily grammatically correct.
[0022] A "conversation log" refers to a record of conversations linked to a user account. In other words, a conversation log is a record of the conversation text. The conversation log accurately records the content of the conversation down to the last word. However, if there are errors in speech recognition, the conversation log will record the incorrectly recognized utterances as they are.
[0023] Conversation logs, by their very nature as data that records spoken content verbatim, may be discarded at any time (for example, at midnight every day) from a privacy standpoint, or they may be stored in a database for a predetermined period. A "summary generated from a conversation log" has a single conversation as its smallest unit. Hereafter, a summary of a conversation unit will be referred to as "conversation unit summary data." In addition, "summaries generated from conversation logs" also include summaries generated from multiple summaries.
[0024] A conversational text refers to a series of interactions between a person and an AI character. Technically, a conversation is considered to be the entire exchange from the first utterance to the last utterance. The first utterance is, for example, an utterance that has a predetermined period of silence (e.g., 5 minutes or more) between it and the preceding utterance. On the other hand, the last utterance is, for example, an utterance that has a predetermined period of silence (e.g., 5 minutes or more) between it and the next utterance. Therefore, even if there is a 1-minute period of silence between an utterance by a person or AI character, the utterances before and after the 1-minute period of silence are treated as utterances within a single conversation.
[0025] The subjects of a "summary" are determined, for example, by a rule-based system. These rules may include specifying time periods (e.g., conversation units, days, months, or years). Other rules may include grouping (i.e., classifying) utterances or summaries with similar content. A conversation model is used to generate summaries. Instructions given to the conversation model are called "prompts." For example, a conversation log and prompts can be given to the conversation model to generate a summary. Alternatively, a summary generated from a conversation log and prompts can be given to the conversation model to generate a new summary. The generated summaries are also used to generate responses (i.e., response text) to human utterances.
[0026] <Embodiment 1> <System Configuration> Figure 1 is a diagram showing an example of the overall configuration of the conversation system 1 assumed in Embodiment 1. The conversation system 1 is an example of an information processing system. The conversation system 1 includes an owner terminal 10 connected via a network N, a voice / text conversion server 20, a front server 30, and a conversation server 40.
[0027] The owner terminal 10 is an example of a terminal operated by an owner using the conversation service. However, the owner terminal 10 does not need to be a dedicated device for the owner registered with the conversation service. For example, the owner terminal 10 may be shared by multiple people, including the owner. The owner terminal 10 is an example of an information processing device.
[0028] The voice / text conversion server 20, the front server 30, and the conversation server 40 are terminals (i.e., servers) on the provider side of the conversation service. Incidentally, the conversation server 40 is also an example of an information processing device. Network N includes, for example, the Internet, LAN (= Local Area Network), and mobile communication systems such as 4G and 5G. Network N may be a wired network, a wireless network, or a hybrid of wired and wireless networks.
[0029] The owner terminal 10 is a terminal that transmits voice and other information to the voice / text conversion server 20, while outputting data received from the cloud. The voice that the owner terminal 10 transmits to the voice / text conversion server 20 may include the voice of a person other than the owner. In this embodiment, the owner terminal 10 transmits all sounds detected by the microphone 15 to the voice / text conversion server 20.
[0030] The owner terminal 10 may have a function to identify the speaker. If the owner terminal 10 can identify the speaker, it sends the audio with speaker information to the voice / text conversion server 20. However, it is also possible to operate without sending the identified speaker information as supplementary information to the audio. If the speaker cannot be identified, the audio is considered to be from an unknown person.
[0031] Data related to the conversation service is transmitted and received in streaming format. That is, data related to the conversation service is sent from the owner terminal 10 to the cloud in real time, and the processing results on the cloud side are sent back to the owner terminal 10 in real time.
[0032] The data streamed from the owner terminal 10 to the cloud may include, for example, the user's voice data (i.e., audio data), output data from touch sensors and accelerometers (i.e., sensor data), and image data of the user, etc. Image data of the user, etc. may also include, for example, the user's background image.
[0033] The data streamed from the cloud to the owner terminal 10 includes, for example, audio data, as well as control data and image data. The control data includes, for example, instructions for outputting sound effects and melodies, instructions for outputting vibrations, instructions for movement to movable mechanisms, and instructions for outputting predetermined images. The image data includes, for example, images that the cloud instructs to output according to the content and progress of the conversation.
[0034] In Figure 1, the owner terminal 10 is a so-called AI speaker. As mentioned above, AI speakers include computer terminals equipped with voice assistance functions and dedicated devices specialized for conversation. In this embodiment, we assume that the owner terminal 10 is a dedicated device specialized for conversation. More specifically, we assume a robot-type owner terminal 10 with a movable head relative to the main body. In the case of Figure 1, there is one owner terminal 10 connected to the network N. However, in reality, multiple owner terminals 10 are connected to the network N.
[0035] In this embodiment, one owner terminal 10 is provided for each owner who uses the conversation service. In other words, the owner terminal 10 is linked to one user account. The user (i.e., owner) linked to the user account is the subscriber to the conversation service.
[0036] However, it is also possible for multiple people to use a single owner terminal 10 linked to a single user account. In this case, the utterances of multiple people will be recorded and linked to a single user account. If it is possible to identify the person who made the utterance, each utterance may be recorded separately for each person linked to a single user account.
[0037] It is also possible to link multiple owner devices 10 to a single user account. For example, one owner may be linked to a conversation-focused conversational robot and a computer terminal such as a smartphone. By linking multiple owner devices 10 to a single user account, it becomes possible to switch between using AI speakers at home and while out and about.
[0038] In addition, it is possible to link multiple user accounts to a single owner terminal 10. In other words, a single owner terminal 10 may be shared by multiple users. For example, if it is possible to identify or distinguish the user using the owner terminal 10 (e.g., the speaker or the person being talked to) using speech recognition technology or image recognition technology, then it is possible to share a single owner terminal 10 with multiple users. When a single owner terminal 10 is shared by multiple users, the content of the conversation is managed separately for each user.
[0039] The voice / text conversion server 20 is a server that converts voice data received from the owner terminal 10 into text. Hereinafter, the text converted by the voice / text conversion server 20 will be referred to as "spoken text". In the case of Figure 1, the spoken text is sent to the front server 30. The voice / text conversion server 20 may also send the spoken text to both the front server 30 and the conversation server 40. Sensor data and image data received from the owner terminal 10 are transferred to the front server 30 separately from the spoken text. However, sensor data and image data may also be sent directly from the owner terminal 10 to the front server 30.
[0040] The front server 30 is a server that instructs the owner terminal 10 to output simple responses and replies to input from the owner terminal 10. Simple responses and replies are expected to serve the purpose of detecting the start and end of speech. Simple responses and replies include, for example, the output of sound effects or melodies, the movement of movable parts, the output of vibrations, and changes in the screen. In the case of Figure 1, the front server 30 transfers the speech text received from the speech / text conversion server 20 to the conversation server 40 in real time.
[0041] The conversation server 40 is a server that generates responses in accordance with spoken text, etc. The conversation server 40 generates response text using not only the latest spoken text, but also past conversations (i.e., memories) related to the content of the latest spoken text. In this embodiment, the past conversation is summarized (summary DB 43C, described later) and used as past conversations. However, as past conversations, for example, past conversations within the last month may be used as the conversation log itself (conversation log 43B, described later), or both past conversations within the last month and summarized past conversations may be used. By summarizing conversations, less important information is lost, while more important information remains.
[0042] Furthermore, the summary DB43C may generate new summaries (re-summarization) from multiple summaries belonging to each predetermined period. For example, summaries generated from conversations in the same month of utterance may be re-summarized to generate monthly summaries. When re-summarizing, the importance level of each summary may be utilized. For example, the content of summaries with high importance may be more likely to be retained after re-summarization. In other words, the content of summaries with low importance may be more likely to be lost through re-summarization. As a result, the information stored in the summary DB43C will be closer to the nature of human memory, where "trivial matters are forgotten over time, and important matters are remembered (even after time has passed)."
[0043] For example, a person might forget what they had for dinner a month ago, but remember the name of a prefecture they visited on a family trip three months ago. However, if the conversation is stored mechanically, it will be impossible to forget such trivial details, which could cause the user to feel uncomfortable during the conversation. However, as mentioned above, the response text generated using the summary DB 43C is more likely to match the user's memory. The response text generated by the conversation server 40 is sent to the user terminal 10.
[0044] In addition to the above, the conversation server 40 may also have a function of generating a response to an input (e.g., sensor data or image data) from the owner terminal 10. The generated control data and image data are transmitted from the conversation server 40 to the owner terminal 10 together with the response text. The conversation server 40 is also an example of an information processing apparatus. Furthermore, the conversation server 40 according to the present embodiment may have a function of starting a conversation without the owner's utterance. That is, the conversation server 40 may have a function of generating utterance text independently of the owner's utterance and initiating a conversation with the owner. However, the content of the utterance is changed according to the situation of the conversation partner (i.e., the owner).
[0045] FIG. 2 is a diagram explaining an example in which a conversation is started from an utterance (i.e., an opening utterance) of the conversation server 40. In FIG. 2, parts corresponding to those in FIG. 1 are denoted by corresponding reference numerals. However, in the conversation system 1 shown in FIG. 2, the internal configurations of the owner terminal 10, the voice / text conversion server 20, and the front-end server 30 are omitted. When the owner terminal 10 satisfies a predetermined condition, the conversation server 40 shown in FIG. 2 generates utterance text corresponding to the predetermined condition.
[0046] The predetermined conditions include, for example, detection of a pre-registered face image, detection of a pre-registered voiceprint, detection of a pre-registered event or schedule, and detection of an image related to a past conversation. Hereinafter, conversation text transmitted to the owner terminal 10 as a response from the conversation server 40 to the owner's utterance is referred to as "response text", and conversation text when the conversation server 40 initiates a conversation is referred to as "utterance text".
[0047] <Hardware Configuration of Terminal> Hereinafter, the hardware configuration of each terminal will be described with reference to FIG. 1. <Owner Terminal> The owner terminal 10 includes a processor 11, a semiconductor memory 12, an auxiliary storage device 13, a camera 14, a microphone 15, a speaker 16, a display 17, a movable mechanism 18, and a communication interface 19. The processor 11 and each device are connected to each other via a bus or other signal lines.
[0048] The processor 11 is, for example, a CPU. The semiconductor memory 12 includes a ROM (=Read Only Memory) storing UEFI (=Unified Extensible Firmware Interface) and the like, and a RAM (=Random Access Memory) used as a work area for the processor 11. The processor 11 and the semiconductor memory 12 operate as a computer.
[0049] The auxiliary storage device 13 is, for example, a hard disk device or a semiconductor storage. The auxiliary storage device 13 stores programs and data for realizing functions as an AI speaker. The programs herein include a program for converting response text and utterance text received from a cloud side into audio data, a program for controlling output in accordance with control data received from the cloud side, and a program for displaying image data received from the cloud side.
[0050] FIG. 3 is a diagram illustrating an example of a personal identification table 13A stored in the auxiliary storage device 13. The personal identification table 13A is used when identifying a person on the side of the owner terminal 10 (see FIG. 1). The personal identification table 13A shown in FIG. 3 includes a user ID 13A1, face information 13A2, and voiceprint information 13A3. The user ID 13A1 stores, for example, management IDs (=identifiers) of registered users (e.g., an owner and the owner's family members) who use the owner terminal 10 (see FIG. 1).
[0051] In the case of Figure 3, the personal identification table 13A stores information for three people. User ID 13A1 is the same as User ID 43A1 (see Figure 4) managed by the conversation server 40 (see Figure 1). User ID 13A1 and User ID 43A1 are managed on a per-owner terminal 10 basis used for registration. Therefore, even if it is the same person, a different ID will be assigned if the owner terminal 10 used for registration is different. Note that User ID 13A1 does not have to be related to User ID 43A1. In other words, User ID 13A1 may be a local ID used only on the owner terminal 10. Face information 13A2 stores face images captured by the camera 14 (see Figure 1). Voiceprint information 13A3 stores voice and extracted features recorded by the microphone 15 (see Figure 1).
[0052] Returning to the explanation of Figure 1, the camera 14 is an imaging device, and for example, a CMOS (Complementary Metal Oxide Semiconductor) image sensor is used. The camera 14 may be integrated with the main body of the owner terminal 10, or it may be attached externally to the main body of the owner terminal 10. The camera 14 may be detachable from the owner terminal 10. The camera 14 may also be wirelessly connected to the owner terminal 10. The microphone 15 is a device that converts the voice of the user operating the owner terminal 10 into an electrical signal (i.e., voice data).
[0053] The speaker 16 is a device that converts voice data or sound data received from the conversation server 40 into air vibrations (i.e., sound). However, the voice data or sound data to be converted may be provided by various programs executed on the owner terminal 10. The display 17 is, for example, a liquid crystal display or an organic EL (= Electro Luminescence) display. The display 17 displays, for example, an image of an AI character's face. A capacitive touch sensor may be laminated on the surface of the display 17. A display 17 with this type of touch sensor integrated is called a touch panel.
[0054] The movable mechanism 18 is provided when the owner terminal 10 is a conversational robot modeled after a character or living creature. However, even if it is a conversational robot, the movable mechanism 18 may not be provided. The movable mechanism 18 includes a power source (e.g., a motor) and a drive mechanism that transmits power to realize a predetermined operation. For example, if the conversational robot consists of a torso and a head, the movable mechanism 18 realizes the movement of rotating the head up and down or left and right. Also, for example, if the conversational robot has arms and legs attached to the torso, the movable mechanism 18 realizes the movement of the arms and legs.
[0055] The communication interface 19 is a device that enables communication with a server that provides conversation services. The communication interface 19 is equipped with communication functions compatible with the network N. The owner terminal 10 may also include a touch sensor and an accelerometer. The touch sensor is used, for example, to detect operations such as stroking the conversation robot. The accelerometer is used to detect operations such as lifting the conversation robot. In addition, the owner terminal 10 may be equipped with, for example, a temperature sensor, a humidity sensor, and an illuminance sensor. The outputs of these sensors may be utilized in the conversation service.
[0056] <Voice / Text Conversion Server> The voice / text conversion server 20 includes a processor 21, semiconductor memory 22, auxiliary storage device 23, and communication interface 24. The processor 21 and each device are connected to each other via buses and other signal lines. The processor 21 is, for example, a CPU. The semiconductor memory 22 includes ROM that records UEFI, etc., and RAM used as a work area for the processor 21. The processor 21 and the semiconductor memory 22 operate as a computer.
[0057] The auxiliary storage device 23 is, for example, a hard disk drive or semiconductor storage. The auxiliary storage device 23 stores a program for converting voice data into text. The communication interface 24 is a device that enables communication with other servers that provide conversation services. The communication interface 24 is equipped with communication functions that are compatible with the network N.
[0058] <Front Server> The front server 30 includes a processor 31, semiconductor memory 32, auxiliary storage device 33, and communication interface 34. The processor 31 and each device are connected to each other via buses and other signal lines. The processor 31 is, for example, a CPU. The semiconductor memory 32 includes ROM that records UEFI, etc., and RAM used as a work area for the processor 31. The processor 31 and the semiconductor memory 32 operate as a computer.
[0059] The auxiliary storage device 33 is, for example, a hard disk drive or semiconductor storage. The auxiliary storage device 33 stores programs and data tables for outputting simple reactions and responses. Through the execution of the programs, sound effects, etc., are output from the owner terminal 10 when the start or end of speech is detected. The communication interface 34 is a device that enables communication with other servers providing conversation services and the owner terminal 10. The communication interface 34 is equipped with communication functions compatible with the network N.
[0060] <Conversation Server> The conversation server 40 includes a processor 41, semiconductor memory 42, auxiliary storage device 43, and communication interface 44. The processor 41 and each device are connected to each other via buses and other signal lines. The processor 41 is, for example, a CPU. The semiconductor memory 42 includes ROM that records UEFI, etc., and RAM used as a work area for the processor 41. The processor 41 and the semiconductor memory 42 operate as a computer.
[0061] The auxiliary storage device 43 is, for example, a hard disk drive or semiconductor storage. The auxiliary storage device 43 stores programs and various data for realizing conversational services. The programs here include, for example, a program called a conversational engine. The conversational engine includes, for example, a program that generates response text to spoken text. The conversational engine is also called a chatbot. Conversational engines are classified into, for example, function-specific, general-purpose rule-based, and AI-type types.
[0062] In this embodiment, the auxiliary storage device 43 stores multiple conversation engines. However, some or all of the conversation engines may be provided as a conversation service independent of the conversation server 40. Other programs include a program for recording conversation logs and a program for summarizing conversation content. However, a generation AI independent of the conversation server 40 may be used for generating summaries.
[0063] The auxiliary storage device 43 stores a user database (Data Base) 43A, a conversation log 43B, and a summary database 43C. The user database 43A is a database that stores information on all users (the owner and their family) who use the conversation service. In this embodiment, users other than the owner who have registered themselves are considered "the owner's family." Users other than the owner who have registered themselves are also called "family users." The conversation log 43B is a database that stores the history of conversations between the user and the AI character (i.e., conversation text). The conversation text is managed in conjunction with the user account.
[0064] The summary DB 43C is a database that stores summaries of the content of conversations (i.e., conversation text). In this embodiment, the summary DB 43C stores multiple summaries that differ in the time elapsed since the conversation (e.g., one day, one week, one month, one year). The summary DB 43C also stores vectors representing the content of the conversation, which are used to detect summaries similar to the dialogue text. In this embodiment, the vectors include, for example, the end time of the conversation, a summary of the conversation text (i.e., summary text), and importance. The dimensions of the vectors can be, for example, several hundred or more.
[0065] Importance is given as a score that represents the importance of the content of the summary. For example, the score is given as a numerical value (e.g., on a scale of 1 to 10) according to the topics included in the summary. For example, in a conversation on the day of a school sports day, the importance of the sports day would be higher than that of food. On the other hand, if someone is highly interested in health, the topic of food is likely to come up frequently in the conversation. Topics that appear frequently in conversations like this tend to be topics of high importance to the user.
[0066] To generate a score indicating importance, for example, a generation AI independent of the conversation server 40 is used. The prompts given to the generation AI include, for example, the scoring rules and the conversation text or summary to be summarized. In this embodiment, the scoring rules are set in advance by the provider of the conversation service. These scoring rules are common to all users.
[0067] However, the scoring rules may be made customizable for each user. For example, the conversation server 40 (see Figure 1) may switch the scoring rules applied to a user according to attribute information 43A4 (see Figure 4). Alternatively, the user may be able to set or switch the scoring rules applied to themselves via the owner terminal 10. Switching the applied scoring rules allows for adjustments to topics that remain after summarization or topics that appear in the response text.
[0068] The conversation text includes text representing what the user uttered (i.e., utterance text) and text representing what the AI character uttered (i.e., response text). In this embodiment, conversation text containing three or more texts is subject to summarization. However, conversations consisting of two or more texts may also be included in the summarization. In this embodiment, previously generated summary data is further summarized (i.e., re-summarized) to generate new summary data. For example, daily summary data for one month is re-summarized together to generate monthly summary data. After generating the new summary data, the previously generated summary data used for re-summarization is deleted.
[0069] The communication interface 44 is a device that enables communication with other servers providing conversation services and with the owner terminal 10. The communication interface 44 is equipped with communication functions adapted to the network N. In this embodiment, the communication interface 44 is used to transmit response text, control data, and image data to the owner terminal 10.
[0070] <Data Example 1> <User DB> Figure 39 is a diagram illustrating an example of data stored in the user DB 43A. The user DB 43A shown in Figure 39 includes user ID 43A11, user account 43A12, username 43A13, and user attributes 43A14.
[0071] User ID 43A11 is the management ID (identifier) of a user using the conversation service. In this embodiment, User ID 43A11 is assigned to user account 43A12. The contents of User ID 43A11 shown in Figure 39 are different from the contents of User ID 13A1 shown in Figure 3. In this case, a table is prepared to link the same user. Note that, as with Data Example 2 described later, the contents of User ID 43A11 may be the same as the contents of User ID 13A1 shown in Figure 3.
[0072] User account 43A12 is registration information used, for example, to identify a user. An email address is an example of registration information. Username 43A13 is the name of the user used in the conversation service. For example, the response text of the conversation service uses the information of username 43A13, such as "Hello A."
[0073] User attribute 43A14 is user attribute information, including, for example, age, gender, region of residence, language, interests, and concerns. User attribute 43A14 may be used, for example, to generate response text. For example, by including user attribute 43A14 in the prompt, it becomes possible to generate different response text for each user, even if the utterance text is the same. For example, it becomes possible to adjust the tone and manner of speaking of the response text according to age and gender.
[0074] <Data Example 2> <User DB> Figure 4 is a diagram illustrating an example of data stored in the user DB 43A. The user DB 43A shown in Figure 4 includes a user ID 43A1, a "How you address" field 43A2, a "Relationship with the owner" field 43A3, and attribute information 43A4. In addition, if necessary, the management ID, email address, password, gender, place of residence, date of birth, conversation service usage course, and other information of the owner terminal 10 are also stored.
[0075] User ID 43A1 stores the management ID of the user using the conversation service. The content of User ID 43A1 shown in Figure 4 is the same as User ID 13A1 shown in Figure 3. In this embodiment, User ID 43A1 is managed for each owner terminal 10 used to register a new user. Therefore, even for the same person, a different User ID 43A1 will be assigned if a different owner terminal 10 is used to register the new user.
[0076] However, if one owner uses multiple owner terminals 10, the user ID 13A1 may be shared across the multiple owner terminals 10 associated with that one owner. If the user ID 13A1 is shared across the multiple owner terminals 10 associated with one owner, it becomes possible to manage the content of conversations exchanged by the same person (for example, "Taro") through different owner terminals 10 as conversations with the same person. Of course, this assumes that the same person is authenticated on all owner terminals 10.
[0077] The "How to Address" field 43A2 stores the name that the conversation service uses to address the registered user. For example, the conversation service's response text uses the information in the "How to Address" field 43A2, such as "Hello Ai-chan". The "Relationship with Owner" field 43A3 stores the relationship between the registered user and the owner. In Figure 4, "Taro," whose user ID is "3432," is the owner himself (i.e., the contracted user of the conversation service). "Hanako," whose user ID is "3433," is the wife of "3432." "Ai-chan," whose user ID is "3521," is the daughter of "3432." "Hanako" and "Ai-chan" are family users of "Taro."
[0078] Attribute information 43A4 stores information representing the attributes of each user. Attribute information 43A4 can be registered or edited by, for example, the owner. However, after registration by the owner, family users may also register or edit their own attribute information. Attribute information 43A4 may also be registered by the conversation server 40. For example, the conversation server 40 may extract and store specific information (e.g., hobbies, preferences, likes and dislikes, desires, interests, concerns) that appears in the user's conversation.
[0079] For example, the conversation server 40 may store information of high importance extracted from the summary DB 43C (see Figure 1) as attribute information 43A4. The information extracted and stored by the conversation server 40 may overlap with the contents of the conversation log 43B (see Figure 1) and the summary DB 43C, but extracting it separately makes it easier to refer to when creating conversation text.
[0080] Figure 5 illustrates an example of an owner screen 500 used for registering owner information. The owner screen 500 shown in Figure 5 corresponds to the "Owner Information" screen of "MyRoom" provided by the conversation service shown in Figure 5. The owner screen 500 is displayed by accessing the conversation server 40 (see Figure 1). The owner screen 500 shown in Figure 5 includes a name field 501, an email address field 502, a password field 503, a gender field 504, a place of residence field 505, and a date of birth field 506. All or part of this information is also stored in the owner terminal 10 (see Figure 1).
[0081] Figure 6 illustrates an example of the family registration screen 510. The family registration screen 510 shown in Figure 6 corresponds to the "Family Registration" screen of "MyRoom" in the conversation service. The family registration screen 510 includes a "Remember to recognize owner" field 511, a "Remember to recognize family" field 512, and an "Add Family" button 513.
[0082] In Figure 6, the "Memory for Owner Recognition" section 511 includes a face information section 511A and a voice information section 511B. In Figure 6, both the face and voice are already stored. By tapping the face information section 511A, it is possible to delete or re-register a registered face. Similarly, by tapping the voice information section 511B, it is possible to delete or re-register a registered voice.
[0083] In Figure 6, the "Memory for Recognizing Family" section 512 includes an information update button 512A and a section 512B for registered family information. In Figure 6, two people, "Hanako" and "Ai-chan," are registered as family users. However, only their names 43A2 (see Figure 4) are registered for "Hanako" and "Ai-chan," and their facial images and voiceprints are not registered. Therefore, the camera icon and speaker icon corresponding to "Hanako" and "Ai-chan" are both displayed in gray. Incidentally, when either a face or voice is registered, the display of the corresponding icon switches from gray to color.
[0084] Furthermore, pressing the ">" button switches to the registration screen 520 (see Figure 7) corresponding to each user. Pressing the "Add Family" button 513 displays the new user registration screen. In this embodiment, up to three family users can be registered. Family users (hereinafter also referred to as "family") do not include the owner user.
[0085] Figure 7 illustrates an example of a family registration screen 520. The family registration screen 520 is used as a new user registration screen, as well as a registration screen for already registered users. The registration screen 520 shown in Figure 7 is provided with a name input field 521, a gender input field 522, a date of birth input field 523, a relationship to the owner input field 524, a face memory button 525, a voice memory button 526, a save button 527, and a "cancel" button 528.
[0086] The registration screen 520 shown in Figure 7 is an example of the registration screen for "Ai-chan," the owner's daughter. In Figure 7, the daughter's face is to be registered, so the face memory button 525 is displayed in color. On the other hand, the daughter's voice is not to be registered, so the voice memory button 526 is grayed out. When the save button 527 is tapped, the registration of the information on the screen is confirmed. If the "cancel" button 528 is pressed, the input is canceled and the user returns to, for example, the family registration screen 510 (see Figure 5).
[0087] <Conversation Log> Figure 8 illustrates an example of data stored in the conversation log 43B. The conversation log 43B shown in Figure 8 includes User ID 43B1, Conversation ID 43B2, Start Date and Time 43B3, End Date and Time 43B4, and Conversation Text 43B5. User ID 43B1 is the management ID of the user using the conversation service and is the same as User ID 43A1 (see Figure 4).
[0088] Conversation ID 43B2 is the conversation management ID. As shown in Figure 8, conversation ID 43B2 is managed on a per-user account basis (i.e., per-owner basis). However, conversation ID 43B2 may also be used to manage the conversation history on a per-user ID basis, including family users registered by the owner.
[0089] "Family user" is used to mean a user registered as a family by the owner. However, "family" here does not require a socially accepted family relationship, blood relationship, or cohabitation; it may also include friends and acquaintances. Start date and time 43B3 is information about the date and time the conversation started. For example, the date on which user ID 43B1 started a conversation with user "10001" is "February 11, 2025," and the time the conversation started is "11:03:42."
[0090] The end date and time 43B4 is information about the date and time the conversation ended. For example, the date on which the conversation with user ID 43B1, "10001", ended was "February 11, 2025", and the time the conversation started was "11:05:13". The conversation text 43B5 is a record of the conversation text. In the case of Figure 8, in order to enable speaker identification, the beginning of the text spoken by the user is marked with "Mr. A", etc., and the beginning of the text spoken by the AI character is marked with "AI character". However, in terms of data, it is sufficient as long as the recording format allows for distinction between the two.
[0091] <Summary DB> Figure 9 illustrates an example of data stored in the summary DB 43C. In Figure 9, the summary DB 43C is managed on a per-user account basis. In Figure 9, Person A, Person B, and Person C can all be owners or family users. The data structure of the summary DB 43C is common among users. In Figure 9, an example of the data in the summary DB 43C is shown, with Person A as a representative example.
[0092] The summary DB 43C shown in Figure 9 includes summary data 43C1 at the conversation level, summary data 43C2 at the daily level, summary data 43C3 at the monthly level, and summary data 43C4 at the yearly level. In other words, the summary DB 43C includes four types of summary data with different target periods. These four summary data sets 43C1, 43C2, 43C3, and 43C4 are examples of first and second summary data generated using different processing methods.
[0093] For example, if the conversation-based summary data 43C1 is considered the first summary data, then the other summary data 43C2, 43C3, and 43C4 are examples of the second summary data. In this case, the daily summary data 43C2 is an example of the third summary data in that it summarizes multiple conversation-based summary data 43C1 (i.e., multiple first summary data). The monthly summary data 43C3 is an example of the fourth summary data in that it summarizes multiple daily summary data 43C2 (i.e., multiple third summary data).
[0094] In other words, summary data 43C2, as the second summary data, is generated from first summary data where the elapsed time since the utterance satisfies a predetermined condition (for example, one month has passed). Summary data 43C4, as the second summary data, is generated from third summary data where the elapsed time since the utterance satisfies a predetermined condition (for example, two years have passed). Note that the second summary data may also be generated from first summary data where the elapsed time since the generation of first summary data satisfies a predetermined condition (for example, one month has passed since the generation of first summary data), rather than recording the utterance time of the conversation that is the subject of first summary data.
[0095] The conversation-unit summary data 43C1 is a summary of the conversation text 43B5 (see Figure 8). In this embodiment, one conversation contains three or more texts. Therefore, the conversation-unit summary data 43C1 is a summary of the content of three or more texts. In this embodiment, the summary data is recorded in bullet point format. When summary data 43C2 is generated, the used summary data 43C1 is deleted; when summary data 43C3 is generated, the used summary data 43C2 is deleted; and when summary data 43C4 is generated, the used summary data 43C3 is deleted.
[0096] Figure 10 illustrates an example of conversation-based summary data 43C1. The summary data 43C1 shown in Figure 10 includes a user ID 43C11, a conversation ID 43C12, a start date and time 43C13, an end date and time 43C14, a summary text 43C15, an importance score 43C16, and a vector value 43C17. User ID 43C11 is the management ID of the user using the conversation service and is common to user IDs 43A1 (see Figure 4) and 43B1 (see Figure 8).
[0097] Conversation ID 43C12 is the conversation management ID and is the same as conversation ID 43B2 (see Figure 8). The start date and time 43C13 and end date and time 43C14 are the same as start date and time 43B3 (see Figure 8) and end date and time 43B4 (see Figure 8), respectively. The summary text 43C15 is a summary of conversation text 43B5 (see Figure 8). For example, if conversation ID 43C12 is "A1001", the summary text 43C15 is "The weather is nice, but the wind is cold."
[0098] Importance 43C16 stores a score representing the importance of the summary's content. Incidentally, topics related to daily life are assigned a lower importance. For example, "The owner took a bath" is assigned a score of "0". On the other hand, special events are assigned a higher importance. For example, "A grandchild was born" is assigned a score of "9". Furthermore, even for topics related to daily life, scoring rules may be established so that topics that are frequent or appear many times within the target period (e.g., conversation, day, month, year) are assigned a higher importance. Vector value 43C17 stores a vector value of the summary's content. The vector value is stored as a sequence of hundreds of numbers, such as [0.37, 0.91, 0.35, ..., 0.77].
[0099] Returning to the explanation of Figure 9, the daily summary data 43C2 is data that summarizes the content of the summary data 43C1 of one or more conversation units whose end times belong to the same day. Note that the day for summary generation does not have to be from 0:00:00 to 23:59:59 on the same day. For example, the day for summary generation may be from 3:00:00 to 2:59:59 on the following day. In any case, the content of the summary data 43C1 of conversation units whose end times belong to the 24-hour period is re-summarized to obtain the daily summary data 43C2.
[0100] Figure 40 illustrates an example of daily summary data 43C2. The summary data 43C2 shown in Figure 40 includes user ID 43C21, conversation date 43C22, summary text 43C23, importance 43C24, and vector value 43C25. User ID 43C21 is the management ID of the user using the conversation service and is the same as user ID 43C11 (see Figure 10).
[0101] The conversation date 43C22 indicates the date to which the summary data corresponds. For example, the conversation date 43C22 is the date to which the end date and time 43C14 (see Figure 10) is associated. Alternatively, the date corresponding to the start date and time 43C13 (see Figure 10) can also be used. The summary text 43C23 is a summary of the content of one or more summary data 43C1 (see Figure 10) that share the conversation date 43C22. The importance score 43C24 stores a score representing the importance of the content summarized from one or more summary data 43C1 (see Figure 10) that share the conversation date 43C22. The vector value 43C25 stores the vector value of the content summarized from one or more summary data 43C1 (see Figure 10) that share the conversation date 43C22.
[0102] Returning to the explanation of Figure 9, the monthly summary data 43C3 is data that summarizes the content of one or more daily summary data 43C2 belonging to the same month. Figure 41 is a diagram illustrating an example of monthly summary data 43C3. The summary data 43C3 shown in Figure 41 includes user ID 43C31, conversation month 43C32, summary text 43C33, importance 43C34, and vector value 43C35. In the case of monthly summary data 43C3, the unit of the management period is changed to "month".
[0103] In the case of Figure 41, the summary text 43C33 for the daily summary data 43C2 in February 2025 is stored as "My grandchild was born." This summary indicates that there were many topics related to the birth of grandchildren in that month. The importance score 43C34 stores a score representing the importance of the content summarized from one or more summary data 43C2 (see Figure 40) that are common to the conversation month 43C32. The vector value 43C35 stores the vector value of the content summarized from one or more summary data 43C2 (see Figure 40) that are common to the conversation month 43C32.
[0104] Returning to the explanation of Figure 9, similarly, the annual summary data 43C4 is data that summarizes the content of one or more monthly summary data 43C3 belonging to the same year. In other words, the annual summary data 43C4 is a summary of the content of 12 months of summary data 43C3. To put it another way, the annual summary data 43C4 is also a summary of the content of 365 days of summary data 43C2.
[0105] Figure 42 illustrates an example of annual summary data 43C4. The summary data 43C4 shown in Figure 42 includes user ID 43C41, conversation year 43C42, summary text 43C43, importance 43C44, and vector value 43C45. In the case of annual summary data 43C44, the unit of the management period is changed to "year". In Figure 42, the summary text 43C43 for 2025 is recorded as "My grandchild was born in February." and "I went to Hawaii in March."
[0106] Furthermore, the summary text 43C43 for 2026 contains the statement, "In May, I fractured my right wrist and lived a life of hardship." The importance score 43C44 stores a score representing the importance of the content summarized from one or more summary data 43C3 (see Figure 41) that share the same conversation year 43C42. The vector value 43C45 stores the vector value of the content summarized from one or more summary data 43C3 (see Figure 41) that share the same conversation year 43C42.
[0107] Furthermore, the generated daily summary data 43C2, monthly summary data 43C3, and yearly summary data 43C4 do not need to be limited to one per set of summary data to be summarized; there may be multiple. For example, if a set can be classified into subsets (clusters) for each topic, summaries may be generated on a subset basis. For example, there may be two or more summary data 43C2 for person A on February 11, 2025. Similarly, there may be two or more summary data 43C3 for person A in February 2025.
[0108] Figure 43 illustrates the timing of generation of summary data at the conversation, daily, and monthly levels. In Figure 43, the summary data is referred to as "memory." This is because the information remaining during the summarization process is considered to be human memory. Humans cannot remember everything in the conversation text 43B5 (see Figure 8), and they forget unnecessary memories. The aforementioned summary data at the conversation, daily, monthly, and yearly levels are employed as mechanisms to mimic or reproduce the mechanisms of human memory. For example, on January 1, 2024, there are eight conversation-level summary data 43C1 (i.e., memories).
[0109] In the case of Figure 43, on February 1, 2024, one month after the conversation date, the eight memories from January 1, 2024 are summarized. As a result, the summary data 43C2 for January 1, 2024 is generated. In Figure 43, two summary data 43C2 are generated for each day. Note that if there are three or more topics, three or more summary data 43C2 may be generated, and if there is only one topic, one summary data 43C2 may be generated. In Figure 43, the summary data 43C2 corresponding to January 29, January 30, and January 31 are generated on March 1, but they may also be generated on February 28, the last day of February. Incidentally, in the case of a leap year, the summary data 43C2 corresponding to January 30 and January 31 may also be generated on February 29, the last day of February.
[0110] In the case of Figure 43, the monthly summary data 43C3 is generated one year after the month in question. Specifically, on January 1, 2025, the content summary data 43C3 is generated based on the daily summary data 43C2 from January 1 to January 31, 2024. In the case of Figure 43, two summary data 43C3 for January 2024 are also generated.
[0111] Figure 44 illustrates the timing of the generation of monthly and yearly summary data. In Figure 44, today (the current date) is set to January 1, 2026. In this case, the summary data 43C3 for December 2025, one month prior, is generated. Also on the same day, the summary data 43C4 for 2024, two years prior, is generated. In Figure 44, there is one summary data 43C3 for each month. Therefore, the 2024 summary data 43C4 is generated from 12 summary data 43C3s. Note that the yearly summary data may be generated 10 years after the target year, for example.
[0112] <Conversation Sequence 1> The following describes the provision of conversation services through conversation system 1 (see Figure 1). This conversation sequence is an example using camera 14 (see Figure 1).
[0113] <When a conversation is initiated by user utterance> <Overview> Figure 11 is a diagram illustrating the overview of a conversation sequence initiated by user utterance. The symbol S in the figure represents a step. Note that the conversation server 40 in Figure 11 is a collective term for the speech / text conversion server 20 (see Figure 1), the front server 30 (see Figure 1), and the conversation server 40.
[0114] Step 101: The owner (including family users) speaks to the owner terminal 10. The owner's voice is converted into an electrical signal (i.e., voice data) through the microphone of the owner terminal 10 (see Figure 1).
[0115] Step 102: The owner terminal 10 sends the voice data to the conversation server 40.
[0116] Step 103: The conversation server 40 analyzes the audio data. If the audio data does not contain an imaging instruction, the conversation server 40 starts the process of generating a response text. If the audio data contains an imaging instruction, it sends an imaging instruction to the owner terminal 10. The conversation server 40 suspends the process of generating the response text until it obtains image data from the owner terminal 10.
[0117] Step 104: Upon receiving the imaging instruction, the owner terminal 10 starts imaging with the camera 14 (see Figure 1). The owner terminal 10 also transmits the image data captured by the camera 14 to the conversation server 40. The image data may be a still image or a moving image.
[0118] Step 105: Upon receiving the image data, the conversation server 40 begins the process of generating response text. In this embodiment, the conversation server 40 generates response text using the content spoken by the owner, the received image data, and information related to the owner. Information related to the owner includes, for example, the person, gender, relationship with the owner 43A3 (see Figure 4), conversation history (e.g., conversation log 43B (see Figure 1), summary DB 43C (see Figure 1)), event information, schedule information, and attribute information 43A4 (see Figure 4). Here, "person" includes, for example, the "way of addressing" the owner set by the owner themselves (e.g., see Figures 4 to 7). The "way of addressing" the owner is used when the AI character speaks to the owner.
[0119] In other words, information related to the owner is different from image data captured by camera 14 and audio data acquired by microphone 15. Information related to the owner may include information related not only to the owner themselves but also to family users. Conversation history is one example of stored information related to the owner.
[0120] Event information and schedule information may be included in attribute information 43A4, or they may be information independent of attribute information 43A4. It is not necessary to use all the information included in the owner-related information to generate the response text; one or more of them may be used. For example, only the person, gender, relationship to the owner 43A3 (see Figure 4), conversation history, and summary DB 43C may be used. Once the response text is generated, the conversation server 40 converts the response text into speech synthesis data. The conversation server 40 then transmits the generated speech synthesis data to the owner terminal 10.
[0121] Step 106: The owner terminal 10 outputs the speech synthesis data as sound from the speaker 16 (see Figure 1). This enables the owner terminal 10 to respond to the owner's speech. In other words, a conversation begins between the owner and the owner terminal 10. From here on, the processing operations of steps 101 to 106 are repeated.
[0122] <Example of Processing Operation> The following describes an example of processing operation performed through the cooperation between the owner terminal 10 and the conversation server 40. Needless to say, the processing operation example described below is merely an example. For example, it is possible to have the conversation server 40 perform some or all of the steps performed on the owner terminal 10. Conversely, it is also possible to have the owner terminal 10 perform some or all of the steps performed on the conversation server 40.
[0123] Furthermore, the processing operation examples shown in Figures 11, 12, 15, 20, 25, and 38, which will be described later, can be combined with each other. In addition, the combinations may include one or more steps from Figures 26, 27, 30, 33, and 37, which will be described later.
[0124] <Example 1> Figure 12 illustrates one example of a processing operation that initiates a conversation based on user utterance. Note that the conversation server 40 in Figure 12 is also a collective term for the speech / text conversion server 20, the front server 30, and the conversation server 40.
[0125] Step 111: The owner terminal 10 accepts voice input. Incidentally, voice input is possible not only with the owner's voice but also with the voices of family users.
[0126] Step 112: The owner terminal 10 identifies the person who spoke. To identify the person who spoke, for example, a registered facial image or voiceprint is used. In the mechanism shown in Figure 12, it is assumed that the user's facial image or voiceprint is stored in the owner terminal 10. In the case of Figure 12, the camera 14 (see Figure 1) is constantly capturing images. For example, it captures several images per second, regardless of voice input. When voice input is received, the owner terminal 10 uses the captured images to identify the person who spoke. Note that the image capture in step 112 is used only to identify the person who spoke. After the person who spoke is identified, the captured images are immediately deleted from the semiconductor memory 12 (see Figure 1), etc.
[0127] If voice input is not being received, the captured image is immediately deleted from the semiconductor memory 12 (see Figure 1), etc. Regardless of whether voice input is detected or not, the image data at this stage is not transmitted to external parties, including the conversation server 40. This prevents images unintended by the owner from being transmitted externally.
[0128] In addition, the camera 14 may initiate image capture triggered by the detection of voice input. The voice input detection here does not involve natural language processing. Therefore, the content of the voice input is not analyzed. Person identification is performed on a person-by-person basis, for example. For example, whether it is the owner or a family user (e.g., wife, daughter) is identified using registration data. Alternatively, the person who spoke may be identified by distinguishing between the owner and others (i.e., family users). Furthermore, the person who spoke may be identified by distinguishing between registered users (including owners and family users) and unregistered individuals. The identified person's information is used as person information to generate the response text.
[0129] Step 113 The owner terminal 10 transmits the person information and voice data to the conversation server 40. As mentioned above, image data is not transmitted to the conversation server 40. This is because the image capture in step 112 is used solely for identifying the person who spoke. If using the personal identification table 13A (see Figure 3), the owner terminal 10 may transmit the user ID of the identified user to the conversation server 40.
[0130] Step 114: The conversation server 40 converts the received audio data into text. Specifically, the audio / text conversion server 20 (see Figure 1) performs the text conversion of the audio data. The audio / text conversion server 20 converts the audio data into text data using, for example, an audio analysis AI. Hereinafter, the text data converted from the audio data will be referred to as the spoken text.
[0131] Step 115 The conversation server 40 determines whether or not the spoken text contains an imaging instruction. If the spoken text contains a predetermined keyword, the conversation server 40 determines that the spoken text contains an imaging instruction. In this case, a positive result is obtained in step 115. On the other hand, if the spoken text does not contain a predetermined keyword, the conversation server 40 determines that the spoken text does not contain an imaging instruction. In this case, a negative result is obtained in step 115. The predetermined keyword includes, for example, "look" or "visible". Alternatively, the determination may be made using the spoken text converted into a regular expression. Also, if a combination of multiple predetermined keywords ("this" + "look") is included in the spoken text, it may be determined that an imaging instruction is included.
[0132] Step 116 Step 116 is executed if a positive result is obtained in step 115. The conversation server 40 obtains the latest image data from the owner terminal 10. This process corresponds to steps 103 and 104 (see Figure 11). That is, the conversation server 40 instructs the owner terminal 10 to take an image, and in response to that instruction, the owner terminal 10 sends the image data to the conversation server 40.
[0133] Step 117: The conversational server 40 passes the image data to the conversational AI and obtains the image description text (i.e., image description text). The conversational AI is an example of a generative AI. The conversational AI here uses natural language processing and machine learning models to output text that describes the content of the input image data.
[0134] Step 118: The conversation server 40 provides the conversational AI with the spoken text, image description text, and person information to generate a response text. The conversational AI used here may be the same as or different from the conversational AI used in step 117. For example, if the conversational AI in step 117 is good at handling image data, then in step 118, a conversational AI that is good at generating conversational text will be used. The conversation server 40 reads the person information corresponding to the user ID notified from the owner terminal 10 from the user DB 43A (see Figure 4). In the case of Figure 4, the person information is provided as, for example, "Taro", "Hanako", and "Ai-chan".
[0135] Step 119 Step 119 is performed if a negative result is obtained in Step 115. The conversation server 40 provides the conversational AI with the spoken text to generate a response text. Unlike Step 118, neither image description text nor person information is used.
[0136] Step 120: The conversation server 40 generates speech synthesis data from the response text. The response text here is generated in step 118 or step 119. The conversation server 40 sends the generated speech synthesis data to the owner terminal 10. This process corresponds to step 105 (see Figure 11). Step 121: The owner terminal 10 outputs speech. This process corresponds to step 106 (see Figure 11).
[0137] Figure 13 illustrates an example of response text output when the identified person is "Hanako". The spoken text is, "Take a look, what do you think?". The image description text is, "A woman with long hair is wearing a white knit hat with 'AAA' written on it," and "The woman is wearing a brown patterned knit and is facing sideways." The person information is, "The owner is Hanako." Therefore, the response text output is, "Hanako, that white knit hat looks great on you!"
[0138] In the case of Figure 13, the captured image includes the person who spoke to the owner terminal 10 (see Figure 1). Therefore, a response text is generated that reflects the person information (i.e., information about the owner) identified in step 112 (see Figure 12). For example, if the user who spoke to the owner terminal 10 (see Figure 1) is the "owner himself," then a response such as "Taro, who is the woman who looks good in a white knit hat?" or "Taro, that's a woman who looks good in a white knit hat!" will be output.
[0139] Thus, even if the spoken text and image description text are the same, different response texts will be output if the information of the person speaking is different. Note that if person information is not used to generate the response text, the content of the response text will be generated based on the spoken text and image description text. Therefore, it will be limited to general expressions such as "That white knit hat is lovely!" or "That hat suits you!"
[0140] Figure 14 illustrates another example of response text output when the identified person is "Ai-chan". The spoken text is, "Look, there's fruit!". The image description text is, "Pineapple, red grapes, white grapes, and strawberries are arranged in the basket" and "The pineapple is cut into bite-sized pieces". The person information is, "The owner is Ai-chan.". Therefore, the response text output is, "Ai-chan, there are so many different kinds of fruit!"
[0141] In the case of Figure 14, the person who spoke to the owner terminal 10 (see Figure 1) is not visible in the captured image. However, a response text is generated that reflects the person information (i.e., information about the owner) identified in step 112 (see Figure 12). For example, if the user who spoke to the owner terminal 10 (see Figure 1) is the "owner himself," a response text such as "Taro, there are so many different kinds of fruit!" will be output.
[0142] Furthermore, if the user who speaks to the owner terminal 10 (see Figure 1) is "Hanako," a response text such as "Hanako, there are so many different kinds of fruit!" will be output. In this way, even if the spoken text and image description text are the same, the content of the response text can be changed for each owner. Note that if person information is not used to generate the response text, the content of the response text will be generated based only on the spoken text and image description text. Therefore, it will be limited to general expressions such as "There are so many different kinds of fruit!"
[0143] <Example 2> Figure 15 illustrates another example of a processing operation that starts a conversation based on user utterance. Figure 15 is denoted with reference numerals corresponding to the parts that correspond to those in Figure 12. The only difference between the processing operation shown in Figure 15 and the processing operation shown in Figure 12 is steps 131 and 132. Therefore, only steps 131 and 132 will be explained below. Step 131 is inserted after step 117. Step 132 is executed instead of step 118 (see Figure 12).
[0144] Step 131 The conversation server 40 refers to the memory information of the identified person. Figure 16 is a diagram illustrating an example of the owner's memory information. In Figure 16, for illustrative purposes, the conversation log 43B (see Figure 8) and the summary DB 43C1 (see Figure 10) are integrated. Note that in Figure 16, the conversation start date and time 43B3 or end date and time 43B4 is used as the conversation date and time.
[0145] In Figure 16, we can see that the owner's "favorite fruit is strawberries," and the importance of this topic is "7." We can also see that the owner "owns a dog named Pochi," that "Pochi is 7 years old," and that "Pochi's birthday is October 6th," and the importance of this topic is "8." Furthermore, we can see that the owner "wants a white knit hat," and the importance of this topic is "3." Returning to the explanation of Figure 15.
[0146] Step 132: The conversation server 40 provides the conversational AI with the spoken text, image description text, and stored information (i.e., conversation history) to generate a response text. In this case, the conversation server 40 may include information indicating the importance level of the topic (e.g., a numerical value) as stored information. The information indicating the importance level can be used to narrow down the topics to be included in the response text. For example, if the description text of an image captured by the owner terminal 10 corresponds to multiple stored items, it becomes possible to use the topic with a higher importance level compared to other topics as information related to the owner.
[0147] Figure 17 illustrates an example of response text output when the identified person is the "owner." Figure 17 includes corresponding symbols for parts that correspond to those in Figure 13. In Figure 17, the spoken text is "Take a look, what do you think?", and the image description text is the same as in Figure 13. However, in Figure 17, the identified person's memory information is used instead of person information.
[0148] Figure 17 assumes that the owner is identified as "Hanako." Therefore, the memory information entered includes "The owner is female," "The owner wants a white knit hat," and "The owner's favorite fruit is strawberries." The memory information also includes information about the pet dog, Pochi. In step 118 (see Figure 12), only the identified person's information was used, but in step 132 (see Figure 15), the identified person's memory information is also used to generate the response text. Therefore, the response text output is "That white knit hat is lovely!", "You said you wanted one, didn't you?", and "Have you finally found something you like?".
[0149] Since the memory information of the identified person is used to generate the response text, even if the spoken text and the image description text are the same, if the memory information is different, the content of the response text will be different. For example, if "Taro's" memory information includes "My wife wants a white knit hat," then possible outputs include "That white knit hat is lovely!", "Hanako wanted one, didn't she?", and "Did you give it to her as a present?".
[0150] Figure 18 illustrates another example of the owner's memory information. In Figure 18, corresponding parts with those in Figure 16 are indicated with corresponding symbols. The difference between the conversation text 43B5 in Figure 18 and the conversation text 43B5 in Figure 16 is that the content of the first and third lines of the conversation text 43B5 in Figure 18 concerns the wife.
[0151] Therefore, the content of the summary text 43C15 has also been changed to "Taro's wife's favorite fruit is strawberries" and "Hanako wants a white knit hat." Incidentally, "Hanako" is "Taro's wife." The example of response text output mentioned above is based on the memory information shown in Figure 18. Note that if memory information is not used to generate the response text, the content of the response text will be limited to general expressions such as "That white knit hat is lovely!" or "You're a woman who looks good in hats!"
[0152] Figure 19 illustrates another example of response text output when the identified person is the "owner." Figure 19 is denoted with corresponding symbols for parts corresponding to those in Figure 14. In Figure 19, the spoken text is "Look, there's fruit," and the image description text is the same as in Figure 14. However, in Figure 19, memory information is used instead of person information. This memory information belongs to the identified person.
[0153] Figure 19 assumes that the owner is identified as "Hanako." Therefore, the memory information entered includes "The owner is female," "The owner wants a white knit hat," and "The owner's favorite fruit is strawberries." The memory information also includes information about the pet dog, Pochi. In step 118 (see Figure 12), only the identified person's information was used, but in step 132 (see Figure 15), the memory information of the identified person is also used to generate the response text.
[0154] Therefore, the response text outputs, "There are lots of different fruits!" and "There are strawberries, the owner's favorite." In other words, the fact that the person likes strawberries is included in the response text. As with the case in Figure 17, since the memory information of the identified person is used to generate the response text, even if the spoken text and the image description text are the same, if the memory information is different, the content of the response text will be different.
[0155] For example, if the memory information for "Taro" shown in Figure 18 includes "Taro's wife's favorite fruit is strawberries," then it becomes possible to output things like "There are lots of different fruits!", "There are strawberries, Hanako's favorite!", and "Hanako will be happy." Note that if memory information is not used to generate the response text, the content of the response text will be limited to general expressions such as "There are lots of different fruits!".
[0156] <Example 3> Figure 20 illustrates another example of a processing operation that starts a conversation based on user utterance. Figure 20 is denoted with reference numerals corresponding to the parts that correspond to those in Figure 12. The only difference between the processing operation shown in Figure 20 and the processing operation shown in Figure 12 is steps 141 and 142. Therefore, only steps 141 and 142 will be explained below. Step 141 is inserted after step 117. Step 142 is executed instead of step 118 (see Figure 12).
[0157] Step 141 The conversation server 40 retrieves events or schedules associated with person information. Figure 21 is a diagram illustrating an example of the event / schedule DB 43D. The event / schedule DB 43D is stored in the auxiliary storage device 43 (see Figure 1). The event / schedule DB 43D shown in Figure 21 is a dedicated DB used for managing event information or schedule information.
[0158] The event schedule DB 43D shown in Figure 21 includes conversation date and time 43D1, user ID 43D2, conversation text 43D3, event schedule information 43D4, start date and time 43D5, and end date and time 43D6. In this embodiment, event and schedule information is extracted from the conversation text 43D3 and stored. If event and schedule information is registered independently of the conversation, the conversation text 43D3 will be blank.
[0159] The conversation text 43D3 is used by the user to confirm and modify the event schedule information 43D4. The event schedule information 43D4 uses information from, for example, the summary text 43C15 (see Figure 10). The content of the event schedule information 43D4 can also be confirmed and modified by the user. If the user modifies the summary text 43C15, the corresponding event schedule information 43D4 will also be modified.
[0160] The start date and time 43D5 and end date and time 43D6 are registered if they can be extracted from the conversation text. It is also possible that only the date or only the time may be stored. In Figure 21, for events or schedules where the time cannot be specified, a time range from "0:00" to "24:00" is stored. Furthermore, if an approximate time period (e.g., AM, PM, breakfast, lunch, dinner) can be identified during the conversation, that time period is stored as the start and end time.
[0161] In Figure 21, for "Hanako," whose user ID is "3433," the time range from "12:00" to "24:00" is stored for the statement "Tomorrow I'm going to the arcade with my daughter from noon." Also in Figure 21, for "Hanako," whose user ID is "3433," the time range from "3 / 2" to "3 / 9" is stored for the statement "My hair has grown long, so I want to go to the hair salon sometime this week."
[0162] For events and schedules where the start date and time are publicly available, such as concerts, the conversation server 40 may supplement the information. The start date and time 43D5 and end date and time 43D6 can also be confirmed and modified by the user. Event information or schedule information may also be managed in attribute information 43A4 (see Figure 4).
[0163] Furthermore, the event / schedule information 43D4 may be made directly registrable, for example, from the owner screen 500 (see Figure 5). The event / schedule information 43D4 may also be linked with information from an event or schedule registration site provided by the conversation service. Additionally, the conversation server 40 may obtain user-related event and schedule information from an external server linked through a user account, etc. Returning to the explanation of Figure 20.
[0164] Step 142: The conversation server 40 provides the conversational AI with the spoken text, image description text, and event information or schedule information associated with the person information to generate a response text. For example, if the person information is "Taro," then the event information or schedule information associated with "Taro" is used. "Taro's" user ID is "3432." Therefore, if the event / schedule information DB 43D shown in Figure 21 is used, the conversational AI will be given the schedule to "go to a live concert of my favorite artist" on "March 4th."
[0165] In this example 3, the event or schedule information provided to the conversational AI can have start and end times that are either before or after the current time. Including events or schedules that took place before the current time makes it possible to generate response text associated with completed events or schedules. On the other hand, including events or schedules that will take place after the current time makes it possible to generate response text associated with planned events or schedules.
[0166] In the case of Figure 21, there is only one event or schedule related to "Taro," but if multiple events or schedules are found, all of them are provided to the conversational AI. Incidentally, the importance level 43C16 (see Figure 10) may be obtained from the attribute information 43A4 (see Figure 4) and summary text 43C15 (see Figure 10) used in the event / schedule information 43D4 and provided to conversational A1. On the other hand, if no events or schedules related to "Taro" are stored, the conversation server 40 may provide the conversational AI with only the spoken text, image description text, and person information.
[0167] Figure 22 illustrates an example of response text output using an event or schedule. In Figure 22, corresponding parts with those in Figure 14 are indicated with corresponding numerals. In Figure 22, the spoken text is "Look, there's fruit," and the image description text is the same as in Figure 14. However, in Figure 22, the registered event information used is "The owner plans to go shopping at the supermarket in front of the station."
[0168] Therefore, the response text output is, "I bought lots of fruit at the supermarket in front of the station!" This response text is possible because the conversational AI is given information that there is a plan to go shopping at the supermarket in front of the station. Note that if event schedule information 43D4 (see Figure 21) is not used to generate the response text, the content of the response text will be limited to general expressions such as, "There are lots of different kinds of fruit!"
[0169] Figure 23 illustrates another example of response text output using an event or schedule. The spoken text is "Look." The image description texts are "A girl with her hair in two pigtails is holding a large stuffed dog" and "There is a crane game machine in the background."
[0170] Figure 23 assumes that the identified person is "Hanako". The registered event information for "Hanako" is "The owner will go out to play with a friend and her daughter on March 5th". Therefore, the response text output is, "Is that your friend's daughter? She's holding a stuffed animal! Where are you playing?"
[0171] Figure 24 illustrates another example of response text output using an event or schedule. In Figure 24, corresponding parts with those in Figure 23 are indicated by corresponding reference numerals. Note that the owner terminal 10 in Figure 24 (see Figure 1) is assumed to be a portable device such as a smartphone. In Figure 24 as well, the spoken text is "Look," and the image description text is "A girl with her hair tied in two pigtails is holding a large stuffed dog," and "There is a crane game machine in the background."
[0172] The registered event information is "The owner will go to the arcade on March 5th." Therefore, the response text outputs "Did you go to the arcade?", "I'm glad you got the stuffed animal!", and "Who is the girl in the picture with you?". The image used in Figure 24 is the same as in Figure 23, but the content of the registered event information is different. Consequently, even if the spoken text, the captured image, and the image description text are the same, the content of the response text is different.
[0173] <Example 4> Figure 25 illustrates another example of a processing operation that starts a conversation based on a user's utterance. Figure 25 is denoted with reference numerals corresponding to the parts that correspond to those in Figure 20. Due to space limitations, Figure 25 only shows the processing operation on the conversation server 40 side. Therefore, the operation in step 114 starts when person information and voice data are received from step 113 (see Figure 20).
[0174] The only difference between the processing operation shown in Figure 25 and the processing operation shown in Figure 20 is steps 151 and 152. Therefore, only steps 151 and 152 will be explained below. Step 151 is inserted after step 141. Step 152 is executed if a negative result is obtained in step 151.
[0175] Step 151: The conversation server 40, having obtained an event or schedule associated with the person information in Step 141, determines whether there is an event or schedule that corresponds to the current time. "Corresponding to the current time" means, for example, that it has been executed or is scheduled to be executed within a predetermined period including the current date and time. The predetermined period includes, for example, one week, one day, 12 hours, 6 hours, 2 hours, 1 hour, etc.
[0176] The designated time may be set according to the content of the event or schedule. For example, for topics of interest such as concerts or trips, the time may be set over a period of several days. On the other hand, for items with specific dates and times, such as shopping, the time may be set to within one day. Furthermore, the designated time may be changed depending on whether the event has already been completed or is scheduled to be completed. In this case as well, the designated time may not be uniform regardless of the content of the event or schedule, but rather may be set according to the content of the event or schedule.
[0177] If the event or schedule obtained in step 141 corresponds to the current time, a positive result is obtained in step 151. In this case, the conversation server 40 proceeds to step 142. On the other hand, if the event or schedule obtained in step 141 does not correspond to the current time, a negative result is obtained in step 151. In this case, the conversation server 40 proceeds to step 152.
[0178] By the way, in step 151, it may be determined whether or not the event or schedule is used in a conversation that includes the current time. Here, "conversation that includes the current time" refers to a conversation within a predetermined period that includes the current date and time. By providing this determination function, it becomes possible to change the content of the response text even if the same spoken text or image description text is given to the conversational AI. As a result, the problem of the same event or schedule being repeated can be reduced. Note that the additional determination may be performed only if an event or schedule corresponding to the current time is found. If a positive result is obtained in the additional determination, proceed to step 142; if a negative result is obtained in the additional determination, proceed to step 152.
[0179] Step 152 The conversation server 40 provides the conversational AI with the spoken text and image description text to generate a response text. In this case, the owner's or family user's events or schedule are not taken into consideration, so the response content will be as follows: For example, in the example in Figure 22, it will be something like "There are lots of different fruits!" and will not include information about "supermarket in front of the station" or "shopping". For example, in the example in Figure 23, it will be something like "That's a big stuffed animal." and will not include information about "friend's daughter" or "playing". For example, in the example in Figure 24, it will be something like "That's a big stuffed animal." and will not include information about "game center".
[0180] <Summary> According to the aforementioned conversation system 1 (see Figure 1), even if the content of the utterance and the image captured by the owner terminal 10 are the same, different responses can be output by reflecting the speaking user's personal information, memory information, event information, schedule information, etc.
[0181] <When a conversation is initiated by an utterance from the conversation server> <Overview> Figure 26 is a diagram illustrating the overview of a conversation sequence initiated by an utterance from the conversation server 40. In Figure 26, the conversation server 40 is a collective term for the speech / text conversion server 20 (see Figure 1), the front server 30 (see Figure 1), and the conversation server 40.
[0182] Step 201: In the case of Figure 26, the owner terminal 10 is constantly capturing images with the camera 14 (see Figure 1). For example, regardless of the owner's speech or instructions, the camera 14 captures several images per second. However, the owner terminal 10 does not upload the captured images to the conversation server 40. In this embodiment, the owner terminal 10 determines whether or not the captured images contain the face image of a person who has been registered in advance. In other words, the owner terminal 10 is constantly performing face recognition.
[0183] For example, if a pre-registered face image is included in the captured image, the owner terminal 10 sends a message to the conversation server 40 indicating that face recognition has occurred. This message may also include information about the recognized person. However, the message to the conversation server 40 may only be sent if a specific face image from the pre-registered face images is captured. For example, the system may be configured to send a message indicating face recognition if the owner's face image is captured, but not if a family user's face image is recognized.
[0184] On the other hand, if a pre-registered face image is not included in the captured image, the owner terminal 10 does not send the face recognition to the conversation server 40. Furthermore, regardless of whether a pre-registered face image is included in the captured image or not, the owner terminal 10 deletes the image used for face recognition from the semiconductor memory 12 (see Figure 1), etc. This prevents images unintended by the owner from being transmitted externally.
[0185] Step 202: When the conversation server 40 receives notification of face recognition, it sends an image capture instruction to the owner terminal 10. Alternatively, the conversation server 40 may be configured to send the image capture instruction only when a specific person's face image is recognized.
[0186] For example, if the facial image recognized is that of the owner, an imaging command is sent to the owner terminal 10, but if the facial image recognized is that of a family user, the imaging command is not sent to the owner terminal 10. In other words, the system can control whether or not to send an imaging command to the owner terminal 10 based on each user's registration information. By adopting this mechanism, the conversation service can be provided only to users who wish to receive spontaneous speech from the conversation server 40.
[0187] Step 203: Upon receiving the imaging instruction, the owner terminal 10 starts imaging with the camera 14 (see Figure 1). The owner terminal 10 also transmits the image data captured by the camera 14 to the conversation server 40. The image data may be a still image or a moving image.
[0188] Step 204: Upon receiving the image data, the conversation server 40 begins the process of generating response text. In this embodiment, the conversation server 40 generates spoken text using the captured image, regardless of the owner's instructions. In other words, the conversation server 40 generates spoken text independently of the owner's utterances. Once the spoken text is generated, the conversation server 40 converts the spoken text into speech synthesis data. The conversation server 40 then transmits the generated speech synthesis data to the owner terminal 10.
[0189] Step 205: The owner terminal 10 outputs the speech synthesis data as sound from the speaker 16 (see Figure 1). This enables the owner terminal 10 to spontaneously initiate a conversation with the user. After this, if the user responds, a conversation is established between the owner terminal 10 and the user. From here on, the processing operations in steps 101 to 106 (see Figure 11) are repeated.
[0190] <Example of Processing Operation> The following describes an example of processing operation performed through the cooperation between the owner terminal 10 and the conversation server 40. Needless to say, the processing operation example described below is merely an example. For example, it is possible to have the conversation server 40 perform some or all of the steps performed on the owner terminal 10. Conversely, it is also possible to have the owner terminal 10 perform some or all of the steps performed on the conversation server 40.
[0191] Furthermore, the processing operation examples in Figures 27, 30, 33, and 37, which will be described later, can be combined with each other. In addition, the combinations may include one or more steps from Figures 11, 12, 15, 20, 25, and Figure 38, which will be described later.
[0192] <Example 1> Figure 27 illustrates one example of a processing operation in which a conversation is initiated by an utterance from the conversation server 40. Note that in Figure 27, the conversation server 40 is also a collective term for the speech / text conversion server 20, the front server 30, and the conversation server 40.
[0193] Step 211: The owner terminal 10 determines whether the facial recognition meets predetermined conditions. As mentioned above, one of the predetermined conditions is the detection of an image containing the face image of a person registered as the owner. If the facial recognition meets the predetermined conditions, a positive result is obtained in step 211. In this case, the owner terminal 10 proceeds to step 212. If the facial recognition does not meet the predetermined conditions, a negative result is obtained in step 211. In this case, the owner terminal 10 repeats the determination in step 211. If voice input is received during this determination (for example, in step 111 in Figure 12), the owner terminal 10 proceeds to step 112 (see Figure 12).
[0194] Step 212: The owner terminal 10 sends a message to the conversation server 40 indicating that facial recognition has been performed. Step 212 corresponds to step 201 (see Figure 26). The message indicating facial recognition may also include information about the user (including the owner) who was authenticated from an image including a facial image.
[0195] For user authentication, voice other than utterances directed to the owner terminal 10 may be used. Voice other than utterances directed to the owner terminal 10 includes, for example, utterances that do not include the nickname assigned to the owner terminal 10 (e.g., Romi). This type of utterance may include sighs such as "phew" or monologues such as "I'm tired." The facial recognition message does not need to include information about the authenticated user (including the owner). This is because the transmission in step 212 is merely a trigger to initiate utterances from the conversation server 40, and it is possible to operate the system so that the content of the utterance text is entirely assumed to be that of the owner.
[0196] Step 213: The conversation server 40 obtains the latest image data from the owner terminal 10. This process corresponds to steps 202 and 203 (see Figure 26). The conversation server 40 may authenticate the user who is captured in the image obtained from the owner terminal 10. Similarly, the user may be authenticated using the voice obtained from the owner terminal 10. However, in step 213, user authentication is optional.
[0197] Step 214: The conversational server 40 passes the image data to the conversational AI and obtains the image description text (i.e., image description text). The conversational AI here is also an example of a generative AI. The conversational AI here uses natural language processing and machine learning models to output text that describes the content of the input image data.
[0198] Step 215: The conversation server 40 provides the image description text to the conversational AI to generate spoken text. The conversational AI used here may be the same as the one used in step 214, or a different one. For example, if the conversational AI in step 214 is skilled at handling image data, then in step 215, a conversational AI skilled at generating conversational text will be used. The spoken text here may be based on the owner (e.g., "Taro"), regardless of the authenticated person. However, the conversational AI may also be given identified person information to generate the spoken text.
[0199] Step 216: The conversation server 40 generates speech synthesis data from the spoken text. The conversation server 40 sends the generated speech synthesis data to the owner terminal 10. This process corresponds to step 204 (see Figure 26). Step 217: The owner terminal 10 outputs the audio. This process corresponds to step 205 (see Figure 26).
[0200] Figure 28 illustrates an example of speech text output when the identified person is "Hanako". In Figure 28, there is no spoken text; only image description text and person information are provided to the conversational AI.
[0201] Incidentally, the image description text is: "A woman with long hair is wearing a white knit hat with 'AAA' written on it." and "The woman is wearing a brown patterned knit and is facing sideways." Also, the person information is: "The owner is Hanako." Therefore, the output text is: "Hanako, that white knit hat looks great!"
[0202] If only image description text is provided to the interactive AI, the generated speech text will be indistinguishable from a person's, such as "That white knit hat is lovely." In any case, the owner terminal 10 will be able to initiate a conversation with the user in front of it. If the user responds to this conversation with something like, "Really? I liked it so I bought it!" or "It's cold outside," a conversation between the user and the owner terminal 10 will begin. The owner terminal 10's response to the user's response will follow the conversation sequence shown in Figure 11.
[0203] Figure 29 illustrates another example of speech text output when a person is not identified. In Figure 29, there is no spoken text. Also in Figure 29, only image description text is provided to the conversational AI. The image description text is, "Pineapple, red grapes, white grapes, and strawberries are arranged in the basket" and "The pineapple has been cut into bite-sized pieces." Therefore, the spoken text output is, "There are so many different kinds of fruit!" In this case as well, a conversation between the user and the owner terminal 10 can be initiated by a spoken message from the owner terminal 10.
[0204] <Example 2> Figure 30 illustrates another example of processing operation in which a conversation is initiated by an utterance from the conversation server 40. Figure 30 is denoted with reference numerals corresponding to the parts that correspond to those in Figure 27. The only difference between the processing operation shown in Figure 30 and the processing operation shown in Figure 27 is steps 221 and 222. Therefore, only steps 221 and 222 will be explained below. Step 221 is inserted after step 214. Step 222 is executed instead of step 215.
[0205] Step 221 The conversation server 40 refers to the owner's stored information (see, for example, Figure 16). In Figure 16, it is found that the owner's "favorite fruit is strawberries" and the importance of this topic is "7". It is also found that the owner "owns a dog named Pochi", "Pochi is 7 years old", "Pochi's birthday is October 6th", and the importance of this topic is "8".
[0206] Furthermore, the owner "wants a white knit hat," indicating that the importance of this topic is "3." If the user is identified through the owner terminal 10, the identified user's memory information may be referenced. Needless to say, referencing the identified user's memory information reduces topic mismatches.
[0207] Step 222 The conversation server 40 provides the interactive AI with the image description text and the stored information (i.e., the conversation history) to generate the utterance text. In this case, the conversation server 40 may include information indicating the importance level of the topic (e.g., a numerical value) as stored information. The information indicating the importance level can be used to narrow down the topics to be included in the response text. For example, if the description text of an image captured by the owner terminal 10 corresponds to multiple stored items, it becomes possible to use the topic with a higher importance level compared to other topics as information related to the owner.
[0208] Figure 31 illustrates an example of response text output when the identified person is the "owner." Figure 31 includes corresponding symbols for parts corresponding to those in Figure 28. In Figure 31, there is no spoken text; only image description text and the owner's memory information are provided to the conversational AI. In Figure 31, the image description text is: "A woman with long hair is wearing a white knit hat with 'AAA' written on it." and "The woman is wearing a brown patterned knit and is facing sideways." In Figure 31, the memory information includes: "The owner is a woman," "The owner wants a white knit hat," and "The owner's favorite fruit is strawberries." The memory information also includes information about the pet dog, Pochi.
[0209] Therefore, the output text of the utterance reads, "Isn't that the white knit hat you wanted?" and "It's lovely!" In Example 1, the conversation was based on information that could be read from the image, but in Example 2, the utterance reflects the memory information. Specifically, "wanted" is added. As a result, the owner experiences the feeling of being spoken to by an acquaintance who knows them well.
[0210] If there is no memory information related to the image description text, an utterance that can be identified from the appearance, such as "Going out? That knit hat is lovely," will be uttered. Incidentally, the conversation server 40 may provide the conversational AI with memory information (including conversations initiated by the user) that has not been used within a predetermined period, including the current date and time. This function prevents the same conversation based on the same memory information from being repeated multiple times. For example, if the owner terminal 10 says "Isn't that the white knit hat you wanted?" and "It's lovely!" when the user goes out, the AI can avoid repeating the same lines "Isn't that the white knit hat you wanted?" and "It's lovely!" when the user returns home. As a result, a more human-like conversation can be achieved.
[0211] Figure 32 illustrates another example of response text output when the identified person is the "owner." Figure 32 is denoted with corresponding symbols for parts corresponding to those in Figure 29. In Figure 32, there is no spoken text; only image description text and the owner's stored information are provided to the conversational AI.
[0212] In the case of Figure 32, the image description text is "Pineapple, red grapes, white grapes, and strawberries are arranged in a basket" and "The pineapple is cut into bite-sized pieces." Meanwhile, the memory information includes "The owner is a woman," "The owner wants a white knit hat," and "The owner's favorite fruit is strawberries." The memory information also includes information about the pet dog, Pochi. Therefore, in the spoken text, the focus is placed on the strawberries, which are the owner's favorite fruit, among the pineapple, red grapes, white grapes, and strawberries that appear in the image description text, resulting in a more human-like conversation such as "There are so many strawberries! What's that?"
[0213] <Example 3> Figure 33 illustrates another example of processing operation in which a conversation is initiated by an utterance from the conversation server 40. Figure 33 is denoted with reference numerals corresponding to the parts that correspond to those in Figure 27. The only difference between the processing operation shown in Figure 33 and the processing operation shown in Figure 27 is steps 231 and 232. Therefore, only steps 231 and 232 will be explained below. Step 231 is inserted after step 214. Step 232 is executed instead of step 215.
[0214] Step 231: The conversation server 40 retrieves events or schedules associated with person information. Person information refers to information about the person to be spoken to. The person to be spoken to is assumed to be the owner or a user whose face was recognized in step 212. For example, events or schedules of the person to be spoken to are retrieved from the event / schedule DB 43D (see Figure 21).
[0215] Step 232: The conversation server 40 provides the interactive AI with image description text and event information or schedule information associated with the person information to generate spoken text. For example, if the person information is "Taro," then the event information or schedule information associated with "Taro" will be used. In this example 3, the event information or schedule information provided to the interactive AI can have a start date and time or an end date and time that is before or after the current date and time.
[0216] Figure 34 illustrates an example of speech text output using an event or schedule. In Figure 34, corresponding parts with those in Figure 29 are indicated with corresponding numerals. In Figure 34, there is no spoken text; only image description text and registered event information are provided to the conversational AI. In Figure 34, the image description text is "Pineapple, red grapes, white grapes, and strawberries are arranged in the basket" and "The pineapple has been cut into bite-sized pieces."
[0217] By the way, in the case of Figure 34, the conversational AI is given the registered event information that "the owner plans to go shopping at the supermarket in front of the station." Therefore, the output text is "I bought lots of fruit at the supermarket in front of the station!" This output text is possible because the conversational AI is given that the owner speaking to it plans to go shopping at the supermarket in front of the station. Note that if event schedule information 43D4 (see Figure 21) is not used to generate the output text, the content of the output text will be limited to general expressions such as "There are lots of different kinds of fruit!"
[0218] Figure 35 illustrates another example of response text output using an event or schedule. In Figure 35, the owner terminal 10 (see Figure 1) is assumed to be a portable device such as a smartphone. The image description text is, "A girl with her hair tied in two pigtails is holding a large stuffed dog," and "There is a crane game machine in the background." The registered event information is, "The owner will go out with a friend and his daughter on March 5th."
[0219] Therefore, the output text reads, "Is that your friend's daughter? That's a big stuffed animal!" This type of conversational text is made possible when the interactive AI is given event information, such as going out to play with a friend and their daughter. In other words, it becomes possible to speak in a way that is similar to a human speaking to you. However, if the registered event information is not provided to the interactive AI, the output text may end up being something like, "The girl is holding a big stuffed animal," or "Ai-chan is holding a big stuffed animal," which could lead to a misunderstanding of the friend's daughter and the owner's daughter.
[0220] Figure 36 illustrates another example of response text output using an event or schedule. Figure 36 is denoted with corresponding reference numerals for parts corresponding to those in Figure 35. In Figure 36 as well, the image description text is "A girl with her hair tied in two pigtails is holding a large stuffed dog" and "There is a crane game machine in the background."
[0221] However, in the case of Figure 36, the registered event information is "The owner will go to the arcade on March 5th." In the event information in Figure 36, destination information is included, but information about who the playmates are is missing. Therefore, the output text is "Looks like you're having fun at the arcade!" In other words, even if the same image description text is given to the conversational AI, the output content will be different.
[0222] <Example 4> Figure 37 illustrates another example of a processing operation initiated by an utterance from the conversation server 40. Figure 37 is denoted with reference numerals corresponding to the parts that correspond to those in Figure 33. The only difference between the processing operation shown in Figure 37 and the processing operation shown in Figure 33 is steps 241 and 242. Therefore, only steps 241 and 242 will be explained below. Step 241 is inserted after step 231. Step 242 is executed if a negative result is obtained in step 231.
[0223] Step 241: The conversation server 40, having obtained an event or schedule associated with the person information in Step 231, determines whether there is an event or schedule that corresponds to the current time. "Corresponding to the current time" means, for example, that has been completed or is scheduled to be completed within a predetermined period including the current date and time. The predetermined period includes, for example, one week, one day, 12 hours, 6 hours, 2 hours, 1 hour, etc.
[0224] The designated time may be set according to the content of the event or schedule. For example, for topics of interest such as concerts or trips, the time may be set over a period of several days. On the other hand, for items with specific dates and times, such as shopping, the time may be set to within one day. Furthermore, the designated time may be changed depending on whether the event has already been completed or is scheduled to be completed. In this case as well, the designated time may not be uniform regardless of the content of the event or schedule, but rather may be set according to the content of the event or schedule.
[0225] If the event or schedule obtained in step 241 corresponds to the current time, a positive result is obtained in step 241. In this case, the conversation server 40 proceeds to step 222. On the other hand, if the event or schedule obtained in step 241 does not correspond to the current time, a negative result is obtained in step 241. In this case, the conversation server 40 proceeds to step 242.
[0226] By the way, in step 241, it may be determined whether or not the event or schedule is used in a conversation that includes the current time. Here, "conversation that includes the current time" refers to a conversation within a predetermined period that includes the current date and time. By providing this determination function, it becomes possible to change the content of the spoken text even if the same image description text is given to the conversational AI. As a result, the problem of repeated spoken texts about the same event or schedule can be reduced. Note that the additional determination may be performed only if an event or schedule corresponding to the current time is found. If a positive result is obtained in the additional determination, proceed to step 222; if a negative result is obtained in the additional determination, proceed to step 242.
[0227] Step 242: The conversation server 40 provides the image description text to the interactive AI to generate spoken text. In other words, the conversation is based only on the image captured by the owner terminal 10. For example, in the example in Figure 34, the speech would be something like "There are lots of different kinds of fruit!" and information such as "supermarket in front of the station" or "shopping" would not be included. For example, in the example in Figure 35, the speech would be something like "That's a big stuffed animal." and information such as "my friend's daughter" would not be included. For example, in the example in Figure 36, the speech would be something like "That's a big stuffed animal." and information such as "game center" would not be included.
[0228] <Summary> The aforementioned conversation system 1 (see Figure 1) enables the AI character to initiate conversations spontaneously, even when the user is not speaking.
[0229] <Other Embodiments> The embodiments of the invention are not limited to those described above. For example, elements of each embodiment can be combined as appropriate. Combinations here include the deletion of elements from each embodiment.
[0230] (1) In the above-described embodiment, the conversation server 40 (see Figure 1) assigns a new user ID each time a family user is registered. Therefore, even for the same person, a different user ID will be assigned depending on whether they are registered as a family user by owner A or by owner B. For example, a different user ID will be assigned depending on whether they are registered as owner A's child or owner B's grandchild. That is, the same person will be assigned a user ID as a child and a user ID as a grandchild. Also, when owner A is registered as a family member of owner B, owner A will be assigned a user ID as the owner and a user ID as a family user of owner B.
[0231] However, if it is confirmed that the users are the same person through facial images, voiceprints, etc., the user IDs used to manage the same person may be consolidated into one. Alternatively, the multiple user IDs may remain as they are, and a process (hereinafter referred to as "name matching process") may be performed to determine that the multiple user IDs belong to the same person. However, it is desirable to require confirmation from the owner before consolidating or matching user IDs. Consolidating into a single user ID or performing the name matching process makes it possible to respond and speak in a way that reflects past conversations, regardless of the owner associated with the owner terminal 10 (see Figure 1) used for the conversation, as long as the conversation is with the same person.
[0232] (2) In the above-described embodiment, the owner terminal 10 identifies the user with whom it is speaking. However, the identification of the user with whom it is speaking may be performed on the conversation server 40 side. For example, the user with whom it is speaking may be identified by providing the image data captured by the owner terminal 10, along with the face image and voiceprint information registered in association with the owner terminal 10 which is the target of processing, to the conversational AI or the like.
[0233] (3) Step 131 (see Figure 15) assumes the memory information of the identified person. However, it is also possible to include the memory information of family users in addition to the memory information of the identified person. In this case, the conversation server 40 provides the conversational AI with the information of the person identified in step 112 (see Figure 15) and the memory information of the identified person and their family users. For example, even if the identified person is a male named "Taro" and there is no memory information about a "white knit hat" in his memory information, it is possible to output response text such as "Have you found the hat that Hanako wanted?" by using the memory information of family user "Hanako".
[0234] (4) In step 117 (see Figure 12) described above, image data is given to the interactive AI to obtain image description text. However, step 117 may be omitted, and in step 118 (see Figure 12) or step 132 (see Figure 15), image data obtained from the owner terminal 10 may also be given to the interactive AI to generate response text. Figure 38 is a diagram illustrating another example of processing operation in which a conversation is started by a user's utterance. Figure 38 is indicated with reference numerals corresponding to parts that correspond to Figure 12.
[0235] The difference between Figure 38 and Figure 12 is that the processing operation shown in Figure 38 does not include step 117 (see Figure 12), and step 118A is executed instead of step 118. In step 118A, the conversational AI is provided with the spoken text, image data, and person information. In other words, the image description text is replaced with image data.
[0236] In the processing operation shown in Figure 38, step 117, which involves acquiring image description text, is omitted, and all image information is provided to the interactive AI that generates the response text. As a result, the interactive AI generates the response text including information that was missing during the generation of the image description text. Consequently, the variety of content in the response text increases compared to when the image description text is provided to the interactive AI. This mechanism of providing the image data itself to the interactive AI instead of the image description text can also be applied to Figures 15, 20, and 25. Furthermore, this mechanism of providing the image data itself to the interactive AI instead of the image description text can also be applied to processing operations where a conversation is initiated by an utterance from the conversation server 40.
[0237] (5) In step 142 (see Figure 20) described above, event information or schedule information linked to person information is provided to the conversational AI. However, the conversation server 40 may also provide the conversational AI with the content of events or schedules of one or more people registered for the owner as information related to the owner. For example, if the person identified in step 112 (see Figure 20) is "Taro", the conversational AI may be provided with the content of events or schedules of not only "Taro" but also "Hanako" and "Ai-chan" who are registered as family members.
[0238] In this case, the conversational AI is given information about "Taro" as a person to identify the person being spoken to, as well as information about who each event or schedule relates to. In this case, even if event or schedule information linked to "Taro" cannot be found, it becomes possible to reflect events or schedules related to people associated with the owner, such as "Oh, that reminds me, Hanako had an appointment to go to the hair salon," in the response text.
[0239] Furthermore, the event information or schedule information of one or more people registered for the owner may be used in a mechanism that does not have event or schedule information associated with the person identified in step 112 (see Figure 20) (for example, Taro). This mechanism prioritizes the generation of response text that reflects the owner's event or schedule information, and secondarily generates response text that reflects the event or schedule information of other people such as family members.
[0240] (6) In step 201 described above (see Figure 26), if facial recognition by the owner terminal 10 is successful, a message to that effect is sent to the conversation server 40. However, if facial recognition fails, a message to that effect is also sent to the conversation server 40, and image capture is initiated.
[0241] (7) In step 211 described above (see Figure 27), facial recognition on the owner terminal 10 is assumed, but the owner terminal 10 may also notify the conversation server 40 that it has met other predetermined conditions besides facial recognition. For example, the owner terminal 10 may notify the conversation server 40 that it has detected the turning on of a lighting fixture that was previously off. The detection of this event information can be considered as the owner entering the room where the stationary owner terminal 10 is located, and is a preferred timing for the conversation server 40 to initiate contact.
[0242] For example, the owner terminal 10 may notify the conversation server 40 that it has detected a noise while silence has continued for a predetermined period of time (e.g., 10 minutes) or longer. The detection of this event information can also be considered as the owner entering the room where the stationary owner terminal 10 is located, and is a preferred timing for the conversation server 40 to initiate contact. In addition, the owner terminal 10 may notify the conversation server 40 of other information, such as the release of a locked front door, the detection of a person by a motion sensor, the detection of operating sounds from home equipment or appliances, notifications from a thermometer or hygrometer, or other information.
[0243] (8) In step 201 described above (see Figure 26), it is assumed that the user's face is recognized while the owner terminal 10 is constantly capturing images regardless of the owner's instructions. However, the activation of the camera function of another application or the like may be used as a trigger event for the conversation server 40 to initiate a conversation.
[0244] In existing devices, activating the camera function does not involve user speech, nor is it intended to initiate a conversation with the owner terminal 10. Therefore, a mechanism may be adopted in which the activation of the camera function is considered a trigger event, and the image data being captured by the camera 14 (see Figure 1) is sent to the conversation server 40. In this case, there is no need for the conversation server 40 to issue an imaging instruction to the owner terminal 10. Furthermore, it becomes possible for the user to spontaneously speak to an object that they are interested in and have pointed the camera 14 at.
[0245] (9) In the above-described embodiment, the response text and spoken text generated by the conversation server 40 (see Figure 1) are directly transmitted to the owner terminal 10 (see Figure 1). However, a mechanism may be adopted in which the text is provided to a dedicated server that converts it into voice data, and the voice data generated by the dedicated server is transmitted to the owner terminal 10.
[0246] <Provision of Conversation Service (Conversation Sequence 2)> The following describes the provision of conversation services through conversation system 1 (see Figure 1).
[0247] <Overview of Processing Sequence> Figure 45 is a diagram illustrating the processing sequence of the conversation service. Incidentally, the symbol S shown in the figure represents a step. In this embodiment, conversation is initiated by the user's speech or operation (for example, stroking or lifting the user terminal 10). The user terminal 10 transmits voice data, etc. to the voice / text conversion server 20 (step 101A). The voice data, etc. includes not only voice data but also image data and sensor data.
[0248] When the voice / text conversion server 20 receives voice data etc. from the user terminal 10, it generates spoken text (step 102A). After this, the voice / text conversion server 20 transmits the spoken text etc. to the front server 30 and the conversation server 40 (step 103A). The spoken text etc. may include not only spoken text converted from voice data, for example, but also image data and sensor data. The front server 30 generates a simple response based on the received voice data etc. (step 104A). As mentioned above, the simple response is, for example, a response (i.e., action) to inform the user of the detection of the start or end of speech.
[0249] When a simple response is generated, the front server 30 sends the simple response to the user terminal 10 (step 105A). The user terminal 10, having received the simple response, outputs the received simple response (step 106A). For example, the user terminal 10 outputs a sound effect or melody. Also, if the user terminal 10 includes a movable mechanism 18 (see Figure 1), it outputs a response that moves, for example, its head or limbs. Furthermore, if the AI character's facial expression is displayed on the user terminal 10's display 17 (see Figure 1), the AI character's facial expression changes as a response. All of these serve to inform the user without delay that their speech has been received.
[0250] The conversation server 40 generates a response text based on the received audio data, etc. (step 107A). In addition to the response text, control data and image data related to the response may also be generated. Once the response text is generated, the conversation server 40 sends the response text to the user terminal 10 (step 108A). The time required to generate the response text is longer than the time required to generate a simple response. For this reason, as shown in Figure 45, the transmission of the response text occurs after the transmission of the simple response.
[0251] Upon receiving the response text, the user terminal 10 converts the received voice text into voice data (step 109A). The user terminal 10 then plays the voice data (step 110A). If the user terminal 10 also receives control data or image data from the conversation server 40, it outputs those as well. For example, it changes the display on the user terminal 10's movable mechanism 18 or display 17 in conjunction with the voice playback. This adds a sense of life to the AI character's response.
[0252] <Response Text Generation Process> Figure 46 is a flowchart illustrating a detailed example of the response text generation process. Figure 46 is denoted with corresponding symbols for parts corresponding to Figure 45. The process shown in Figure 46 corresponds to step 107A (see Figure 45). When the speech / text conversion server 20 (see Figure 45) receives the utterance text, the processor 41 (see Figure 1) performs preprocessing (step 111A). The processor 41 performs, for example, normalization of the utterance text, morphological analysis, removal of unnecessary words, and vectorization. In addition, the processor 41 stores the utterance text in the conversation log 43B (see Figure 1).
[0253] Next, the processor 41 provides the conversation engine with pre-processed utterance text. In the case of Figure 46, three conversation engines are provided. The three conversation engines can be classified into, for example, a general-purpose rule-based conversation engine, a generative AI-based conversation engine, and a function-specific conversation engine. For example, the processor 41 generates response text using the general-purpose rule-based conversation engine (step 112A). An example of a general-purpose rule-based conversation engine is "Scenario Graph". The processor 41 also generates response text using the generative AI-based conversation engine (step 113A). An example of a generative AI-based conversation engine is ChatGPT or Gemini.
[0254] Furthermore, the processor 41 generates response text using a function-specific conversation engine (step 114A). A function-specific conversation engine could be, for example, a conversation engine specialized for a specific function such as word games. Note that there may be more than one conversation engine operating in steps 112A, 113A, and 114A. In that case, each conversation engine generates response text in parallel.
[0255] The processor 41, having provided the utterance text to the conversation engine, selects one of the response texts based on predetermined rules (step 115A). For example, the processor 41 prioritizes response texts generated by a group of conversation engines with a higher priority. For instance, it prioritizes the response text generated by step 114A over the response texts generated by step 112A and step 113A. It also prioritizes the response text generated by step 113A over the response text generated by step 112A.
[0256] If multiple conversation engines with the same priority can generate response text, the first response text generated may be selected, or one response text may be selected randomly. Once a response text is selected, the processor 41 performs post-processing (step 116A). Post-processing includes, for example, the addition of emotion. For the addition of emotion, for example, the presence and type of sound effects, the presence and type of melody, the playback speed of the response text, the tone of voice, the type of animation, and the movement of the movable mechanism 18 (see Figure 1) are specified.
[0257] <Detailed Operation of Step 113A> Figure 47 is a flowchart illustrating an example of a method for generating response text using a generative AI-type conversation engine. Figure 47 is denoted with reference numerals corresponding to parts in Figure 46. The process shown in Figure 47 corresponds to step 113A (see Figure 46). In this embodiment, the processor 41 acquires four pieces of information when generating a prompt to be given to the conversation model (generative AI).
[0258] As part of the information, the processor 41 obtains user attributes and context (step 113A1). The processor 41 initiates step 113A1 triggered by the receipt of spoken text. Incidentally, user attributes are obtained from the user DB 43A (see Figure 2). Context includes information such as date and time, weather, and location.
[0259] As another piece of information, the processor 41 retrieves the most recent conversation text (step 113A2). The processor 41 initiates step 113A2 triggered by the reception of utterance text. The most recent conversation text may be, for example, conversation text within the same conversation (i.e., utterance text and response text), or conversation text from the previous conversation.
[0260] As another piece of information, the processor 41 obtains the user's spoken text (step 113A3). This spoken text is the most recent (i.e., last) spoken text notified to the speech / text conversion server 20 (see Figure 45) in step 103A (see Figure 45). As another piece of information, the processor 41 obtains summary data (i.e., memory) related to the content of the spoken text. First, the processor 41 vectorizes the spoken text (step 113A4). For example, the processor 41 generates vectors using the preprocessing results from step 111A (see Figure 46).
[0261] Once a vector representing the content of the utterance text is generated, the processor 41 retrieves one or more summary data related to the utterance text from the summary DB 43C (see Figures 10, 40-42) (step 113A5). The one or more summary data retrieved in step 113A5 is an example of a related summary. As mentioned above, the summary DB 43C stores the content of conversations exchanged with the AI character as summary data on a conversational, daily, monthly, and yearly basis. In other words, the vector DB 43D stores conversation texts that have been affected by the forgetting effect over time.
[0262] In step 113A5, vectors similar to the vector generated in step 113A4 are retrieved from the summary DB 43C. This retrieves summary data (i.e., memories) related to what the user uttered. Cosine similarity is used to retrieve similar vectors. Cosine similarity is calculated by dividing the dot product of two vectors by the magnitudes of the two vectors. Cosine similarity is given as a value between +1 and -1, with values closer to 1 indicating high similarity and values closer to -1 indicating low similarity. A value of 0 means that the similarity is neither high nor low (i.e., it cannot be said that they are similar or dissimilar).
[0263] The processor 41 performs an approximate nearest neighbor search based on the cosine similarity between the vectors stored in the summary DB 43C and the vectors generated in step 113A4, and identifies one or more vectors (summary data). The methods for identifying one or more summary data include, for example, a method for detecting the vector (summary data) that gives the maximum cosine similarity, a method for detecting vectors (summary data) that exceed a predetermined threshold, and a method for detecting the vector (summary data) that gives the maximum cosine similarity for each utterance time.
[0264] Figure 48 illustrates the process for obtaining summary data (first summary) given to a prompt. First, four types of summary data 43C1, 43C2, 43C3, and 43C4 (see Figure 9), each with a different utterance time, are generated from the conversation log 43B, which stores the conversation text between the user and the AI character, and stored in the summary DB 43C.
[0265] For example, the conversation-based summary data 43C1 is an example of the first summary data. The daily summary data 43C2 is an example of the third summary data. The monthly summary data 43C3 is an example of the fourth summary data. The yearly summary data 43C4 is an example of the fourth summary data. Also, the daily summary data 43C2, the monthly summary data 43C3, and the yearly summary data 43C4 are examples of the second summary data.
[0266] Of these summary data, summary data related to the utterance content is extracted as related summaries. That is, a portion of the summaries generated from the conversation log 43B is extracted as related summaries. Note that the extracted summaries may include summaries with different utterance times or generation times. Note that not all of the extracted summary data is used to generate the response text; only a portion may be used. In that case, differences may be made in the topics of the summary data used to generate the response text based on the differences in utterance times.
[0267] Figure 49 illustrates an example of extracting summary data based on differences in utterance timing. In Figure 49, corresponding parts are indicated with corresponding labels in Figure 9. Figure 49 shows Person A's summary DB 43C. As mentioned above, the summary DB 43C stores four sets of data with different utterance timings. Specifically, the summary DB 43C includes summary data 43C1 at the conversational level, summary data 43C2 at the daily level, summary data 43C3 at the monthly level, and summary data 43C4 at the yearly level.
[0268] In the case of Figure 49, summary data with high importance related to the utterance content (i.e., above the first threshold) is more easily extracted from the monthly summary data 43C3. Summary data is extracted from the daily summary data 43C2 regardless of its importance related to the utterance content. Since extraction is performed regardless of importance, summary data with low importance is also extracted. Furthermore, for the daily summary data 43C2, it is also possible to more actively extract summary data with low importance (i.e., below the first threshold).
[0269] Incidentally, the summary data 43C1, which is closest to the utterance time, is conversational summary data 43C2, which is next closest is monthly summary data 43C3, and the furthest is yearly summary data 43C4. In the case of Figure 49, monthly summary data 43C3 corresponds to the first period, and daily summary data 43C2 corresponds to the second period, which is closer to the utterance time than the first period.
[0270] This extraction method mimics human memory. Human memory retains various conversational content, regardless of importance, as long as it is close to the time of utterance. However, as time passes, information of lower importance to the user is forgotten, and relatively more important information remains. Similarly, the daily summary data C2 contains more relatively less important information compared to the monthly summary data 43C3.
[0271] Therefore, according to this embodiment, as summary data containing less important content is eliminated during the memory compression process, summary data containing more important content is preferentially retained and extracted for longer periods of time since the utterance. On the other hand, for shorter periods of time since the utterance, summary data on topics of relatively lower importance is also included in the extraction target, thereby broadening the variety of topics. As a result, even if less important content is included in the response text, the user is likely to remember the content of the conversation, thus achieving a conversation that is not heavily reliant on long-term memory. This conversation realizes an interaction that is closer to a conversation between two people.
[0272] Returning to the explanation of Figure 47, once the four pieces of information have been obtained through the above procedure, the processor 41 generates a prompt to give to the conversation model (generating AI) (step 113A6). Next, the processor 41 gives the prompt to the conversation model and obtains the response text (step 113A7). In other words, the processor 41 generates a response by providing the conversation model with the user's utterance and summary data related to the utterance. Thus, in this embodiment, instead of a conversation log, the response text is generated using summary data related to the user's utterance text.
[0273] As mentioned above, the conversation log 43B (see Figure 1) may accurately store the content of the conversation down to the last word, regardless of how many years ago it occurred. However, due to its nature as data that records the content of the utterances exactly as they were spoken, it may be discarded at any time (for example, at midnight every day) from a privacy standpoint, or it may be stored in the database for a predetermined period. In any case, unlike the conversation log, the summary data 43C stores only information of high importance and frequency, similar to the user's memory. Therefore, the content of the response text to the user's utterance text does not become unnecessarily detailed. As a result, it becomes possible to make conversations with AI characters closer to conversations between humans.
[0274] <Summary Data Generation Process> As mentioned above, in this embodiment, when generating response text as utterances of the AI character, summary data generated from the conversation text stored in the conversation log 43B (see Figure 1) is used in addition to the user's utterance text. The generation process for four types of summary data will be described below.
[0275] <Conversation-Unit Summary Data> Figure 50 is a flowchart illustrating the procedure for generating conversation-unit summary data. The processing operations shown in Figure 50 are executed by the processor 41 of the conversation server 40. The processor 41 obtains utterance text and response text that satisfy predetermined conditions from the conversation log (step 201A). The predetermined conditions here are, for example, that the recording is made after the time of the last memory created previously and a predetermined fixed time, whichever is more recent. The fixed time is, for example, 24 hours. However, the fixed time is not limited to 24 hours; it may be 1 hour, 6 hours, or 12 hours.
[0276] Next, the processor 41 selects one of the texts in the order in which they occur (step 202A). The text here is either an utterance text corresponding to a user's utterance or a response text corresponding to an AI character's utterance. Next, the processor 41 determines whether or not there are any unselected texts (step 203A).
[0277] If there are unselected texts remaining, a positive result is obtained in step 203A. For example, if 50 texts were acquired in step 201A, a positive result will be obtained until the 50th text is selected. In this case, the processor 41 determines whether the elapsed time since the end time of the previous text is greater than or equal to a threshold (step 204A).
[0278] Step 204A is a step for determining which text belongs to one conversation and which belongs to the next conversation. In this embodiment, 5 minutes is used as the threshold. Note that 5 minutes is just an example, and it may be 1 minute or 10 minutes. In this embodiment, the threshold is set on the conversation server 40 side. However, it may be possible to allow the user to adjust the threshold.
[0279] If the elapsed time is less than the threshold, a negative result is obtained in step 204A. In this case, the processor 41 returns to step 202A and selects the next text. On the other hand, if the elapsed time is greater than or equal to the threshold, a positive result is obtained in step 204A. In this case, the processor 41 determines that the texts up to the previous one constitute a single conversation (step 205A). For example, if there is a silence period of 5 minutes or more between the third text and the fourth text, the first three texts are determined to have been spoken within a single conversation.
[0280] Next, the processor 41 removes text containing NG (= No Good) words from the conversation (step 206A). In this embodiment, NG words are removed at this stage so that they are not generated in the summary data containing them. Note that NG words include, for example, words that apply to all users and words that are set individually for each user. The former includes words that contain discriminatory, immoral, or antisocial expressions. The latter includes words that the user personally dislikes, such as marriage, lovers, and grades. The latter words are set from the user terminal 10 (see Figure 1). Alternatively, the system may store all text containing NG words as summary data without removing them, and then remove them when generating the robot's spoken text.
[0281] Next, the processor 41 determines whether there are three or more texts in the conversation (step 207A). This operation is for generating summary data for each conversation, so it requires the recording of three or more texts. However, three or more is just an example; there could be two or more, or four or more. Four or more texts mean that there were two or more exchanges between the user and the AI character. If the number of texts is less than three (i.e., two or fewer texts), a negative result is obtained in step 207A. In this case, the processor 41 returns to step 202A. In other words, conversations with two or fewer texts will not be generated as summary data for each conversation. In other words, they will not be retained in memory.
[0282] On the other hand, if there are three or more texts, a positive result is obtained in step 207A. In this case, the processor 41 determines whether or not a conversation has taken place (step 208A). This determination process determines whether the exchange between the user's spoken text and the AI character's response text makes sense. This determination process is provided to exclude conversations where, for example, television or radio audio is mistakenly recognized as the user's speech. This is because memories of this type of conversation would be meaningless to the user.
[0283] If a conversation is not established, a negative result is obtained in step 208A. In this case, the processor 41 returns to step 202A. That is, no conversation-specific summary data is generated from the conversation in question. On the other hand, if a conversation is established, a positive result is obtained in step 208A. In this case, the processor 41 generates conversation-specific summary data (step 209A). Note that the aforementioned importance level may be assigned at the stage of generating the conversation-specific summary data. When summary data summarizing the content of a conversation that satisfies specific conditions is generated, the processor 41 returns to step 202A.
[0284] Subsequently, the processor 41 determines whether or not there are any unselected texts (step 203A). If there are no texts to select, a negative result is obtained in step 203A. For example, if 50 texts were obtained in step 201A, and the text selected in the previous step 202A was the 50th text, a negative result is obtained in step 203A. In this case, the processor 41 terminates the process.
[0285] <Daily Summary Data> <Generation Procedure 1> Figure 51 is a flowchart illustrating the procedure for generating daily summary data. The processing operations shown in Figure 51 are also executed by the processor 41 of the conversation server 40. The processor 41 refers to the conversation date of the conversation-unit summary data (step 221A). The conversation date is identified, for example, from the conversation end date and time 43C14 (see Figure 10). As mentioned above, if the administrative day does not start at 0:00:00, the conversation date is managed by the configured day.
[0286] Next, the processor 41 picks up summary data for conversation units that are one month old since the conversation date (step 222A). For example, in the example in Figure 43, summary data 43C1 for a conversation unit whose conversation date was January 1, 2024 is picked up on February 1, 2024. Next, the processor 41 re-summarizes the picked-up summary data 43C1 for conversation units and stores it as daily summary data 43C2 (step 223A). In this daily summary data 43C2, the date one month prior (i.e., January 1, 2024) is stored as the conversation date.
[0287] <Generation Procedure 2> Figure 52 is a flowchart illustrating another generation procedure for daily summary data. Figure 52 is denoted with corresponding symbols for parts that correspond to Figure 51. The processing operations shown in Figure 52 are also executed by the processor 41 of the conversation server 40. In generation procedure 2, steps 221A and 222A are executed in order. That is, the processor 41 refers to the conversation date of the conversation-unit summary data and picks up summary data for conversation units where one month has passed since the conversation date.
[0288] Next, the processor 41 determines whether the number of summary data in the set exceeds a threshold (step 224A). If the number of summary data exceeds the threshold, a positive result is obtained in step 224A. In this case, the processor 41 reads out the summary data for each conversation unit from the summary DB 43C (see Figures 10, 40-42) (step 225A). Subsequently, the processor 41 clusters the summary data 43C1 (see Figure 9) for conversation units belonging to the same day (step 226A).
[0289] In other words, the processor 41 classifies the summary data 43C1 of conversation units belonging to the same day into subsets (clusters) of summary data with similar vectors. In this embodiment, the cosine distance between vectors is calculated, and clustering is performed using a density-based method with the resulting distance matrix. "Density-based" means that nearby data is grouped together. That is, if there are many data points together, there is a high probability that they will form a cluster, and if there is little data around them, there is a low probability that they will form a cluster. Alternatively, a process such as classifying groups of vectors with a cosine similarity of 0.5 or higher into the same cluster may be used. In this case, the number of clusters is not predetermined.
[0290] Figure 53 illustrates the clustering of summary data 43C1 from multiple conversation units with the same conversation date. In Figure 53, the conversation date is February 10, 2024, and it includes summary data 43C1 (memories A, B, C, and D) from corresponding conversation units. In Figure 53, the four summary data 43C1 are classified into two subsets (clusters). For example, cluster A contains memories A and C, and cluster B contains memories B and D. Here, clusters correspond to topics, for example. Topics include school life, meals, work, childcare, and fashion. That is, the conversations from February 10, 2024, are classified into two topics.
[0291] In the case of generation procedure 1 described above, these two topics are also combined into a single summary data 43C2 (see Figure 9). However, in this generation procedure 2, the summarization of the conversation text is judged on a unit of two topics (i.e., clusters). In Figure 53, a number is attached to the upper left of the summary data 43C1 (i.e., memory). This number represents the amount of information or importance calculated according to predetermined rules. In Figure 53, the number for memory A is the highest, and the number for memory C is the lowest. Also, the average value of the numbers of the classified summary data 43C1 is attached to the lower left of the cluster. For example, cluster A is attached with "4", and cluster B is attached with "4.5".
[0292] Returning to the explanation of Figure 52, the processor 41 re-summarizes the summary data in the cluster with the lowest amount of information (step 227A). In the case of Figure 53, it re-summarizes memory A and memory C contained in cluster A. After re-summarization, the processor 41 disassembles the cluster and records it as daily summary data (step 228A). Then the processor 41 returns to step 224A.
[0293] Figure 54 illustrates the re-summarization of summary data within a cluster and the decomposition of the cluster after re-summarization. In Figure 54, memory E is generated by re-summarizing memory A and memory C of cluster A. The value calculated for memory E according to a predetermined rule is "6". As a result, cluster A is assigned the number "6". Subsequently, when the cluster is decomposed, the summary data for February 10, 2024 consists of three memories: B, D, and E.
[0294] Returning to the explanation of Figure 52, the processing operation described above is repeated as long as a positive result is obtained in step 224A. Note that if the number of summary data in the set falls below a threshold, a negative result is obtained in step 224A. In this case, the processor 41 terminates the generation process of daily summary data 43C2. As described above, in the processing operation shown in Figure 52, the summarization process is repeated based on the importance of the content until the number of summary data 43C1 for conversation units belonging to the same day falls below a predetermined number. This makes it possible to retain summary data 43C2 for conversation units with high importance, unlike when summarizing summary data for conversation units belonging to the same day into a single summary data.
[0295] <Monthly Summary Data> Figure 55 is a flowchart illustrating the procedure for generating monthly summary data 43C3 (see Figure 9). The processing operations shown in Figure 55 are also executed by the processor 41 of the conversation server 40. The processor 41 checks the current year, month, and day (step 231A). Next, the processor 41 determines whether it is currently the 1st or not (step 232A). In this embodiment, the first day of each month is used as the determination day, but the day used for determination can be the 2nd, the last day of the month, or any other day.
[0296] If the current date is not the 1st, a negative result is obtained in step 232A. In this case, the processor 41 returns to step 231A. On the other hand, if the current date is the 1st, a positive result is obtained in step 232A. In this case, the processor 41 picks up daily summary data 43C2 (see Figure 9) for the same month of the previous year (step 233A). For example, 31 summary data 43C2 (see Figure 9) are picked up. The target of the pick-up can be freely set as long as daily summaries have already been created for each interval.
[0297] Next, the processor 41 re-summarizes the selected summary data and records it as summary data for the corresponding month (step 234A). This generates monthly summary data 43C3. Note that the same processing as in "Generation Procedure 2" for daily summary data 43C2 may also be applied to the monthly summary data 43C3. By applying the same processing as in "Generation Procedure 2," it becomes easier for information (i.e., memories) on multiple topics to be retained in the monthly summary data 43C3.
[0298] <Annual Summary Data> Figure 56 is a flowchart illustrating the procedure for generating annual summary data 43C4 (see Figure 9). The processing operations shown in Figure 56 are also executed by the processor 41 of the conversation server 40. The processor 41 checks the current year, month, and day (step 241A). Next, the processor 41 determines whether it is currently January 1st or not (step 242A). In this embodiment, January 1st is used as the determination date, but the day used for determination can be January 2nd or any other day.
[0299] If the current date is not January 1st, a negative result is obtained in step 242A. In this case, the processor 41 returns to step 241A. On the other hand, if the current date is January 1st, a positive result is obtained in step 242A. In this case, the processor 41 picks up monthly summary data from two years ago (step 243A). For example, 12 summary data 43C3 (see Figure 9) are picked up. The period to be picked up does not necessarily have to be two years ago; it can be freely set as long as monthly summary data has already been created for that period.
[0300] Next, the processor 41 re-summarizes the selected summary data and records it as summary data for the corresponding year (step 244A). This generates the yearly summary data 43C4. Note that the same processing as in "Generation Procedure 2" for the daily summary data 43C2 may also be applied to the yearly summary data 43C4. By applying the same processing as in "Generation Procedure 2", it becomes easier for information (i.e., memories) of multiple topics to be retained in the yearly summary data 43C4 as well.
[0301] <Size Management of the Summary Database> The summary database 43C (see Figure 1) stores summary data 43C1 to 43C4 (see Figure 9) for each user who uses the conversation service. This information grows in proportion to the duration of use of the conversation service. Therefore, in this embodiment, the data size of the summary database 43C is monitored on a per-user basis, and when the data size reaches a standard level, the data size is compressed.
[0302] Here, data size compression refers to organizing memory. Note that monitoring is not limited to individual users; it can also apply to predetermined groups designated for monitoring, or to all users. Figure 57 is a flowchart illustrating the data compression process in the summary DB43C (see Figure 1).
[0303] First, the processor 41 determines whether the size of the summary DB 43C (see Figure 1) exceeds a threshold (step 251A). This determination process is executed at a predetermined time, for example, at a time specified in the maintenance schedule. If the size of the summary DB 43C is less than or equal to the threshold, a negative result is obtained in step 251A. In this case, the processor 41 terminates the process that was started. On the other hand, if the size of the summary DB 43C exceeds the threshold, a negative result is obtained in step 251A. In this case, the processor 41 reads the summary data from the summary DB 43C in units (step 252A).
[0304] Next, the processor 41 clusters the summary data in units of read data (step 253A). For example, it clusters the summary data 43C1 of conversations belonging to the same day. For example, it clusters the summary data 43C2 of days belonging to the same month. For example, it clusters the summary data 43C3 of months belonging to the same year. For the clustering process, the method described in "Generation Procedure 2" for the daily summary data 43C2 (see Figure 9) is used.
[0305] Next, the processor 41 re-summarizes the summary data in the cluster with the lowest amount of information (step 254A). Subsequently, the processor 41 decomposes the clusters and records them as summary data for each unit (step 255A). After that, the processor 41 returns to step 251A. In the case of Figure 57, the processor 41 repeatedly performs the series of processing operations as long as the size of the summary DB 43C exceeds the threshold. When the size of the summary DB 43C falls below the threshold, a negative result is obtained in step 251A. In this case, the processor 41 terminates the series of processing operations.
[0306] <Example User Screen> The following describes user screens related to the conversation service. <Setting and Changing the Time Period> The above explanation described the case where summary data is generated when the elapsed time since the utterance meets predetermined conditions. For example, daily summary data 43C2 (see Figure 9) is generated one month after the conversation date. However, it is conceivable that some users may want to individually set or change the timing of summary data generation.
[0307] Figure 58 is a diagram illustrating an example of an operation screen 300 displayed on the user terminal 10 (see Figure 1). The operation screen 300 shown in Figure 58 includes a summary mode setting field 301, a summary data generation timing setting field 302, a "back" button 303, and a "confirm" button 304.
[0308] The summary mode setting section 301 displays "On" and "Off" buttons in radio button format. In Figure 58, the "On" button is selected, meaning the summary mode is enabled. When the summary mode is enabled, the aforementioned summary data 43C1-4 (see Figure 9) are generated, and these summary data 43C1-4 are referenced to generate the response text. On the other hand, when the "Off" button is selected, the summary mode is disabled. In this case, the response text is generated using the conversation log 43B (see Figure 1). Using the conversation log 43B means that every word of a conversation from 10 years ago may be referenced to generate the response text.
[0309] The summary data generation timing setting field 302 is displayed when the "On" button in the summary mode setting field 301 is selected, or when it is operable. In Figure 58, the "Daily" button is selected from the four buttons corresponding to the time periods. The four buttons corresponding to the time periods are displayed as radio buttons. Therefore, the description reads, "Please select the shortest time period for which to generate daily summary data." Of course, the content of the description changes when other buttons are selected.
[0310] In addition, the summary data generation timing setting section 302 displays buttons corresponding to the current setting period (i.e., 1 month) and selectable periods (in this case, 7 days, 1 month, and 6 months). In Figure 58, the "6 months" button is selected. When the "Back" button 303 is pressed, the current selection is canceled and the user returns to the previous screen. When the "Confirm" button 304 is pressed, the selection is confirmed and used to generate new summary data.
[0311] <Memory Verification> As mentioned above, in this embodiment, summary data related to the user's spoken text is referenced to generate the response text. This summary data is generated using a conversation log that records the conversation between the user and the AI character. The spoken text, which is the content spoken by the user within the conversation log, is generated by the speech / text conversion server 20 (see Figure 1).
[0312] However, there is a non-zero possibility that errors may be present in the text conversion performed by the speech / text conversion server 20. As a result, errors may remain in the summary data, potentially reducing the accuracy of the response text to the spoken text. To address this, a function is needed that allows users to delete or modify the content of recorded conversations. There is also a need among users to review the content of recorded conversations. Therefore, a function is needed to review the content of the recorded conversation log.
[0313] Figure 59 illustrates the summary data confirmation screens 310A and 310B. The confirmation screens 310A and 310B shown in Figure 59 display "Your (Person A's) memories of February 23, 2025," etc. The title display is optional and could be "AI Character's Memories," etc. The confirmation screens 310A and 310B shown in Figure 59 include an item field 311, an information field 312, a "Delete" button 313, an "Exclude" button 314, a "Modify" button 315, and a "Close" button 316.
[0314] The item column 311 displays items such as "Select," "Exclude," "Importance," "Time," and "Content," corresponding to the content displayed in the information column 312. "Select" is associated with a checkbox. "Exclude" is associated with a checkbox. "Importance" represents the column displaying the importance level assigned to each memory. "Time" represents the column displaying the date and time of the conversation, etc. "Content" represents the column displaying the corresponding conversation text or summary data. The content column displays the summary data in bullet points.
[0315] In the confirmation screen 310A shown in Figure 59, the second conversation text from the top, "It looks like the roads are congested. The weather is bad, I'm worried," is selected. This is because the user did not want to retain this information in their memory. The confirmation screen 310B shown in Figure 59 represents the display after the "Delete" button 313 is pressed on confirmation screen 310A. In confirmation screen 310B, "It looks like the roads are congested. The weather is bad, I'm worried," which was selected on confirmation screen 310A, has been deleted from the screen. Therefore, the deleted conversation content will no longer be included in the monthly or yearly summary data. As a result, the deleted memory will not be reflected in the generation of response text.
[0316] Figure 60 illustrates the summary data confirmation screens 310A and 310C. In Figure 60, corresponding parts with those in Figure 59 are indicated with corresponding symbols. In confirmation screen 310A shown in Figure 60, the second conversation text from the top, "Ms. B's son, who lives nearby, has been hospitalized," is selected. The "Exclude" button 314 is selected when you want to keep the data in memory but do not want to use it to generate the response text. In other words, the "Exclude" button 314 is selected when you want to exclude data from being input to the conversation model.
[0317] The confirmation screen 310C shown in Figure 60 represents the display after the "Exclude" button 314 has been pressed on confirmation screen 310A. On confirmation screen 310C, the message "Ms. B's son, who lives nearby, has been hospitalized," which was selected on confirmation screen 310A, remains displayed. However, the corresponding checkbox is still displayed as selected. This setting is also carried over to the monthly and yearly summary data that inherits the content. As a result, the excluded memory is no longer reflected in the generation of the response text.
[0318] Figure 61 illustrates the summary data confirmation screens 310A and 310D. In Figure 61, corresponding parts with reference numerals are shown in relation to Figure 59. In confirmation screen 310A shown in Figure 61, the third conversation text from the top, "I need to call the patent restaurant to change my reservation time," is selected. Also, the "Edit" button 315 is selected. Therefore, in confirmation screen 310D shown in Figure 61, the edit box 312A is displayed as a pop-up. The edit box 312A displays the conversation text selected by the checkbox as editable.
[0319] In Figure 61, the typographical error has been corrected to "I need to call the Tokyo restaurant to change my reservation time." When the "OK" button 312B is pressed in this state, the display changes to the corrected content. Specifically, the third conversation text from the top on the confirmation screen 310A changes to "I need to call the Tokyo restaurant to change my reservation time." The "Select" checkbox is deselected. If you want to keep the text in memory but do not want to use it to generate the response text, you can select the "Exclude" button 314. As another example, you could select genres you don't want to remember in advance on a settings screen that is not shown, and then have conversations of the selected genres automatically excluded.
[0320] <Topic Exclusion> As mentioned above, the confirmation screen 310 allows you to specify the exclusion of topics from the summary data used to modify conversation logs and generate response text. On the other hand, topics that you do not want to include in the response text may be specified from other screens. Figure 62 illustrates an example of a screen 320 for setting topics that you do not want to include in the conversation.
[0321] The settings screen 320 shown in Figure 62 includes the title "Unwanted Topics," as well as an item field 321, an information field 322, a "Close" button 323, and a "Confirm" button 324. In the item field 321, for example, "Select" and "Topic" are displayed, corresponding to the content displayed in the information field 322.
[0322] In Figure 62, two checkboxes corresponding to "illness" and "marriage" are selected. Note that in the settings screen 320 shown in Figure 62, undesirable topic candidates are listed as options, but it may also be possible to input free text. The entered text may be used as a prompt for the conversation model. If text is used, it will also be possible to specify conditions for topics to avoid. When the "Close" button 323 is operated, the current selection is canceled. Also, when the "Confirm" button 324 is operated, the settings are confirmed.
[0323] Figure 63 illustrates an example of a settings screen 330 for topics that users do not want to remember. The settings screen 330 shown in Figure 63 is titled "Activity Log". The settings screen 330 has an "Activity Log" tab and a "Conversation Log" tab. Figure 63 shows the contents of the "Conversation Log" tab. Utterances 331A to 331E made between the owner and the conversation server 40 (see Figure 1) are arranged on the timeline 331 in chronological order. The settings screen 330 here is an example of a screen that presents the contents of individual conversations to the user. Each utterance in the settings screen 330 shown in Figure 63 includes a "good" button 332, a "bad" button 333, and an "edit" button 334. These are icons for accepting specific actions from the user.
[0324] For example, the "good" button 332 is used to increase the importance of the utterance or topic of the utterance being manipulated. For example, the "bad" button 333 is used to decrease the importance of the utterance or topic of the utterance being manipulated. The "edit" button 334 is used to exclude the utterance or topic of the utterance being manipulated from the summary related to the content of the utterance (i.e., the related summary).
[0325] Through these button operations, users can, for example, exclude or include specific utterances or topics from the summary data. Furthermore, through these button operations, users can, for example, prevent summaries generated from specific utterances or topics from being extracted as related summaries, or make them more likely to be extracted as related summaries. In this way, users can reflect their own ideas in the generation of summary data and response text through the settings screen 330.
[0326] <Summary> The conversation system 1 described in this embodiment (see Figure 1) generates the AI character's response text by providing a conversation model with the user's utterances and a summary of the conversation data as prompts. By adopting this generation method, it becomes possible to make the AI character's responses closer to those of a human. In other words, it becomes possible to make the AI character's responses closer to the characteristics of human memory. Here, memory characteristics refer to the fact that memories deteriorate and become vague over time since the conversation.
[0327] <Embodiment 2> In the embodiment described above, the conversation server 40, etc., generates the AI character's response to the user's utterance. In this embodiment, the case in which the AI character initiates a topic with the user will be described. Figure 64 is a diagram showing an example of the overall configuration of the conversation system 1 assumed in Embodiment 2. Figure 64 is denoted by reference numerals corresponding to the parts that correspond to those in Figure 1. The conversation system 1 shown in Figure 64 is also an example of an information processing system.
[0328] The hardware configuration of the user terminal 10 and other servers that constitute the conversation system 1 shown in Figure 64 is the same as in Embodiment 1. In Figure 64, the starting point of speech is the conversation server 40, so the descriptions of the user terminal 10, speech / text conversion server 20, and front server 30 are simplified. The functions described in this embodiment are also realized through the execution of programs by the processor 41.
[0329] Figure 65 is a flowchart illustrating the speech processing from the AI character. The processor 41 determines whether or not it has detected an event that satisfies predetermined conditions (step 301). Events that satisfy predetermined conditions include, for example, when the predetermined conditions are met during or after a conversation with the user has ended.
[0330] One of the predetermined conditions that must be met during a conversation with a user is that, for example, three minutes or more have passed since the user's last utterance. In Embodiment 1, a conversation is considered to have ended when five minutes or more have passed since the user's utterance or the AI character's response, so the AI character may offer a new topic before the conversation is considered to have ended.
[0331] In addition, the specified conditions may include, for example, the detection of specific facial expressions or gestures of the user in camera images uploaded from the user terminal 10. For example, a user waving at the camera may be considered a specific gesture. Incidentally, waving is an example of a user action that does not involve speech. However, the detection of a user shedding tears or closing their eyes in camera images of a user during a conversation may also be used as a speech event initiated by the AI character.
[0332] The predetermined conditions that must be met when a conversation with the user has ended may include, for example, the detection of events such as a pre-set date and time, news updates, weather updates, traffic updates, etc. The pre-set date and time may be, for example, 0:00 AM on January 1st, 7:00 AM on the user's birthday, 7:00 AM on the 1st of each month, or 12:00 PM, 3:00 PM, and 6:00 PM every day. These dates and times are just examples of events that the user has set in advance. The predetermined conditions that must be met when a conversation with the user has ended may also include the fact that the conversation function is running on the user terminal 10.
[0333] If no event that meets the predetermined conditions is detected, a negative result is obtained in step 301. In this case, the processor 41 repeats the determination process in step 301. On the other hand, if an event that meets the predetermined conditions is detected, a positive result is obtained in step 301. In this case, the processor 41 extracts summary data with an importance higher than the threshold value (step 302). The summary data with an importance higher than the threshold value here is an example of a relevant summary that satisfies a predetermined rule.
[0334] The reason for extracting summary data with an importance level higher than the threshold is that it is more likely to be a topic remembered by the user being spoken to. Summary data with an importance level higher than the threshold includes important events such as entrance ceremonies, trips, and weddings. The threshold level may be lower for summary data with a more recent memory, such as conversational summary data 43C1 (see Figure 9) and daily summary data 43C2 (see Figure 9), than for summary data with a longer memory, such as monthly summary data 43C3 (see Figure 9) and daily summary data 43C4 (see Figure 9). By changing the importance threshold according to the elapsed time since the utterance, the bias in the topics extracted can be reduced.
[0335] Next, the processor 41 instructs the conversation model to generate question text related to the extracted summary data (step 303). The prompt here may include the date and time, day of the week, and the user's camera image (or the situation as perceived from the camera image). Including information other than the extracted summary data in the prompt makes it possible to generate question text that is appropriate for the timing of speaking to the user. After this, the processor 41 performs post-processing such as adding emotion (step 304) and sends the question text, etc. to the user terminal 10 (step 305).
[0336] <Summary> The conversation system 1 described in this embodiment (see Figure 64) can output question text from the user terminal 10 as speech from the AI character, using memories of past conversations, when predetermined conditions are met during or after a conversation with the user has ended. By adopting this technology, it is possible to make questions using memories of past conversations with the user that are of high importance. In other words, it becomes possible to make the speech from the AI character sound more like that of a human.
[0337] <Other Embodiments> The embodiments of the invention are not limited to those described above. For example, elements of each embodiment can be combined as appropriate. Combinations here include the deletion of elements from each embodiment.
[0338] (1) In the above-described embodiment, a silence period lasting for a certain period of time (e.g., 5 minutes) is detected as a boundary between conversations. However, the duration of a single conversation may be set mechanically. For example, a conversation may be considered to have ended each time a certain amount of time has elapsed.
[0339] (2) In the above-described embodiment, text containing NG words is excluded at the stage of generating summary data for conversation units. However, text containing NG words may also be excluded when selecting the response text to be output from the AI character. For example, it may be excluded at any stage in step 107A (see Figure 45). For example, NG words may be excluded at the selection stage in step 115A (see Figure 46) or at the post-processing stage in step 116A (see Figure 46).
[0340] (3) In the above-described embodiment, text containing NG words is excluded at the stage of generating summary data for each conversation unit. However, instead of excluding NG words, a low priority (or importance) may be assigned to text or summary data containing NG words. In this case, text or summary data containing NG words will remain. However, by assigning a low priority (importance) to text or summary data containing NG words, it is possible to ensure that NG words are not included in the response text.
[0341] (4) In the above-described embodiment, the importance is displayed numerically on the summary data confirmation screen 310A (see Figure 60, etc.), but it may also be displayed as a symbol such as a heart mark.
[0342] (5) In the above-described embodiment, a method is employed in which summary data is not generated when the conversation text contains two or fewer texts. However, summary data may be generated even if the conversation text contains only one text.
[0343] (6) In the above-described embodiment, the response text and question text generated by the conversation server 40 (see Figure 45) are sent directly to the user terminal 10 (see Figure 45). However, a mechanism may be adopted in which the text is provided to a dedicated server that converts it into audio data, and the audio data generated by the dedicated server is sent to the user terminal 10.
[0344] (7) In the above-described embodiment, the conversation log 43B (see Figure 1) and the summary DB 43C (see Figure 1) are managed in association with the user account. Therefore, when one user account is shared by a family, conversations between the husband and the AI character, conversations between the wife and the AI character, and conversations between the child and the AI character are all managed under one user account. In this case, it is unavoidable that responses to the wife's utterances will be made based on the memory of the conversation with the husband. As a result, the conversation may become unnatural.
[0345] Therefore, if the speaker can be identified using speech recognition or image recognition technology, the conversation log 43B and summary DB 43C may be managed separately for each identified speaker. This enables natural interaction between the speaker and the AI character even when using a single user account. If speaker identification fails, the AI character may be asked, "Who is speaking now?" or similar to identify the speaker.
[0346] (8) In the above-described embodiment, summaries and vectors are managed in the summary DB 43C (see Figure 1), but the summary DB 43C may manage summaries and vectors in a separate DB (for example, vector DB 43D).
[0347] <Summary 1> The main features of the information processing system, information processing method, program, and information processing device described in the embodiments are shown below. [General Problem] One of the objectives of this disclosure is to provide an information processing system, information processing method, program, and information processing device that can personalize the response content for each owner, even if the content of the utterance and the captured image are the same.
[0348] [Note 1] One of the objectives of this disclosure is to change the content of the response text for each owner, even if the content of the utterance and the captured image are the same. [Note 1] The information processing system relating to this disclosure includes one or more processors, which generate a response text using the content uttered by the owner, the image captured by the owner terminal, and information related to the owner, and output audio from the owner terminal corresponding to the generated response text. According to the above information processing system, the content of the response text can be changed for each owner, even if the content of the utterance and the captured image are the same.
[0349] [Appendix 2] One of the purposes of this disclosure is to limit the generation of response text using images captured by the owner terminal to cases where the owner has given an instruction to capture an image. [Appendix 2] The information processing system described in Appendix 1, wherein when the content spoken by the owner includes an instruction to capture an image, one or more processors acquire an image from the owner terminal and use it to generate response text. This makes it possible to limit the generation of response text using images captured by the owner terminal to cases where the owner has given an instruction to capture an image.
[0350] [Appendix 3] One of the purposes of this disclosure is to reflect information related to the identified person in the content of the response text. [Appendix 3] An information processing system as described in Appendix 1, wherein one or more processors identify the person who spoke as the owner and use information related to the identified person as information related to the owner. This makes it possible to reflect information related to the identified person in the content of the response text.
[0351] [Appendix 4] One of the purposes of this disclosure is to reflect the content of events or schedules registered for an identified person in the content of the response text. [Appendix 4] An information processing system as described in Appendix 3, in which one or more processors use the content of events or schedules registered for an identified person as information related to the owner. This makes it possible to reflect the content of events or schedules registered for an identified person in the content of the response text.
[0352] [Appendix 5] One of the purposes of this disclosure is to reflect in the content of the response text the content of events or schedules that are highly relevant to the current date and time. [Appendix 5] An information processing system as described in Appendix 4, wherein one or more processors generate a response text using events or schedules that have been completed or are scheduled to be completed within a predetermined period including the current date and time, from among the events or schedules registered for an identified person. This makes it possible to reflect in the content of the response text the content of events or schedules that are highly relevant to the current date and time.
[0353] One of the purposes of this disclosure, corresponding to the problem described in [Appendix 6], is to prevent the same event or schedule content from appearing repeatedly in the response text. [Appendix 6] An information processing system as described in Appendix 4, wherein one or more processors generate a response text using events or schedules registered for an identified person that have not been used in a conversation within a predetermined period including the current date and time. This prevents the same event or schedule content from appearing repeatedly in the response text.
[0354] [Appendix 7] One of the purposes of this disclosure is to reflect the content of events or schedules of one or more people registered with the owner in the content of the response text. [Appendix 7] An information processing system as described in Appendix 1, in which one or more processors use the content of events or schedules of one or more people registered with the owner as information related to the owner. This makes it possible to reflect the content of events or schedules of one or more people registered with the owner in the content of the response text.
[0355] One of the purposes of this disclosure, corresponding to the problem described in [Appendix 8], is to reflect the history of conversations with an identified person in the content of the response text. [Appendix 8] An information processing system as described in Appendix 3, wherein one or more processors use the history of conversations recorded about an identified person as information related to the owner. This makes it possible to reflect the history of conversations with an identified person in the content of the response text.
[0356] One of the purposes of this disclosure, corresponding to the problem described in [Appendix 9], is to reflect topics of high importance to the identified person in the content of the response text. [Appendix 9] The information processing system described in Appendix 8, wherein the conversation history includes information indicating the level of importance of topics, and one or more processors use topics with a higher level of importance than other topics as information related to the owner. This makes it possible to reflect topics of high importance to the identified person in the content of the response text.
[0357] [Appendix 10] One of the purposes of this disclosure is to facilitate the generation of response text by natural language processing by using information that describes the content of an image. [Appendix 10] An information processing system as described in Appendix 1, wherein one or more processors use information that describes the content of an image captured by the owner terminal instead of the image captured by the owner terminal. This makes it possible to facilitate the generation of response text by natural language processing by using information that describes the content of an image.
[0358] [Appendix 11] One of the purposes of this disclosure is to address the diverse range of images captured by the owner terminal. [Appendix 11] An information processing system as described in Appendix 10, comprising one or more processors, which acquire information describing the content of images captured by the owner terminal through a generating AI. This enables the system to address the diverse range of images captured by the owner terminal.
[0359] [Appendix 12] One of the purposes of this disclosure is to generate a response text based on the content of the owner's speech when the owner does not give an imaging instruction. [Appendix 12] An information processing system as described in Appendix 1, wherein one or more processors generate a response text using the content of the owner's speech when the content of the owner's speech does not include an image imaging instruction. This makes it possible to generate a response text based on the content of the owner's speech when the owner does not give an imaging instruction.
[0360] [Appendix 13] One of the objectives of this disclosure is to change the content of the response text for each owner, even if the content of the utterance and the captured image are the same. [Appendix 13] The information processing method relating to this disclosure involves one or more processors performing the following processes: generating a response text using the content uttered by the owner, an image captured by the owner terminal, and information related to the owner; and outputting audio from the owner terminal corresponding to the generated response text. According to the above information processing method, the content of the response text can be changed for each owner, even if the content of the utterance and the captured image are the same.
[0361] [Appendix 14] One of the objectives of this disclosure is to change the content of the response text for each owner, even if the content of the utterance and the captured image are the same. [Appendix 14] The program of this disclosure causes one or more processors to generate a response text using the content uttered by the owner, the image captured by the owner terminal, and information related to the owner, and to output audio from the owner terminal corresponding to the generated response text. According to the above program, the content of the response text can be changed for each owner, even if the content of the utterance and the captured image are the same.
[0362] [Appendix 15] One of the objectives of this disclosure is to change the content of the response text for each owner, even if the content of the utterance and the captured image are the same. [Appendix 15] The information processing device of this disclosure includes one or more processors, which generate a response text using the content uttered by the owner, the image captured by the owner terminal, and information related to the owner, and output audio from the owner terminal corresponding to the generated response text. According to the above information processing device, the content of the response text can be changed for each owner, even if the content of the utterance and the captured image are the same.
[0363] <Summary 2> [General Issues] One of the purposes of this disclosure is to provide an information processing system, information processing method, program, and information processing device that can spontaneously initiate conversations while understanding the surrounding situation, including the owner.
[0364] [Note 1] One of the objectives of this disclosure is to enable spontaneous speech that reflects the content of autonomously captured images. [Note 1] The information processing system relating to this disclosure includes one or more processors, which generate speech text using captured images without instructions from the owner, and output voice corresponding to the generated speech text from the owner terminal. According to the above information processing system, spontaneous speech that reflects the content of autonomously captured images can be enabled.
[0365] [Appendix 2] One of the purposes of this disclosure is to reflect information about an identified person in the content of spontaneous speech. [Appendix 2] The information processing system described in Appendix 1, wherein when an image containing a facial image of a person registered as the owner is detected, one or more processors generate speech text for the person identified from the image. This makes it possible to reflect information about an identified person in the content of spontaneous speech.
[0366] [Appendix 3] One of the purposes of this disclosure is to reflect information about an identified person in the content of spontaneous speech. [Appendix 3] An information processing system as described in Appendix 1, in which one or more processors generate speech text for a person identified from a voice. This makes it possible to reflect information about an identified person in the content of spontaneous speech.
[0367] One of the purposes of this disclosure, corresponding to the problem described in [Appendix 4], is to reflect the content of events or schedules registered for the owner in the content of spontaneous utterances. [Appendix 4] The information processing system described in Appendix 1, wherein if an image contains information related to the content of events or schedules registered for the owner, one or more processors generate utterance text for the detected information. This makes it possible to reflect the content of events or schedules registered for the owner in the content of spontaneous utterances.
[0368] [Appendix 5] One of the purposes of this disclosure is to reflect the content of events or schedules that are highly relevant to the current date and time in the content of spontaneous utterances. [Appendix 5] An information processing system as described in Appendix 4, in which one or more processors generate utterance text using events or schedules that have been completed or are scheduled to be completed within a predetermined period including the current date and time, from among the events or schedules registered for the owner. This makes it possible to reflect the content of events or schedules that are highly relevant to the current date and time in the content of spontaneous utterances.
[0369] [Appendix 6] One of the purposes of this disclosure is to prevent the same event or schedule from appearing repeatedly in the content of spontaneous utterances. [Appendix 6] An information processing system as described in Appendix 4, wherein one or more processors generate utterance text using events or schedules registered for the owner that have not been used in conversations within a predetermined period including the current date and time. This prevents the same event or schedule from appearing repeatedly in the content of spontaneous utterances.
[0370] One of the purposes of this disclosure, corresponding to the problem described in [Appendix 7], is to reflect the history of conversations with the owner in the content of spontaneous utterances. [Appendix 7] The information processing system described in Appendix 1, wherein if an image contains information related to the history of conversations recorded about the owner, one or more processors generate speech text corresponding to the information. This makes it possible to reflect the history of conversations with the owner in the content of spontaneous utterances.
[0371] One of the purposes of this disclosure, corresponding to the problem described in [Appendix 8], is to reflect topics of high importance in the content of spontaneous utterances. [Appendix 8] The information processing system described in Appendix 7, wherein the conversation history includes information indicating the level of importance of topics, and one or more processors generate utterance text using topics that have a higher level of importance compared to other topics. This makes it possible to reflect topics of high importance in the content of spontaneous utterances.
[0372] One of the purposes of this disclosure is to address the challenges described in [Appendix 9] and to support the diverse range of images captured by the owner terminal. [Appendix 9] An information processing system as described in Appendix 1, comprising one or more processors that acquire information describing the content of images captured by the owner terminal through a generating AI. This enables support for the diverse range of images captured by the owner terminal.
[0373] [Appendix 10] One of the purposes of this disclosure is to enable spontaneous speech with content appropriate to the identified person. [Appendix 10] An information processing system as described in Appendix 9, wherein one or more processors provide the generating AI with information describing the content of the acquired image, as well as information of the person identified from the image or sound, to generate speech text. This enables spontaneous speech with content appropriate to the identified person.
[0374] One of the objectives of this disclosure, corresponding to the problem described in [Appendix 11], is to enable spontaneous speech that reflects the content of autonomously captured images. [Appendix 11] An information processing system as described in Appendix 1, in which one or more processors generate speech text independently of the owner's speech. This enables spontaneous speech that reflects the content of autonomously captured images.
[0375] [Appendix 12] One of the objectives of this disclosure is to enable spontaneous speech that reflects the content of autonomously captured images. [Appendix 12] The information processing method relating to this disclosure involves one or more processors performing the following processes: generating speech text using images captured without instructions from the owner; and outputting audio corresponding to the generated speech text from the owner terminal. According to the above information processing method, spontaneous speech that reflects the content of autonomously captured images can be enabled.
[0376] [Appendix 13] One of the objectives of this disclosure is to enable spontaneous speech that reflects the content of autonomously captured images. [Appendix 13] The program relating to this disclosure causes one or more processors to generate speech text using images captured without instructions from the owner, and to output audio corresponding to the generated speech text from the owner terminal. According to the above program, spontaneous speech that reflects the content of autonomously captured images can be enabled.
[0377] [Appendix 14] One of the objectives of this disclosure is to enable spontaneous speech that reflects the content of autonomously captured images. [Appendix 14] The information processing device relating to this disclosure includes one or more processors, which generate speech text using captured images without instructions from the owner, and output voice corresponding to the generated speech text from the owner terminal. According to the above information processing device, spontaneous speech that reflects the content of autonomously captured images can be enabled.
[0378] <Summary 3> [General Issues] One of the purposes of this disclosure is to provide an information processing device, an information processing method, a program, and an information processing system that can make the response content closer to that of a human.
[0379] [Note 1] One of the objectives of this disclosure is to realize a response that closely resembles the characteristics of human memory. [Note 1] The information processing device relating to this disclosure has one or more processors acquire relevant summaries that are related to the user's utterances from summaries generated from conversation logs linked to an account, and provides the acquired relevant summaries and utterances to a conversation model to output a response to the utterances. According to the above information processing device, a response that closely resembles the characteristics of human memory can be realized.
[0380] [Appendix 2] One of the objectives of this disclosure is to produce a response that is closer to that of a human than when conversation logs are directly provided to a conversation model. [Appendix 2] The information processing device described in Appendix 1 includes first summary data and second summary data generated by different processing methods. This makes it possible to produce a response that is closer to that of a human than when conversation logs are directly provided to a conversation model.
[0381] [Appendix 3] One of the objectives of this disclosure is to make the information provided to a conversation model more similar to the characteristics of human memory. [Appendix 3] The information processing device described in Appendix 2, wherein the first summary data includes summary data that summarizes individual conversations, and the second summary data includes at least one of a third summary data that summarizes a plurality of first summary data and a fourth summary data that further summarizes a plurality of third summary data. This makes it possible to make the information provided to a conversation model more similar to the characteristics of human memory.
[0382] [Appendix 4] One of the purposes of this disclosure is to make the generation of summary data closer to human memory. [Appendix 4] The second summary data is generated from a set of first summary data or a set of third summary data, where the elapsed time from utterance or the elapsed time from summary generation satisfies predetermined conditions. This makes it possible to make the generation of summary data closer to human memory.
[0383] [Appendix 5] One of the purposes of this disclosure is to enable customization of the output response. [Appendix 5] An information processing device as described in Appendix 4, wherein one or more processors accept the setting or modification of predetermined conditions through the operation screen of a terminal operated by a user. This makes it possible to customize the output response.
[0384] [Appendix 6] One of the purposes of this disclosure is to prevent the loss of information or memory due to excessive summarization. [Appendix 6] The second summary data is an information processing device described in Appendix 3, which is generated multiple times for each set of first summary data or for each set of third summary data. This makes it possible to prevent the loss of information or memory due to excessive summarization.
[0385] One of the purposes of this disclosure, corresponding to the problem described in [Appendix 7], is to ensure that dissimilar content is not lost. [Appendix 7] An information processing device described in Appendix 6, wherein one or more processors classify a set into multiple subsets based on the similarity of their content. This ensures that dissimilar content is not lost.
[0386] [Appendix 8] One of the purposes of this disclosure is to retain summary data of content of high importance. [Appendix 8] An information processing device as described in Appendix 7, wherein one or more processors repeatedly perform summarization processing based on the importance of the content until the number of summaries in the set falls below a predetermined number. This makes it possible to retain summary data of content of high importance.
[0387] [Appendix 9] One of the purposes of this disclosure is to reduce responses based on incorrect memories. [Appendix 9] An information processing device as described in Appendix 1, in which one or more processors display a confirmation screen containing the content of the summary on a terminal operated by the user. This reduces responses based on incorrect memories.
[0388] [Appendix 10] One of the purposes of this disclosure is to reduce responses based on false memories. [Appendix 10] An information processing device as described in Appendix 9, wherein one or more processors accept the deletion of a specific summary selected by the user, or its exclusion from input to the conversation model, or the modification of the content of the summary, through a confirmation screen. This reduces responses based on false memories.
[0389] [Appendix 11] One of the purposes of this disclosure is to facilitate the prediction of stored conversation content. [Appendix 11] One or more processors are the information processing device described in Appendix 9, which displays importance information attached to the summary on a confirmation screen. This makes it easier to predict stored conversation content.
[0390] [Appendix 12] One of the purposes of this disclosure is to reduce responses that contain content the user does not want. [Appendix 12] An information processing device as described in Appendix 1, wherein one or more processors display a screen that presents the content of individual conversations to the user, and display icons on the screen for receiving requests from the user to individually exclude the presented content from the subject of related summaries, or to individually change the importance of the presented content. This reduces responses that contain content the user does not want.
[0391] [Appendix 13] One of the purposes of this disclosure is to reduce responses that include unwanted topics. [Appendix 13] An information processing device as described in Appendix 1, wherein one or more processors receive topics to be excluded from the conversation model through the operation screen of a terminal operated by the user. This reduces responses that include unwanted topics.
[0392] [Appendix 14] One of the purposes of this disclosure is to reduce responses that include unwanted topics. [Appendix 14] An information processing device as described in Appendix 1, wherein one or more processors exclude conversations that include a topic selected by the user from the target of summary generation. This reduces responses that include unwanted topics.
[0393] [Appendix 15] One of the objectives of this disclosure is to improve the accuracy of the response content. [Appendix 15] An information processing device according to Appendix 2, wherein one or more processors, when outputting a response, provide a conversation model with at least one of a plurality of first summary data and a plurality of second summary data generated by a processing method different from that of the first summary data. This makes it possible to improve the accuracy of the response content.
[0394] [Appendix 16] One of the objectives of this disclosure is to output a response that reflects memories with different utterance times. [Appendix 16] The information processing device described in Appendix 15, wherein multiple first summary data or second summary data have different utterance or generation times. This makes it possible to output a response that reflects memories with different utterance times.
[0395] [Appendix 17] One of the purposes of this disclosure is to make the response content more human-like. [Appendix 17] The program described in Appendix 16, wherein one or more processors acquire first summary data or second summary data in which the importance of the conversation content is equal to or greater than a first threshold from a first period, and acquire first summary data or second summary data in which the importance of the conversation content is at least lower than the first threshold from a second period which is closer to the speech tense than the first period. This makes the response content more human-like.
[0396] [Appendix 18] One of the purposes of this disclosure is to enable speech based on memory of past conversations. [Appendix 18] An information processing device as described in Appendix 2, wherein one or more processors, when predetermined conditions are met during or after a conversation with a user, extract first summary data or second summary data that satisfy predetermined rules from a summary, and provide the extracted first summary data or second summary data to a conversation model to output utterances to the user. This enables speech based on memory of past conversations.
[0397] One of the objectives of this disclosure, corresponding to the problem described in [Appendix 19], is to enable speech using high-importance memories. [Appendix 19] The rule is the information processing device described in Appendix 18, which includes the importance of the summary being higher than the threshold value. This enables speech using high-importance memories.
[0398] [Appendix 20] One of the objectives of this disclosure is to realize a response that is close to the characteristics of human memory. [Appendix 20] The information processing method relating to this disclosure involves one or more processors performing the following processes: obtaining relevant summaries that are related to the user's utterances from summaries generated from conversation logs linked to an account; and providing the obtained relevant summaries and utterances to a conversation model to output a response to the utterances. According to the above information processing method, a response that is close to the characteristics of human memory can be realized.
[0399] [Note 21] One of the purposes of this disclosure is to realize a response that closely resembles the characteristics of human memory. [Note 21] The program relating to this disclosure causes one or more processors to obtain relevant summaries that are relevant to the user's utterances from summaries generated from conversation logs associated with an account, and to provide the obtained relevant summaries and utterances to a conversation model to output a response to the utterances. According to the above program, a response that closely resembles the characteristics of human memory can be realized.
[0400] [Note 22] One of the purposes of this disclosure is to realize a response that closely resembles the characteristics of human memory. [Note 22] The information processing system relating to this disclosure comprises a server and a terminal device. The server obtains relevant summaries from the summaries generated from conversation logs linked to an account that are relevant to the user's utterances, and provides the obtained relevant summaries and utterances to a conversation model to output a response to the utterances. According to the above information processing system, a response that closely resembles the characteristics of human memory can be realized.
[0401] 1...Conversation system 10...Owner terminal 11, 21, 31, 41...Processor 12, 22, 32, 42...Semiconductor memory 13, 23, 33, 43...Auxiliary storage device 14...Camera 15...Microphone 16...Speaker 17...Display 18...Movement 19, 24, 34, 44...Communication interface 20...Voice / text conversion server 30...Front server 40...Conversation server 43A...User DB 43B...Conversation log 43C...Summary DB
Claims
Includes one or more processors, The one or more processors mentioned above are: Using the owner's spoken words, images captured by the owner's device, and information related to the owner, a response text is generated. The owner terminal outputs audio corresponding to the generated response text. Information processing system. If the owner's spoken words include instructions for image capture, The one or more processors described above acquire an image from the owner terminal and use it to generate the response text. The information processing system according to claim 1. The one or more processors mentioned above are: Identify the person who spoke as the owner, and use the information related to the identified person as information related to the said owner. The information processing system according to claim 1. The one or more processors mentioned above are: The content of events or schedules registered for an identified person will be used as information related to the said owner. The information processing system according to claim 3. The one or more processors mentioned above are: The system generates the response text using the events or schedules registered for the identified person that have been executed or are scheduled to be executed within a predetermined period including the current date and time. The information processing system according to claim 4. The one or more processors mentioned above are: The system generates the response text using the events or schedules registered for the identified person that have not been used in a conversation within a predetermined period including the current date and time. The information processing system according to claim 4. The one or more processors mentioned above are: The owner shall use the details of events or schedules of one or more individuals registered with the owner as information related to the said owner. The information processing system according to claim 1. The one or more processors mentioned above are: The conversation history recorded for the identified person is used as information related to the owner. The information processing system according to claim 3. The aforementioned conversation history includes information indicating the level of importance of the topics, The one or more processors mentioned above are: Topics with a higher level of importance compared to other topics are used as information related to the owner. The information processing system according to claim 8. The one or more processors mentioned above are: Instead of using the image captured by the owner terminal, information describing the content of the image captured by the owner terminal is used. The information processing system according to claim 1. The one or more processors mentioned above are: Information describing the content of the image captured by the aforementioned owner terminal is obtained through a generating AI. The information processing system according to claim 10. The one or more processors mentioned above are: If the owner's spoken content does not include instructions for image capture, the response text is generated using the owner's spoken content. The information processing system according to claim 1. One or more processors, A process that generates response text using the content spoken by the owner, images captured by the owner's device, and information related to the owner. The process involves outputting audio from the owner terminal corresponding to the generated response text, An information processing method that performs the following. One or more processors, Using the owner's spoken words, images captured by the owner's device, and information related to the owner, a response text is generated. The owner terminal outputs audio corresponding to the generated response text. A program that executes a process. Includes one or more processors, The one or more processors mentioned above are: Using the owner's spoken words, images captured by the owner's device, and information related to the owner, a response text is generated. The owner terminal outputs audio corresponding to the generated response text. Information processing device.