Information processing systems, information processing methods, programs, and information processing devices.
The information processing system personalizes responses by integrating speech, image, and user data, addressing the uniformity issue in existing conversation robots, ensuring varied and engaging interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MIXI INC
- Filing Date
- 2025-04-14
- Publication Date
- 2026-05-27
AI Technical Summary
Existing conversation robots provide uniform responses to the same speech and image content, lacking personalization for different users.
An information processing system that generates personalized response text using speech content, captured images, and user-specific information, enabling varied responses based on individual user data.
Ensures that response content varies for each user, even with identical speech and image inputs, enhancing user engagement and personalization.
Smart Images

Figure 0007866224000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing system, an information processing method, a program, and an information processing apparatus.
Background Art
[0002] Dialogue-type conversation robots have been put into practical use. This type of conversation robot is provided with a function of determining the content of speech based on the conversation log with the user. Hereinafter, a user who uses the conversation service provided by the conversation robot is also referred to as an "owner".
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] There is a conversation robot equipped with a camera. The conversation robot equipped with a camera starts imaging an image in response to an instruction from the owner and outputs an answer to the captured image.
[0005] One object of the present disclosure is to provide an information processing system, an information processing method, a program, and an information processing apparatus in which the response content varies for each owner even when the content of speech and the captured image are the same.
Means for Solving the Problems
[0006] One aspect of the present disclosure includes one or more processors. The one or more processors generate a response text using the content of speech uttered by the owner, the image captured by the owner terminal, and the information related to the owner, and output the voice corresponding to the generated response text from the owner terminal. This is an information processing system.
Effects of the Invention
[0007] According to one form of this disclosure, the content of the response text can be changed for each owner, even if the content of the utterance and the captured image are the same. [Brief explanation of the drawing]
[0008] [Figure 1] This figure shows an example of the overall configuration of the conversation system assumed in Embodiment 1. [Figure 2] This diagram illustrates an example of initiating a conversation from a conversation server's utterance (i.e., a message). [Figure 3] This diagram illustrates an example of a personal identification table stored in an auxiliary storage device. [Figure 4] This diagram illustrates an example of data stored in the user database. [Figure 5] This diagram illustrates an example of the owner screen used for registering owner information. [Figure 6] This is a diagram illustrating an example of a family registration screen. [Figure 7] This is a diagram illustrating an example of a family registration screen. [Figure 8] This diagram illustrates an example of data stored in a conversation log. [Figure 9] This diagram illustrates an example of data stored in the summary database. [Figure 10] This diagram illustrates an example of summary data at the conversational level. [Figure 11] This diagram illustrates the outline of a conversation sequence initiated by a user's utterance. [Figure 12] This diagram illustrates one example of a process that initiates a conversation based on user utterance. [Figure 13] This diagram illustrates an example of the response text output when the identified person is "Hanako". [Figure 14] This diagram illustrates other output examples of response text when the identified person is "Ai-chan". [Figure 15] This diagram illustrates other processing examples that start a conversation based on user utterances. [Figure 16] This is a diagram for explaining an example of the owner's memory information. [Figure 17] This is a diagram for explaining an output example of a response text when the specified person is the "owner". [Figure 18] This is a diagram for explaining another example of the owner's memory information. [Figure 19] This is a diagram for explaining another output example of a response text when the specified person is the "owner". [Figure 20] This is a diagram for explaining another example of a processing operation for starting a conversation based on the user's utterance. [Figure 21] This is a diagram for explaining an example of an event schedule database. [Figure 22] This is a diagram for explaining an output example of a response text using an event or schedule. [Figure 23] This is a diagram for explaining another output example of a response text using an event or schedule. [Figure 24] This is a diagram for explaining another output example of a response text using an event or schedule. [Figure 25] [[ID=二十九]]This is a diagram for explaining another example of a processing operation for starting a conversation based on the user's utterance. [Figure 26] This is a diagram for explaining an overview of a conversation sequence started by the utterance of a conversation server. [Figure 27] This is a diagram for explaining an example of a processing operation for starting a conversation by the utterance of a conversation server. [Figure 28] This is a diagram for explaining an output example of an utterance text when the specified person is "Hanako". [Figure 29] This is a diagram for explaining another output example of an utterance text when the person is not specified. [Figure 30] This is a diagram for explaining another example of a processing operation for starting a conversation by the utterance of a conversation server. [Figure 31] This is a diagram for explaining an output example of a response text when the specified person is the "owner". [Figure 32] This diagram illustrates other output examples of response text when the identified person is the "owner." [Figure 33] This diagram illustrates other processing examples that start a conversation based on an utterance from the conversation server. [Figure 34] This diagram illustrates an example of speech text output using an event or schedule. [Figure 35] This figure illustrates other examples of response text output using events or schedules. [Figure 36] This figure illustrates other examples of response text output using events or schedules. [Figure 37] This diagram illustrates another example of a processing operation initiated by an utterance from a conversational server. [Figure 38] This diagram illustrates other processing examples that start a conversation based on user utterances. [Modes for carrying out the invention]
[0009] Embodiments of this disclosure will be described below with reference to the drawings. The embodiments described below are merely examples of forms for implementing the present disclosure, and the embodiments of the present disclosure are not limited to the examples described below. Therefore, the technical scope of this disclosure is not limited to the embodiments described below. For example, various modifications or improvements to the embodiments described are also included in the technical scope of this disclosure. Furthermore, the various functional components described later are implemented through the execution of programs by the processor.
[0010] <Terminology> First, let's explain the terminology used in the embodiments. The term "computer" includes not only processors as hardware, but also combinations of software programs and hardware processors. Computers may also include general-purpose computers, computers designed for specific purposes, workstations, or other systems capable of performing each type of processing.
[0011] A "processor" includes an integrated circuit that performs various processes in cooperation with a program. The processor can function as the various parts (Units) and means (Means) described in the embodiments below. The execution order of the processes performed by the processor is not limited to the order described in the embodiments and can be changed as necessary. A processor can be configured with one or more hardware components. The types of hardware that make up the processor are not limited to any particular type.
[0012] The processor can be, for example, a CPU (=Central Processing Unit), an MPU (=Micro Processing Unit), a programmable logic device such as an FPGA (=Field Programmable Gate Array), a dedicated circuit for executing specific processing such as an ASIC (=Application Specific Integrated Circuit), a GPU (=Graphics Processing Unit), or hardware such as an NPU (=Neural Processing Unit).
[0013] A processor can be a combination of multiple hardware components of different types, not just multiple hardware components of the same type. When multiple hardware components are configured to perform one or more processes of a given processor, these components may reside in physically separate devices or in the same device. Hardware is composed of circuits, which are combinations of circuit elements such as semiconductor devices.
[0014] "Program" includes software such as the OS (Operating System), firmware, application programs, and microcode. The program may be, for example, a group of program modules. Each function constituting the group of program modules may be implemented by a processor configured to execute each function. The program in each embodiment may be program code or multiple code segments stored in one or more non-temporary computer-readable media (e.g., semiconductor memory, magnetic or optical storage media, or other storage).
[0015] A program may be divided and stored on multiple non-temporary computer-readable media located on devices that are physically separated from each other. Program code and multiple code segments can be represented as any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, instructions, data structures, and program statements. Program code and multiple code segments may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents. The program may run on multiple terminals located on devices that are physically separated from each other.
[0016] "Conversation" includes exchanges using natural language, either spoken or written. Written exchanges are also called "chat." Incidentally, "exchange" can also be called "response." Conversation may also include dialogues, discussions, and debates. "Conversation" may include non-natural language reactions. Reactions may include, for example, the output of sound effects, light emission, changes in light emission color, changes in light emission pattern, vibration, and the movement of movable parts.
[0017] A "topic" refers to the main subject of a conversation. Topics can include, for example, events from everyday life. Conversations that focus on everyday events are sometimes called "small talk." However, the "topic" doesn't have to be limited to events in everyday life. For example, in conversations, discussions, and debates, a specific topic is designated in advance. This specific topic could include, for example, problems, issues, solutions, or likes and dislikes. The "topic" can also be determined by the words included in the conversation. For example, if a user says, "I went to a friend's wedding in my hometown today," possible topics could include "wedding," "marriage," "hometown," and "friend."
[0018] "Speaker" refers to a person participating in a conversation. In this specification, one of the participants is a natural person (hereinafter referred to as "person"). There may be one or more people. The other participant is a virtual character based on AI (=Artificial Intelligence) technology (hereinafter also referred to as "AI character"). The AI character in question includes application programs and conversation models that understand the content of natural language spoken or input by the other party and generate appropriate response sentences. The AI character described later includes application programs and conversation models that generate appropriate question sentences without using the other party's speech or input.
[0019] The content of what the AI character speaks is determined based on rule-based or machine learning models (hereinafter also referred to as "conversation models"). The term "AI character" can refer to, for example, one of the objects displayed on a screen, a conversational service, or a physical AI speaker. Many physical AI speakers are used within the owner's home. The term "conversational model" can refer to the machine learning model itself, or it can refer to a conversational service that utilizes a machine learning model.
[0020] An "AI speaker" is a device that includes a microphone that converts human speech into electrical signals and a speaker that converts text data into speech. An AI speaker may be a computer terminal such as a smartphone equipped with voice assistant functionality, or it may be a dedicated device specifically designed for conversation. Dedicated devices specifically designed for conversation include, for example, smart speakers. The external shape of a dedicated device may include, for example, a cylinder, a robot, a pet, or a fictional creature.
[0021] "Conversation text" includes both text representing human speech recognized by speech recognition technology and text representing speech by an AI character. AI character speech includes speech as a response to user speech and speech initiated by the AI character. The text corresponding to the user or AI character may include, for example, one or more sentences, interjections, and slurs. Note that sentences spoken in this manner are not necessarily grammatically correct.
[0022] A "conversation log" refers to a record of conversations linked to a user account. In other words, a conversation log is a record of the conversation text. The conversation log accurately records the content of the conversation down to the last word. However, if there are errors in speech recognition, the conversation log will record the incorrectly recognized utterances as they are. Due to the nature of conversation logs as data that records the exact content of speech, they may be discarded at any time (for example, at midnight every day) from a privacy standpoint, or they may be stored in a database for a predetermined period. A "summary generated from a conversation log" uses a single conversation as its smallest unit. Hereafter, a summary at the conversation level will be referred to as "conversation-level summary data." In addition, "summaries generated from conversation logs" also include summaries generated from multiple summaries.
[0023] A conversational text refers to a series of interactions between a person and an AI character. Technically, a "single conversation" is considered to be the entire exchange from the first utterance to the last utterance. The first utterance is, for example, an utterance that has a predetermined period of silence (e.g., 5 minutes or more) between it and the preceding utterance. On the other hand, the last utterance is, for example, an utterance that has a predetermined period of silence (e.g., 5 minutes or more) between it and the next utterance. Therefore, even if there is a one-minute silence between one utterance by a human or AI character, the utterances before and after the one-minute silence will be treated as utterances within a single conversation.
[0024] The subjects of a "summary" are determined, for example, by a rule-based system. The rules may include specifying a time period (e.g., conversation units, days, months, or years). In addition, the rules may include grouping (i.e., classifying) utterances or summaries with similar content. To generate summaries, for example, a conversation model is used. Instructions given to the conversation model are called "prompts." For example, a conversation log and prompts are given to the conversation model to generate a summary. Alternatively, a new summary can be generated by giving the conversation model a summary generated from the conversation log and a prompt. The generated summaries are also used to generate responses (i.e., response text) to human utterances.
[0025] <Embodiment 1> <System Configuration> Figure 1 shows an example of the overall configuration of the conversation system 1 assumed in Embodiment 1. The conversation system 1 is an example of an information processing system. The conversation system 1 includes an owner terminal 10 connected via network N, a voice / text conversion server 20, a front server 30, and a conversation server 40.
[0026] Owner terminal 10 is an example of a terminal operated by an owner using the conversation service. However, owner terminal 10 does not need to be a dedicated device for the owner registered with the conversation service. For example, owner terminal 10 may be shared by multiple people, including the owner. Owner terminal 10 is an example of an information processing device.
[0027] The voice / text conversion server 20, the front server 30, and the conversation server 40 are terminals (i.e., servers) on the provider side of the conversation service. Incidentally, the conversation server 40 is also an example of an information processing device. Network N includes, for example, the Internet, LAN (=Local Area Network), and mobile communication systems such as 4G and 5G. Network N may be a wired network, a wireless network, or a hybrid of wired and wireless networks.
[0028] The owner terminal 10 is a terminal that transmits voice and other information to the voice / text conversion server 20, while outputting data received from the cloud. The audio transmitted by the owner terminal 10 to the voice / text conversion server 20 may include the voices of people other than the owner. In this embodiment, the owner terminal 10 transmits all sounds detected by the microphone 15 to the voice / text conversion server 20.
[0029] The owner terminal 10 may have a function to identify the speaker. If the owner terminal 10 can identify the speaker, it sends the audio with speaker information to the speech / text conversion server 20. However, it is also possible to operate the system without sending the identified speaker information as supplementary information to the audio. If the speaker cannot be identified, the voice is considered to be from an unknown person.
[0030] Data related to the conversation service is transmitted and received in streaming format. That is, data related to the conversation service is sent from the owner terminal 10 to the cloud in real time, and the processing results on the cloud side are sent back to the owner terminal 10 in real time. The data streamed from the owner terminal 10 to the cloud may include, for example, user voice data (i.e., audio data), output data from touch sensors and accelerometers (i.e., sensor data), and image data of the user, etc. Image data of the user, etc. may also include, for example, the user's background image.
[0031] The data streamed from the cloud to the owner terminal 10 includes, for example, audio data, as well as control data and image data. The control data includes, for example, instructions for outputting sound effects and melodies, instructions for outputting vibrations, instructions for movement to movable mechanisms, and instructions for outputting predetermined images. The image data includes, for example, images that the cloud instructs to output according to the content and progress of the conversation.
[0032] In Figure 1, the owner terminal 10 is a so-called AI speaker. As mentioned above, AI speakers include computer terminals equipped with voice assistance functions and dedicated devices specialized for conversation. In this embodiment, we assume that the owner terminal 10 is a dedicated device specialized for conversation. More specifically, we assume a robot-type owner terminal 10 in which the head is movable relative to the main body. In Figure 1, there is one owner terminal 10 connected to network N. However, in reality, multiple owner terminals 10 are connected to network N.
[0033] In this embodiment, one owner terminal 10 is provided for each owner who uses the conversation service. In other words, the owner terminal 10 is linked to one user account. The user (i.e., owner) linked to the user account is the subscriber to the conversation service. However, it is also possible for multiple people to use a single owner terminal 10 linked to a single user account. In this case, the utterances of multiple people will be recorded and linked to a single user account. If it is possible to identify the person who made the utterance, each utterance may be recorded separately for each person linked to a single user account.
[0034] It is also possible to link multiple owner devices 10 to a single user account. For example, one owner may be linked to a conversational robot specialized in conversation and a computer terminal such as a smartphone. By linking multiple owner devices 10 to a single user account, it becomes possible to switch between using the AI speaker at home and while out and about.
[0035] In addition, it is possible to link multiple user accounts to a single owner device 10. In other words, a single owner device 10 can be shared by multiple users. For example, if it is possible to identify or distinguish the user using the owner device 10 (e.g., the speaker or the person being talked to) using speech recognition technology or image recognition technology, then it is possible to share a single owner device 10 with multiple users. When a single owner device 10 is shared by multiple users, the content of conversations is managed separately for each user.
[0036] The speech / text conversion server 20 is a server that converts speech data received from the owner terminal 10 into text. Hereafter, the text converted by the speech / text conversion server 20 will be referred to as "spoken text". In the case of Figure 1, the spoken text is sent to the front server 30. The speech / text conversion server 20 may also send the spoken text to both the front server 30 and the conversation server 40. Sensor data and image data received from the owner terminal 10 are transferred to the front server 30 separately from the spoken text. However, sensor data and image data may also be sent directly from the owner terminal 10 to the front server 30.
[0037] The front server 30 is a server that instructs the owner terminal 10 to output simple responses and replies to input from the owner terminal 10. Simple responses and replies are expected to serve the purpose of detecting the start and end of speech. Simple responses and replies may include, for example, the output of sound effects or melodies, the movement of movable parts, the output of vibrations, and changes in the screen. In Figure 1, the front server 30 forwards the spoken text received from the speech / text conversion server 20 to the conversation server 40 in real time.
[0038] The conversation server 40 is a server that generates responses based on spoken text, etc. The conversation server 40 generates response text using not only the latest spoken text, but also past conversations (i.e., memories) that are related to the content of the latest spoken text. In this embodiment, a summary of past conversations (summary DB43C, described later) is used as the past conversation. However, the past conversation itself (conversation log 43B, described later) from within the last month may be used as the past conversation, or both the conversation itself from within the last month and a summary of the past conversation may be used. By summarizing conversations, less important information is lost, while more important information remains.
[0039] Furthermore, the summary DB43C may generate new summaries (re-summarization) from multiple summaries belonging to each predetermined period. For example, summaries generated from conversations in the same month of utterance may be re-summarized to generate monthly summaries. When re-summarizing, the importance level of each summary may be utilized. For example, the content of high-importance summaries may be more likely to be retained after re-summarization. In other words, the content of low-importance summaries may be more likely to be lost through re-summarization. As a result, the information stored in the summary DB43C will be closer to the nature of human memory, where "trivial matters are forgotten over time, and important matters are remembered (even after time has passed)."
[0040] In addition, the conversation server 40 may also have a function to generate responses to input from the owner terminal 10 (e.g., sensor data or image data). The generated control data and image data are sent from the conversation server 40 to the owner terminal 10 along with the response text. Furthermore, the conversation server 40 in this embodiment may have a function to initiate a conversation without the owner speaking. That is, the conversation server 40 may have a function to generate speech text independently of the owner's speech and speak to the owner. However, the content of the speech will be changed according to the situation of the person being spoken to (i.e., the owner).
[0041] Figure 2 illustrates an example of initiating a conversation from an utterance (i.e., a speech) by the conversation server 40. Figure 2 is denoted with corresponding symbols for parts that correspond to those in Figure 1. However, in the conversation system 1 shown in Figure 2, the internal configurations of the owner terminal 10, the voice / text conversion server 20, and the front server 30 are omitted. The conversation server 40 shown in Figure 2 generates spoken text according to predetermined conditions when the owner terminal 10 meets predetermined conditions.
[0042] The specified conditions include, for example, the detection of pre-registered facial images, pre-registered voiceprints, pre-registered events or schedules, and images related to past conversations. In the following, the conversation text sent to the owner terminal 10 as a response from the conversation server 40 to the owner's utterance will be referred to as the "response text," and the conversation text when the conversation server 40 initiates a conversation will be referred to as the "utterance text."
[0043] <Device Hardware Configuration> The following describes the hardware configuration of each terminal, referring to Figure 1. <Owner's terminal> The owner terminal 10 includes a processor 11, semiconductor memory 12, auxiliary storage device 13, camera 14, microphone 15, speaker 16, display 17, movable mechanism 18, and communication interface 19. The processor 11 and each device are connected to each other via buses and other signal lines.
[0044] The processor 11 is, for example, a CPU. The semiconductor memory 12 includes ROM (Read Only Memory) which stores UEFI (Unified Extensible Firmware Interface), etc., and RAM (Random Access Memory) which is used as the work area of the processor 11. The processor 11 and the semiconductor memory 12 operate as a computer.
[0045] The auxiliary storage device 13 is, for example, a hard disk drive or semiconductor storage. The auxiliary storage device 13 stores programs and data necessary to realize the functions of an AI speaker. The programs here include a program that converts response text and spoken text received from the cloud into voice data, a program that controls the output according to control data received from the cloud, and a program that displays image data received from the cloud.
[0046] Figure 3 illustrates an example of a personal identification table 13A stored in the auxiliary storage device 13. The personal identification table 13A is used on the owner terminal 10 (see Figure 1) side to identify a person. The personal identification table 13A shown in Figure 3 includes user ID 13A1, facial information 13A2, and voiceprint information 13A3. User ID 13A1 stores, for example, the management ID (=identifier) of a registered user (e.g., the owner and the owner's family) who uses the owner terminal 10 (see Figure 1).
[0047] In Figure 3, the personal identification table 13A stores information for three people. User ID 13A1 is the same as User ID 43A1 (see Figure 4), which is managed by the conversation server 40 (see Figure 1). User IDs 13A1 and 43A1 are managed on a per-owner terminal 10 basis used for registration. Therefore, even if it is the same person, if the owner terminal 10 used for registration is different, a different ID will be assigned. Note that User ID 13A1 does not need to be related to User ID 43A1. In other words, User ID 13A1 may be a local ID used only on owner terminal 10. The facial information 13A2 stores the facial image captured by the camera 14 (see Figure 1). The voiceprint information 13A3 stores the voice and extracted features recorded by the microphone 15 (see Figure 1).
[0048] Returning to the explanation of Figure 1. Camera 14 is an imaging device, and for example, a CMOS (Complementary Metal Oxide Semiconductor) image sensor is used. Camera 14 may be integrated with the main body of the owner terminal 10, or it may be attached externally to the main body of the owner terminal 10. Camera 14 may be detachable from the owner terminal 10. Camera 14 may also be wirelessly connected to the owner terminal 10. Microphone 15 is a device that converts the voice of the user operating the owner terminal 10 into an electrical signal (i.e., voice data).
[0049] The speaker 16 is a device that converts voice data or sound data received from the conversation server 40 into air vibrations (i.e., sound). However, the voice data or sound data to be converted may be provided by various programs or the like running on the owner terminal 10. The display 17 is, for example, a liquid crystal display or an organic EL (=Electro-Luminescence) display. The display 17 displays, for example, the face image of an AI character. A capacitive touch sensor may be laminated on the surface of the display 17. A display 17 with this type of touch sensor integrated is called a touch panel.
[0050] The movable mechanism 18 is provided when the owner terminal 10 is a conversational robot modeled after a character or living creature. However, even if it is a conversational robot, the movable mechanism 18 may not be provided. The movable mechanism 18 includes a power source (e.g., a motor) and a drive mechanism that transmits power to realize a predetermined operation. For example, if the conversational robot consists of a torso and a head, the movable mechanism 18 enables the head to rotate up and down or left and right. Also, if the conversational robot has arms and legs attached to its torso, the movable mechanism 18 enables the movement of the arms and legs.
[0051] The communication interface 19 is a device that enables communication with a server that provides conversation services. The communication interface 19 is equipped with communication functions that are compatible with network N. The owner terminal 10 may also include a touch sensor and an accelerometer. The touch sensor is used, for example, to detect an action such as stroking the conversational robot. The accelerometer is used to detect an action such as lifting the conversational robot. In addition, the owner terminal 10 may be equipped with, for example, a temperature sensor, a humidity sensor, and an illuminance sensor. The outputs of these sensors may be used for the conversation service.
[0052] <Speech / Text Conversion Server> The speech / text conversion server 20 includes a processor 21, semiconductor memory 22, auxiliary storage device 23, and a communication interface 24. The processor 21 and each device are connected to each other via buses and other signal lines. The processor 21 is, for example, a CPU. The semiconductor memory 22 includes ROM that stores UEFI, etc., and RAM used as the work area of the processor 21. The processor 21 and the semiconductor memory 22 operate as a computer.
[0053] The auxiliary storage device 23 is, for example, a hard disk drive or semiconductor storage. The auxiliary storage device 23 stores a program for converting audio data into text. The communication interface 24 is a device that enables communication with other servers providing conversation services. The communication interface 24 is equipped with communication functions compatible with network N.
[0054] <Front Server> The front server 30 includes a processor 31, semiconductor memory 32, auxiliary storage device 33, and a communication interface 34. The processor 31 and each device are connected to each other via buses and other signal lines. The processor 31 is, for example, a CPU. The semiconductor memory 32 includes ROM that stores UEFI, etc., and RAM used as a work area for the processor 31. The processor 31 and the semiconductor memory 32 operate as a computer.
[0055] The auxiliary storage device 33 is, for example, a hard disk drive or semiconductor storage. The auxiliary storage device 33 stores programs and data tables for outputting simple reactions and responses. Through the execution of the program, sound effects and the like are output from the owner terminal 10, for example, when the start or end of speech is detected. The communication interface 34 is a device that enables communication with other servers and owner terminals 10 that provide conversation services. The communication interface 34 is equipped with communication functions that are compatible with network N.
[0056] <Conversation Server> The conversation server 40 includes a processor 41, semiconductor memory 42, auxiliary storage device 43, and a communication interface 44. The processor 41 and each device are connected to each other via buses and other signal lines. The processor 41 is, for example, a CPU. The semiconductor memory 42 includes ROM that stores UEFI, etc., and RAM used as a work area for the processor 41. The processor 41 and the semiconductor memory 42 operate as a computer.
[0057] The auxiliary storage device 43 is, for example, a hard disk drive or semiconductor storage. The auxiliary storage device 43 stores programs and various data for realizing conversational services. The programs here include, for example, a program called a conversational engine. The conversational engine includes, for example, a program that generates response text to spoken text. The conversational engine is also called a chatbot. Conversational engines are classified into, for example, function-specific, general-purpose rule-based, and AI-type.
[0058] In this embodiment, the auxiliary storage device 43 stores multiple conversation engines. However, some or all of the conversation engines may be provided as a conversation service independent of the conversation server 40. Other programs include a program for recording conversation logs and a program for summarizing conversation content. However, a generation AI independent of the conversation server 40 may be used for generating summaries.
[0059] The auxiliary storage device 43 stores the user DB (=Data Base) 43A, the conversation log 43B, and the summary DB 43C. User DB43A is a database that stores information on all users (the owner and their family) who use the conversation service. In this embodiment, users other than the owner who have registered themselves are considered "the owner's family." Users other than the owner who have registered themselves are also called "family users." Conversation Log 43B is a database that stores the history of conversations between the user and the AI character (i.e., conversation text). The conversation text is managed in conjunction with the user account.
[0060] The summary DB43C is a database that stores summaries of the content of conversations (i.e., conversation text). In this embodiment, the summary DB43C stores multiple summaries that are different in length of time since the conversation (e.g., 1 day, 1 week, 1 month, 1 year). Furthermore, the summary DB43C stores vectors representing the content of the conversation, which are used to detect summaries similar to the dialogue text. In this embodiment, the vectors include, for example, the end time of the conversation, a summary of the conversation text (i.e., the summary text), and importance. The dimensions of the vectors can be, for example, several hundred or more.
[0061] Importance is given as a score that represents the importance of the content of the summary. For example, the score is given as a numerical value (e.g., on a scale of 1 to 10) corresponding to the topics included in the summary. For example, in a conversation on the day of a school sports day, the importance of the sports day will be higher than that of food. On the other hand, if someone is highly interested in health, the topic of food is likely to come up frequently in conversation. Topics that appear frequently in conversations like this tend to be topics of high importance to the user.
[0062] To generate a score indicating importance, a generation AI independent of the conversation server 40 is used, for example. The prompts given to the generation AI include, for example, the scoring rules and the conversation text or summary to be summarized. In this embodiment, the scoring rules are pre-configured by the conversation service provider. These scoring rules are common to all users.
[0063] However, the scoring rules may be changed for each user. For example, the conversation server 40 (see Figure 1) may switch the scoring rules applied to a user according to the user attribute 43A4 (see Figure 4). Furthermore, the owner terminal 10 may allow the user to instruct the system to set or switch the scoring rules applied to themselves. By switching the scoring rules applied, you can adjust, for example, the topics that remain after summarization or the topics that appear in the response text.
[0064] The conversation text includes text representing what the user said and text representing what the AI character said. In this embodiment, conversational texts containing three or more texts are subject to summarization. However, conversations consisting of two or more texts may also be included in the summarization. In this embodiment, previously generated summary data is further summarized (i.e., re-summarized) to generate new summary data. For example, daily summary data for one month is re-summarized together to generate monthly summary data. After the new summary data is generated, the previously generated summary data used for re-summarization is deleted.
[0065] The communication interface 44 is a device that enables communication with other servers providing conversation services and with the owner terminal 10. The communication interface 44 is equipped with communication functions adapted to the network N. In this embodiment, the communication interface 44 is used to transmit response text, control data, and image data to the owner terminal 10.
[0066] <Data Example> <User DB> Figure 4 illustrates an example of data stored in the user database 43A. The user database 43A shown in Figure 4 includes the user ID 43A1, the "How you address" field 43A2, the "Relationship to the owner" field 43A3, and attribute information 43A4. In addition, if necessary, the management ID, email address, password, gender, place of residence, date of birth, conversation service usage plan, and other information of the owner terminal 10 are also stored.
[0067] User ID 43A1 stores the management ID of the user using the conversation service. The content of User ID 43A1 shown in Figure 4 is the same as User ID 13A1 shown in Figure 3. In this embodiment, User ID 43A1 is managed for each owner terminal 10 used to register a new user. Therefore, even for the same person, a different User ID 43A1 will be assigned if a different owner terminal 10 was used to register the new user.
[0068] However, if one owner uses multiple owner terminals 10, the user ID 13A1 associated with the single owner may be shared across the multiple owner terminals 10. If the user ID 13A1 is common across multiple owner terminals 10 linked to a single owner, it becomes possible to manage conversations exchanged by the same person (for example, "Taro") through different owner terminals 10 as conversations with the same person. Of course, this assumes that the same person is authenticated on all owner terminals 10.
[0069] The "How to Address" field 43A2 stores the name that the conversation service uses to address the registered user. For example, the conversation service's response text will use the information in the "How to Address" field 43A2, such as "Hello, Ai-chan". The "Relationship with Owner" column 43A3 stores the relationship between the registered user and the owner. In Figure 4, "Taro," whose user ID is "3432," is the owner himself (i.e., the contracted user of the conversation service). "Hanako," whose user ID is "3433," is the wife of "3432." "Ai-chan," whose user ID is "3521," is the daughter of "3432." "Hanako" and "Ai-chan" are family users of "Taro."
[0070] Attribute information 43A4 stores information representing the attributes of each user. Attribute information 43A4 can be registered or edited, for example, by the owner. However, after registration by the owner, family users may also register or edit their own attribute information. Attribute information 43A4 may be registered by the conversation server 40. For example, the conversation server 40 may extract and store specific information (e.g., hobbies, preferences, likes and dislikes, desires, interests, concerns) that appears in the user's conversation. For example, the conversation server 40 may store highly important information extracted from the summary DB 43C (see Figure 1) as attribute information 43A4. The information extracted and stored by the conversation server 40 may overlap with the contents of the conversation log 43B (see Figure 1) and the summary DB 43C, but extracting it separately makes it easier to refer to when creating conversation text.
[0071] Figure 5 illustrates an example of the owner screen 500 used for registering owner information. The owner screen 500 shown in Figure 5 corresponds to the "Owner Information" screen of "MyRoom" provided by the conversation service shown in Figure 5. The owner screen 500 is displayed by accessing the conversation server 40 (see Figure 1). The owner screen 500 shown in Figure 5 includes a name field 501, an email address field 502, a password field 503, a gender field 504, a place of residence field 505, and a date of birth field 506. All or part of this information is also stored on the owner terminal 10 (see Figure 1).
[0072] Figure 6 illustrates an example of the family registration screen 510. The family registration screen 510 shown in Figure 6 corresponds to the "Family Registration" screen of "MyRoom" in the conversation service. The family registration screen 510 includes a "Remember owner" field 511, a "Remember family" field 512, and an "Add family" button 513.
[0073] In Figure 6, the "Memory for Owner Recognition" section 511 includes a face information section 511A and a voice information section 511B. In Figure 6, both the face and voice are already memorized. By tapping the face information section 511A, it is possible to delete or re-register a registered face. Similarly, by tapping the voice information section 511B, it is possible to delete or re-register a registered voice.
[0074] In Figure 6, the "Memory for Recognizing Family" section 512 includes an information update button 512A and a section 512B for registered family information. In Figure 6, two people, "Hanako" and "Ai-chan," are registered as family users. However, only their names 43A2 (see Figure 4) are registered for "Hanako" and "Ai-chan," and their facial images and voiceprints are not registered. Therefore, the camera icon and speaker icon corresponding to "Hanako" and "Ai-chan" are both grayed out. Incidentally, when either a face or voice is registered, the corresponding icon's display switches from grayed out to color.
[0075] Furthermore, pressing the ">" button will switch to the registration screen 520 (see Figure 7) corresponding to each user. By clicking the "Add Family" button 513, the new user registration screen is displayed. In this embodiment, up to three family users can be registered. Family users (hereinafter also referred to as "family") do not include the owner user.
[0076] Figure 7 illustrates an example of the family registration screen 520. The family registration screen 520 is used as a new user registration screen, as well as a registration screen for already registered users. The registration screen 520 shown in Figure 7 includes a name input field 521, a gender input field 522, a date of birth input field 523, a relationship to the owner input field 524, a face memory button 525, a voice memory button 526, a save button 527, and a "cancel" button 528.
[0077] The registration screen 520 shown in Figure 7 is an example of the registration screen for "Ai-chan," the owner's daughter. In Figure 7, the face of the daughter is being registered, so the face memory button 525 is displayed in color. On the other hand, the voice memory button 526 is grayed out because the daughter's voice is not being registered. When the save button 527 is tapped, the registration of the information on the screen is confirmed. If the "cancel" button 528 is pressed, the input is canceled and the user returns to, for example, the family registration screen 510 (see Figure 5).
[0078] <Conversation Log> Figure 8 illustrates an example of data stored in the conversation log 43B. The conversation log 43B shown in Figure 8 includes user ID 43B1, conversation ID 43B2, start date and time 43B3, end date and time 43B4, and conversation text 43B5. User ID 43B1 is the management ID for users of the conversation service and is the same as User ID 43A1 (see Figure 4).
[0079] Conversation ID 43B2 is the conversation management ID. As shown in Figure 8, conversation ID 43B2 is managed on a per-user account basis (i.e., per-owner basis). However, conversation ID 43B2 may also be used to manage the conversation history on a per-user ID basis, including family users registered by the owner. The term "family user" is used to refer to a user registered as part of the owner's family. However, "family" here does not necessarily mean a family relationship in the conventional sense, blood relation, or cohabitation; it may also include friends and acquaintances. The start date and time 43B3 is information about when the conversation started. For example, the date on which user ID 43B1 started a conversation with user "10001" was "February 11, 2025", and the time the conversation started was "11:03:42".
[0080] The end date and time 43B4 is information about the date and time the conversation ended. For example, the date on which user ID 43B1 ended the conversation with user "10001" was "February 11, 2025", and the time the conversation started was "11:05:13". Conversation text 43B5 is a record of the conversation. In Figure 8, to enable speaker identification, the text spoken by the user is marked with "Mr. / Ms. A," etc., and the text spoken by the AI character is marked with "AI character" at the beginning of the sentence. However, from a data perspective, it is sufficient if the recording is in a format that allows for distinction between the two.
[0081] <Summary DB> Figure 9 illustrates an example of data stored in the summary DB43C. In Figure 9, the summary DB43C is managed on a per-user account basis. In the case of Figure 9, Person A, Person B, and Person C could all be owners or family users. The data structure of Summary DB43C is common among users. Figure 9 shows an example of Summary DB43C data, with Person A as a representative example.
[0082] The summary DB43C shown in Figure 9 includes summary data 43C1 at the conversation level, summary data 43C2 at the daily level, summary data 43C3 at the monthly level, and summary data 43C4 at the yearly level. In other words, the summary DB43C contains four types of summary data with different time periods. In this embodiment, summary data 43C2 is generated from one or more summary data 43C1 with the same conversation date. Summary data 43C3 is generated from one or more summary data 43C2 with the same conversation month. Summary data 43C4 is generated from one or more summary data 43C3 with the same conversation year.
[0083] By repeatedly summarizing in a hierarchical manner, information of relatively low importance or infrequent occurrence tends to be lost. In other words, information of relatively high importance or frequent occurrence tends to remain. In this embodiment, when new summary data is generated, the summary data used for its generation is deleted. For example, after the generation of summary data 43C2, one or more summary data 43C1 with the same conversation date are deleted. After the generation of summary data 43C3, one or more summary data 43C2 with the same conversation month are deleted. After the generation of summary data 43C4, one or more summary data 43C3 with the same conversation year are deleted.
[0084] Figure 10 illustrates an example of summary data 43C1 for a conversation unit. The summary data 43C1 shown in Figure 10 includes a user ID 43C11, a conversation ID 43C12, a start date and time 43C13, an end date and time 43C14, a summary text 43C15, an importance score 43C16, and a vector value 43C17. User ID 43C11 is the management ID for users utilizing the conversation service, and is shared with User IDs 43A1 (see Figure 4) and 43B1 (see Figure 8).
[0085] Conversation ID 43C12 is the management ID for the conversation and is the same as conversation ID 43B2 (see Figure 8). The start date and time 43C13 and the end date and time 43C14 are the same as the start date and time 43B3 (see Figure 8) and the end date and time 43B4 (see Figure 8), respectively. Summary text 43C15 is a summary of conversation text 43B5 (see Figure 8). For example, if conversation ID 43C12 is "A1001", the summary text 43C15 is "The weather is nice, but the wind is cold."
[0086] Importance 43C16 stores a score representing the importance of the summary's content. Incidentally, topics related to daily life are given lower importance. For example, "The owner took a bath" is assigned a score of "0". On the other hand, special events are given higher importance. For example, "A grandchild was born" is assigned a score of "9". Note that even for topics related to daily life, scoring rules may be established so that topics that are frequent or appear many times within the target period (e.g., conversation, day, month, year) are given higher importance. Vector value 43C17 stores the vector value of the summary's content. The vector value is stored as a sequence of hundreds of numbers, such as [0.37, 0.91, 0.35, ..., 0.77].
[0087] Note that the generation process for summary data 43C1, 43C2, 43C3, and 43C4 (see Figure 9) is performed independently of the conversation sequence described later. For example, summary data 43C2 is generated from summary data 43C1 one month after the utterance date. For example, summary data 43C3 is generated from summary data 43C2 one year after the utterance month. For example, summary data 43C4 is generated from summary data 43C3 more than one year after the utterance year.
[0088] <Conversation Sequence> The following describes the provision of conversational services through conversational system 1 (see Figure 1).
[0089] <When the conversation is initiated by the user's utterance> <Overview> Figure 11 illustrates an overview of a conversation sequence initiated by user utterance. The symbol S in the figure represents a step. Note that the conversation server 40 in Figure 11 is a collective term for the speech / text conversion server 20 (see Figure 1), the front server 30 (see Figure 1), and the conversation server 40.
[0090] Step 101 The owner (including family users) speaks to the owner terminal 10. The owner's voice is converted into an electrical signal (i.e., voice data) through the microphone of the owner terminal 10 (see Figure 1).
[0091] Step 102 The owner terminal 10 sends voice data to the conversation server 40.
[0092] Step 103 The conversation server 40 analyzes the audio data. If the audio data does not contain imaging instructions, the conversation server 40 starts the process of generating response text. In response, if the audio data includes an imaging instruction, the imaging instruction is sent to the owner terminal 10. The conversation server 40 suspends the generation of the response text until it has acquired image data from the owner terminal 10.
[0093] Step 104 Upon receiving the imaging command, the owner terminal 10 starts imaging with the camera 14 (see Figure 1). The owner terminal 10 also transmits the image data captured by the camera 14 to the conversation server 40. The image data may be a still image or a moving image.
[0094] Step 105 Upon receiving the image data, the conversation server 40 begins the process of generating a response text. In this embodiment, the conversation server 40 generates a response text using the content spoken by the owner, the received image data, and information related to the owner. Information related to the owner includes, for example, the person, gender, relationship with the owner 43A3 (see Figure 4), conversation history (e.g., conversation log 43B (see Figure 1), summary DB 43C (see Figure 1)), event information, schedule information, and attribute information 43A4 (see Figure 4). Here, "person" includes, for example, the way the owner is addressed by the owner themselves (e.g., see Figures 4-7). The way the owner is addressed is used when the AI character speaks to the owner.
[0095] In other words, information related to the owner is different from the image data captured by camera 14 and the audio data acquired by microphone 15. Information related to the owner may include information related not only to the owner themselves but also to family users. The conversation history is an example of memory information related to the owner.
[0096] Event information and schedule information may be included in attribute information 43A4, or they may be information independent of attribute information 43A4. To generate response text, it is not necessary to use all the information contained in the owner-related information; one or more of that information may be used. For example, only the person, gender, relationship to the owner 43A3 (see Figure 4), conversation history, and summary DB 43C may be used. Once the response text is generated, the conversation server 40 converts the response text into speech synthesis data. The conversation server 40 sends the generated speech synthesis data to the owner terminal 10.
[0097] Step 106 The owner terminal 10 outputs speech synthesis data as sound from the speaker 16 (see Figure 1). This enables the owner terminal 10 to respond to the owner's speech. In other words, a conversation begins between the owner and the owner terminal 10. From this point onward, the processing operations in steps 101 to 106 are repeated.
[0098] <Example of processing operation> The following describes an example of processing operations performed through the cooperation between the owner terminal 10 and the conversation server 40. Needless to say, the processing examples described later are merely examples. For example, it is possible to execute some or all of the steps performed on the owner terminal 10 on the conversation server 40. Conversely, it is also possible to perform some or all of the steps performed on the conversation server 40 on the owner terminal 10.
[0099] Furthermore, the processing operation examples shown in Figures 11, 12, 15, 20, 25, and 38, which will be described later, can be combined with each other. In addition, the combinations may include one or more of the steps shown in Figures 26, 27, 30, 33, and 37, which will be described later.
[0100] <Example 1> Figure 12 illustrates one example of a process that initiates a conversation based on user utterance. In Figure 12, the conversation server 40 is also a collective term for the speech / text conversion server 20, the front server 30, and the conversation server 40.
[0101] Step 111 The owner terminal 10 accepts voice input. Incidentally, voice input is possible not only with the owner's voice but also with the voices of family users.
[0102] Step 112 The owner terminal 10 identifies the person who spoke. For example, registered facial images or voiceprints are used to identify the speaker. In the mechanism shown in Figure 12, it is assumed that the user's facial image or voiceprint is stored in the owner terminal 10. In the case of Figure 12, camera 14 (see Figure 1) is constantly capturing images. For example, it captures several images per second, regardless of audio input. When voice input is received, the owner terminal 10 uses the captured image to identify the person who spoke. Note that the image capture in step 112 is used solely for identifying the person who spoke. After identifying the person who spoke, the captured image is immediately deleted from the semiconductor memory 12 (see Figure 1), etc.
[0103] If voice input is not being received, the captured image is immediately deleted from the semiconductor memory 12 (see Figure 1), etc. Regardless of whether voice input is detected or not, image data at this stage is not transmitted externally, including to the conversation server 40. This prevents images unintended by the owner from being transmitted externally.
[0104] In addition, the camera 14 may initiate image acquisition triggered by the detection of audio input. This audio input detection does not involve natural language processing; therefore, the content of the audio input is not analyzed. Person identification can be done on an individual basis, for example. For instance, whether someone is the owner or a family user (e.g., wife, daughter) can be determined from their registration data. Furthermore, the person who spoke can be identified by distinguishing between the owner and others (i.e., family users). Furthermore, the person who spoke can be identified by distinguishing between registered users (including owners and family users) and unregistered individuals. The information of the identified individual is used as person information to generate the response text.
[0105] Step 113 The owner terminal 10 transmits person information and voice data to the conversation server 40. As mentioned above, image data is not transmitted to the conversation server 40. This is because the image capture in step 112 is used solely for identifying the person who spoke. When using the personal identification table 13A (see Figure 3), the owner terminal 10 may send the user ID of the identified user to the conversation server 40.
[0106] Step 114 The conversation server 40 converts the received audio data into text. Specifically, the audio / text conversion server 20 (see Figure 1) performs this conversion. The audio / text conversion server 20 converts the audio data into text data, for example, using an audio analysis AI. Hereafter, the text data converted from the audio data will be referred to as the spoken text.
[0107] Step 115 The conversation server 40 determines whether or not the spoken text contains an imaging instruction. If the spoken text contains predetermined keywords, the conversation server 40 determines that the spoken text contains an imaging instruction. In this case, a positive result is obtained in step 115. On the other hand, if the predetermined keywords are not included in the spoken text, the conversation server 40 determines that the spoken text does not contain an imaging instruction. In this case, a negative result is obtained in step 115. The specified keywords include, for example, "look" or "visible." Alternatively, regular expressions may be used, or the presence of a combination of multiple specified keywords (e.g., "this" + "look") may be used to determine that an imaging instruction is included.
[0108] Step 116 Step 116 is performed if a positive result is obtained in Step 115. The conversation server 40 obtains the latest image data from the owner terminal 10. This process corresponds to steps 103 and 104 (see Figure 11). That is, the conversation server 40 instructs the owner terminal 10 to take an image, and in response to that instruction, the owner terminal 10 sends the image data to the conversation server 40.
[0109] Step 117 The conversational server 40 passes the image data to the conversational AI and obtains the image description text (i.e., image description text). The conversational AI is an example of a generative AI. Here, the conversational AI uses natural language processing and machine learning models to output text that describes the content of the input image data.
[0110] Step 118 The conversation server 40 provides the conversational AI with the spoken text, image description text, and person information to generate response text. The conversational AI here may be the same as the one used in step 117, or a different one. For example, if the conversational AI in step 117 is good at handling image data, then in step 118, a conversational AI that is good at generating conversational text will be used. The conversation server 40 reads person information corresponding to the user ID notified from the owner terminal 10 from the user database 43A (see Figure 4). In Figure 4, the person information is given as, for example, "Taro," "Hanako," and "Ai-chan."
[0111] Step 119 Step 119 is performed if a negative result is obtained in Step 115. The conversation server 40 provides the spoken text to the conversational AI to generate a response text. Unlike Step 118, no image caption text or person information is used.
[0112] Step 120 The conversation server 40 generates speech synthesis data from the response text. The response text here is generated in step 118 or step 119. The conversation server 40 sends the generated speech synthesis data to the owner terminal 10. This process corresponds to step 105 (see Figure 11). Step 121 The owner terminal 10 outputs audio. This process corresponds to step 106 (see Figure 11).
[0113] Figure 13 illustrates an example of the output of the response text when the identified person is "Hanako". The spoken text is, "Take a look, what do you think?" The image captions read: "A woman with long hair is wearing a white knit hat with 'AAA' written on it." and "The woman is wearing a brown patterned knit top and is facing sideways." The personal information states, "The owner is Hanako." Therefore, the response text output is, "Hanako, that white knit hat looks lovely!"
[0114] In the case of Figure 13, the captured image shows the person who spoke to the owner terminal 10 (see Figure 1). Therefore, a response text is generated that reflects the person information (i.e., information about the owner) identified in step 112 (see Figure 12). For example, if the user who speaks to the owner terminal 10 (see Figure 1) is the "owner himself," the output might be something like, "Taro, who is the woman who looks good in a white knit hat?" or "Taro, that's a woman who looks good in a white knit hat!"
[0115] Thus, even if the spoken text and image description text are the same, different response texts will be output if the information about the person who spoke is different. If person information is not used to generate the response text, the content of the response text will be generated based on the spoken text and the image description text. Therefore, it will be limited to general expressions such as "That white knit hat is lovely!" or "That hat suits you!"
[0116] Figure 14 illustrates another example of response text output when the identified person is "Ai-chan". The spoken text is, "Look, there's fruit." The image caption text reads, "Pineapple, red grapes, white grapes, and strawberries are arranged in a basket," and "The pineapple is cut into bite-sized pieces." The person's information is, "The owner is Ai-chan." Therefore, the response text output is, "Ai-chan, there are so many different kinds of fruit!"
[0117] In the case of Figure 14, the person who spoke to the owner terminal 10 (see Figure 1) is not visible in the captured image. However, a response text is generated that reflects the person information identified in step 112 (see Figure 12) (i.e., information about the owner). For example, if the user speaking to the owner terminal 10 (see Figure 1) is the "owner himself," a response text such as "Taro, there are so many different kinds of fruit!" will be output.
[0118] Furthermore, if the user speaking to the owner terminal 10 (see Figure 1) is named "Hanako," a response text such as "Hanako, there are so many different kinds of fruit!" will be output. In this way, even if the spoken text and image description text are the same, the content of the response text can be changed for each owner. If person information is not used to generate the response text, the content of the response text will be generated based only on the spoken text and the image description text. Therefore, it will be limited to general expressions such as "There are so many different kinds of fruit!"
[0119] <Example 2> Figure 15 illustrates another example of a processing operation that initiates a conversation based on user utterance. Figure 15 is denoted with corresponding reference numerals for parts that correspond to those in Figure 12. The only difference between the processing operation shown in Figure 15 and the processing operation shown in Figure 12 is steps 131 and 132. Therefore, only steps 131 and 132 will be explained below. Note that step 131 is inserted after step 117. Step 132 is executed instead of step 118.
[0120] Step 131 The conversation server 40 accesses the memory information of the identified person. Figure 16 illustrates an example of the owner's memory information. For illustrative purposes, Figure 16 integrates the conversation log 43B (see Figure 8) and the summary DB 43C1 (see Figure 10). Note that in Figure 16, the start date and time 43B3 or the end date and time 43B4 of the conversation are used as the conversation date and time.
[0121] In Figure 16, we can see that the owner's favorite fruit is "strawberries," and the importance of this topic is "7." Furthermore, we learn that the owner "owns a dog named Pochi," that "Pochi is 7 years old," that "Pochi's birthday is October 6th," and that the importance of this topic is "8." Furthermore, the owner "wants a white knit hat," indicating that the importance of this topic is rated "3."
[0122] Step 132 The conversation server 40 provides the conversational AI with the spoken text, image description text, and stored information (i.e., conversation history) to generate response text. In this case, the conversation server 40 may include information (e.g., a numerical value) indicating the importance level of the topic as stored information. This information indicating the importance level can be used to narrow down the topics to be included in the response text. For example, if the descriptive text of an image captured by the owner terminal 10 corresponds to multiple stored items, it becomes possible to use the topic with a higher importance level compared to other topics as information related to the owner.
[0123] Figure 17 illustrates an example of the output of the response text when the identified person is the "owner." Figure 17 is denoted with corresponding symbols for parts that correspond to those in Figure 13. In the case of Figure 17, the spoken text is "Take a look, what do you think?", and the content of the image description text is the same as in Figure 13. However, in the case of Figure 17, memory information of a specific person is used instead of person information.
[0124] Figure 17 assumes that the owner is identified as "Hanako." Therefore, the memory information includes statements such as "The owner is female," "The owner wants a white knit hat," and "The owner's favorite fruit is strawberries." Information about the pet dog, Pochi, is also included in the memory information. In step 118 (see Figure 12), only the identified person's information was used, but in step 132 (see Figure 15), the identified person's memory information is also used to generate the response text. Therefore, the response text includes phrases like, "That white knit hat looks great!", "You said you wanted one, didn't you?", and "Did you finally find something you like?".
[0125] Since the memory information of the identified person is used to generate the response text, even if the spoken text and the image description text are the same, if the memory information is different, the content of the response text will be different. For example, if Taro's memory contains the information, "My wife wants a white knit hat," then it would be possible to output things like, "A white knit hat would be lovely!", "Hanako wanted one, didn't she?", or "Did you give it to her as a present?".
[0126] Figure 18 illustrates another example of the owner's stored information. Figure 18 is denoted with corresponding reference numerals for parts that correspond to those in Figure 16. The difference between the conversation text 43B5 shown in Figure 18 and the conversation text 43B5 shown in Figure 16 is that the content of the first and third lines of the conversation text 43B5 shown in Figure 18 concerns the wife.
[0127] Therefore, the content of summary text 43C15 has also been changed to "Taro's wife's favorite fruit is strawberries" and "Hanako wants a white knit hat." Incidentally, "Hanako" is "Taro's wife." The example response text output mentioned above is based on the stored information shown in Figure 18. If memory information is not used to generate the response text, the content of the response text will be limited to general expressions such as "That white knit hat looks great!" or "That hat suits you!"
[0128] Figure 19 illustrates another example of response text output when the identified person is the "owner." Figure 19 is denoted with corresponding symbols for parts that correspond to those in Figure 14. In Figure 19, the spoken text is "Look, there's fruit," and the image description text is the same as in Figure 14. However, in the case of Figure 19, memory information is used instead of person information. It is, however, memory information of a specific person.
[0129] Figure 19 assumes that the owner is identified as "Hanako." Therefore, the memory information includes statements such as "The owner is female," "The owner wants a white knit hat," and "The owner's favorite fruit is strawberries." Information about the pet dog, Pochi, is also included in the memory information. In step 118 (see Figure 12), only the identified person's information was used, but in step 132 (see Figure 15), the identified person's memory information is also used to generate the response text.
[0130] Therefore, the response text outputs, "There are lots of different fruits!" and "There are strawberries, the owner's favorite." In other words, the response text includes the fact that the owner likes strawberries. Similar to the case in Figure 17, since the memory information of the identified person is used to generate the response text, even if the spoken text and the image description text are the same, if the memory information is different, the content of the response text will be different.
[0131] For example, if the memory information for "Taro" shown in Figure 18 includes "Taro's wife's favorite fruit is strawberries," then it becomes possible to output things like "There are lots of different fruits!", "There are strawberries, which Hanako likes," and "Hanako will be happy." If memory information is not used to generate the response text, the content of the response text will be limited to general expressions such as "There are lots of different kinds of fruit!"
[0132] <Example 3> Figure 20 illustrates another example of a processing operation that initiates a conversation based on user utterance. Figure 20 is denoted with corresponding reference numerals for parts that correspond to those in Figure 12. The only difference between the processing operation shown in Figure 20 and the processing operation shown in Figure 12 is steps 141 and 142. Therefore, only steps 141 and 142 will be explained below. Step 141 is inserted after step 117. Step 142 is executed instead of step 118.
[0133] Step 141 The conversation server 40 retrieves events or schedules associated with person information. Figure 21 illustrates an example of the event schedule DB 43D. The event schedule DB 43D is stored in the auxiliary storage device 43 (see Figure 1). The event / schedule DB43D shown in Figure 21 is a dedicated database used for managing event information or schedule information.
[0134] The event schedule DB43D shown in Figure 21 includes the conversation date and time 43D1, the user ID 43D2, the conversation text 43D3, the event schedule information 43D4, the start date and time 43D5, and the end date and time 43D6. In this embodiment, event and schedule information is extracted from the conversation text 43D3 and stored. If event and schedule information is registered independently of the conversation, the conversation text 43D3 will be blank.
[0135] The conversation text 43D3 is used by the user to review and modify event schedule information 43D4. The event schedule information 43D4 uses information from, for example, the summary text 43C15 (see Figure 10). The content of the event schedule information 43D4 can also be viewed and modified by the user. If the user modifies the summary text 43C15, the corresponding event schedule information 43D4 will also be modified.
[0136] The start date and time (43D5) and end date and time (43D6) are registered if they can be extracted from the conversation text. It is also possible that only the date or only the time may be recorded. In Figure 21, for events or schedules where the time cannot be specified, a time range from "0:00" to "24:00" is stored. Additionally, if an approximate time period (e.g., AM, PM, breakfast, lunch, dinner) can be identified during a conversation, that time period is stored as the start and end time.
[0137] In Figure 21, for "Hanako," whose user ID is "3433," the time range from "12:00" to "24:00" is stored for the statement "Tomorrow I'm going to the arcade with my daughter from noon." Furthermore, in Figure 21, the date "3 / 2" to "3 / 9" is stored in response to "Hanako," whose user ID is "3433," saying "My hair has grown long, so I want to go to the hair salon this week."
[0138] For events and schedules where the start date and time are publicly available, such as concerts, the conversation server 40 may provide additional information. The start date and time 43D5 and the end date and time 43D6 can also be checked and modified by the user. Event information or schedule information may also be managed in attribute information 43A4 (see Figure 4).
[0139] Furthermore, event schedule information 43D4 may be made directly registrable, for example, from the owner screen 500 (see Figure 5). Furthermore, the event / schedule information 43D4 may be linked with information from event or schedule registration sites provided by the conversation service. Furthermore, the conversation server 40 may obtain user-related event information and schedule information from external servers with which it collaborates through user accounts, etc. Returning to the explanation of Figure 20.
[0140] Step 142 The conversation server 40 provides the conversational AI with the spoken text, image description text, and event information or schedule information associated with the person's information to generate response text. For example, if the person's information is "Taro," then event or schedule information associated with "Taro" will be used. "Taro's" user ID is "3432." Therefore, in the event / schedule information DB43D shown in Figure 21, the conversational AI will be given the plan to "go to a live concert of a favorite artist" on "March 4th."
[0141] In this Example 3, the event or schedule information provided to the conversational AI can have start and end dates that are either before or after the current date. Including events or schedules that took place before the current date makes it possible to generate response text associated with completed events or schedules. On the other hand, including events or schedules that will take place after the current date makes it possible to generate response text associated with planned events or schedules.
[0142] In the case of Figure 21, there is only one event or schedule related to "Taro," but if multiple events or schedules are found, all of them will be provided to the conversational AI. Incidentally, the importance level 43C16 (see Figure 10) can also be obtained from the attribute information 43A4 (see Figure 4) and summary text 43C15 (see Figure 10) used in the event / schedule information 43D4 and provided to conversational A1. On the other hand, if no events or schedules related to "Taro" are remembered, the conversation server 40 may provide the conversational AI with only the spoken text, image description text, and person information.
[0143] Figure 22 illustrates an example of response text output using an event or schedule. Figure 22 is denoted with corresponding reference numerals for parts corresponding to those in Figure 14. In the case of Figure 22, the spoken text is "Look, there's fruit," and the content of the image description text is the same as in Figure 14. However, in the case of Figure 22, the registered event information used is "The owner plans to go shopping at the supermarket in front of the station."
[0144] Therefore, the response text output is, "I bought lots of fruit at the supermarket in front of the station!" This response text is possible because the conversational AI is given information that there are plans to go shopping at the supermarket in front of the station. If event schedule information 43D4 (see Figure 21) is not used to generate the response text, the content of the response text will be limited to general expressions such as "There are lots of different kinds of fruit!".
[0145] Figure 23 illustrates other examples of response text output using events or schedules. The spoken text is "Look." The image description text is "A girl with her hair tied in two pigtails is holding a large stuffed dog." and "There is a crane game machine in the background."
[0146] Figure 23 assumes that the identified person is "Hanako." Hanako's registered event information is "The owner will go out with a friend and her daughter on March 5th." Therefore, the response text output is, "Is that your friend's daughter? She's holding a stuffed animal! Where are you playing?"
[0147] Figure 24 illustrates another example of response text output using an event or schedule. Figure 24 uses corresponding reference numerals to indicate parts that correspond to those in Figure 23. Note that the owner terminal 10 in Figure 24 (see Figure 1) is assumed to be a portable device such as a smartphone. In the case of Figure 24, the spoken text is "Look," and the image description text is "A girl with her hair tied in two pigtails is holding a large stuffed dog" and "There is a crane game machine in the background."
[0148] The registered event information is "The owner will go to the arcade on March 5th." Therefore, the response text includes phrases like, "Did you come to the arcade?", "I'm glad you were able to get a stuffed animal!", and "Who is the girl in the picture with you?". The image used in Figure 24 is the same as in Figure 23, but the content of the registered event information is different. Therefore, even if the spoken text, the captured image, and the image description text are the same, the content of the response text will be different.
[0149] <Example 4> Figure 25 illustrates another example of a processing operation that initiates a conversation based on user utterance. Figure 25 is denoted with corresponding reference numerals for parts that correspond to those in Figure 20. Note that, due to space limitations, Figure 25 only shows the processing operations on the conversation server 40 side. Therefore, the operation in step 114 begins when person information and voice data are received from step 113 (see Figure 20).
[0150] The only difference between the processing operation shown in Figure 25 and the processing operation shown in Figure 20 is steps 151 and 152. Therefore, only steps 151 and 152 will be explained below. Step 151 is inserted after step 141. Step 152 is executed if a negative result is obtained in step 151.
[0151] Step 151 In step 141, the conversation server 40, having obtained an event or schedule associated with the person's information, determines whether or not there is an event or schedule corresponding to the current time. "Applicable to the current time" means, for example, that something has been completed or is scheduled to be completed within a specified period including the current date and time. This specified period includes, for example, one week, one day, 12 hours, 6 hours, 2 hours, 1 hour, etc.
[0152] The designated time may be set according to the content of the event or schedule. For example, for topics of interest such as concerts or trips, the time may be set over a period of several days. On the other hand, for matters with a specific date and time, such as shopping, the time may be set within one day. Furthermore, the specified time may be changed depending on whether the event or schedule has already been completed or is scheduled to be completed. In this case as well, the specified time may not be uniform regardless of the content of the event or schedule, but may be set according to the content of the event or schedule.
[0153] If the event or schedule obtained in step 141 corresponds to the current time, a positive result is obtained in step 151. In this case, the conversation server 40 proceeds to step 142. Conversely, if the event or schedule obtained in step 141 does not correspond to the current time, a negative result is obtained in step 151. In this case, the conversation server 40 proceeds to step 152.
[0154] Incidentally, in step 151, it may further be determined whether the event or schedule is being used in a conversation that includes the current time. Here, "conversation that includes the current time" refers to a conversation within a predetermined period that includes the current date and time. By providing this determination function, it becomes possible to change the content of the response text even if the same spoken text or image description text is given to the conversational AI. As a result, the problem of the same event or schedule being repeated can be reduced. Note that the additional check may only be performed if an event or schedule corresponding to the current time is found. If the additional check also yields a positive result, proceed to step 142; if the additional check also yields a negative result, proceed to step 152.
[0155] Step 152 The conversational server 40 provides the conversational AI with the spoken text and image description text to generate a response text. In this case, the owner's or family user's events or schedule are not taken into consideration, so the response will be as follows: For example, in the example shown in Figure 22, the message would become something like "There are so many different kinds of fruit!", and information such as "supermarket in front of the station" or "shopping" would no longer be included. For example, in the example shown in Figure 23, the comment would be something like "That's a big stuffed animal," and information such as "a friend's daughter" or "playing" would no longer be included. For example, in the example in Figure 24, the description would become something like "That's a big stuffed animal," and the information about "game center" would no longer be included. <Summary> According to the aforementioned conversation system 1 (see Figure 1), even if the content of the utterance and the image captured by the owner terminal 10 are the same, different responses can be output by reflecting the speaking user's personal information, memory information, event information, schedule information, etc.
[0156] <When the conversation is initiated by an utterance from the conversation server> <Overview> Figure 26 is a diagram illustrating the outline of a conversation sequence initiated by an utterance from the conversation server 40. In Figure 26, the conversation server 40 is a collective term for the speech / text conversion server 20 (see Figure 1), the front server 30 (see Figure 1), and the conversation server 40.
[0157] Step 201 In the case of Figure 26, the owner terminal 10 is constantly capturing images with the camera 14 (see Figure 1). For example, regardless of the owner's speech or instructions, the camera 14 captures several images per second. However, the owner terminal 10 does not upload the captured images to the conversation server 40. In this embodiment, the owner terminal 10 determines whether or not the captured image contains a face image of a person who has been registered in advance. In other words, the owner terminal 10 is always performing face recognition.
[0158] For example, if a pre-registered facial image is included in the captured image, the owner terminal 10 sends a message to the conversation server 40 indicating that facial recognition has occurred. This message may also include information about the recognized person. However, a message indicating face recognition may be sent to the conversation server 40 only if a specific face image from the pre-registered face images is captured in the image. For example, the system may be configured to send a message indicating face recognition if the owner's face image is captured, but not if a family user's face image is recognized.
[0159] On the other hand, if a pre-registered face image is not included in the captured image, the owner terminal 10 does not send face recognition information to the conversation server 40. Furthermore, the owner terminal 10 deletes the images used for face recognition from the semiconductor memory 12 (see Figure 1), etc., regardless of whether or not the captured image contains a pre-registered face image. This prevents images unintended by the owner from being transmitted externally.
[0160] Step 202 When the conversation server 40 receives a notification that face recognition has been performed, it sends an image capture instruction to the owner terminal 10. Alternatively, the system may be configured to send an imaging command only when a specific person's face image is recognized.
[0161] For example, if the facial image recognized is that of the owner, an imaging command is sent to the owner terminal 10, but if the facial image recognized is that of a family user, it is possible to operate in a way that does not send an imaging command to the owner terminal 10. In other words, the system may control whether or not to send imaging instructions to the owner terminal 10 based on each user's registration information. By adopting this mechanism, the conversation service can be provided only to users who wish to receive spontaneous speech from the conversation server 40.
[0162] Step 203 Upon receiving the imaging command, the owner terminal 10 starts imaging with the camera 14 (see Figure 1). The owner terminal 10 also transmits the image data captured by the camera 14 to the conversation server 40. The image data may be a still image or a moving image.
[0163] Step 204 Upon receiving the image data, the conversation server 40 begins the process of generating response text. In this embodiment, the conversation server 40 generates utterance text using the captured image, regardless of the owner's instructions. In other words, the conversation server 40 generates utterance text independently of the owner's utterances. Once the spoken text is generated, the conversation server 40 converts the spoken text into speech synthesis data. The conversation server 40 sends the generated speech synthesis data to the owner terminal 10.
[0164] Step 205 The owner terminal 10 outputs speech synthesis data as voice from the speaker 16 (see Figure 1). This enables the owner terminal 10 to spontaneously initiate a conversation with the user. After the user responds, a conversation is established between the owner terminal 10 and the user. From this point onward, the processing operations in steps 101 to 106 (see Figure 11) are repeated.
[0165] <Example of processing operation> The following describes an example of processing operations performed through the cooperation between the owner terminal 10 and the conversation server 40. Needless to say, the processing examples described later are merely examples. For example, it is possible to execute some or all of the steps performed on the owner terminal 10 on the conversation server 40. Conversely, it is also possible to perform some or all of the steps performed on the conversation server 40 on the owner terminal 10.
[0166] Furthermore, the processing operation examples shown in Figures 26, 27, 30, 33, and 37, which will be described later, can be combined with each other. In addition, the combinations may include one or more of the steps shown in Figures 11, 12, 15, 20, and 25, as well as Figure 38, which will be described later.
[0167] <Example 1> Figure 27 illustrates one example of a processing operation in which a conversation is initiated by an utterance from the conversation server 40. In Figure 27, the conversation server 40 is also a collective term for the speech / text conversion server 20, the front server 30, and the conversation server 40.
[0168] Step 211 The owner terminal 10 determines whether facial recognition meets predetermined conditions. As mentioned above, one of the predetermined conditions is the detection of an image containing the face image of a person registered as the owner. If facial recognition meets the predetermined conditions, a positive result is obtained in step 211. In this case, the owner terminal 10 proceeds to step 212. If facial recognition does not meet the predetermined conditions, a negative result is obtained in step 211. In this case, the owner terminal 10 repeats the determination in step 211. If voice input is received during this determination (for example, in step 111 in Figure 12), the owner terminal 10 proceeds to step 112 (see Figure 12).
[0169] Step 212 The owner terminal 10 sends a message to the conversation server 40 indicating that facial recognition has been performed. Step 212 corresponds to step 201 (see Figure 26). Furthermore, the term "facial recognition" may include information about the user (including the owner) authenticated from an image containing a facial image.
[0170] For user authentication, voices other than those directed at the owner terminal 10 may be used. Voices other than those directed at the owner terminal 10 include, for example, utterances that do not include the nickname assigned to the owner terminal 10 (e.g., Romi). This type of utterance may also include sighs such as "phew" or monologues such as "I'm tired." The facial recognition process does not necessarily require the inclusion of information about the authenticated user (including the owner). This is because the transmission in step 212 is merely a trigger to initiate speech from the conversation server 40, and the content of the speech text can be entirely assumed to be that of the owner.
[0171] Step 213 The conversation server 40 retrieves the latest image data from the owner terminal 10. This process corresponds to steps 202 and 203 (see Figure 26). The conversation server 40 may authenticate users who are captured in images obtained from the owner terminal 10. Similarly, it may authenticate users using audio obtained from the owner terminal 10. However, in the case of step 213, user authentication is optional.
[0172] Step 214 The conversational server 40 passes the image data to the conversational AI and obtains descriptive text for the image (i.e., image description text). The conversational AI here is also an example of a generative AI. The conversational AI here uses natural language processing and machine learning models to output text that describes the content of the input image data.
[0173] Step 215 The conversational server 40 provides the image description text to the conversational AI to generate spoken text. The conversational AI used here may be the same as the one used in step 214, or a different one. For example, if the conversational AI in step 214 is skilled at handling image data, then in step 215, a conversational AI skilled at generating conversational text will be used. The spoken text here may assume the owner (e.g., "Taro"), regardless of the authenticated person. Alternatively, the identified person information may be provided to an interactive AI to generate spoken text.
[0174] Step 216 The conversation server 40 generates speech synthesis data from the spoken text. The conversation server 40 sends the generated speech synthesis data to the owner terminal 10. This process corresponds to step 204 (see Figure 26). Step 217 The owner terminal 10 outputs audio. This process corresponds to step 205 (see Figure 26).
[0175] Figure 28 illustrates an example of the output of spoken text when the identified person is "Hanako". In the case of Figure 28, there is no spoken text; only image description text and person information are provided to the conversational AI.
[0176] Incidentally, the image captions read: "A woman with long hair is wearing a white knit hat with 'AAA' written on it." and "The woman is wearing a brown patterned knit top and is facing sideways." Furthermore, the personal information states, "The owner is Hanako." Therefore, the output text of the speech reads, "Hanako, that white knit hat looks lovely!"
[0177] Furthermore, if only the image description text is provided to the conversational AI, the generated spoken text will be non-personally specific, such as "That white knit hat is lovely." In any case, the owner terminal 10 will be able to initiate a conversation with the user in front of it. If the user responds to this conversation with something like, "Really? I liked it so I bought it!" or "It's cold outside," a conversation will begin between the user and the owner terminal 10. The owner terminal 10's response to the user's response will follow the conversation sequence shown in Figure 11.
[0178] Figure 29 illustrates another example of speech text output when the person is not identified. In the case of Figure 29, there is no spoken text. Also, in Figure 29, only the image description text is provided to the conversational AI. The image description text is "Pineapples, red grapes, white grapes, and strawberries are arranged in a basket" and "The pineapples are cut for easy eating". Therefore, the utterance text outputs "There are so many different fruits!" In this case as well, the conversation between the user and the owner terminal 10 can be started triggered by the call from the owner terminal 10.
[0179] <Example 2> FIG. 30 is a diagram for explaining another example of a processing operation for starting a conversation by the utterance of the conversation server 40. In FIG. 30, the corresponding parts to those in FIG. 27 are denoted by the same reference numerals. The difference between the processing operation shown in FIG. 30 and the processing operation shown in FIG. 27 is only in steps 221 and 222. Therefore, only steps 221 and 222 will be described below. Note that step 221 is inserted after step 214. Step 222 is executed instead of step 215.
[0180] · Step 221 The conversation server 40 refers to the owner's storage information (see, for example, FIG. 16). In the case of FIG. 16, it can be seen that the owner's favorite fruit is strawberries and the importance of this topic is "7". Also, it can be seen that the owner has a dog named Pochi, Pochi is 7 years old, Pochi's birthday is October 6th, and the importance of this topic is "8".
[0181] Also, it can be seen that the owner wants a white knitted hat and the importance of this topic is "3". When the user is identified through the owner terminal 10, the storage information of the identified user may be referred to. Needless to say, referring to the storage information of the identified user reduces the topic mismatch.
[0182] · Step 222 The conversation server 40 provides the conversational AI with image description text and stored information (i.e., conversation history) to generate spoken text. In this case, the conversation server 40 may include information (e.g., a numerical value) indicating the importance level of the topic as stored information. This information indicating the importance level can be used to narrow down the topics to be included in the response text. For example, if the descriptive text of an image captured by the owner terminal 10 corresponds to multiple stored items, it becomes possible to use the topic with a higher importance level compared to other topics as information related to the owner.
[0183] Figure 31 illustrates an example of the output of the response text when the identified person is the "owner." Figure 31 is denoted with corresponding symbols for parts that correspond to those in Figure 28. In the case of Figure 31, there is no spoken text; instead, the interactive AI is provided with image description text and the owner's memory information. In the case of Figure 31, the image caption text is "A woman with long hair is wearing a white knit hat with 'AAA' written on it." and "The woman is wearing a brown patterned knit and is facing sideways." In the case of Figure 31, the memory information entered includes "The owner is female," "The owner wants a white knit hat," and "The owner's favorite fruit is strawberries." The memory information also includes information about the pet dog, Pochi.
[0184] Therefore, the output text of the speech reads, "Isn't that the white knit hat you wanted?" and "That's lovely!" In Example 1, the conversation was based on information gleaned from the image, but in Example 2, the utterance reflects memory information. Specifically, the phrase "wanted" is added. As a result, the owner experiences the feeling of being spoken to by an acquaintance who knows them well.
[0185] If there is no memory information related to the image description text, an utterance that identifies the image from its appearance, such as "Going out? That knit hat is lovely," will be uttered. Incidentally, the conversation server 40 may provide the conversational AI with memory information (including conversations initiated by the user) that has not been used within a predetermined period, including the current date and time. This function prevents the same conversation based on the same memory information from being repeated multiple times. For example, if the owner terminal 10 says, "Isn't that the white knit hat you wanted?" and "It's lovely!" when the user goes out, the AI can avoid repeating the same lines, "Isn't that the white knit hat you wanted?" and "It's lovely!" when the user returns home. As a result, it is possible to achieve a conversation that is closer to that of a human.
[0186] Figure 32 illustrates another example of response text output when the identified person is the "owner." Figure 32 is denoted with corresponding symbols for parts corresponding to those in Figure 29. In the case of Figure 32, there is no spoken text; instead, the interactive AI is provided with image description text and the owner's memory information.
[0187] In the case of Figure 32, the image caption text reads, "Pineapple, red grapes, white grapes, and strawberries are arranged in a basket," and "The pineapple has been cut into bite-sized pieces." On the other hand, the memory information includes statements such as, "The owner is female," "The owner wants a white knit hat," and "The owner's favorite fruit is strawberries." The memory information also includes details about the pet dog, Pochi. Therefore, in the spoken text, the focus is on the strawberries, which are the owner's favorite among the pineapple, red grapes, white grapes, and strawberries that appear in the image description text, resulting in a more human-like conversation such as, "There are so many strawberries! What happened?"
[0188] <Example 3> Figure 33 illustrates another example of processing operation in which a conversation is initiated by an utterance from the conversation server 40. In Figure 33, parts corresponding to those in Figure 27 are indicated with corresponding reference numerals. The only difference between the processing operation shown in Figure 33 and the processing operation shown in Figure 27 is steps 231 and 232. Therefore, only steps 231 and 232 will be explained below. Step 231 is inserted after step 214. Step 232 is executed instead of step 215.
[0189] Step 231 The conversation server 40 retrieves events or schedules associated with person information. Person information refers to information about the person being spoken to. The person being spoken to is assumed to be the owner or a user whose face was recognized in step 212. For example, the event or schedule of the person being spoken to is retrieved from the event / schedule DB 43D (see Figure 21).
[0190] Step 232 The conversational server 40 provides the conversational AI with image description text and event information or schedule information associated with person information to generate spoken text. For example, if the person's information is "Taro," then event information or schedule information associated with "Taro" will be used. In this example 3, the event or schedule information provided to the conversational AI can have start and end dates that are either before or after the current date and time.
[0191] Figure 34 illustrates an example of speech text output using an event or schedule. Figure 34 is denoted with corresponding reference numerals for parts corresponding to those in Figure 29. In the case of Figure 34, there is no spoken text; instead, the conversational AI is provided with image description text and registered event information. In the case of Figure 34, the image caption text reads, "Pineapple, red grapes, white grapes, and strawberries are arranged in a basket," and "The pineapple has been cut into bite-sized pieces."
[0192] By the way, in the case of Figure 34, the conversational AI is given the registered event information that "the owner plans to go shopping at the supermarket in front of the station." Therefore, the uttered text "You bought a lot of fruits at the supermarket near the station!" is output. This uttered text becomes possible because the interactive AI is provided with the information that the owner who makes the call plans to go shopping at the supermarket near the station. Note that when the event / schedule information 43D4 (see FIG. 21) is not used for generating the uttered text, the content of the uttered text remains a general expression such as "There are a lot of different fruits!".
[0193] FIG. 35 is a diagram for explaining another output example of the response text using an event or schedule. The owner terminal 10 (see FIG. 1) in FIG. 35 assumes a portable device such as a smartphone. The image description text is "A girl with her hair tied in two has a large stuffed dog." and "There is a crane game machine in the back." Also, the registered event information is "The owner will go play with a friend and the friend's daughter on 3 / 5."
[0194] Therefore, the uttered text "Is it the friend's daughter? A big stuffed animal!" is output. This greeting based on the uttered text becomes possible because the interactive AI is provided with the event information that the owner will go play with a friend and the friend's daughter. That is, it becomes possible to make an utterance similar to a greeting from a person. Note that when the registered event information is not provided to the interactive AI, the content of the uttered text may be a misspoken utterance such as "The girl has a big stuffed animal" or "Aichan has a big stuffed animal", confusing the friend's daughter and the owner's daughter.
[0195] FIG. 36 is a diagram for explaining another output example of the response text using an event or schedule. In FIG. 36, the corresponding parts to those in FIG. 35 are marked with corresponding reference numerals. Also in the case of FIG. 36, the image description text is "A girl with her hair tied in two has a large stuffed dog." and "There is a crane game machine in the back."
[0196] However, in the case of Figure 36, the registered event information is "The owner will go to the arcade on March 5th." In Figure 36, the event information includes destination information, but the information about the playmates is missing. Therefore, the output text is "It looks like you're having fun at the arcade!". In other words, even if the same image description text is given to the conversational AI, the output will be different.
[0197] <Example 4> Figure 37 illustrates another example of a processing operation initiated by an utterance from the conversation server 40. Figure 37 is denoted with reference numerals corresponding to the parts that correspond to those in Figure 33. The only difference between the processing operation shown in Figure 37 and the processing operation shown in Figure 33 is steps 241 and 242. Therefore, only steps 241 and 242 will be explained below. Step 241 is inserted after step 231. Step 242 is executed if a negative result is obtained in step 231.
[0198] Step 241 In step 231, the conversation server 40, having obtained an event or schedule associated with the person's information, determines whether or not there is an event or schedule corresponding to the current time. "Applicable to the current time" means, for example, that something has been completed or is scheduled to be completed within a specified period including the current date and time. This specified period includes, for example, one week, one day, 12 hours, 6 hours, 2 hours, 1 hour, etc.
[0199] The designated time may be set according to the content of the event or schedule. For example, for topics of interest such as concerts or trips, the time may be set over a period of several days. On the other hand, for matters with a specific date and time, such as shopping, the time may be set within one day. Furthermore, the specified time may be changed depending on whether the event or schedule has already been completed or is scheduled to be completed. In this case as well, the specified time may not be uniform regardless of the content of the event or schedule, but may be set according to the content of the event or schedule.
[0200] If the event or schedule obtained in step 241 corresponds to the current time, a positive result is obtained in step 241. In this case, the conversation server 40 proceeds to step 222. Conversely, if the event or schedule obtained in step 241 does not correspond to the current time, a negative result is obtained in step 241. In this case, the conversation server 40 proceeds to step 242.
[0201] By the way, in step 241, it may also be determined whether the event or schedule is being used in a conversation that includes the current time. Here, "conversation that includes the current time" refers to a conversation within a predetermined period that includes the current date and time. By providing this determination function, it becomes possible to change the content of the spoken text even if the same image description text is given to the conversational AI. As a result, the problem of repeated spoken texts about the same event or schedule can be reduced. Note that the additional check may only be performed if an event or schedule corresponding to the current time is found. If the additional check also yields a positive result, proceed to step 222; if the additional check yields a negative result, proceed to step 242.
[0202] Step 242 The conversation server 40 provides image description text to the conversational AI to generate spoken text. In other words, the conversation is based solely on the images captured by the owner terminal 10. For example, in the example in Figure 34, the message would become something like "There are so many different kinds of fruit!", and information such as "supermarket in front of the station" or "shopping" would no longer be included. For example, in the example shown in Figure 35, the comment would be something like "That's a big stuffed animal," and the information "it's a friend's daughter" would no longer be included. For example, in the example in Figure 36, the description would become something like "That's a big stuffed animal," and the information about "game center" would no longer be included.
[0203] <Summary> The aforementioned conversation system 1 (see Figure 1) enables the AI character to initiate conversations spontaneously, even when the user is not speaking.
[0204] <Other Embodiments> The embodiments of the invention are not limited to those described above. For example, elements of each embodiment can be combined as appropriate. Combinations here also include the deletion of elements from each embodiment.
[0205] (1) In the above-described embodiment, the conversation server 40 (see Figure 1) assigns a new user ID each time a family user is registered. Therefore, even for the same person, a different user ID will be assigned depending on whether the person is registered as a family user by owner A or by owner B. For example, a different user ID will be assigned depending on whether the user is registered as owner A's child or owner B's grandchild. In other words, the same person will be assigned both a user ID as a child and a user ID as a grandchild. Furthermore, when owner A is registered as a family member of owner B, owner A will be assigned both a user ID as the owner and a user ID as a family user of owner B.
[0206] However, if it is confirmed that the user IDs belong to the same person through facial images, voiceprints, etc., the user IDs used to manage that person may be consolidated into one. Alternatively, the multiple user IDs may remain as they are, and a process (hereinafter referred to as "name matching process") may be performed to determine that all multiple user IDs belong to the same person. However, it is desirable to require confirmation from the owner before merging user IDs or performing the name matching process. By integrating or matching user IDs into a single user ID, it becomes possible to make responses and pronouncements that reflect past conversations, regardless of the owner associated with the owner terminal 10 (see Figure 1) used for the conversation, as long as the conversation is with the same person.
[0207] (2) In the above-described embodiment, the owner terminal 10 identifies the user with whom it is speaking. However, the identification of the user being spoken to may be performed on the conversation server 40 side. For example, the user being spoken to may be identified by providing the conversational AI with the image data captured by the owner terminal 10, and the face image and voiceprint information registered in association with the owner terminal 10 that is being processed.
[0208] (3) Step 131 (see Figure 15) assumes the memory information of the identified person. However, the memory information may include not only the memory information of the identified person but also the memory information of their family users. In this case, in step 112 (see Figure 15), the conversation server 40 provides the conversational AI with the information of the identified person and the memory information of the identified person and their family users. For example, even if the identified person is a male named "Taro" and there is no memory of a "white knit hat" in his memory information, it is possible to output response text such as "Have you found the hat that Hanako wanted?" by using the memory information of family user "Hanako".
[0209] (4) In step 117 above (see Figure 12), image data is provided to the conversational AI to obtain image description text. However, step 117 may be omitted, and in step 118 (see Figure 12) and step 132 (see Figure 15), image data acquired from the owner terminal 10 may be provided to the conversational AI to generate response text. Figure 38 illustrates another example of a processing operation that initiates a conversation based on user utterance. Figure 38 is denoted with corresponding reference numerals for parts that correspond to those in Figure 12.
[0210] The difference between Figure 38 and Figure 12 is that the processing operation shown in Figure 38 does not include step 117 (see Figure 12), and step 118A is executed instead of step 118. In step 118A, the conversational AI is provided with the spoken text, image data, and person information. In other words, the image description text is replaced with image data.
[0211] In the processing operation shown in Figure 38, step 117, which involves acquiring image description text, is omitted, and all image information is provided to the conversational AI that generates the response text. As a result, the conversational AI generates the response text including information that was missing when generating the image description text. Consequently, the variation in the content of the response text increases compared to when the image description text is provided to the conversational AI. Furthermore, the mechanism of providing the image data itself to the interactive AI instead of image description text can also be applied to Figures 15, 20, and 25. Furthermore, the mechanism of providing the interactive AI with the image data itself instead of image description text can also be applied to examples of processing operations where a conversation is initiated by an utterance from the conversation server 40.
[0212] (5) In step 142 above (see Figure 20), event information or schedule information linked to the person's information is provided to the conversational AI. However, the conversation server 40 may provide the conversational AI with the content of events or schedules of one or more people registered with the owner as information related to the owner. For example, if the person identified in step 112 (see Figure 20) is "Taro," the interactive AI may be given not only information about "Taro" but also about events or schedules for "Hanako" and "Ai-chan," who are registered as family members.
[0213] In this process, the conversational AI is provided with information about "Taro" as a person, as well as information about who each event or schedule relates to, in order to identify the person it is talking to. In this case, even if no event or schedule information linked to "Taro" can be found, it becomes possible to reflect events or schedules related to people associated with the owner, such as "Oh, that reminds me, Hanako was planning to go to the hair salon," in the response text.
[0214] Furthermore, the event or schedule information of one or more people registered for the owner may be used if there is no event or schedule information associated with the person identified in step 112 (see Figure 20) (for example, Taro). This mechanism prioritizes the generation of response text that reflects the owner's event or schedule information, and secondarily generates response text that reflects the event or schedule information of other people such as family members.
[0215] (6) In step 201 described above (see Figure 26), if facial recognition by the owner terminal 10 is successful, a message to that effect is sent to the conversation server 40. However, if facial recognition fails, a message to that effect is sent to the conversation server 40, and image capture is initiated.
[0216] (7) In step 211 described above (see Figure 27), facial recognition on the owner terminal 10 is assumed, but the owner terminal 10 may also notify the conversation server 40 that it has met other predetermined conditions besides facial recognition. For example, the owner terminal 10 may notify the conversation server 40 that it has detected the activation of a lighting fixture that was previously off. The detection of this event information can be considered as the owner entering the room where the stationary owner terminal 10 is located, and this is a favorable timing for the conversation server 40 to initiate contact.
[0217] For example, the owner terminal 10 may notify the conversation server 40 that it has detected a noise while silence has continued for a predetermined period of time (e.g., 10 minutes) or longer. The detection of this event information can also be considered as the owner entering the room where the stationary owner terminal 10 is located, and is a preferred timing for the conversation server 40 to initiate contact. In addition, the owner terminal 10 may notify the conversation server 40 of other information, such as the release of the locked front door, the detection of a person by a motion sensor, the detection of operating sounds from household equipment or appliances, notifications from a thermometer or hygrometer, and other detected information.
[0218] (8) Step 201 described above (see Figure 26) assumes that the user's face is recognized while the owner terminal 10 is constantly capturing images regardless of the owner's instructions. However, the activation of the camera function of another app or other application may be used as a trigger event for the conversation server 40 to initiate a conversation.
[0219] In the case of existing devices, activating the camera function does not involve user speech, nor is it intended to involve conversation with the owner device 10. Therefore, a mechanism may be adopted in which the activation of the camera function is considered a trigger event, and the image data being captured by camera 14 (see Figure 1) is sent to the conversation server 40. In this case, it becomes unnecessary for the conversation server 40 to issue an imaging instruction to the owner terminal 10. In addition, it becomes possible for the user to spontaneously speak to an object that they are interested in and have pointed camera 14 at.
[0220] (9) In the above-described embodiment, the response text and spoken text generated by the conversation server 40 (see Figure 1) are sent directly to the owner terminal 10 (see Figure 1). However, a mechanism may be adopted in which the text is provided to a dedicated server that converts it into voice data, and the voice data generated by the dedicated server is sent to the owner terminal 10.
[0221] <Summary 1> The main features of the information processing system, information processing method, program, and information processing device described in the embodiments are shown below. [General tasks] One of the purposes of this disclosure is to provide an information processing system, information processing method, program, and information processing device that can personalize the response content for each owner, even if the content of the utterance and the captured images are the same.
[0222] Issues corresponding to [Appendix 1] One of the purposes of this disclosure is to change the content of the response text for each owner, even if the content of the utterance and the captured images are the same. [Note 1] The information processing system relating to this disclosure includes one or more processors, which generate response text using the content spoken by the owner, images captured by the owner terminal, and information related to the owner, and output audio from the owner terminal corresponding to the generated response text. According to the information processing system described above, even if the content of the speech and the images captured are the same, the content of the response text can be changed for each owner.
[0223] Issues corresponding to [Appendix 2] One of the purposes of this disclosure is to limit the generation of response text using images captured by the owner terminal to cases where the owner has instructed the device to capture an image. [Note 2] If the owner's spoken content includes an instruction to capture an image, one or more processors in the information processing system described in Appendix 1 acquire the image from the owner's terminal and use it to generate response text. This allows the generation of response text using images captured on the owner's device to be limited to cases where the owner has instructed the device to capture an image.
[0224] Issues corresponding to [Appendix 3] One of the purposes of this disclosure is to reflect information related to the identified person in the content of the response text. [Note 3] An information processing system as described in Appendix 1, comprising one or more processors, which identifies the person who spoke as the owner and uses information related to the identified person as information related to the owner. This allows information related to the identified person to be reflected in the response text.
[0225] Issues corresponding to [Appendix 4] One of the purposes of this disclosure is to reflect in the response text the content of events or schedules registered for the identified person. [Note 4] An information processing system as described in Appendix 3, comprising one or more processors, which uses the content of registered events or schedules concerning an identified person as information related to the owner. This allows the content of the response text to reflect the registered events or schedules for the identified person.
[0226] Issues corresponding to [Appendix 5] One of the purposes of this disclosure is to reflect in the response text the content of events or schedules that are highly relevant to the current date and time. [Note 5] An information processing system as described in Appendix 4, wherein one or more processors generate response text using events or schedules registered for an identified person that have been executed or are scheduled to be executed within a predetermined period including the current date and time. This allows the response text to reflect events or schedules that are highly relevant to the current date and time.
[0227] Issues corresponding to [Appendix 6] One of the purposes of this disclosure is to ensure that the same event or schedule content does not appear repeatedly in the response text. [Note 6] The information processing system described in Appendix 4, wherein one or more processors generate response text using events or schedules registered for an identified person that have not been used in a conversation within a predetermined period including the current date and time. This prevents the same event or schedule content from appearing repeatedly in the response text.
[0228] Issues corresponding to [Appendix 7] One of the purposes of this disclosure is to reflect in the response text the details of events or schedules of one or more individuals registered as the owner. [Note 7] An information processing system as described in Appendix 1, in which one or more processors use the content of events or schedules of one or more persons registered with the owner as information related to the owner. This allows the content of the response text to reflect the event or schedule details of one or more people registered with the owner.
[0229] Issues corresponding to [Appendix 8] One of the purposes of this disclosure is to reflect the conversation history with the identified person in the content of the response text. [Note 8] An information processing system as described in Appendix 3, comprising one or more processors, which uses the conversation history recorded about an identified person as information related to the owner. This allows the conversation history with the identified person to be reflected in the response text.
[0230] Issues corresponding to [Appendix 9] One of the purposes of this disclosure is to reflect topics of high importance to the identified person in the content of the response text. [Note 9] The information processing system described in Appendix 8 includes a conversation history that contains information indicating the level of importance of topics, and one or more processors use topics with a higher level of importance compared to other topics as information related to the owner. This allows the response text to reflect topics that are highly important to the identified individual.
[0231] Issues corresponding to [Appendix 10] One of the purposes of this disclosure is to facilitate the generation of response text using natural language processing by using information that describes the content of an image. [Note 10] An information processing system as described in Appendix 1, wherein one or more processors use information describing the content of an image captured by the owner terminal instead of the image captured by the owner terminal. This makes it easier to generate response text using natural language processing by using information that describes the content of the image.
[0232] Issues corresponding to [Appendix 11] One of the purposes of this disclosure is to support the wide variety of images captured by the owner's terminal. [Note 11] An information processing system as described in Appendix 10, comprising one or more processors, which acquires information describing the content of an image captured by the owner terminal through a generating AI. This allows the system to handle a wide variety of images captured by the owner's device.
[0233] Issues corresponding to [Appendix 12] One of the purposes of this disclosure is to generate response text based on the content of the owner's speech when the owner does not give an imaging instruction. [Note 12] An information processing system as described in Appendix 1, wherein one or more processors generate response text using the content spoken by the owner if the content spoken by the owner does not include instructions for capturing an image. This allows the system to generate response text based on the owner's speech if the owner does not issue an imaging instruction.
[0234] Issues corresponding to [Appendix 13] One of the purposes of this disclosure is to change the content of the response text for each owner, even if the content of the utterance and the captured images are the same. [Note 13] The information processing method relating to this disclosure involves one or more processors performing the following processes: generating response text using the content spoken by the owner, an image captured by the owner terminal, and information related to the owner; and outputting audio from the owner terminal corresponding to the generated response text. According to the information processing method described above, even if the content of the utterance and the captured image are the same, the content of the response text can be changed for each owner.
[0235] Issues corresponding to [Appendix 14] One of the purposes of this disclosure is to change the content of the response text for each owner, even if the content of the utterance and the captured images are the same. [Note 14] The program disclosed herein causes one or more processors to generate response text using the content spoken by the owner, images captured by the owner terminal, and information related to the owner, and to output audio from the owner terminal corresponding to the generated response text. According to the program described above, even if the content of the speech and the images captured are the same, the content of the response text can be changed for each owner.
[0236] Issues corresponding to [Appendix 15] One of the purposes of this disclosure is to change the content of the response text for each owner, even if the content of the utterance and the captured images are the same. [Note 15] The information processing device disclosed herein includes one or more processors, which generate response text using the content spoken by the owner, an image captured by the owner terminal, and information related to the owner, and output audio from the owner terminal corresponding to the generated response text. According to the above information processing device, even if the content of the utterance and the captured image are the same, the content of the response text can be changed for each owner.
[0237] <Summary 2> [General tasks] One of the purposes of this disclosure is to provide an information processing system, information processing method, program, and information processing device that can spontaneously initiate conversations while being aware of the surrounding circumstances, including the owner.
[0238] Issues corresponding to [Appendix 1] One of the purposes of this disclosure is to enable spontaneous speech that reflects the content of autonomously captured images. [Note 1] The information processing system relating to this disclosure includes one or more processors, which generate spoken text using images captured without instruction from the owner, and output audio corresponding to the generated spoken text from the owner's terminal. According to the above information processing system, spontaneous speech that reflects the content of autonomously captured images can be enabled.
[0239] Issues corresponding to [Appendix 2] One of the purposes of this disclosure is to reflect information about an identified person in the content of spontaneous speech. [Note 2] The information processing system described in Appendix 1, wherein if an image containing a facial image of a person registered as the owner is detected, one or more processors generate spoken text for the person identified from the image. This allows information about the identified person to be reflected in the content of spontaneous speech.
[0240] Issues corresponding to [Appendix 3] One of the purposes of this disclosure is to reflect information about an identified person in the content of spontaneous speech. [Note 3] An information processing system as described in Appendix 1, comprising one or more processors, which generates spoken text for a person identified from their voice. This allows information about the identified person to be reflected in the content of spontaneous speech.
[0241] Issues corresponding to [Appendix 4] One of the purposes of this disclosure is to reflect the content of events or schedules registered for the owner in the content of spontaneous utterances. [Note 4] If the image contains information related to the content of an event or schedule registered for the owner, one or more processors generate spoken text for the detected information, as described in Appendix 1 of the information processing system. This allows the content of spontaneous speech to reflect the events or schedules registered for the owner.
[0242] Issues corresponding to [Appendix 5] One of the purposes of this disclosure is to reflect the content of events or schedules that are highly relevant to the current date and time in the content of spontaneous utterances. [Note 5] An information processing system as described in Appendix 4, comprising one or more processors, which generates spoken text using events or schedules registered for the owner that have been executed or are scheduled to be executed within a predetermined period including the current date and time. This allows users to reflect events or schedules that are highly relevant to the current date and time in their spontaneous speech.
[0243] Issues corresponding to [Appendix 6] One of the purposes of this disclosure is to prevent the same event or schedule from appearing repeatedly in the content of spontaneous utterances. [Note 6] The information processing system described in Appendix 4, wherein one or more processors generate spoken text using events or schedules registered for the owner that have not been used in conversations within a predetermined period including the current date and time. This prevents the same event or schedule from appearing repeatedly in spontaneous speech.
[0244] Issues corresponding to [Appendix 7] One of the purposes of this disclosure is to reflect the history of conversations with the owner in the content of spontaneous utterances. [Note 7] If the image contains information related to the history of conversations recorded about the owner, one or more processors generate speech text corresponding to the information, as described in Appendix 1 of the information processing system. This allows the history of conversations with the owner to be reflected in the content of spontaneous speech.
[0245] Issues corresponding to [Appendix 8] One of the purposes of this disclosure is to reflect high-priority topics in the content of spontaneous speech. [Note 8] The information processing system described in Appendix 7 includes a conversation history that contains information indicating the importance level of topics, and one or more processors generate spoken text using topics that have a higher importance level compared to other topics. This allows users to reflect important topics in their spontaneous speech.
[0246] Issues corresponding to [Appendix 9] One of the purposes of this disclosure is to support the wide variety of images captured by the owner's terminal. [Note 9] An information processing system as described in Appendix 1, comprising one or more processors, which acquires information describing the content of an image captured by the owner terminal through a generating AI. This allows the system to handle a wide variety of images captured by the owner's device.
[0247] Issues corresponding to [Appendix 10] One of the purposes of this disclosure is to enable spontaneous speech with content tailored to the identified person. [Note 10] The information processing system described in Appendix 9, wherein one or more processors provide the generating AI with information describing the content of the acquired image, as well as information about a person identified from the image or audio, to generate spoken text. This enables spontaneous speech tailored to the identified individual.
[0248] Issues corresponding to [Appendix 11] One of the purposes of this disclosure is to enable spontaneous speech that reflects the content of autonomously captured images. [Note 11] An information processing system as described in Appendix 1, comprising one or more processors, which generates spoken text independently of the owner's utterances. This enables spontaneous speech that reflects the content of autonomously captured images.
[0249] Issues corresponding to [Appendix 12] One of the purposes of this disclosure is to enable spontaneous speech that reflects the content of autonomously captured images. [Note 12] The information processing method relating to this disclosure involves one or more processors performing the following processes: generating spoken text using images captured without the owner's instructions; and outputting audio corresponding to the generated spoken text from the owner terminal. According to the above information processing method, spontaneous speech that reflects the content of autonomously captured images can be enabled.
[0250] Issues corresponding to [Appendix 13] One of the purposes of this disclosure is to enable spontaneous speech that reflects the content of autonomously captured images. [Note 13] The program relating to this disclosure causes one or more processors to generate spoken text using images captured without the owner's instructions, and to output audio corresponding to the generated spoken text from the owner's terminal. According to the program described above, spontaneous speech that reflects the content of autonomously captured images can be enabled.
[0251] Issues corresponding to [Appendix 14] One of the purposes of this disclosure is to enable spontaneous speech that reflects the content of autonomously captured images. [Note 14] The information processing device relating to this disclosure includes one or more processors, which generate spoken text using images captured without instruction from the owner, and output audio corresponding to the generated spoken text from the owner terminal. According to the above information processing device, spontaneous speech that reflects the content of autonomously captured images can be enabled. [Explanation of Symbols]
[0252] 1...Conversation System 10…Owner terminal 11, 21, 31, 41… processors 12, 22, 32, 42... Semiconductor memory 13, 23, 33, 43…Auxiliary storage device 14…Camera 15... Mike 16…Speaker 17…Display 18...Movable mechanism 19, 24, 34, 44… Communication Interfaces 20…Speech / Text Conversion Server 30…Front server 40... Conversation Server 43A...User DB 43B...Conversation Log 43C…Summary DB
Claims
1. Includes one or more processors, The one or more processors mentioned above are: Using the owner's spoken words, images captured by the owner's device, and information related to the owner, a response text is generated. The owner terminal outputs audio corresponding to the generated response text. It is an information processing system, The one or more processors mentioned above are: The system identifies the person who spoke as the owner, extracts content from the registered scheduled events or schedules for the identified person that falls within a predetermined time period determined according to the content of the scheduled events or schedules, and uses the extracted content as information related to the owner. Information processing system.
2. If the owner's spoken words include instructions for image capture, The one or more processors described above acquire an image from the owner terminal and use it to generate the response text. The information processing system according to claim 1.
3. The one or more processors mentioned above are: The system generates the response text using the scheduled events or schedules registered for the identified person that have not been used in a conversation within a predetermined period including the current date and time. The information processing system according to claim 1.
4. The one or more processors mentioned above are: The content of events or schedules planned by one or more registered individuals related to the owner will be used as information related to the said owner. The information processing system according to claim 1.
5. The one or more processors mentioned above are: The conversation history recorded for the identified person is used as information related to the owner. The information processing system according to claim 1.
6. The aforementioned conversation history includes information indicating the level of importance of the topics, The one or more processors mentioned above are: Topics with a higher level of importance compared to other topics are used as information related to the owner. The information processing system according to claim 5.
7. The one or more processors mentioned above are: Instead of using the image captured by the aforementioned owner terminal, information describing the content of the image captured by the aforementioned owner terminal is used. The information processing system according to claim 1.
8. The one or more processors mentioned above are: Information describing the content of the image captured by the aforementioned owner terminal is obtained through a generating AI. The information processing system according to claim 7.
9. The one or more processors mentioned above are: If the owner's spoken content does not include instructions for image capture, the response text is generated using the owner's spoken content. The information processing system according to claim 1.
10. The aforementioned owner includes users registered as family members, The information processing system according to claim 1.
11. One or more processors, A process that generates response text using the content spoken by the owner, images captured by the owner's device, and information related to the owner. The process involves outputting audio from the owner terminal corresponding to the generated response text, The process involves identifying the person who spoke as the owner, extracting content from the registered scheduled events or schedules for the identified person that falls within a predetermined time period determined according to the content of the scheduled events or schedules, and using the extracted content as information related to the owner. An information processing method that performs the following.
12. One or more processors, Using the owner's spoken words, images captured by the owner's device, and information related to the owner, a response text is generated. The owner terminal outputs audio corresponding to the generated response text. The system identifies the person who spoke as the owner, extracts content from the registered scheduled events or schedules for the identified person that falls within a predetermined time period determined according to the content of the scheduled events or schedules, and uses the extracted content as information related to the owner. A program that executes a process.
13. Includes one or more processors, The one or more processors mentioned above are: Using the owner's spoken words, images captured by the owner's device, and information related to the owner, a response text is generated. The owner terminal outputs audio corresponding to the generated response text. The system identifies the person who spoke as the owner, extracts content from the registered scheduled events or schedules for the identified person that falls within a predetermined time period determined according to the content of the scheduled events or schedules, and uses the extracted content as information related to the owner. Information processing device.