Information processing system, information processing method, and information processing device
The information processing system enhances virtual agent presence by synchronizing its display and responses with its schedule, addressing the perceived lack of engagement in existing systems.
Patent Information
- Application Number
- JP2025066447
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2026-01-08
- Estimated Expiration
- 2045-04-14
AI Technical Summary
Existing virtual agents are perceived as mere software and machines, lacking a sense of presence, making users reluctant to engage in conversations.
An information processing system that determines the display mode of a virtual agent based on its schedule, incorporates user utterances and schedule information into prompts for a generative AI model, and outputs the agent's display mode and conversation content to enhance its perceived presence.
The system enables users to recognize the virtual agent as a closer presence, making it more familiar and engaging through synchronized display and responsive interactions.
Smart Images

Figure 0007796275000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing system, an information processing method, and an information processing device. [Background technology]
[0002] Conventionally, there are applications that allow users to converse with AI agents (virtual agents) (for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2024-028844 Summary of the Invention [Problem to be solved by the invention]
[0004] However, existing virtual agents only provide the ability to hold conversations, and even if they display characters, the characters only function as decoration.
[0005] For this reason, existing virtual agents are ultimately nothing more than software and machines to users, and they tend to feel reluctant to talk to them, making it difficult to build a relationship with them.
[0006] In view of the above background, the present invention aims to provide an information processing system etc. that enables a user to recognize a virtual agent as a closer presence living at the same time as the user, and that makes the virtual agent a more familiar presence to the user. [Means for solving the problem]
[0007] The information processing system of the present invention comprises: a mode determination unit that determines the display mode of the virtual agent based on a schedule obtained from a schedule memory unit that stores the schedule of the virtual agent; an utterance information acquisition unit that acquires utterance information from a user; a prompt generation unit that includes the utterance information from the user and the schedule in a prompt for generating conversation content of the virtual agent in a generative artificial intelligence model; a transmission unit that transmits the prompt to the generative artificial intelligence model; a receiving unit that receives conversation content generated by the generative artificial intelligence model as a response to the prompt; and an output unit that outputs information on the display mode of the virtual agent determined by the mode determination unit and the conversation content received by the receiving unit as a response to the prompt.
[0008] The information processing method of the present invention includes the steps of determining the display mode of the virtual agent based on a schedule obtained from a schedule memory unit that stores the schedule of the virtual agent; obtaining utterance information from a user; including the utterance information from the user and the schedule in a prompt for generating the conversation content of the virtual agent by a generative artificial intelligence model; transmitting the prompt to the generative artificial intelligence model; receiving the conversation content generated by the generative artificial intelligence model as a response to the prompt; and outputting information on the determined display mode of the virtual agent and the conversation content as a response to the received prompt.
[0009] The information processing device of the present invention comprises: a mode determination unit that determines the display mode of the virtual agent based on a schedule obtained from a schedule memory unit that stores the schedule of the virtual agent; an utterance information acquisition unit that acquires utterance information from a user; a prompt generation unit that includes the utterance information from the user and the schedule in a prompt for generating conversation content of the virtual agent in a generative artificial intelligence model; a transmission unit that transmits the prompt to the generative artificial intelligence model; a receiving unit that receives conversation content generated by the generative artificial intelligence model as a response to the prompt; and an output unit that outputs information on the display mode of the virtual agent determined by the mode determination unit and the conversation content received by the receiving unit as a response to the prompt. [Effects of the Invention]
[0010] According to the present invention, the user can recognize the virtual agent as a closer presence living at the same time as the user, and the virtual agent can become a more familiar presence to the user. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram showing an example of a system configuration of an information processing system S according to the present embodiment. [Figure 2] FIG. 2 is a diagram showing an example of information processing implemented by the information processing system S according to this embodiment. [Figure 3] FIG. 3 is a diagram showing an example of the hardware configuration of the user terminal 10 according to the present embodiment. [Figure 4] FIG. 4 is a diagram showing an example of the hardware configuration of distribution server 20 according to the present embodiment. [Figure 5] FIG. 5 is a diagram showing the functional configuration of the distribution server 20. As shown in FIG. [Figure 6] FIG. 6 is a diagram showing an example of a schedule of the virtual agent VA stored in the schedule storage unit 211 according to this embodiment. [Figure 7] FIG. 7 is a diagram showing an example of information stored in background graphic storage unit 212 according to this embodiment. [Figure 8] FIG. 8 is a diagram showing an example of information stored in the motion storage unit 213 according to this embodiment. [Figure 9] FIG. 9 is a diagram showing an example of a conversation history stored in the conversation history storage unit 214 according to this embodiment. [Figure 10] FIG. 10 is a diagram showing the functional configuration of the user terminal 10. As shown in FIG. [Figure 11] FIG. 11 is a flowchart showing an example of processing executed in the information processing system S according to this embodiment. [Figure 12] FIG. 12 is a flowchart showing an example of processing performed by the distribution server 20 from determining the display mode of the virtual agent VA in step S103 of FIG. 11 to generating a prompt to cause the artificial intelligence system 30 to generate the conversation content of the virtual agent VA to the user in step S108 of FIG. 11. DETAILED DESCRIPTION OF THE INVENTION
[0012] An information processing system according to the present embodiment will be described below with reference to the drawings. Note that the following description merely shows one example of a preferred embodiment, and is not intended to limit the invention described in the claims. In all figures describing the present embodiment, common components are designated by the same reference numerals, and repeated description will be omitted.
[0013] [Outline of Information Processing System S] Fig. 1 is a diagram showing an example of the system configuration of an information processing system S according to this embodiment. As shown in Fig. 1, the information processing system S according to this embodiment is an information processing system that uses an artificial intelligence system (generative artificial intelligence model) 30 to generate conversation content for a virtual agent (AI agent) to a user.
[0014] The information processing system S according to this embodiment includes a user terminal 10, a distribution server 20, and an artificial intelligence system 30. The information processing system S may also include other devices such as servers and terminals. The user terminal 10 and the distribution server 20 are communicatively connected to each other via a network NW. The distribution server 20 and the artificial intelligence system 30 are communicatively connected to each other via the network NW. The network NW is configured with the Internet, a local area network (LAN), various mobile communication systems constructed using wireless base stations, etc. For example, the network includes 3G, 4G, and 5G mobile communication systems, LTE (Long Term Evolution), and wireless networks (e.g., Wi-Fi (registered trademark)) that can connect to the Internet via a predetermined access point. For wireless connections, communication protocols include, for example, Z-Wave (registered trademark), ZigBee (registered trademark), and Bluetooth (registered trademark). For wired connections, the network also includes networks directly connected using a USB (Universal Serial Bus) cable, etc.
[0015] The user terminal 10 is an information processing device used by a user. The user terminal 10 accepts operations by the user. The user terminal 10 launches an application according to this embodiment (hereinafter may be simply referred to as an "app") in response to a user operation. The app according to this embodiment is a virtual agent app that provides a conversation service with a virtual agent. In response to a user operation, the user terminal 10 requests content to be displayed or played by the app from the distribution server 20. Note that the user terminal 10 may display content using any app as long as it can display or play content distributed by the distribution server 20. For example, the user terminal 10 may display or play content using a dedicated application such as the virtual agent app according to this embodiment, or may display or play content using a browser app (hereinafter may be simply referred to as a "browser").
[0016] The user terminal 10 is realized by, for example, a smartphone, a tablet terminal, a notebook PC (Personal Computer), a desktop PC, a mobile phone, a PDA (Personal Digital Assistant), etc. In this embodiment, the case where the user terminal 10 is a smartphone having a touch panel function is shown.
[0017] In this embodiment, the user terminal 10 has a function of a voice response system (voice recognition), detects voice information (hereinafter also referred to as "utterance") uttered by a user, and accepts the detected utterance as input information. For example, the user terminal 10 is a terminal device that performs processing on the user's utterance. When the user terminal 10 detects the user's utterance, it transmits input information based on the user's utterance to the distribution server 20, receives content corresponding to the user's utterance, and presents it to the user.
[0018] The distribution server 20 accepts a distribution request from the user terminal 10 of the user, and distributes content related to a virtual agent that interacts with the user in a virtual agent app launched on the user terminal 10. The distribution server 20 uses an artificial intelligence system 30 to generate conversation content for the virtual agent to the user, and distributes the generated conversation content to the user terminal 10. The distribution server 20 is configured by a computer having at least a communication function and a program execution function. The distribution server 20 may be realized by one information processing device or by multiple information processing devices.
[0019] The distribution server 20 may also have a voice recognition function. The distribution server 20 may also be able to acquire information from a voice recognition server that provides a voice recognition service. In this case, the information processing system S may include the voice recognition server. In this embodiment, the user terminal 10, the distribution server 20, and the voice recognition server recognize the user's utterance and identify the user who made the utterance by appropriately using various conventional technologies, and therefore, description thereof will be omitted.
[0020] The artificial intelligence system 30 is an information processing device that has an internal LLM (Large Language Model) and is capable of outputting an output sentence (a response to a prompt) in response to an input sentence called a prompt. Examples of the artificial intelligence system 30 include ChatGPT, OpenAI GPT, PerplexityAsk, BingAI, and BERT. The artificial intelligence system 30 may be operated by a business that provides the virtual agent app according to the present embodiment, or may be used via an API (Application Programming Interface) provided by an external artificial intelligence system 30 operated by another business.
[0021] The artificial intelligence system 30 has a dialogue response (chat) function, and by giving any inquiry or command in text to the artificial intelligence system 30, a response to the inquiry or command can be obtained. In this embodiment, by sending a prompt to the artificial intelligence system 30 to cause the artificial intelligence system 30 to generate the content of a conversation between the virtual agent and the user, the content of the conversation can be obtained as a response.
[0022] Furthermore, in this embodiment, the artificial intelligence system 30 is not limited to text-based dialogue responses. For example, the artificial intelligence system 30 may be capable of voice-based dialogue responses. In this case, the artificial intelligence system 30 may receive voice data of an utterance from a user (for example, the voice data of the utterance itself, or voice data of the utterance with noise, etc., canceled out) along with the user's identification information, etc., and generate the content of the conversation for the virtual agent. Furthermore, the artificial intelligence system 30 may be an image generation AI system, such as Midjourney or Stable Diffusion, that is capable of generating image data.
[0023] An example of information processing realized by the information processing system S according to this embodiment will be described below with reference to Fig. 2. Fig. 2 is a diagram showing an example of information processing realized by the information processing system S according to this embodiment. Fig. 2 shows a case where a user is identified by a user ID "U1" (hereinafter, may be referred to as "user U1"). In Fig. 1, user U1 is assumed to be a user currently using a conversation service with a virtual agent VA provided by distribution server 20.
[0024] In step S1, user U1 operates the user terminal 10 to launch the virtual agent app. Specifically, in response to the operation of user U1, the user terminal 10 launches the virtual agent app provided by the distribution server 20. For example, user U1 launches the virtual agent app by tapping the icon AppI of the virtual agent app displayed on the user terminal 10. User U1 may also launch the virtual agent app by making a predetermined utterance to launch the virtual agent app, such as "launch the virtual agent app." Note that launching the virtual agent app also includes returning the virtual agent app from the background to the foreground.
[0025] In step S2, the user terminal 10 transmits a distribution request for the display mode of the virtual agent VA to the distribution server 20. Here, the information on the display mode of the virtual agent VA includes the motion of the virtual agent VA and the background graphic of the virtual agent VA.
[0026] In step S3, the distribution server 20, which has received a request for distribution of the display mode of the virtual agent VA from the user terminal 10, acquires the schedule of the virtual agent VA and determines the display mode of the virtual agent VA based on the acquired schedule. The schedule of the virtual agent VA defines the location and behavior of the virtual agent VA for each time period (see FIG. 6). For example, if the schedule of the virtual agent VA at the time of launching the virtual agent app is to clean (behavior) the living room (location) of the home, the distribution server 20 determines the display mode of the virtual agent VA to be a living room image as the background graphic and a cleaning motion as the motion. Details of the process of determining the display mode of the virtual agent VA will be described later.
[0027] In step S4, the distribution server 20 outputs to the user terminal 10 information on the display mode of the virtual agent VA determined in step S3 (for example, a living room image as background graphics and a cleaning motion as motion).
[0028] In step S5, the user terminal 10 receives from the distribution server 20 information about the display mode of the virtual agent VA determined in step S3, and displays / plays the virtual agent VA in the determined display mode on the display unit of the user terminal 10. In the example shown in Fig. 2, in step S3, the distribution server 20 determines the display mode of the virtual agent VA to be a living room image as the background graphic and a cleaning motion as the motion, so that an image of the virtual agent VA cleaning the living room of his / her home is displayed on the display unit of the user terminal 10.
[0029] In step S6, the user U1 recognizes the image displayed on the display unit of the user terminal 10 and, as an example, utters "What are you doing now?" into the microphone icon MI displayed on the display unit of the user terminal 10. As a result, the user terminal 10 acquires text data corresponding to "What are you doing now?" through voice recognition.
[0030] In step S7, the user terminal 10 transmits to the distribution server 20 a request for distribution of the conversation content of the virtual agent VA to the user U1, together with the text data corresponding to the acquired "What are you doing now?".
[0031] In step S8, the distribution server 20 receives a distribution request including text data corresponding to "What are you doing now?" and generates a prompt by including the text data as speech information of the user U1 and the schedule of the virtual agent VA obtained in step S3 in the prompt.
[0032] In step S9, the distribution server 20 transmits the prompt generated in step S8 to the artificial intelligence system 30.
[0033] In step S10, the artificial intelligence system 30 receives a prompt from the distribution server 20 and generates a conversation content as a response to the prompt. The conversation content generated by the artificial intelligence system 30 is text data, but may also be voice data.
[0034] In step S11, the distribution server 20 receives from the artificial intelligence system 30 the conversation content generated by the artificial intelligence system 30 in step S10.
[0035] In step S12, the distribution server 20 receives the conversation content generated by the artificial intelligence system 30 and outputs the received conversation content to the user terminal 10. At this time, if the received conversation content is text data, the distribution server 20 may convert the text data into voice data and transmit it to the user terminal 10. The conversion from text data to voice data may be performed on the user terminal 10 side.
[0036] In step S13, the user terminal 10, which has received the conversation content from the distribution server 20, outputs the conversation content. Specifically, the user terminal 10 outputs the audio data of the conversation content from a speaker and / or displays the text data of the conversation content on a display unit. In the example shown in FIG. 2, the user terminal 10 plays back the audio data of "I'm cleaning the house!" as the conversation content of the virtual agent VA to the user U1, and displays the text data of "I'm cleaning the house!"
[0037] In this embodiment, by including the utterance information of user U1 and the schedule of virtual agent VA in the prompt generated in step S8, it is possible to make virtual agent VA reply with content that is in line with the schedule of virtual agent VA, such as "I'm cleaning the house!" Note that an instruction to reply with content that is in line with the schedule of virtual agent VA may be added to the prompt generated in step S8. Furthermore, an instruction to reply with content that is in line with the schedule of virtual agent VA may be included in the prompt when the user asks virtual agent VA a question about its current location or activity.
[0038] The subsequent processing is the same as steps S6 to S13, and the processing is repeated until the user U1 stops talking to the virtual agent VA.
[0039] As described above, the information processing system S according to this embodiment can display the virtual agent VA in a display mode that conforms to the schedule of the virtual agent VA, and can cause the virtual agent VA to respond in a manner that conforms to the schedule. Therefore, the information processing system S according to this embodiment can make the user recognize the virtual agent VA as a closer presence who lives the same time as the user, and can make the virtual agent VA a more familiar presence to the user.
[0040] The configurations of the user terminal 10 and the distribution server 20 will be described in detail below.
[0041] 3 is a diagram showing an example of a hardware configuration of a user terminal 10 according to the present embodiment. The user terminal 10 is a computer such as a smartphone, a tablet terminal, a notebook PC (Personal Computer), a desktop PC, a mobile phone, or a PDA (Personal Digital Assistant).
[0042] The user terminal 10 is physically configured as a computer including a processor 101, memory 102, storage 103, communication device 104, input device 105, output device 106, and a bus connecting these. Each of these devices operates using power supplied from a battery (not shown). In the following description, the term "device" can be interpreted as a circuit, device, unit, etc. The hardware configuration of the user terminal 10 may be configured to include one or more of the devices shown in FIG. 2, or may be configured with some of the devices omitted. Furthermore, the user terminal 10 may be configured by communicating and connecting multiple devices each having a different housing.
[0043] Each function of the user terminal 10 is realized by loading specific software (programs) onto hardware such as the processor 101 and memory 102, causing the processor 101 to perform calculations, control communication via the communication device 104, and control at least one of reading and writing data in the memory 102 and storage 103.
[0044] The processor 101 controls the entire computer by running, for example, an operating system. The processor 101 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. Furthermore, for example, a baseband signal processing unit, a call processing unit, etc. may be realized by the processor 101.
[0045] The processor 101 reads programs (program codes), software modules, data, etc. from at least one of the storage 103 and the communication device 104 into the memory 102, and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described below. The functional blocks of the user terminal 10 may be implemented by a control program stored in the memory 102 and running on the processor 101. Various processes may be executed by one processor 101, or may be executed simultaneously or sequentially by two or more processors 101. The processor 101 may be implemented by one or more chips. The programs may be transmitted to the user terminal 10 via a telecommunications line.
[0046] The memory 102 is a computer-readable recording medium and may be configured by, for example, at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 102 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 102 can store executable programs (program codes), software modules, etc. for implementing the method according to this embodiment.
[0047] Storage 103 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray® disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy disk, a magnetic strip, etc. Storage 103 may also be referred to as an auxiliary storage device.
[0048] The communication device 104 enables communication with the distribution server 20. The communication device 104 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc., to realize at least one of Frequency Division Duplex (FDD) and Time Division Duplex (TDD). For example, a transmitting / receiving antenna, an amplifier unit, a transmitting / receiving unit, a transmission path interface, etc. may be realized by the communication device 104. The transmitting / receiving unit may be implemented as a transmitting unit and a receiving unit that are physically or logically separated.
[0049] The input device 105 is an input device that accepts input from the outside (for example, keys, microphones, cameras, switches, buttons, various sensors such as position information sensors and motion sensors, etc.). The output device 106 is an output device that outputs to the outside (for example, a display, speakers, LED lamps, etc.). The input device 105 and the output device 106 may be integrated into one device (for example, a touch panel).
[0050] Each device, such as the processor 101 and the memory 102, is connected by a bus for communicating information. The bus may be configured using a single bus, or different buses may be used between each device.
[0051] The user terminal 10 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 101 may be implemented using at least one of these pieces of hardware.
[0052] FIG. 4 is a diagram illustrating an example of the hardware configuration of distribution server 20 according to this embodiment. The hardware configuration of artificial intelligence system 30 according to this embodiment is similar to that of distribution server 20, and therefore a description thereof will be omitted. Distribution server 20 is an information processing device realized by, for example, a computer, a server device, a cloud system, or the like. Physically, distribution server 20 is configured as a computer device including processor 201, memory 202, storage 203, communication device 204, input device 205, output device 206, and buses connecting these. Since processor 201, memory 202, storage 203, input device 205, and output device 206 are similar in hardware to processor 101, memory 102, storage 103, input device 105, and output device 106 of user terminal 10, a description thereof will be omitted.
[0053] The communication device 204 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also called, for example, a network device, a network controller, a network card, a communication module, etc. The communication device 204 enables communication with the user terminal 10.
[0054] [Functional configuration of distribution server 20] Next, a description will be given of the functional configuration of distribution server 20 according to this embodiment. 5 is a diagram showing the functional configuration of distribution server 20. Distribution server 20 has storage unit 21 and control unit 22.
[0055] The storage unit 21 stores various types of data. For example, the storage unit 21 stores programs executed by the control unit 22. The storage unit 21 stores programs that cause the control unit 22 to function as a delivery request acquisition unit 221, a mode determination unit 222, a prompt generation unit 223, a transmission unit 224, a reception unit 225, and an output unit 226.
[0056] The storage unit 21 includes a schedule storage unit 211 , a background graphic storage unit 212 , a motion storage unit 213 , and a conversation history storage unit 214 .
[0057] (Schedule storage unit 211) The schedule storage unit 211 stores the schedule of the virtual agent VA. Fig. 6 is a diagram showing an example of the schedule of the virtual agent VA stored in the schedule storage unit 211 according to this embodiment. As shown in Fig. 6, the schedule of the virtual agent VA defines the "location" of the virtual agent VA and the "action" of the virtual agent VA for each "time".
[0058] Furthermore, the schedule of the virtual agent VA may include a schedule for holidays and a schedule for weekdays, which allows the virtual agent VA to have the concept of holidays and weekdays, allowing the user to recognize the virtual agent VA as a closer presence living the same hours as the user, making the virtual agent VA a more familiar presence to the user.
[0059] Figure 6 shows an example of a weekday schedule for a virtual agent VA, in which a schedule is set up that defines locations and actions for each time period, such as sleeping (action) in the bedroom (location) at home until 7:00 a.m., having breakfast (action) in the living room (location) at home from 7:00 to 8:00 a.m., and commuting (action) by train (location) from 8:00 to 8:30 a.m.
[0060] As an example of a holiday schedule for the virtual agent VA, a schedule is set in which locations and actions are specified for each time period, such as sleeping (action) in the bedroom (location) at home until 8:00 a.m., having breakfast (action) in the living room (location) at home from 8:00 to 9:00 a.m., and cleaning (action) in the living room (location) at home from 9:00 to 9:30 a.m.
[0061] In addition to the holiday schedule and the weekday schedule, the virtual agent VA's schedule may also include special day schedules, such as a New Year's Eve schedule, a New Year's Day schedule, and a Christmas schedule.
[0062] The schedules of these virtual agents VA are set, for example, by a business providing the virtual agent application according to this embodiment. The business may update the schedules of these virtual agents VA or add new schedules. Note that it may also be possible for a user or a third party to set, update, or add schedules for the virtual agents VA.
[0063] (Background graphic storage unit 212) The background graphic storage unit 212 stores image data for the background graphic of the virtual agent VA included in the display mode of the virtual agent VA. The background graphic storage unit 212 stores the location of the virtual agent VA and the image data of the background graphic corresponding to the location, linking them together. FIG. 7 is a diagram showing an example of information stored in the background graphic storage unit 212 according to this embodiment. As shown in FIG. 7, the information stored in the background graphic storage unit 212 consists of two items: "location" and "background graphic."
[0064] The "location" is the location of the virtual agent VA, and multiple predetermined locations are set.
[0065] The "background graphic" is image data corresponding to the "location" and is the background image of the virtual agent VA.
[0066] In the example of FIG. 7, one bedroom image is linked as a "background graphic" to a bedroom at home, which is a "location." Furthermore, multiple living room images are linked as "background graphics" to a living room at home, which is also a "location." In this embodiment, sunny, rainy, cloudy, and snowy living room images are prepared for the living room at home, which differ depending on the climate or weather. This is because the living room at home is set to have a window, and the outside can be seen through the window. Therefore, the sunny living room image is, for example, an image in which the sun is visible through the window, the cloudy living room image is, for example, an image in which clouds are visible through the window, the rainy living room image is, for example, an image in which rain is visible through the window, and the snowy living room image is, for example, an image in which snow is visible through the window. In this way, the type of image data for the "background graphic" corresponding to the "location" (e.g., sunny, rainy, cloudy, snowy) is determined by a business operator or the like depending on the setting, and the number of "background graphic" images corresponding to the "location" corresponding to the determined type is prepared by the business operator or the like. On the other hand, since the bedroom at home is set to have no windows, there is only one bedroom image corresponding to the bedroom at home.
[0067] Alternatively, living room images for special occasions such as Christmas, New Year's Eve, New Year's Day, etc. may be prepared for the living room of a home. For the living room images for special occasions, images for sunny, rainy, cloudy, and snowy days may be prepared, which differ depending on the climate or weather.
[0068] The type of "location" and the image data of "background graphics" may be added or updated as appropriate by the business operator that provides the virtual agent application according to this embodiment.
[0069] (Motion memory unit 213) The motion storage unit 213 stores motion data for the motion of the virtual agent VA included in the display mode of the virtual agent VA. The motion storage unit 213 stores the actions of the virtual agent VA and the motion data of the motion corresponding to the actions, linking them to each other. FIG. 8 is a diagram showing an example of information stored in the motion storage unit 213 according to this embodiment. As shown in FIG. 8, the information stored in the motion storage unit 213 consists of two items: "action" and "motion."
[0070] "Action" is an action that the virtual agent VA performs, and multiple predetermined actions are set.
[0071] "Motion" is motion data corresponding to "behavior" and defines the behavior of the virtual agent VA. Motion data is data that defines the behavior of the virtual agent VA. Specifically, motion data includes a timeline in which data that causes the virtual agent VA to operate is arranged along a time axis. Motion data may also be a timeline in which image data that causes the virtual agent VA to operate is arranged along a time axis.
[0072] In the example of FIG. 8 , a sleep motion is associated as a "motion" with going to bed as an "action." In addition, in this embodiment, a meal motion and a New Year's Day meal motion are associated as "motions" with breakfast as an "action," a meal motion is associated as a "motion" with lunch as an "action," and a meal motion and a Christmas meal motion are associated as "motions" with dinner as an "action." In this embodiment, a common meal motion is associated with breakfast, lunch, and dinner, but different meal motions may be associated with each. For example, the dinner meal motion may be a slower meal motion than breakfast and lunch. The New Year's Day meal motion is a motion different from a normal meal motion, such as eating ozoni, and is a motion that is used only when current events information, described below, indicates that today is New Year's Day. The Christmas meal motion is a motion different from a normal meal motion, such as holding chicken in one's hand and eating it, and is a motion that is used only when current events information, described below, indicates that today is Christmas.
[0073] Furthermore, in the embodiment, for walking as an "action," walking motions for sunny weather, rainy weather, cloudy weather, and snowy weather are prepared. The walking motion for sunny weather is, for example, a motion in which the virtual agent VA simply walks, the walking motion for cloudy weather is, for example, a motion in which the virtual agent VA walks while holding an umbrella in its hand without holding it, the walking motion for rainy weather is, for example, a motion in which the virtual agent VA walks while holding an umbrella, and the walking motion for snow is, for example, a motion in which the virtual agent VA walks while wearing a scarf and holding an umbrella. Note that the virtual agent VA may be made to perform the same walking motion regardless of the weather.
[0074] In addition, for example, for work as an "action," different work motions may be prepared depending on the time of day, such as a morning work motion, a necessary work motion, and a night work motion.
[0075] The types of "action" and the motion data of "motion" may be added or updated as appropriate by the business operator that provides the virtual agent application according to this embodiment.
[0076] Incidentally, a schedule for a virtual agent VA defines the "location" and "action" of the virtual agent VA for each "time." The "location" in the schedule is selected from a plurality of locations stored in the background graphic storage unit 212, and the "action" in the schedule is selected from a plurality of actions stored in the motion storage unit 213. In other words, the schedule for a virtual agent VA can be created by combining the data stored in the background graphic storage unit 212 and the motion storage unit 213.
[0077] (Conversation history storage unit 214) The conversation history storage unit 214 stores various information related to the conversation history between the user and the virtual agent VA. FIG. 9 is a diagram showing an example of a conversation history stored in the conversation history storage unit 214 according to this embodiment. As shown in FIG. 9, the information stored in the conversation history storage unit 214 consists of two items: "user ID" and "conversation history." The "conversation history" includes items such as "type," "date and time," and "conversation content."
[0078] "User ID" indicates identification information for identifying a user. "Conversation history" indicates the conversation history between a user and a virtual agent VA in the virtual agent app according to this embodiment. "Type" indicates the type of content in the virtual agent app. In the example of FIG. 9, "Type" stores either "Input," which indicates that speech information has been input by the user's speech, or "Output," which indicates that content has been output to the user.
[0079] "Date and time" indicates the date and time when each content was input or output. In the example of Figure 9, "date and time" is illustrated using an abstract symbol such as "date and time DA-1." Furthermore, "conversation content" indicates text data included in the content that was input or output at the corresponding date and time.
[0080] 9, content corresponding to the speech information of user U1 was input from the user terminal 10 of the user (user U1) identified by the user ID "U1" at the date and time "date and time DA-1." For example, the content input to the distribution server 20 at the date and time "date and time DA-1" includes the conversation content "What are you doing now?"
[0081] It also indicates that content was output by the distribution server 20 to the user terminal 10 of user U1 at the date and time "Date and time DA-2." For example, it indicates that the content output by the distribution server 20 at the date and time "Date and time DA-2" includes the conversation "I'm cleaning the house!"
[0082] Returning to Figure 5, the control unit 22 functions as a distribution request acquisition unit 221, a mode determination unit 222, a prompt generation unit 223, a transmission unit 224, a reception unit 225, and an output unit 226 by executing the program stored in the memory unit 21.
[0083] The distribution request acquisition unit 221 acquires a distribution request from the user terminal 10. The distribution request acquisition unit 221 acquires a distribution request for the display mode of the virtual agent VA from the user terminal 10. Here, the information on the display mode of the virtual agent VA includes the motion of the virtual agent VA and background graphics of the virtual agent VA.
[0084] Furthermore, the distribution request acquisition unit 221 acquires a distribution request for the content of the conversation between the virtual agent VA and the user, including the user's utterance information, from the user terminal 10. In this regard, the distribution request acquisition unit 221 can also be said to be an utterance information acquisition unit that acquires the user's utterance information.
[0085] The mode determination unit 222 acquires the schedule of the virtual agent VA from the schedule storage unit 211, and determines the display mode of the virtual agent VA based on the acquired schedule of the virtual agent VA. As will be described later, the schedule of the virtual agent VA is included in the prompt sent to the artificial intelligence system 30 for generating the content of the conversation of the virtual agent VA with the user. Therefore, the answers of the virtual agent VA generated by the artificial intelligence system 30 can be made to conform to the schedule of the virtual agent VA.
[0086] This allows the virtual agent VA to be displayed in a display mode that conforms to the schedule of the virtual agent VA, and allows the virtual agent VA to respond with content that conforms to the schedule of the virtual agent VA, so that the user can recognize the virtual agent VA as a closer presence that lives at the same time as the user, and the virtual agent VA can become a more familiar presence to the user.
[0087] Specifically, the behavior determination unit 222 determines the background graphic of the virtual agent VA so as to correspond to the location defined in the schedule of the virtual agent VA, and determines the motion of the virtual agent VA so as to correspond to the behavior defined in the schedule of the virtual agent VA.
[0088] For example, if the schedule for virtual agent VA when the virtual agent app is launched is cleaning (action) the living room (location) of one's home, the display mode determination unit 222 determines the display mode of virtual agent VA to be a living room image as the background graphic and a cleaning motion as the motion. In this embodiment, as shown in FIG. 7, multiple background graphic images are prepared for the living room of one's home. However, if the display mode is determined using only the schedule, a default background graphic image may be set in advance and the default image may be used. For example, in this case, a cloudy living room image is considered applicable to any climate or weather, and is therefore suitable as the default living room image.
[0089] This allows the display mode of the virtual agent VA to match the scheduled activity and location when the virtual agent VA answers in accordance with the schedule. This prevents the virtual agent VA's answer from being inconsistent with the display mode. Even if the virtual agent VA answers in accordance with the schedule, if the display mode of the virtual agent VA does not match the scheduled activity and location, the inconsistency will cause a sense of incongruity for the user.
[0090] The mode determination unit 222 may also acquire at least one of weather information and current events information for a specific region predetermined for the virtual agent VA, and determine the background graphic of the virtual agent VA's display mode based on the virtual agent VA's schedule and the acquired information. Here, the specific region predetermined for the virtual agent VA is the region where the virtual agent VA lives. For example, it may be the same region as the user's region. Weather information and current events information can be acquired, for example, from an external server. Weather information includes information such as sunny, rainy, cloudy, and snowy. Current events information includes information about what day it is today (Christmas, New Year's Day, New Year's Eve, etc.), information about ongoing events and recent news, and information about the latest events in various fields such as politics, economics, society, international affairs, sports, and culture.
[0091] For example, if the schedule of virtual agent VA at the time of launching the virtual agent app is to clean (action) the living room (place) of one's home and the weather information for the specific region at that time indicates that it will rain, the mode determination unit 222 may determine the display mode of virtual agent VA to be a rainy living room image as the background graphic and a cleaning motion as the motion. Also, for example, if the schedule of virtual agent VA at the time of launching the virtual agent app is to have dinner (action) in the living room (place) of one's home and the current events information for that day in the specific region indicates that today is Christmas, the mode determination unit 222 may determine the display mode of virtual agent VA to be a Christmas living room image as the background graphic and a eating motion as the motion.
[0092] This allows, for example, if a specific area predetermined for the virtual agent VA is set as the area where the user usually lives, the user can imagine that he or she is living in the same area as the virtual agent VA, and can recognize the virtual agent VA as a closer presence who lives at the same time as the user, making the virtual agent VA a more familiar presence to the user.
[0093] In addition, the manner determination unit 222 may acquire at least one of weather information and current affairs information for a specific region predetermined for the virtual agent VA, and determine the display manner (both background graphics and motion) of the virtual agent VA based on the schedule of the virtual agent VA and the information.
[0094] For example, if the schedule of virtual agent VA at the time of launching the virtual agent app is a walk (action) in a city (place) and the weather information in the specific region at that time indicates rain, the mode determination unit 222 may determine the display mode of virtual agent VA to be a rainy city image as the background graphic and a rainy walking motion as the motion. Also, if the schedule of virtual agent VA at the time of launching the virtual agent app is dinner (action) in the living room (place) of one's home, the weather information in the specific region at that time indicates snow, and the current events information for that day in the specific region indicates Christmas, the mode determination unit 222 may determine the display mode of virtual agent VA to be a snowy living room image as the background graphic and a Christmas eating motion as the motion.
[0095] This allows the display mode of the virtual agent VA to be determined based on the weather information and current events information of a specific region that has been predefined for the virtual agent VA, which allows the user to more easily imagine living in the same region as the virtual agent VA and to recognize the virtual agent VA as a closer presence living at the same time as the user, making the virtual agent VA a more familiar presence to the user.
[0096] Furthermore, the mode determination unit 222 may determine the display mode of the virtual agent VA to be a display mode that does not follow the schedule of the virtual agent VA according to a predetermined probability. That is, when determining the display mode of the virtual agent VA, the mode determination unit 222 may determine the display mode to be a display mode that does not follow the schedule of the virtual agent VA once in several tens of times, such as 10 times, 20 times, or 30 times. The predetermined probability can be set as appropriate by the business operator or the like.
[0097] For example, the mode determination unit 222 may randomly select one of several schedules, such as one to five, as a display mode that does not follow the schedule of the virtual agent VA, according to a predetermined probability. Specifically, referring to FIG. 6 , if today is a weekday and it is currently 9:00 a.m., the virtual agent VA would be working at the office according to the schedule. If the mode determination unit 222 determines the display mode of the virtual agent VA according to the schedule, the mode determination unit 222 may select an image of the office as the background graphic and a work motion as the motion. However, if the mode determination unit 222 determines the display mode of the virtual agent VA to be a display mode that does not follow the schedule according to a predetermined probability, the mode determination unit 222 may select, for example, sleeping in the bedroom at home, which is three schedules before the schedule of working at the office, as the display mode to be determined. In this case, the user may believe that the virtual agent VA is at home, even though he or she would normally be at the office, because he or she overslept or is feeling unwell, etc.
[0098] In this way, by having the virtual agent VA sometimes perform actions that are not according to schedule due to oversleeping, poor health, etc., the virtual agent VA can be made to seem more human and become more familiar to the user.
[0099] Also, for example, the mode determination unit 222 may randomly determine one of the schedules for that day as a display mode that does not follow the schedule of the virtual agent VA according to a predetermined probability. Also, the mode determination unit 222 may determine the display mode that does not follow the schedule of the virtual agent VA based on some kind of rule base.
[0100] The predetermined probability may also be varied depending on the time period. For example, the frequency of determining a display mode that does not follow the schedule may be medium in the morning, low in the afternoon, and high in the evening and at night.
[0101] Furthermore, the mode determination unit 222 may be realized using an LLM such as the artificial intelligence system 30, or may be realized using an LLM that can be executed locally provided in the distribution server 20. Examples of LLMs that can be executed locally include Llama 3, Gema 2, and Deep Seek-R1.
[0102] In this case, the prompt that the mode determination unit 222 uses to cause the LLM to determine the display mode of the virtual agent VA includes the schedule of the virtual agent VA, the information stored in the background graphic storage unit 212 (the information in FIG. 7), and the information stored in the motion storage unit 213 (the information in FIG. 8). Furthermore, the prompt includes an instruction to determine the display mode by selecting a background graphic and a motion in accordance with the schedule. Also, in order to determine a display mode that does not follow the schedule of the virtual agent VA according to a predetermined probability, the prompt may include, in addition to the above information, an instruction to determine a display mode that does not follow the schedule of the virtual agent VA according to a predetermined probability.
[0103] Furthermore, when the LLM is to determine a display mode that does not follow the schedule of the virtual agent VA according to a predetermined probability, the prompt may further include an instruction to the virtual agent VA to search for at least one of weather information and current affairs information for a specific region that has been predetermined, and to use that information to determine a display mode that does not follow the schedule of the virtual agent VA.
[0104] The prompt generation unit 223 generates a prompt for causing the artificial intelligence system 30 to generate the content of the conversation between the virtual agent VA and the user. The prompt generation unit 223 generates a prompt for causing the artificial intelligence system 30 to generate the content of the conversation between the virtual agent VA and the user by including utterance information from the user and the schedule of the virtual agent VA in the prompt.
[0105] Furthermore, when the mode determination unit 222 determines the display mode of the virtual agent VA to be a display mode that does not conform to the schedule of the virtual agent VA according to a predetermined probability, the prompt generation unit 223 may include inconsistency information of the executed behavior with respect to the schedule in the prompt for causing the artificial intelligence system 30 to generate the conversation content of the virtual agent VA. Note that the inconsistency information of the executed behavior with respect to the schedule is, for example, information that indicates that an action that should have been performed in the schedule has been performed but has not conformed to the schedule.
[0106] For example, if the manner determination unit 222 determines the display manner of the virtual agent VA to be a display manner that does not follow the schedule (for example, an image of a bedroom at home as the background graphic and a sleeping motion as the motion), instead of determining the display manner of the virtual agent VA to be a display manner that does not follow the schedule (for example, an image of a bedroom at home as the background graphic and a sleeping motion as the motion), the prompt generation unit 223 includes information about the inconsistency of the execution behavior with respect to the schedule in a prompt that causes the artificial intelligence system 30 to generate the conversation content of the virtual agent VA. This allows the artificial intelligence system 30 to generate an appropriate response by utilizing the information about the inconsistency of the execution behavior with respect to the schedule, even if the user utters to the virtual agent VA, "Huh? Why are you still sleeping?"
[0107] In this way, when the virtual agent VA sometimes behaves in a way that is not in line with the schedule due to oversleeping or poor health (for example, sleeping in the bedroom at home when the schedule would have called for working at the office), the virtual agent VA can be made to give a consistent answer to that behavior.
[0108] In addition, when the mode determination unit 222 determines the display mode of the virtual agent VA to be a display mode that does not conform to the schedule of the virtual agent VA according to a predetermined probability, the prompt generation unit 223 may include information on the display mode of the virtual agent VA determined by the mode determination unit 222 in addition to information on the discrepancy between the execution behavior and the schedule in a prompt for causing the artificial intelligence system 30 to generate the conversation content of the virtual agent VA.
[0109] As a result, even if the user utters to the virtual agent VA, "What are you doing now?", the artificial intelligence system 30 can generate an appropriate reply based on the information on the display mode.
[0110] The transmitting unit 224 transmits the prompt to the artificial intelligence system 30. For example, the transmitting unit 224 transmits the prompt to the artificial intelligence system 30 to cause the artificial intelligence system 30 to generate conversation content of the virtual agent VA to the user.
[0111] The receiving unit 225 receives the content of the conversation between the virtual agent VA and the user, which is generated by the artificial intelligence system 30 as a response to the prompt generated by the prompt generating unit 223 .
[0112] The output unit 226 outputs information about the display mode of the virtual agent VA determined by the mode determination unit 222 to the user terminal 10. The display mode information is, for example, information about which background graphic is to be displayed on the display unit 13 of the user terminal 10 and which motion the virtual agent VA is to perform. The display mode information may also be image data of the background graphic or motion data. The output unit 226 outputs the content of the conversation received by the receiving unit 225 as a response to the prompt to the user terminal 10. If the content of the conversation is text data, the output unit 226 may convert the text data into audio data and output it to the user terminal 10 together with or instead of the text data.
[0113] [Functional configuration of user terminal 10] Next, a functional configuration of the user terminal 10 according to this embodiment will be described. Fig. 10 is a diagram showing the functional configuration of the user terminal 10. The user terminal 10 has a storage unit 11, a control unit 12, a display unit 13, and an audio output unit 14.
[0114] The display unit 13 is, for example, a display that displays various types of information, and is an example of the output device 106 shown in FIG. 3. The display unit 13 displays various images related to the virtual agent app according to this embodiment under the control of the display control unit 124. The display unit 13 displays, for example, text data of the conversation content of the virtual agent VA under the control of the display control unit 124. Furthermore, the display unit 13 may also display, under the control of the display control unit 124, an image of the virtual agent VA corresponding to the motion data of the virtual agent VA related to the conversation content of the virtual agent VA.
[0115] The audio output unit 14 is, for example, a speaker that outputs (plays back) various types of sound information, and is an example of the output device 106 shown in Fig. 3. The audio output unit 14 plays back various sounds related to the virtual agent app according to this embodiment under the control of the audio control unit 125. The audio output unit 14 plays back audio data of the conversation content of the virtual agent VA under the control of the audio control unit 125, for example.
[0116] The storage unit 11 stores various types of data. For example, the storage unit 11 stores a program executed by the control unit 12. The storage unit 11 stores a program that causes the control unit 12 to function as a reception unit 121, a transmission unit 122, a reception unit 123, a display control unit 124, and a voice control unit 125. The storage unit 11 stores various image data of the virtual agent VA.
[0117] The control unit 12 executes the programs stored in the storage unit 11 to function as a reception unit 121, a transmission unit 122, a reception unit 123, a display control unit 124, and an audio control unit 125.
[0118] The reception unit 121 receives input from the user. For example, the reception unit 121 receives input from the user via an operation unit such as a touch panel provided on the surface of the display unit 13, a microphone, or the like. When the reception unit 121 receives an operation to start a virtual agent app, it starts the virtual agent app and notifies the transmission unit 122 of the operation content. When the reception unit 121 receives utterance information from the user to the virtual agent VA, it notifies the transmission unit 122 of the utterance information.
[0119] The transmitting unit 122 transmits a content distribution request to the distribution server 20 via the communication device 104. For example, when the transmitting unit 122 is notified by the receiving unit 121 of the operation details related to the operation for launching the virtual agent app, the transmitting unit 122 transmits a distribution request for the display mode of the virtual agent VA to the distribution server 20 via the communication device 104. Furthermore, when the transmitting unit 122 is notified by the receiving unit 121 of the user's utterance information to the virtual agent VA, the transmitting unit 122 transmits to the distribution server 20 a distribution request for the conversation details of the virtual agent VA including text data corresponding to the utterance information.
[0120] The receiving unit 123 receives information on the display mode of the virtual agent VA transmitted from the distribution server 20. The receiving unit 123 receives the content of the conversation of the virtual agent VA transmitted from the distribution server 20.
[0121] The display control unit 124 outputs the information on the display mode of the virtual agent VA received by the receiving unit 123, and displays / plays the virtual agent VA in that display mode. Specifically, the display control unit 124 displays the background graphic included in the display mode on the display unit 13, and plays the motion included in the display mode on the display unit 13. The display control unit 124 also causes the display unit 13 to display text data of the conversation content of the virtual agent VA.
[0122] The voice control unit 125 causes the voice output unit 14 to reproduce the voice data of the conversation content of the virtual agent VA.
[0123] [Operation flow] FIG. 11 is a flowchart showing an example of processing executed in the information processing system S according to this embodiment.
[0124] In step S101, the user terminal 10 starts up a virtual agent application. Note that the start-up of the virtual agent application here also includes the case where the virtual agent application is returned from the background to the foreground.
[0125] In step S102, the transmission unit 122 of the user terminal 10 transmits to the distribution server 20 a request for distribution of the display mode of the virtual agent VA.
[0126] In step S103, the mode determination unit 222 of the distribution server 20, which has received a request for distribution of the display mode of the virtual agent VA from the user terminal 10, determines the display mode of the virtual agent VA based on the schedule of the virtual agent VA. Details of the process for determining the display mode in step S103 will be described later with reference to FIG.
[0127] In step S104, the transmitting unit 224 of the distribution server 20 outputs to the user terminal 10 information on the display mode of the virtual agent VA determined by the mode determining unit 222 in step S103.
[0128] In step S105, the display control unit 124 of the user terminal 10, which has received the information on the display mode of the virtual agent VA from the distribution server 20, outputs the received information on the display mode of the virtual agent VA and displays / plays the virtual agent VA in that display mode. Specifically, the display control unit 124 displays the background graphic included in the display mode on the display unit 13, and plays the motion included in the display mode on the display unit 13.
[0129] The user, who has recognized the display mode of the virtual agent VA, speaks to the virtual agent VA toward the microphone icon MI displayed on the display unit 13 of the user terminal 10. As a result, in step S106, the reception unit 121 of the user terminal 10 acquires, by voice recognition, utterance information, which is text data corresponding to the utterance.
[0130] In step S107, the transmitting unit 122 of the user terminal 10 transmits to the distribution server 20 a request for distribution of the conversation content of the virtual agent VA to the user, including the acquired speech information.
[0131] When the delivery request acquisition unit 221 of the delivery server 20 receives a delivery request including utterance information, in step S108, the prompt generation unit 223 of the delivery server 20 generates a prompt by including the utterance information in the prompt. Details of the process of generating the prompt in step S108 will be described later with reference to FIG.
[0132] In step S109, the transmission unit 224 of the distribution server 20 transmits the prompt generated in step S108 to the artificial intelligence system 30.
[0133] In step S110, the artificial intelligence system 30 receives a prompt from the distribution server 20 and generates a conversation content (a conversation content of the virtual agent VA in response to the user's utterance) as a response to the prompt.
[0134] In step S111, the receiving unit 225 of the distribution server 20 receives from the artificial intelligence system 30 the conversation content generated by the artificial intelligence system 30 in step S110.
[0135] In step S112, the output unit 226 of the distribution server 20 receives the conversation content generated by the artificial intelligence system 30 and outputs the received conversation content to the user terminal 10.
[0136] In step S113, the user terminal 10 that has received the conversation content outputs the conversation content. Specifically, the audio control unit 125 of the user terminal 10 causes the audio output unit 14 to play back the audio data of the conversation content, and / or the display control unit 124 causes the display unit 13 to display the text data of the conversation content.
[0137] Thereafter, the processes of steps S106 to S113 are repeated until the user stops talking to the virtual agent VA.
[0138] Fig. 12 is a flowchart showing an example of processing executed by the distribution server 20 from determining the display mode of the virtual agent VA in step S103 of Fig. 11 to generating a prompt for causing the artificial intelligence system 30 to generate a conversation content of the virtual agent VA to the user in step S108 of Fig. 11. Note that the flowchart of Fig. 12 shows a method for determining the display mode of the virtual agent VA without using an LLM.
[0139] In step S201, the distribution request acquisition unit 221 acquires from the user terminal 10 a distribution request for the display mode of the virtual agent VA.
[0140] In step S202, the mode determining unit 222 acquires the schedule of the virtual agent VA stored in the schedule storage unit 211.
[0141] In step S203, the mode determination unit 222 acquires weather information for a specific region predetermined for the virtual agent VA from an external server or the like. The weather information includes information such as sunny, rainy, cloudy, and snowy. The processing of step S203 is necessary when weather information is taken into consideration when determining the display mode. Therefore, if weather information is not used to determine the display mode, the processing of step S203 can be omitted.
[0142] In step S204, the mode determination unit 222 acquires current events information for a specific region predetermined for the virtual agent VA from an external server or the like. The current events information includes information about what day it is today (Christmas, New Year's Day, New Year's Eve, etc.), information about ongoing events and recent news, and information about the latest events in various fields such as politics, economics, society, international affairs, sports, and culture. The processing of step S204 is necessary when current events information is taken into consideration in determining the display mode. Therefore, if current events information is not used in determining the display mode, the processing of step S204 can be omitted.
[0143] In step S205, the mode determination unit 222 determines a background graphic in the display mode of the virtual agent VA using the information acquired in steps S201 to S204. Specifically, the mode determination unit 222 determines the background graphic of the virtual agent VA by selecting one from various background graphics stored in the background graphic storage unit 212 using the information acquired in steps S201 to S204.
[0144] In step S206, the manner determination unit 222 determines a motion in the display manner of the virtual agent VA using the information acquired in steps S201 to S204. Specifically, the manner determination unit 222 determines the motion of the virtual agent VA by selecting one of various motions stored in the motion storage unit 213 using the information acquired in steps S201 to S204. Note that when determining the display manner of the virtual agent VA in steps S205 and S206, the manner determination unit 222 may determine the display manner of the virtual agent VA to be a display manner that does not follow the schedule of the virtual agent VA according to a predetermined probability. The display manner that does not follow the schedule may be determined randomly, or, for example, several display manners to which the original schedule is changed may be prepared (for example, for work, travel, movies, colds, etc.; for cleaning, TV, reading, walks, etc.). Note also that either step S205 or step S206 may be performed first.
[0145] In step S207, the output unit 226 outputs to the user terminal 10 the display mode (background graphic, motion) of the virtual agent VA determined by the mode determination unit 222 in steps S205 and S206.
[0146] In step S208, the delivery request acquisition unit 221 determines whether or not a delivery request for conversation content including user utterance information has been acquired. If the delivery request acquisition unit 221 has acquired a delivery request, the flow proceeds to step S209. If the delivery request acquisition unit 221 has not acquired a delivery request, the delivery request acquisition unit 221 waits until it acquires a delivery request.
[0147] In step S209, the prompt generation unit 223 determines whether the manner determination unit 222 has determined a display manner that conforms to the schedule of the virtual agent VA or a display manner that does not conform to the schedule. If the manner determination unit 222 has determined a display manner that conforms to the schedule of the virtual agent VA, the flow proceeds to step S210. If the manner determination unit 222 has not determined a display manner that does not conform to the schedule of the virtual agent VA, the flow proceeds to step S211.
[0148] In step S210 (when the display mode according to the schedule has been determined), the prompt generating unit 223 includes, in the prompt, utterance information from the user in addition to the schedule of the virtual agent VA.
[0149] In step S211 (when a display mode that does not conform to the schedule has been determined), the prompt generation unit 223 includes, in the prompt, the schedule of the virtual agent VA, as well as utterance information from the user and information about the discrepancy between the execution behavior and the schedule. In addition to this information, the prompt may also include information about the display mode of the virtual agent VA determined by the mode determination unit 222. That is, it includes information about the display mode regarding behavior that differs from the original schedule. For example, if the schedule shows that the user is working but is actually on vacation at the beach, it includes information about the display mode of playing on the beach. This makes it possible to generate an appropriate conversation based on behavior that has changed from the original schedule.
[0150] As described above, in the distribution server 20 of the information processing system S according to this embodiment, the virtual agent VA can be displayed in a display mode that conforms to the schedule of the virtual agent VA, and the virtual agent VA can be made to respond in a manner that conforms to the schedule. This allows the user to recognize the virtual agent as a closer presence that lives at the same time as the user, making the virtual agent a more familiar presence to the user.
[0151] [Other embodiments] The information processing system S of the present invention has been described above with reference to the present embodiment, but the information processing system S of the present invention is not limited to the above embodiment. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the technical scope of the present invention. Furthermore, systems or devices that combine the separate features included in the present embodiment in any way are also included in the technical scope of the present invention.
[0152] The present invention may also be applied to a system consisting of multiple devices, or to a single device. Furthermore, the present invention is also applicable when an information processing program that realizes the functions of the present embodiment is supplied to a system or device and executed by a built-in processor. The technical scope of the present invention also includes a program installed on a computer to realize the functions of the present invention, a medium storing the program, a server that downloads the program, and a processor that executes the program. In particular, the technical scope of the present invention includes at least a non-transitory computer-readable medium storing a program that causes a computer to execute the processing steps included in the above-described embodiments.
[0153] In the above embodiment, the user terminal 10 interacts with the artificial intelligence system 30 via the distribution server 20, allowing the user to converse with the virtual agent VA. If the user terminal 10 can have the functions of the control unit 22 of the distribution server 20 (the functions shown in FIG. 5 ) and can access the data in the storage unit 21 of the distribution server 20, the user terminal 10 can interact directly with the artificial intelligence system 30, allowing the user to converse with the virtual agent VA. Furthermore, the user terminal 10 may be provided with at least some of the functions of the control unit 22 of the distribution server 20 (the functions shown in FIG. 5 ), allowing the user terminal 10 to access the data in the storage unit 21 of the distribution server 20.
[0154] The following additional notes are provided regarding the above-described embodiments.
[0155] [Appendix 1] The information processing system a mode determining unit that determines a display mode of the virtual agent based on a schedule acquired from a schedule storage unit that stores the schedule of the virtual agent; an utterance information acquisition unit that acquires utterance information from a user; a prompt generation unit that includes the user's speech information and the schedule in a prompt for causing a generative artificial intelligence model to generate a conversation content of the virtual agent; a sending unit that sends the prompt to the generative artificial intelligence model; a receiving unit that receives a conversation content generated by the generative artificial intelligence model as a response to the prompt; The display device includes an output unit that outputs information about the display mode of the virtual agent determined by the mode determination unit and the content of the conversation received by the receiving unit as a response to the prompt.
[0156] With the above configuration, the virtual agent can be displayed in a display mode that conforms to the schedule, and the virtual agent can be made to respond in a manner that conforms to the schedule, so that the user can recognize the virtual agent as a closer presence that lives at the same time as the user, making the virtual agent a more familiar presence to the user.
[0157] [Appendix 2] In the information processing system of Appendix 1, The mode determining unit determines the display mode of the virtual agent to be a display mode that does not follow the schedule of the virtual agent, according to a predetermined probability.
[0158] With the above configuration, it is possible to make the virtual agent perform actions that are not in line with the schedule due to occasional oversleeping or poor health, etc., thereby giving the virtual agent a more human-like presence and making it more familiar to the user.
[0159] [Appendix 3] In the information processing system of Appendix 2, When the mode determination unit determines the display mode of the virtual agent to be a display mode that does not follow the schedule of the virtual agent according to a predetermined probability, The prompt generation unit includes, in the prompt, information on the inconsistency of the execution behavior with respect to the schedule.
[0160] With the above configuration, when the virtual agent behaves in a manner that is not in line with the schedule due to oversleeping, poor health, or the like, the virtual agent can be made to give a consistent response to the behavior.
[0161] [Appendix 4] In any of the information processing systems of Supplementary Notes 1 to 3, the virtual agent's schedule defines the location and behavior of the virtual agent for each time period; the information on the display mode of the virtual agent includes a motion of the virtual agent and a background graphic of the virtual agent; The manner determining unit determines the background graphic in accordance with a location defined in a schedule of the virtual agent, and determines the motion in accordance with an action defined in the schedule of the virtual agent.
[0162] With the above configuration, the background graphic can be determined to correspond to the location defined in the virtual agent's schedule, and the motion can be determined to correspond to the action defined in the virtual agent's schedule. As a result, when the virtual agent's response conforms to the schedule, the display mode of the virtual agent matches the scheduled action and location, so there is no contradiction between the virtual agent's response and the display mode. Even if the virtual agent responds in accordance with the schedule, if the display mode does not match the scheduled action and location, the contradiction will be felt, creating a sense of incongruity.
[0163] [Appendix 5] In the information processing system of Appendix 4, The mode determination unit acquiring at least one of weather information and current events information for a region predetermined for the virtual agent; The background graphic is determined based on the virtual agent's schedule and the obtained information.
[0164] With the above configuration, the virtual agent can be made to behave as if it actually exists in a predetermined area. Furthermore, if the predetermined area for the virtual agent is, for example, the area where the user usually lives, the user can imagine that he or she is living in the same area as the virtual agent, and can recognize the virtual agent as a closer presence that lives at the same time as the user, making the virtual agent a more familiar presence to the user.
[0165] [Appendix 6] In any of the information processing systems of Supplementary Notes 1 to 5, The mode determination unit acquiring at least one of weather information and current events information for a region predetermined for the virtual agent; The display mode is determined based on the schedule of the virtual agent and the acquired information.
[0166] With the above configuration, the display mode of the virtual agent can be determined based on weather information and current events information for a region predetermined for the virtual agent. This allows the virtual agent to behave as if it actually exists in the predetermined region. Furthermore, if the region predetermined for the virtual agent is the region where the user usually lives, for example, the user can more easily imagine living in the same region as the virtual agent, and can recognize the virtual agent as a closer presence living at the same time as the user, making the virtual agent a more familiar presence to the user.
[0167] [Appendix 7] In any of the information processing systems of Supplementary Notes 1 to 6, The virtual agent's schedule includes a schedule for holidays and a schedule for weekdays.
[0168] With the above configuration, the virtual agent can be given the concepts of holidays and weekdays, so that the user can recognize the virtual agent as a closer presence living at the same time as the user, making the virtual agent a more familiar presence to the user.
[0169] [Appendix 8] The information processing method is determining a display mode of the virtual agent based on a schedule acquired from a schedule storage unit storing the schedule of the virtual agent; acquiring utterance information from a user; a step of including the user's speech information and the schedule in a prompt for causing a generative artificial intelligence model to generate a conversation content of the virtual agent; sending the prompt to the generative artificial intelligence model; receiving conversational content generated by the generative artificial intelligence model in response to the prompt; and outputting information on the determined display mode of the virtual agent and the content of the conversation as a response to the received prompt.
[0170] This configuration provides the same operational effects as the information processing system of Supplementary Note 1. The information processing method of Supplementary Note 8 may include various configurations of the information processing systems described in Supplementary Notes 2 to 7.
[0171] [Appendix 9] The information processing device a mode determining unit that determines a display mode of the virtual agent based on a schedule acquired from a schedule storage unit that stores the schedule of the virtual agent; an utterance information acquisition unit that acquires utterance information from a user; a prompt generation unit that includes the user's speech information and the schedule in a prompt for causing a generative artificial intelligence model to generate a conversation content of the virtual agent; a sending unit that sends the prompt to the generative artificial intelligence model; a receiving unit that receives a conversation content generated by the generative artificial intelligence model as a response to the prompt; The display device includes an output unit that outputs information about the display mode of the virtual agent determined by the mode determination unit and the content of the conversation received by the receiving unit as a response to the prompt.
[0172] This configuration provides the same operational effects as the information processing system of Supplementary Note 1. The information processing device of Supplementary Note 9 may include various configurations of the information processing systems described in Supplementary Notes 2 to 7. [Explanation of symbols]
[0173] 10 User terminal 11 Storage section 12 Control Unit 13 Display section 14 Audio output section 121 Reception 122 Transmitter 123 Receiving unit 124 Display control unit 125 Audio control unit 20 Distribution server (information processing device) 21 Memory section 211 Schedule memory section 212 Background Graphics Memory 213 Motion Memory Unit 214 Conversation History Mechanism 22 Control Unit 221 Delivery request acquisition unit 222 Mode Determination Unit 223 Prompt Generation Unit 224 Transmitter 225 Receiving Unit 226 Output section 30 Artificial Intelligence Systems (Generative AI Models)
Claims
1. a mode determining unit that determines a display mode of the virtual agent based on a schedule acquired from a schedule storage unit that stores the schedule of the virtual agent; an utterance information acquisition unit that acquires utterance information from a user; a prompt generation unit that includes the user's speech information and the schedule in a prompt for causing a generative artificial intelligence model to generate a conversation content of the virtual agent; a sending unit that sends the prompt to the generative artificial intelligence model; a receiving unit that receives a conversation content generated by the generative artificial intelligence model as a response to the prompt; An information processing system comprising: an output unit that outputs information on the display mode of the virtual agent determined by the mode determination unit and the conversation content as a response to the prompt received by the receiving unit.
2. The information processing system according to claim 1 , wherein the mode determining unit determines the display mode of the virtual agent to be a display mode that does not follow the schedule of the virtual agent, according to a predetermined probability.
3. When the mode determination unit determines the display mode of the virtual agent to be a display mode that does not follow the schedule of the virtual agent according to a predetermined probability, The information processing system according to claim 2 , wherein the prompt generation unit includes inconsistency information of the execution behavior with respect to the schedule in the prompt.
4. the virtual agent's schedule defines the location and behavior of the virtual agent for each time period; the information on the display mode of the virtual agent includes a motion of the virtual agent and a background graphic of the virtual agent; The information processing system according to claim 1 , wherein the manner determination unit determines the background graphic to correspond to a location defined in the schedule of the virtual agent, and determines the motion to correspond to an action defined in the schedule of the virtual agent.
5. The mode determination unit acquiring at least one of weather information and current events information for a region predetermined for the virtual agent; The information processing system according to claim 4 , wherein the background graphic is determined based on the schedule of the virtual agent and the acquired information.
6. The mode determination unit acquiring at least one of weather information and current events information for a region predetermined for the virtual agent; The information processing system according to claim 1 , wherein the display mode is determined based on the schedule of the virtual agent and the acquired information.
7. The information processing system according to claim 1 , wherein the virtual agent's schedule includes a schedule for holidays and a schedule for weekdays.
8. A computer-implemented information processing method, comprising: determining a display mode of the virtual agent based on a schedule acquired from a schedule storage unit storing the schedule of the virtual agent; acquiring utterance information from a user; a step of including the user's speech information and the schedule in a prompt for causing a generative artificial intelligence model to generate a conversation content of the virtual agent; sending the prompt to the generative artificial intelligence model; receiving conversational content generated by the generative artificial intelligence model in response to the prompt; an information processing method comprising a step of outputting information on the determined display mode of the virtual agent and a conversation content as a response to the received prompt.
9. a mode determining unit that determines a display mode of the virtual agent based on a schedule acquired from a schedule storage unit that stores the schedule of the virtual agent; an utterance information acquisition unit that acquires utterance information from a user; a prompt generation unit that includes the user's speech information and the schedule in a prompt for causing a generative artificial intelligence model to generate a conversation content of the virtual agent; a sending unit that sends the prompt to the generative artificial intelligence model; a receiving unit that receives a conversation content generated by the generative artificial intelligence model as a response to the prompt; an output unit that outputs information about the display mode of the virtual agent determined by the mode determination unit and a conversation content as a response to the prompt received by the receiving unit.
Citation Information
Patent Citations
Agent living virtually
JP2015056177A
System
JP2025048816A
System
JP2025051878A
System and program
JP2024028844A