program

The program enhances user interaction with AI by allowing inquiries through virtual objects only when conditions are met, optimizing communication and reducing unnecessary requests and costs.

JP2026081727AActive Publication Date: 2026-05-19COLOPL
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
COLOPL
Filing Date
2024-11-05
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing systems lack convenience in user interaction with artificial intelligence via virtual objects, particularly in controlling information processing and communication connections.

Method used

A program that allows a user to make inquiries to artificial intelligence through an object associated with AI when a predetermined condition is met, such as proximity and orientation, thereby controlling information communication with external devices.

Benefits of technology

Improves user convenience by ensuring that inquiries are processed only when the user is in an appropriate relative position or orientation, reducing unnecessary requests and charges to the AI server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026081727000001_ABST
    Figure 2026081727000001_ABST
Patent Text Reader

Abstract

This program provides an object equipped with a function to query artificial intelligence, thereby improving the convenience of user inquiries to artificial intelligence. [Solution] The program allows a computer to make inquiries to the artificial intelligence via an object (a virtual space concierge avatar, a guidance display device) if the object associated with the artificial intelligence (a virtual space concierge avatar, a guidance display device) and the user (a user object) using the object meet predetermined conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a program.

Background Art

[0002] Conventionally, a technique for controlling information processing for an input via a virtual object has been known (for example, see Patent Document 1).

Prior Art Document

Patent Document

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The present invention aims to improve the convenience for users.

Means for Solving the Problems

[0005] To solve the above problems, the program according to the present invention causes a computer to permit a user to make an inquiry to the artificial intelligence via the object when an object associated with the artificial intelligence and the user who uses the object satisfy a predetermined condition.

Effects of the Invention

[0006] According to the present invention, the convenience for users is improved.

Brief Description of the Drawings

[0007] [Figure 1] It is a diagram showing an overview of the system according to the present embodiment. [Figure 2] It is a hardware configuration diagram of a server. [Figure 3] It is a hardware configuration diagram of an HMD set which is an example of a user terminal. [Figure 4] This is a hardware configuration diagram of a tablet device, another example of a user terminal. [Figure 5] This is a hardware configuration diagram of a digital signage system, which is an example of a guidance device. [Figure 6] This is a diagram that conceptually represents one aspect of a virtual space. [Figure 7] This diagram shows a YZ cross-section of the field of view in a virtual space, viewed from the X direction. [Figure 8] This diagram shows the XZ cross-section of the field of view in a virtual space, viewed from the Y direction. [Figure 9] This is a functional block diagram of the server and user terminal according to the first embodiment. [Figure 10] This is a flowchart showing the server processing according to the first embodiment. [Figure 11] This is a flowchart showing the processing of the first user terminal according to the first embodiment. [Figure 12] This is an image diagram of the virtual space according to the first embodiment. [Figure 13] This is an image diagram of a guidance counter in a virtual space according to the first embodiment. [Figure 14] This is an image diagram of a guidance counter in a virtual space according to the first embodiment. [Figure 15] This is a functional block diagram of the server and guidance device according to the second embodiment. [Figure 16] This is a flowchart showing the processing of the guide device according to the second embodiment. [Figure 17] This is an external view of the guide device according to the second embodiment. [Figure 18] This is a display image diagram of the guidance device according to the second embodiment. [Figure 19] This is a display image diagram of the guidance device according to the second embodiment. [Modes for carrying out the invention]

[0008] Hereinafter, embodiments according to the present invention will be described with reference to the drawings. In this specification, "this embodiment" shall apply to both the first embodiment and the second embodiment described later. The embodiments of the present invention described below show an example when embodying the present invention, and do not limit the scope of the present invention to the scope of the description of the embodiments. Therefore, the present invention can be implemented with various modifications to the embodiments. First, the outline of the system 1 according to this embodiment will be described.

[0009] [Outline of System 1] FIG. 1 is a diagram showing an outline of a system 1 according to an embodiment of the present invention. The system according to the present invention has a function of controlling the permission of information communication with an external device to which the system makes a communication connection. More specifically, when the relative relationship between an object included in the system according to the present invention and having a communication connection function with an external device and the user of the system satisfies a predetermined condition, control is performed to permit communication connection with the external device via the object. In other words, when the relative relationship between the object and the external device does not satisfy the predetermined condition, control is performed not to permit communication connection with the external device via the object.

[0010] In this embodiment, that the system 1 "permits" information communication with an external device to which it makes a communication connection includes, for example, making operable a configuration that operates when a device constituting the system 1 transmits and receives information to and from the external device. It also includes enabling a processing function provided by the device constituting the system 1 for transmitting and receiving information to and from the external device.

[0011] FIG. 1 is a configuration example of the system 1 according to this embodiment. FIG. 1(A) illustrates the system configuration of a system 1A which is a first example of the system 1. FIG. 1(B) illustrates the system configuration of a system 1B which is a second example of the system 1.

[0012] As shown in FIG. 1(A), the system 1A mainly includes a server 10 and user terminals 20A, 20B, 20C (hereinafter, these may be collectively referred to as "user terminals 20"). Although three user terminals 20 are illustrated in FIG. 1A, the example of the user terminals 20 included in the system 1A is not limited thereto. The server 10 and the user terminals 20 are connected to be able to communicate with each other via a communication network 2. Specific examples of the communication network 2 are not particularly limited, but for example, it is composed of the Internet, a mobile communication system (e.g., 4G, 5G, etc.), a wireless network such as Wi-Fi (registered trademark), or a combination thereof.

[0013] Also, as shown in FIG. 1(B), the system 1B mainly includes a server 10 and a guidance device 30. The server 10 and the guidance device 30 are connected to be able to communicate with each other via the communication network 2. Specific examples of the communication network 2 are not particularly limited, but are the same as in the first example.

[0014] In addition, in both the system 1A and the system 1B, an AI server 40 is connected to be able to communicate with each other via the communication network 2. The AI server 40 is an example of an information processing device equipped with a learned model capable of dialogue in natural language (Artificial Intelligence).

[0015] [Overview of System 1A] The system 1A provides a function to respond to inquiries from users of each of the plurality of user terminals 20A to 20C in the virtual space provided by the server 10. More specifically, a user character (hereinafter, may be referred to as "user avatar") associated with the user using the user terminal 20 is arranged in the virtual space, and an image of the virtual space viewed from a preset viewpoint is displayed on the user terminal 20.

[0016] System 1A provides a function for making inquiries to the AI ​​server 40, which allows the user to obtain a predetermined answer by utilizing an object placed in a virtual space. In System 1A, the "object for making inquiries to the AI ​​server 40" refers, for example, to a concierge avatar (concierge object) with which a user avatar (user object) placed in a virtual space generated on the user terminal 20 interacts.

[0017] In other words, in System 1A, the concierge avatar, which is one of the objects with which the user avatar (user object) generated and deployed on the user terminal 20 interacts, is an example of an object associated with artificial intelligence. The concierge avatar is an example of an object that has the function of making inquiries to artificial intelligence. When a user using the user terminal 20 makes an inquiry to the artificial intelligence, predetermined information processing is performed on the information that the user inputs to the server 10 via the concierge avatar, and the inquiry is generated as a request to the artificial intelligence. For example, when a user asks a question to the artificial intelligence, the user asks the question (which may be non-textual information such as voice, or textual information consisting of a question sentence) to the concierge avatar via the user avatar, and the server 10 generates a request to the artificial intelligence from the input question. The request generated by the server 10 is then notified to the AI ​​server 40 as an API request. The AI ​​server 40 notifies the server 10 of the API response, which includes the answer to the question. The server 10 converts the answer included in the API response into a form that can be recognized by the user corresponding to the user avatar via the concierge avatar, and outputs it to the user terminal 20.

[0018] The user terminal 20 operates the user avatar in the virtual space in conjunction with human operations on the control device. Avatar movements include, for example, moving within the virtual space, moving different parts of the body, changing posture, changing facial expressions, speaking, and moving objects placed within the virtual space.

[0019] Furthermore, the user terminal 20 transmits avatar data indicating changes in the avatar's movements and status to the server 10 via the communication network 2. The server 10 transmits the avatar data received from the user terminal 20 to other user terminals 20 via the communication network 2. The user terminal 20 then updates the movements and status of the corresponding avatar in the virtual space based on the avatar data received from the server 10.

[0020] Furthermore, System 1A provides a virtual space that corresponds to the real space. That is, the virtual space according to this embodiment is a space that mimics the real space. However, the virtual space provided by System 1 does not need to be a perfect match to the real space and may be appropriately stylized. The real space refers to, for example, a city, commercial facilities, shops, theme parks, event venues, parks, train stations, etc. that actually exist.

[0021] System 1A allows server 10 to generate a request corresponding to input information (e.g., inquiry information including a question) or to send the generated request to AI server 40 if the relative relationship between the character (concierge avatar) placed on an object corresponding to an information counter in the virtual space and the user avatar satisfies predetermined conditions.

[0022] In other words, System 1A (Server 10) does not permit the concierge avatar to generate a request corresponding to the input inquiry information, or to send the generated request to the AI ​​Server 40, if the user avatar does not meet the predetermined conditions. Here, "does not permit request generation" includes, for example, accepting the input of inquiry information via the user avatar but not executing the request generation process based on that inquiry information.

[0023] In System 1A, the function to control whether to permit or deny requests to the AI ​​server 40, as described above, is implemented by a computer program (hereinafter simply referred to as "the program") described later.

[0024] Server 10 analyzes the information received from the user terminal 20 and performs a process to determine whether it is possible to generate a request to the AI ​​server 40 and to authorize the request device. If Server 10 does not authorize the generation of a request to the AI ​​server 40, for example, if the source of the information input is the user terminal 20, it may be the case that the user avatar is not in a predetermined relative relationship with the concierge avatar in the virtual space.

[0025] "When a predetermined relative relationship is not maintained in the virtual space" means, for example, that the distance between the user avatar and the concierge avatar in the virtual space is not less than a predetermined distance, that is, that the user avatar is not approaching the concierge avatar. Also, "When a predetermined relative relationship is not maintained in the virtual space" means that the user avatar is not facing the concierge avatar in the virtual space, that is, that the user avatar is approaching the concierge avatar but is not facing it. Furthermore, "When a predetermined relative relationship is not maintained in the virtual space" means that the state in which the distance between the user avatar and the concierge avatar is less than a predetermined distance in the virtual space does not continue for a certain period of time or longer. Similarly, if the user avatar is facing the concierge avatar, and that facing state does not continue for a certain period of time or longer.

[0026] As described above, the virtual space according to this embodiment contains a user avatar and a concierge avatar. The user avatar and concierge avatar are characters that can be operated by a person using the user terminal 20 within the virtual space, such as a person, animal, or plant. The user avatar operates within the virtual space in conjunction with the user's operations through the user terminal 20. In contrast, the appearance of the concierge avatar is determined by the server 10. For example, the server 10 may select from a pre-registered list of concierge avatar candidates. The concierge avatar may also operate in a manner determined by the server 10, or it may remain stationary.

[0027] [Overview of System 1B] System 1B provides a function for making inquiries to the AI ​​server 40, which allows a person to use an object placed in the real world to obtain a predetermined answer. In System 1B, the "object for making inquiries to the AI ​​server 40" is, for example, the guidance device 30 that displays a concierge avatar (concierge object) that responds to information input from a person who wants to make an inquiry to the artificial intelligence. The guidance device 30 also outputs a response from the artificial intelligence to the inquiry.

[0028] In other words, in System 1B, the guidance device 30 displaying the concierge avatar is an example of an object associated with artificial intelligence. The guidance device 30 corresponds to an object that has a function to inquire with artificial intelligence. When a person using the guidance device 30 makes an inquiry to the artificial intelligence, predetermined information processing is performed on the information that the person inputs to the server 10 via the concierge avatar, and the inquiry is generated as a request to the artificial intelligence. For example, when a person asks a question to the artificial intelligence, the person asks the question to the concierge avatar. This question may be non-textual information such as voice, or textual information consisting of a question sentence. The server 10 generates a request to inquire with the artificial intelligence from the question input to the concierge avatar. The server 10 then notifies the AI ​​server 40 of the generated request as an API request. The AI ​​server 40 notifies the server 10 of the API response in response to the API request, which includes the answer to the question. The server 10 converts the answer included in the API response into a form that can be recognized by a person via the concierge avatar and outputs it via the concierge avatar.

[0029] The guidance device 30 is a device that detects the movements and voices of people in the real world. The guidance device 30 also acquires information and voices of people and transmits them to the server 10 via the communication network 2. For example, the type of guidance device 30 connected to system 1 is not limited to one type, and multiple types may be combined.

[0030] As an example, the information device 30 is a stationary display device installed in commercial facilities, shops, theme parks, event venues, parks, train stations, etc. A specific example is a digital signage that displays responses to user inquiries such as the location of a store or the characteristics of a store.

[0031] System 1B does not permit the guidance device 30 to generate a request corresponding to the input inquiry information, or to send the generated request to the AI ​​server 40, if the person does not meet the predetermined conditions.

[0032] In System 1B, the function to control whether to allow or deny requests to the AI ​​server 40, as described above, is implemented by a program described later.

[0033] Server 10 analyzes the information received from the guidance device 30 and performs processing to determine whether it is possible to generate a request to the AI ​​server 40 and whether the request device is permitted. Server 10 does not permit the generation of a request to the AI ​​server 40 if, for example, the person is not in a predetermined relative relationship with the guidance device 30 in real space.

[0034] "When the predetermined relative relationship is not in place in real space" means, for example, when the distance between the person and the guidance device 30 is not below the predetermined distance, that is, when the person is not approaching the guidance device 30. Also, when the person is not facing the guidance device 30, that is, even if the person is close to the guidance device 30, the person is not in a posture to make an inquiry to the guidance device 30. Furthermore, when the sound input to the guidance device 30 does not reach the predetermined volume, that is, even if the person is approaching the guidance device 30 and their face is turned towards the guidance device 30, the information input as an inquiry is unidentifiable.

[0035] [Common to System 1A and System 1B] In this specification, requests from server 10 to AI server 40 are sometimes referred to as "question processing," and responses from AI server 40 to server 10 are sometimes referred to as "answer processing."

[0036] The AI ​​server 40 may be charged based, for example, on the volume of requests sent from the server 10 and the corresponding responses. Requests and responses consist of string data, and the volume of requests and responses is determined by factors such as the number of characters in the string data. Factors that determine the volume of requests and responses include the length of usage time, such as when the data is read aloud, and the amount of data transferred when outputting images or videos. Therefore, regardless of the content of the inquiry information received by the server 10 from the user terminal 20 and the guidance device 30, a request is generated from the data input to the server 10, and when the generated request is sent from the server 10 to the AI ​​server 40, a charge is incurred.

[0037] Both System 1A and System 1B according to this embodiment have a function that, if they determine that the input from the user terminal 20 or guidance device 30 to the server 10 does not constitute an inquiry (question) to artificial intelligence, does not generate a request corresponding to that input, or if a request is generated, stores it without sending it to the AI ​​server 40. These functions prevent the sending of unnecessary requests to the AI ​​server 40 and prevent unnecessary charges.

[0038] [Server 10 Configuration] Figure 2 is a hardware configuration diagram of server 10. Server 10 is implemented using a general-purpose computer such as a workstation or personal computer. As shown in Figure 2, server 10 mainly comprises a processor 11, memory 12, storage 13, input / output interface 14, and communication interface 15. Each component of server 10 is connected to the communication bus 19.

[0039] The processor 11 performs the processing described later by executing a series of instructions contained in the server program 13P stored in memory 12 or storage 13. The processor 11 can be implemented as, for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), an MPU (Micro Processor Unit), an FPGA (Field-Programmable Gate Array), or other device.

[0040] Memory 12 temporarily holds the server program 13P and data. The server program 13P is loaded, for example, from storage 13. The data includes data input to the server 10 and data generated by the processor 11. For example, memory 12 can be implemented as RAM (Random Access Memory) or other volatile memory.

[0041] Storage 13 permanently holds the server program 13P and data. Storage 13 can be implemented as, for example, ROM (Read-Only Memory), a hard disk drive, flash memory, or other non-volatile storage device. Alternatively, storage 13 may be implemented as a removable storage device, such as a memory card. In yet another example, instead of being built into the server 10, storage 13 may be connected to the server 10 as an external storage device. With such a configuration, for example, in a scenario where multiple user terminals 20 are used, such as in an amusement facility, it becomes possible to update the server program 13P and data all at once.

[0042] The input / output interface 14 is an interface for connecting external devices such as monitors, input devices (e.g., keyboards, pointing devices), external storage devices, speakers, cameras, microphones, and sensors to the server 10. The processor 11 communicates with external devices through the input / output interface 14. The input / output interface 14 can be implemented using, for example, USB (Universal Serial Bus), DVI (Digital Visual Interface), HDMI (High-Definition Multimedia Interface, registered trademark), or other terminals.

[0043] The communication interface 15 communicates with other devices connected to the communication network 2 (e.g., the guidance device 30, the user terminal 20). The communication interface 15 can be implemented as a wired communication interface such as a LAN (Local Area Network), or a wireless communication interface such as Wi-Fi (Wireless Fidelity), Bluetooth (registered trademark), or NFC (Near Field Communication).

[0044] [Configuration of User Terminal 20] Figure 3 is a hardware configuration diagram of an HMD (Head Mounted Display) set, which is an example of a user terminal 20. Figure 4 is a hardware configuration diagram of a tablet terminal, which is another example of a user terminal 20. In addition to the HMD set shown in Figure 3 and the tablet terminal shown in Figure 4, the user terminal 20 can be a smartphone, feature phone, laptop computer, desktop computer, or a device with similar functions that can display the virtual space 90, which will be described later, to the user.

[0045] As shown in Figure 3, the user terminal 20, implemented as an HMD set, includes a computer 26 having a processor 21, memory 22, storage 23, input / output interface 24, and communication interface 25. Each component of the computer 26 is connected to a communication bus 29. The basic configuration of the processor 21, memory 22, storage 23, input / output interface 24, communication interface 25, and communication bus 29 is the same as that of the processor 11, memory 12, storage 13, input / output interface 14, communication interface 15, and communication bus 19 shown in Figure 2. The storage 23 also holds the terminal program 23P.

[0046] Furthermore, the user terminal 20, which is implemented as an HMD set, includes an HMD 50, a motion sensor 61, and an operating device 62 as external devices to the computer 26. The HMD 50, motion sensor 61, and operating device 62 are connected to the processor 21 via an input / output interface 24.

[0047] The HMD50 is worn on the user's head and provides the user with a virtual space. More specifically, the HMD50 may include either a so-called head-mounted display equipped with a monitor, or a head-mounted device capable of attaching a smartphone or other device with a monitor. The HMD50 mainly comprises a monitor 51 (display device), a gaze sensor 52, cameras 53 and 54, a microphone 55, and a speaker 56.

[0048] For example, the monitor 51 is implemented as an opaque display device. The monitor 51 is positioned on the main body of the HMD 50, for example, in front of the user's eyes. The opaque monitor 51 can be implemented as, for example, a liquid crystal monitor or an organic EL (Electro-Luminescence) monitor.

[0049] As another example, the monitor 51 is implemented as a transparent display device. In this case, the HMD 50 is not a closed type that covers the user's eyes, but an open type like glasses. The monitor 51 may include a configuration that simultaneously displays a portion of the image that constitutes the virtual space and the real space. For example, the transparent monitor 51 may display an image of the real space captured by a camera mounted on the HMD 50. As yet another example, the transparent monitor 51 may be configured to have adjustable transmittance. The transparent monitor 51 may then set the transmittance of a portion of the display area to be high so that the real space can be directly viewed.

[0050] Furthermore, the monitor 51 may employ the following configurations to allow the person wearing the HMD 50 to view the three-dimensional image. For example, the monitor 51 may include a sub-monitor for displaying the image for the right eye and a sub-monitor for displaying the image for the left eye. As another example, the monitor 51 may be configured to display the image for the right eye and the image for the left eye together. In this case, the monitor 51 includes a high-speed shutter. The high-speed shutter operates to alternately display the image for the right eye and the image for the left eye so that the image is perceived by only one eye at a time.

[0051] The gaze sensor 52 detects the direction in which the user's right and left eyes are looking. In other words, the gaze sensor 52 detects the user's gaze. The gaze sensor 52 is implemented, for example, by a sensor having an eye-tracking function. Preferably, the gaze sensor 52 includes a sensor for the right eye and a sensor for the left eye. The gaze sensor 52 detects the rotation angle of each eyeball by irradiating the user's right and left eyes with infrared light and receiving reflected light from the cornea and iris in response to the irradiated light. Then, the gaze sensor 52 identifies the user's gaze based on the detected rotation angles.

[0052] Camera 53 captures the upper part of the user's face (more specifically, the user's eyes, eyebrows, etc.) while the user is wearing the HMD 50. Camera 54 captures the lower part of the user's face (more specifically, the user's nose, mouth, etc.) while the user is wearing the HMD 50. For example, camera 53 may be mounted on the side of the HMD 50 housing that faces the user, and camera 54 may be mounted on the side opposite to the user. Alternatively, the HMD 50 may be equipped with a single camera that captures the user's entire face instead of the two cameras 53 and 54.

[0053] The microphone 55 converts the user's speech into an audio signal (electrical signal) and outputs it to the computer 26. The speaker 56 converts the audio signal output from the computer 26 back into speech and outputs it to the user. Note that the HMD 50 may include earphones instead of the speaker 56.

[0054] The motion sensor 61 has a position tracking function for detecting the movement of the HMD 50. For example, the motion sensor 61 may read multiple infrared rays emitted by the HMD 50 to detect the position and tilt of the HMD 50 in real space. As another example, the motion sensor 61 may be implemented as a camera. In this case, the motion sensor 61 analyzes the image information of the HMD 50 output from the camera to detect the position and tilt of the HMD 50. As yet another example, the motion sensor 61 may be implemented as an angular velocity sensor, a geomagnetic sensor, or an accelerometer.

[0055] The operating device 62 is connected to the computer 26 by wire or wireless connection. The operating device 62 accepts commands (operations) from the user to the computer 26. For example, the operating device 62 may be a so-called controller that is held and operated by a person wearing the HMD 50. As another example, the operating device 62 may be configured to be attachable to the body or part of the clothing of a person wearing the HMD 50 and detect the person's movement using a motion sensor. However, the specific examples of the operating device 62 are not limited to these, and may also be a keyboard, pointing device, touch panel, etc.

[0056] As shown in Figure 4, the user terminal 20, implemented as a tablet device, mainly comprises a processor 31, memory 22, storage 23, communication interface 25, monitor 51, cameras 53 and 54, microphone 55, speaker 56, motion sensor 61, and operating device 62. Each component of the tablet device is connected to the communication bus 29. The basic configuration of the processor 21, memory 22, storage 23, communication interface 25, monitor 51, cameras 53 and 54, microphone 55, speaker 56, motion sensor 61, and operating device 62 is the same as in the HMD set, so the configuration specific to the tablet device will be described below.

[0057] The monitor 51 is mounted on the surface of a flat casing. Camera 53 is a so-called in-camera mounted on the surface of the flat casing to capture images of the user's face as they view the monitor 51. Camera 54 is a so-called out-camera mounted on the back of the flat casing (the side opposite the monitor 51) to capture images of the surroundings. Motion sensor 61 detects the movement of the casing (for example, rotation around three mutually orthogonal axes). An operating device 62 suitable for a tablet terminal is, for example, a touch panel superimposed on the monitor 51 that accepts various touch operations by the user (for example, tap, slide, flick, pinch in, pinch out, etc.).

[0058] As shown in Figure 5, the guidance device 30, which is implemented as a digital signage device, mainly comprises a processor 31, memory 32, storage 33, communication interface 35, touch panel monitor 34, camera 36, ​​microphone 37, and speaker 38. Each component of the digital signage device is connected to the communication bus 39.

[0059] The basic configuration of the processor 31, memory 32, storage 33, communication interface 35, and communication bus 39 is the same as that of the processor 11, memory 12, storage 13, communication interface 15, and communication bus 19 shown in Figure 2. Furthermore, the storage 33 holds the guidance program 33P.

[0060] The information device 30, implemented as a digital signage system, has a unique configuration that includes a touch panel monitor 34, a camera 36, ​​a microphone 37, and a speaker 38. However, a standard monitor may be used instead of the touch panel monitor 34.

[0061] The touch panel monitor 34 is similar in configuration to that of the user terminal 20, which is implemented as a tablet device, in which an operating device 62 is superimposed on the monitor 51, and operates as a touch panel that accepts various touch operations by the user (e.g., tap, slide, flick, pinch in, pinch out, etc.). The touch panel monitor 34 can display an image equivalent to a keyboard so that the user can input questions as text, and the user can input the string of characters for the question by touching this keyboard. In addition, the touch panel monitor 34 displays images based on advertising information stored in the storage 33 of the guidance device 30.

[0062] Camera 36 is installed on the front of the housing of the touch panel monitor 34 and captures images of the user approaching the guidance device 30. When camera 36 captures an image of a user, if the area occupied by the individual user in the image occupies a certain proportion of the total area of ​​the image, the server 10 functions to identify that the user is approaching the guidance device 30 to ask a question.

[0063] Microphone 37 functions as an input interface for users of the guidance device 30 to input questions to the guidance device 30 by voice. When the user's voice input into microphone 37 is sent to server 10, it is converted into a string and used to generate a request.

[0064] Speaker 38 outputs voice modified from the answer string included in the response sent back by AI Server 40 after a request generated based on a question entered by the user into the guidance device 30 is sent to AI Server 40 via Server 10. Speaker 38 also outputs voice and music based on advertising information stored in the guidance device 30's storage 33.

[0065] [Overview of Virtual Space 90] Figure 6 is a conceptual diagram representing one aspect of the virtual space 90. Figure 7 is a diagram showing a YZ cross-section of the field of view 94 in the virtual space 90 as viewed from the X direction. Figure 8 is a diagram showing an XZ cross-section of the field of view 94 in the virtual space 90 as viewed from the Y direction.

[0066] As shown in Figure 6, the virtual space 90 has a spherical structure that covers the entire 360-degree direction from the center C. To avoid complicating the explanation, Figure 6 illustrates the upper half of the celestial sphere of the virtual space 90. Each mesh is defined in the virtual space 90. The position of each mesh is predetermined as a coordinate value in the XYZ coordinate system, which is the global coordinate system defined in the virtual space 90. Each partial image that constitutes a panoramic image 91 (still image, video, etc.) that can be displayed in the virtual space 90 is associated with the corresponding mesh in the virtual space 90.

[0067] For example, the coordinate system of virtual space 90 is defined as an XYZ coordinate system with the origin at center C. It is assumed that the XYZ coordinate system is parallel to the real coordinate system. In the XYZ coordinate system, the horizontal, vertical (up and down), and front-to-back directions are defined as the X-axis, Y-axis, and Z-axis, respectively. Therefore, the X-axis (horizontal) of the XYZ coordinate system is parallel to the x-axis of the real coordinate system, the Y-axis (vertical) of the XYZ coordinate system is parallel to the y-axis of the real coordinate system, and the Z-axis (front-to-back direction) of the XYZ coordinate system is parallel to the z-axis of the real coordinate system.

[0068] A virtual camera 92, associated with the user terminal 20, is placed in the virtual space 90. The position of the virtual camera 92 in the virtual space 90 corresponds to the viewpoint of a person (a person wearing the HMD set) in the virtual space 90. The orientation of the virtual camera 92 corresponds to the person's line of sight (reference line of sight 93) in the virtual space 90. The processor 21 then defines the field of view area 94 (the field of view of the virtual camera 92) in the virtual space 90 based on the position and orientation of the virtual camera 92.

[0069] As shown in Figure 7, the field of view region 94 includes region 95 in the YZ section. Region 95 is the range of the polar angle α centered on the reference line of sight 93 in the vertical section (YZ section) that includes the reference line of sight 93 in the virtual space 90. As shown in Figure 8, the field of view region 94 includes region 96 in the XZ section. Region 96 is the range of the azimuth angle β centered on the reference line of sight 93 in the horizontal section (XZ section) that includes the reference line of sight 93 in the virtual space 90.

[0070] The processor 21 generates (extracts) a portion of the panoramic image 91 deployed in the virtual space 90 that is included in the field of view 94, as a virtual space image 97 captured by the virtual camera 92. The processor 21 then displays the generated virtual space image 97 on the monitor 51. In other words, the field of view 94 corresponds to the field of view of a person wearing the HMD set within the virtual space 90. Furthermore, the field of view 94 moves in accordance with changes in the position and orientation of the virtual camera 92 within the virtual space 90, and the virtual space image 97 displayed on the monitor 51 is updated. In other words, the person's field of view moves.

[0071] For example, the processor 21 moves the virtual camera 92 within the virtual space 90 in conjunction with the user's operation received by the operating device 62. The processor 21 also changes the orientation of the virtual camera 92 (i.e., the reference line of sight 93) in conjunction with the movement of the user terminal 20 detected by the motion sensor 61 (for example, rotation around three mutually orthogonal axes). Furthermore, the processor 21 displays the virtual space image 97 captured by the virtual camera 92 after the change in position and orientation on the monitor 51.

[0072] [First Embodiment] [Functional block diagram of server 10 and user terminal 20] The operation of System 1A according to the first embodiment will be explained with reference to Figures 9 to 14. Figure 9 is a functional block diagram of System 1A. The functional block of Server 10 is such that the server program 13P loaded into the memory 12 of Server 10 causes Server 10 to function as an avatar generation means 110, a request feasibility determination means 120, a request control means 130, an AI dialogue means 140, an input format conversion means 160, an output format conversion means 170, and a request storage means 150.

[0073] Furthermore, the terminal program 23P loaded into the memory 22 of the user terminal 20 causes the user terminal 20 (computer 26) to function as a virtual space definition means 210, a placement means 220, a camera movement means 230, an image generation means 240, an image display means 250, a user input means 260, and a user output means 270.

[0074] The avatar generation means 110 generates a user avatar corresponding to the user terminal 20 and transmits (relays) the avatar data received from one of the user terminals 20 via the communication network 2 to the other user terminals 20. The avatar generation means 110 also generates a concierge avatar and transmits (relays) the avatar data to be placed on the information counter provided in the virtual space 90 to all user terminals 20A to 20C. As a result, all user terminals 20A to 20C hold the same avatar data. Consequently, the same avatar is placed in the same position in the virtual space 90 defined for each user terminal 20A to 20C. In other words, the avatars are synchronized across all user terminals 20A to 20C.

[0075] The request approval / denial means 120 determines whether to approve or deny the request based on whether the relative relationship between each user avatar and the concierge avatar obtained from the avatar generation means 110 satisfies predetermined conditions.

[0076] The request eligibility determination means 120 notifies the request control means 130 that it permits sending a request using input from the user terminal 20 to the AI ​​server 40 if the relative relationship between the user avatar and the concierge avatar satisfies predetermined conditions. On the other hand, if the request eligibility determination means 120 does not satisfy predetermined conditions, it notifies the request control means 130 that it does not permit (denies) sending a request using input from the user terminal 20 to the AI ​​server 40.

[0077] The request control means 130, based on a notification from the request approval / denial means 120, instructs the AI ​​dialogue means 140 to approve or deny the request to the AI ​​server 40. If the notification from the request control means 130 is "request approved," the AI ​​dialogue means 140 sends a request based on the input data to the AI ​​server 40. If the notification from the request control means 130 is "request denied," the AI ​​dialogue means 140 does not send a request based on the input data to the AI ​​server 40. The request control means 130 may also stop the operation of the guidance device 30 if the notification from the request approval / denial means 120 is "request denied." The operation of the guidance device 30 may be stopped by stopping the processor 31 or by stopping the power supply to the guidance device 30.

[0078] The input data to the AI ​​dialogue means 140 is based on "question information" input from the input format conversion means 160. The question information consists of character data. If the notification from the request control means 130 is "request permitted," the AI ​​dialogue means 140 generates a request based on the character data input from the input format conversion means 160 and sends the generated request to the AI ​​server 40. If the notification from the request control means 130 is "request not permitted," the AI ​​dialogue means 140 does not generate a request based on the character data input from the input format conversion means 160, or if it does generate a request, it does not send it to the AI ​​server 40. When the AI ​​dialogue means 140 receives character data from the input format conversion means 160 in the case of request not permitted, it generates a request and stores the generated request in the request storage means 150. Then, when the AI ​​dialogue means 140 changes from request not permitted to request permitted, it sends the request stored in the request storage means 150 to the AI ​​server 40.

[0079] If the question entered from the user input means 260 of the user terminal 20 is text data, the input format conversion means 160 inputs the entered text data to the AI ​​dialogue means 140. If the data entered from the user input means 260 of the user terminal 20 is voice data, the input format conversion means 160 converts the entered voice data into text data and inputs the converted text data to the AI ​​dialogue means 140.

[0080] The output format conversion means 170 converts the response received by the AI ​​dialogue means 140 from the message processing means 410 of the AI ​​server 40 into audio data and outputs it to the user output means 270. When audio data is output to the user output means 270, the speaker 56 outputs the audio of the response included in the response.

[0081] If the request is denied, the request storage means 150 stores the request generated by the AI ​​dialogue means 140 when input data is received from the input format conversion means 160 to the AI ​​dialogue means 140.

[0082] [Functional blocks of user terminal 20] The virtual space definition means 210 defines a virtual space 90 that corresponds to the real space. More specifically, the virtual space definition means 210 expands virtual space data representing the virtual space 90 into memory 22. The virtual space data includes, for example, a panoramic image 91, the shapes and positions of virtual objects such as buildings and plants placed within the virtual space 90, and the shapes and positions of information counters (information desks) placed within the virtual space 90. The virtual space data may be downloaded in advance from server 10 and stored in storage 23, or it may be downloaded from server 10 when defining the virtual space 90. The specific process for defining the virtual space 90 is already well known, so a detailed explanation is omitted.

[0083] The placement means 220 places characters in the virtual space 90 defined by the virtual space definition means 210. Here, we will explain the case where the terminal program 23P is executed on the user terminal 20A. The characters that the placement means 220 places in the virtual space 90 include the user avatar of user terminal 20A (first user terminal) (hereinafter referred to as "user avatar"), the avatars of users of other user terminals 20B and 20C (second user terminals) (hereinafter referred to as "other users' avatars"), and the concierge avatar.

[0084] First, the placement means 220 places the user's avatar in the virtual space 90 based on avatar data pre-stored in the storage 23. Then, the placement means 220 transmits the avatar data of the user's avatar to the server 10 via the communication network 2. Furthermore, when the placement means 220 receives a user operation to instruct the user's avatar's movements via the operation device 62, it operates the user's avatar in the virtual space 90 according to that operation. Finally, the placement means 220 updates the avatar data to indicate the state of the user's avatar after its operation and transmits the updated avatar data to the server 10 via the communication network 2.

[0085] Furthermore, the placement means 220 receives avatar data of other users' avatars from the server 10 via the communication network 2. Based on the avatar data received from the server 10, the placement means 220 places the other users' avatars in the virtual space 90. In addition, the placement means 220 repositions the other users' avatars each time their avatar data is updated, so that the other users' avatars operate within the virtual space 90.

[0086] Furthermore, the concierge avatar is assumed to always be positioned in a predetermined location within the virtual space 90. That is, the concierge avatar is positioned in advance in a predetermined location within the virtual space 90, similar to other objects placed in the virtual space 90 that do not move within the virtual space 90.

[0087] The camera movement means 230 moves the virtual camera 92 within the virtual space 90 based on the movement of the user terminal 20 detected by the motion sensor 61 or on the operation instructions of the avatar via the operation device 62. Moving the user terminal 20 and instructing the avatar's movements via the operation device 62 are examples of movement operations. If the relative position of the virtual camera 92 and the user's avatar does not change, the camera movement means 230 moves the virtual camera 92 in accordance with the user's avatar moved by the placement means 220. Alternatively, the camera movement means 230 may move the virtual camera 92 independently of the user's avatar according to instructions via the motion sensor 61 or the operation device 62.

[0088] The image generation means 240 generates a virtual space image 97 by capturing images of the virtual space 90 with the virtual camera 92. More specifically, the image generation means 240 extracts an image corresponding to the field of view 94 from the panoramic image 91 as the virtual space image 97 and stores the virtual space image data representing the virtual space image 97 in the memory 22. The image generation means 240 also controls how characters are included in the virtual space image 97.

[0089] As an example, the image generation means 240 may generate a virtual space image 97 by changing the display mode of the concierge avatar included in the field of view 94 (i.e., the field of view of the virtual camera 92) based on the relative distance from the user avatar. That is, when the user avatar moves to a position below a certain distance from the concierge avatar, the image generation means 240 may generate a virtual space image 97 that includes the concierge avatar and the input screen for questions to the concierge avatar. As another example, the image generation means 240 may output a message that has been pre-registered and associated with the concierge avatar based on a selection operation of the user avatar performed on the operating device 62. As yet another example, the image generation means 240 may display a question entered by the user avatar in the virtual space image 97 based on an operation performed on the operating device 62. As yet another example, the image generation means 240 may switch the display mode of the avatar based on an operation performed on the operating device 62. The image generation means 240 may generate a virtual space image 97 in such a manner that the user can identify the response notified from the server 10 to the user output means 270.

[0090] The image display means 250 displays the virtual space image 97 on the monitor 51. More specifically, the image display means 250 loads the virtual space image data generated by the image generation means 240 into the graphics memory of the monitor 51. Then, each time the character in the virtual space 90 is updated, the image generation means 240 generates a new virtual space image 97, and the image display means 250 displays the newly generated virtual space image 97 on the monitor 51. As a result, the user of the user terminal 20 can view the image showing the view of the character operating in the virtual space 90 through the monitor 51.

[0091] The user input means 260 transmits the input data to the input format conversion means 160 when a user character enters a question to the concierge avatar within the virtual space 90.

[0092] The user output means 270 displays the question asked to the concierge avatar within the virtual space 90, and if the output format of the response corresponding to the question is text data, it also displays the answer included in the response. If the output format of the response is audio data output from the output format conversion means 170, the user output means 270 outputs the audio data to the user as audio.

[0093] [Functional blocks of AI Server 40] The AI ​​server 40 is not part of the configuration included in System 1 according to this embodiment, but it has a message processing means 410 that receives requests from System 1, generates responses to those requests, and sends the generated responses to Server 10. The message processing means 410 includes a machine learning model that generates optimal answers to questions using natural language, and has the function of enabling mutual data communication with the AI ​​dialogue means 140 of Server 10 via an API. When a question is input to the artificial intelligence from the user terminal 20 or guidance device 30, and a request related to this question is sent to the message processing means 410, the message processing means 410 generates and outputs a response that includes an answer corresponding to this request.

[0094] Next, we will explain the operation of System 1A. System 1A assumes that the user avatar of user terminal 20A is in the virtual space 90. System 1A then provides the user of user terminal 20A with an information guidance service in the virtual space 90.

[0095] Figure 10 is a flowchart showing the processing of the server 10 according to the first embodiment. Figure 11 is a flowchart showing the processing of the user terminal 20A according to the first embodiment. Figure 12 is an image of an information counter placed in the virtual space according to the first embodiment. Figures 13 and 14 are image of a concierge avatar placed in the virtual space according to the first embodiment. The processing by the server 10 in Figure 10 and the processing by the user terminal 20A in Figure 11 are executed in parallel.

[0096] Server 10 repeatedly executes the process shown in Figure 10 at predetermined time intervals. First, Server 10 (request permission determination means 120) determines whether the input data from any of the user terminals 20A to 20C satisfies the request permission conditions (S1001).

[0097] If the server 10 (request permission determination means 120) determines in step S1001, which is repeatedly executed at predetermined time intervals, that the request permission conditions are not met (S1001: No), then it denies the request from the AI ​​dialogue means 140 to the AI ​​server 40 (message processing means 410) (S1002). After that, the process ends and is executed again at the next execution time.

[0098] In step S1001, for example, from the state shown in Figure 12, the user moves through the virtual space 90 and approaches the guidance counter object 910 within a certain distance, and when the server 10 (avatar generation means 110) detects that the user avatar has approached the concierge avatar, the server 10 (request permission determination means 120) determines that the request permission conditions are met (S1001: Yes).

[0099] In this case, the server 10 (avatar generation means 110) generates avatar data as shown in Figure 13(A) and sends it to the placement means 220 of the user terminal 20A. The avatar data sent here displays the concierge avatar 920, the dialogue display area 930, and the question input area 940 together in the user's field of view area 94, as illustrated in Figure 13(A).

[0100] As shown in Figure 13(A), when the user avatar is facing the concierge avatar 920, a question can be entered through the user avatar. Therefore, based on the conditions derived from the relative relationship between the user avatar and the concierge avatar 920, the server 10 (request permission determination means 120) determines that the request permission conditions are met (S1001: Yes).

[0101] Next, the server 10 (request control means 130) permits the request from the AI ​​dialogue means 140 to the AI ​​server 40 (message processing means 410) (S1003). After that, the server 10 waits to execute the processes from step S1006 onwards until it receives input from user terminals 20A to 20C (S1004: Yes) and until a certain period of time has elapsed without input (S1004: NO, S1005: NO). The processes from S1001 onwards are executed at arbitrary timings and frequencies based on the user avatars corresponding to each of the user terminals 20A to 20C.

[0102] When the server 10 (avatar generation means 110) receives input data from one of the user terminals 20A to 20C via the communication network 2 (S1004: Yes), it determines in step S1005 that a certain amount of time has elapsed until the request input is completed (S1006: No). For example, as shown in Figure 12, when the guidance counter object 910 is in the user's field of view in the virtual space 90, and then voice input is received from the microphone 55 and voice data is input to the server 10 (S1004: Yes), if a certain amount of time has elapsed until the request input is completed (S1006: No) (S1005: Yes), the server 10 (request control means 130) denies the request from the AI ​​dialogue means 140 to the AI ​​server 40 (message processing means 410) (S1012). In this case, the server 10 (AI dialogue means 140) generates a request based on the partially input voice data (S1013), but does not send it to the AI ​​server 40, and instead stores the generated request in the storage 13 using the request storage means 150 (S1014).

[0103] If it is determined in step S1006 that the input of the request has been completed (S1006: Yes), the server 10 (AI dialogue means 140) reads the request generated in the storage 13 (S1007: Yes) and sends it to the AI ​​server 40 (S1008). If there is no request generated in the storage 13 (S1007: No), the server 10 (AI dialogue means 140) generates a request using the string that has been entered in the question input area 940, as shown in Figure 13(A) (S1011), and sends the request to the AI ​​server 40 (message processing means 410) (S1008).

[0104] When a request is sent from server 10 (AI dialogue means 140) to AI server 40 (message processing means 410), the message processing means 410 generates a response to the request and sends it to server 10 (AI dialogue means 140) (S1008). As a result, as illustrated in Figure 13(B), the user terminal 20 (user output means 270) updates the display in the dialogue display area 930 according to the response received by server 10 (AI dialogue means 140) (S1009).

[0105] Furthermore, for example, even if the guidance counter object 910 and the concierge avatar 920 shown in Figure 14(A) are displayed and the request is permitted (S1005), if, as in the judgment process in step S1006, no valid question is entered and a message prompting the user to enter a question is displayed in the dialogue display area 930 for a predetermined period of time (S1006: Yes), the server 10 (request control means 130) controls the AI ​​dialogue means 140 to deny the request (S1012). Also, for example, if the request is denied while the concierge avatar 920 is displayed (S1006: Yes, S1012), the displayed input icon 950 may be grayed out, as shown in Figure 14(B), so that the user can recognize that the request has been denied.

[0106] Next, the processing of the user terminal 20 in system 1 according to the first embodiment will be explained using the flowchart in Figure 11. In the first embodiment, as shown in Figure 11, the user terminal 20A (virtual space definition means 210) defines the virtual space 90 by expanding the virtual space data stored in storage 23 into memory 12 (S1101). Next, the user terminal 20A (placement means 220 and camera movement means 230) places the user avatar and virtual camera 92 in the virtual space 90 defined in step S1101 and transmits the avatar data to the server 10 via the communication network 2 (S1102).

[0107] Next, the user terminal 20A (placement means 220) determines whether the user avatar is approaching the guidance counter object 910, which is placed at a predetermined location in the virtual space 90 (S1103). This determination is based on whether the virtual distance between the user avatar and the guidance counter object in the virtual space 90 is shorter than a predetermined distance. If the user avatar corresponding to the user terminal 20A moves in the virtual space and the distance to the guidance counter object 910 becomes shorter than the predetermined distance (S1103: Yes), then, for example, as shown in Figure 13(A), the guidance counter object 910 and the concierge avatar 920 are placed in the virtual space 90 (S1104).

[0108] In step S1103, if the distance to the guidance counter object 910 is not shorter than a predetermined distance (S1103: No), a virtual space 90 containing only the user avatar is generated (S1110). Then, the generated virtual space 90 is displayed on the user terminal 20A (S1111).

[0109] Next, the user terminal 20A (image generation means 240 and image display means 250) performs virtual space image display processing in order to reflect the guidance counter object 910 and concierge avatar 920 placed in the virtual space 90 onto the virtual space image 97 (S1105).

[0110] Next, if a question is entered on the user terminal 20A (S1106: Yes), the input data is sent to the server 10 (S1107). Subsequently, if a response is received from the server 10 (S1108: Yes), a virtual space image including the answer contained in the response is displayed, as illustrated in Figure 13(B) (S1109).

[0111] [Effects of the First Embodiment] According to the first embodiment, if voice input is received from the microphone 55 while the user avatar is not approaching the guidance counter object 910 within the virtual space 90, it is possible to control the system so that a request is not sent to the AI ​​server 40 based on that voice input. In other words, if the input can be determined not to be a question, it is possible to avoid sending an unnecessary request to the AI ​​server 40 and receiving a response.

[0112] Furthermore, according to the first embodiment, when the user avatar meets predetermined conditions with the concierge avatar 920, the sending of a request to the AI ​​server 40 is permitted, and otherwise the request is denied. This allows the AI ​​server 40 to be used to provide a service that responds to questions in a conversational format, and to control the system so that a request is not sent if there is input unrelated to the question.

[0113] According to the first embodiment, since AI servers 40 generally use a pay-as-you-go system based on the volume of requests and responses, System 1 can suppress unnecessary charges.

[0114] Furthermore, according to the first embodiment, the display to the user can be changed to indicate whether or not communication with the AI ​​server 40 is prohibited, as illustrated in Figure 14(B). This allows the user to recognize that when they attempt to ask a question, processing time is required to transition from a request-prohibited state to a request-allowed state in order to obtain an answer. Therefore, the user can be aware in advance that there will be a time lag between entering a question and receiving an answer, thereby improving user convenience.

[0115] [Second Embodiment] Next, the operation of system 1B according to the second embodiment will be described with reference to Figures 9, 10, and 15-19. Detailed explanations of commonalities with the first embodiment will be omitted, and the explanation will focus on processes specific to the second embodiment. It should be noted that the first embodiment and the second embodiment can be combined in whole or in part.

[0116] [Functional block diagram of server 10 and guidance device 30] Figure 15 is a functional block diagram of System 1B. The functional block of Server 10 has the same configuration as in the first embodiment, as shown in Figure 9, but the connection between the information input source and output destination is different. Therefore, a detailed explanation of Server 10 will be omitted, and the functional configuration of Server 10 related to System 1B will be explained while describing the functional block of the guidance device 30.

[0117] [Functional blocks of the guidance device 30] As shown in Figure 15, the guidance program 33P loaded into the memory 32 of the guidance device 30 causes the guidance device 30 to function as an image acquisition means 310, a guidance input means 320, a guidance output means 330, an image display means 340, and an operation control means 350.

[0118] The image acquisition means 310 acquires an image of a person using the guidance device 30 using the camera 36 provided by the guidance device 30 and transmits the image to the request eligibility determination means 120. The image transmitted by the image acquisition means 310 is image data including a person asking a question to the AI ​​server 40 via the guidance device 30. The request eligibility determination means 120 analyzes the input image data to determine whether a person is approaching the guidance device 30 and is in a state to input a question. For example, the request eligibility determination means 120 analyzes the input image data to determine whether a person is approaching the guidance device 30 within a predetermined distance, and if the person is approaching within the predetermined distance, it notifies the request control means 130 to authorize the operation of sending a request to the AI ​​server 40 using input from the guidance input means 320.

[0119] The guidance input means 320 has the function of transmitting the input string to the server 10 (input format conversion means 160) when a person touches the touch panel monitor 34 and inputs a question. In addition, when a person inputs a question by voice to the guidance device 30, the guidance input means 320 transmits the voice collected by the microphone 37 of the guidance device 30 to the input format conversion means 160. If the data format input to the input format conversion means 160 is voice data, the input format conversion means 160 converts the voice data into string data and outputs it.

[0120] Furthermore, the guidance input means 320 also outputs the audio collected by the microphone 37 to the request eligibility determination means 120. For example, the request eligibility determination means 120 determines whether the volume of the input audio exceeds a predetermined threshold. If it exceeds the predetermined volume, it notifies the request control means 130 to permit the operation of sending a request to the AI ​​server 40 using the input from the guidance input means 320. If it does not exceed the predetermined volume, it notifies the request control means 130 to disallow the operation of sending a request to the AI ​​server 40 using the input from the guidance input means 320.

[0121] The guidance output means 330 outputs the response received by the AI ​​dialogue means 140 to the touch panel monitor 34 of the guidance device 30. Furthermore, when the guidance output means 330 receives voice data from the voice conversion means 180, it outputs a response from the speaker 38 as voice based on the voice data.

[0122] The image display means 340 displays the data input by the guidance input means 320 and the data from the guidance output means 330 as images on the touch panel monitor 34 of the guidance device 30. The image display means 340 also displays a concierge avatar that accepts questions on the touch panel monitor 34. The image display means 340 changes the display mode of the concierge avatar according to the input state in the guidance input means 320 and the output state in the guidance output means 330. Specifically, if no data input is detected from the guidance input means 320 for a certain period of time, the image display means 340 displays the concierge avatar in a state where it is displaying a message prompting input. Also, if no data input is detected from the guidance input means 320 for a certain period of time, the image display means 340 displays the concierge avatar stopping its operation in a manner that allows the user to identify that the request has been denied.

[0123] The operation control means 350 controls the operation of the guidance device 30 based on notifications from the server 10 (request approval / rejection determination means 120). For example, if the request approval / rejection determination means 120 notifies the request control means 130 that a request to the AI ​​server 40 is not permitted based on input from the guidance input means 320 or the image acquisition means 310, the request control means 130 sends a notification to the operation control means 350 instructing it to stop operation. Upon receiving the notification instructing it to stop operation, the operation control means 350 stops the input of questions from the guidance device 30 to the server 10. Furthermore, upon receiving the notification instructing it to stop operation, the operation control means 350 puts the power supply of the guidance device 30 into suspend mode or cuts off the operating power, in any case, to prevent the input of questions from the guidance device 30 to the server 10. This prevents requests to the AI ​​server 40 from being generated by inadvertent input from the guidance device 30 to the server 10.

[0124] In the server 10 according to this embodiment, if the data input from the guidance input means 320 of the guidance device 30 is character data, the input format conversion means 160 inputs the input character data to the AI ​​dialogue means 140. If the data input from the guidance input means 320 of the guidance device 30 is voice data, the input format conversion means 160 converts the input voice data into character data and inputs the converted character data to the AI ​​dialogue means 140.

[0125] Furthermore, the output format conversion means 170 of the server 10 converts the response received by the AI ​​dialogue means 140 from the message processing means 410 of the AI ​​server 40 into audio data and outputs it to the guidance output means 330. When audio data is output to the guidance output means 330, the speaker 38 outputs the audio of the response included in the response.

[0126] The functional blocks of the AI ​​server 40 are the same as those in the first embodiment, so their explanation will be omitted.

[0127] Next, the operation of system 1B will be described. Server 10, as in the first embodiment, repeatedly executes the process shown in Figure 10 at predetermined time intervals. First, server 10 waits to execute the process from step S1004 onward, either by receiving input from the guidance device 30 (S1001) or until a certain period of time has elapsed without input (S1001: NO, S1012: NO). The input data from the guidance device 30 is transmitted irregularly. In other words, the process from step S1004 onward is executed at any timing and frequency.

[0128] The server 10 (request permission determination means 120) determines whether the input from the guidance device 30 satisfies the request permission conditions (S1004). The subsequent processing is the same as in the first embodiment, except that the information input source and output destination is the guidance device 30 instead of the user terminal 20. In the second embodiment, the functions of the server 10 may be implemented in the guidance device 30, and the server 10 may be omitted.

[0129] In the second embodiment, as illustrated in Figure 16, a guidance device 30 used as digital signage is assumed. The guidance device 30 is, for example, a device in which a touch panel monitor 34 is mounted on a vertically oriented housing. A camera 36, ​​a microphone 37, and a speaker 38 are mounted on the housing.

[0130] In the second embodiment, the guidance device 30 displays a concierge avatar 920 on a touch panel monitor 34 (S1601). For example, if a person at a predetermined distance from the touch panel monitor 34 inputs a question by voice to the concierge avatar 920 illustrated in Figure 16(A), and the system determines that a person is inputting a question to the touch panel monitor 34 (S1602: Yes), the guidance input means 320 transmits the voice collected by the microphone 37 to the server 10 (input format conversion means 160) (S1603).

[0131] In step 1602, the determination of whether a person is at a predetermined distance from the touch panel monitor 34 is made, for example, by checking whether the volume of the voice uttered by the person as input to the touch panel monitor 34 exceeds a predetermined threshold. In this case, if the volume of the voice picked up by the microphone 37 exceeds the predetermined volume, it is assumed that the person has asked a question towards the touch panel monitor 34, and the process proceeds to step 1603.

[0132] Furthermore, the determination in step 1603 can also be made by estimating the approach of a person using, for example, the camera 36 of the touch panel monitor 34, regardless of the volume of sound collected by the microphone 37. That is, the determination is made by detecting a person within the field of view captured by the camera 36 and determining whether the proportion of the area occupied by that person within the field of view exceeds a threshold. For example, if a person is approaching the touch panel monitor 34, the person within the field of view captured by the camera 36 will be captured larger, increasing the proportion of the area occupied within the field of view. In this case, assuming that the person has asked a question to the touch panel monitor 34, the system proceeds to step 1603. Alternatively, by using a camera 36 equipped with a distance detection function, the distance detected by the camera 36 can be used to determine, in the same manner as above, whether a person is approaching the touch panel monitor 34 and asking a question.

[0133] Furthermore, in step 1602, if the camera 36 captures a person and it is determined whether or not that person is making a specific facial expression, it may be determined that the person is about to input a question into the touch panel monitor 34.

[0134] Furthermore, in step 1602, if a person touches the touch panel monitor 34, it can be determined that the person is attempting to input a question into the touch panel monitor 34. In this case, the touch operation may be sent to the server 10 as input data.

[0135] If it is determined in step 1602 that "a situation where a question is being entered" then in step S1001, the server 10 determines that "input has been received". Consequently, a request based on the entered voice is then sent to the AI ​​server 40 (message processing means 410), the server 10 (AI dialogue means 140) receives the corresponding response, and the guidance output means 330 receives it via the output format conversion means 170 (S1604: Yes).

[0136] The guidance device 30 (guidance output means 330) outputs the answer using the speaker 38 (S1605). The guidance device 30 also outputs the answer as text using the image display means 340, as illustrated in Figure 18(A), and changes the display of the concierge avatar 920 (S1606).

[0137] Furthermore, in the second embodiment, if no input from the user is received after a certain period of time (S1602: No, S1607: Yes), the guidance device 30 changes the display of the concierge avatar 920 so that the user can recognize that the request from the server 10 to the AI ​​server 40 has been denied (S1608). For example, as illustrated in Figure 18(B), the display of the concierge avatar 920 is changed to a state where the user can recognize that it is in an input waiting state and that there has been no input for some time.

[0138] Furthermore, when the guidance device 30 is in an authorized state for requests to the AI ​​server 40, it displays a message that allows the user to identify that it is accepting questions, as illustrated in Figure 17(A). If no questions are entered for a certain period of time, the display changes to one that allows the user to recognize that the concierge avatar 920 is in a waiting state for questions, as shown in Figure 17(B).

[0139] [Effects of the second embodiment] According to the second embodiment, AI can be used to provide answers to facility guidance questions using digital signage (information device 30) installed in commercial facilities and the like. In this case, if the microphone 37 of the information device 30 picks up ambient sounds, it is controlled not to send a request to the AI ​​server unless the input corresponds to a question. This makes it possible to avoid sending unnecessary requests and receiving responses due to inputs that do not correspond to questions.

[0140] Furthermore, according to the second embodiment, when a person meets predetermined conditions with respect to the guidance device 30, the system can permit the sending of a request to the AI ​​server 40, and otherwise, it can deny the request. For example, the server 10 (request permission determination means 120) determines whether a person is facing the guidance device 30 based on the image captured by the camera 36. This allows the system to control the sending of a request when using the AI ​​server 40 to provide a service that responds to questions in an interactive format, provided that there is input unrelated to the question.

[0141] According to the second embodiment, since AI servers 40 generally use a pay-as-you-go system based on the volume of requests and responses, System 1 can suppress unnecessary charges.

[0142] Furthermore, according to the second embodiment, the display of the guidance device 30 can be easily recognized by a person viewing the display, by changing the display mode as illustrated in Figures 17 and 18, to indicate whether or not communication with the AI ​​server 40 is prohibited. For example, in the display mode shown in Figure 18(B), when a user tries to ask a question, the user can recognize in advance that it will take processing time to transition from a request-disallowed state to a request-allowed state in order to obtain an answer. Therefore, the user can recognize in advance that there will be a time lag between entering a question and receiving an answer, thereby improving user convenience.

[0143] Furthermore, the program according to the present invention is not limited to a single program, but may be a collection of multiple programs. Also, the program according to the present invention is not limited to being executed on a single device, but may be executed by multiple devices in a shared manner. Moreover, the division of roles between the server 10 and the user terminal 20 is not limited to the examples described above. That is, part of the processing of the server 10 may be executed by the user terminal 20, and part of the processing of the user terminal 20 may be executed by the server 10.

[0144] Furthermore, some or all of the means implemented by the program can also be implemented by hardware such as integrated circuits. Additionally, the program may be provided by being recorded on a non-transient recording medium readable by a computer. Recording mediums include, for example, hard disks, SD cards, DVDs, and servers on the internet.

[0145] [Note] The present invention is summarized below. [assignment] For example, the present invention aims to improve user convenience. [Solution] (1) On the computer, If an object associated with artificial intelligence and the user who uses the object meet certain conditions, A program that allows the user to make inquiries to the artificial intelligence through the object. (2) In the program described in (1) above, A program that allows the query if the object and the user are at a predetermined distance from each other. (3) In the program described in (2) above, A program that allows the user to make the query if the user is facing the object. (4) In any of the programs described in (1) to (3) above, A program that causes the computer to permit the query in response to the user's image input to the object. (5) In the programs described in (1) through (4) above, A program that causes the computer to query the artificial intelligence using information stored in the storage means when the state in which the object and the user do not satisfy a predetermined condition changes to a state in which they satisfy a predetermined condition. The solutions described in the above program may be applied as appropriate to other categories such as systems, methods, media, and devices. [Effects and Effects] According to the above solution (1), the convenience of users using artificial intelligence is improved by allowing an object in a virtual space or an object installed in the real world that has the function of making inquiries (questions) to the artificial intelligence, i.e., a concierge avatar in a virtual space, a digital information display device installed in the real world, and a user in a virtual space (user avatar) or a user in the real world (person) using it, to make inquiries to the artificial intelligence when predetermined conditions are met.

[0146] According to solution (2) above, in the case of an avatar placed in a virtual space, when the distance between the user avatar and the concierge avatar in the virtual space is shorter than the virtual distance set as a threshold, inquiries (questions) to the artificial intelligence are permitted. In the case of a guidance display device placed in the real world, when the distance between the person and the guidance display device is shorter than the distance set as a threshold, inquiries (questions) to the artificial intelligence are permitted. This improves the convenience for users utilizing the artificial intelligence.

[0147] According to the above solution (3), in the case of an avatar placed in a virtual space, when the user avatar is facing the concierge avatar, and the distance in the virtual space is shorter than the virtual distance, the AI ​​is allowed to make a request (question). In the case of a guidance display device placed in the real world, when the person is facing the guidance device, and the distance between the person and the guidance display device is shorter than a threshold, the AI ​​is allowed to make a request (question). This improves the convenience for users utilizing the AI.

[0148] According to the above solution (4), if the image of a person captured by the guidance display device meets predetermined conditions, permission is granted to make an inquiry (question) to the artificial intelligence. This improves the convenience for users utilizing the artificial intelligence.

[0149] According to the above solution (5), if an inquiry to the artificial intelligence is not permitted, the temporarily recorded information is retrieved when an inquiry to the artificial intelligence is permitted and used for the inquiry (question) to the artificial intelligence. This improves the convenience for users of the artificial intelligence. [Explanation of symbols]

[0150] 1, 1A, 1B…System, 2…Communication Network, 10…Server, 11, 21, 31…Processor, 12, 22, 21…Memory, 13, 23, 33…Storage, 13P…Server Program, 14, 24…Input / Output Interface, 15, 25, 35…Communication Interface, 19, 29, 39…Communication Bus, 20, 20A, 20B, 20C…User Terminal, 23P…Terminal Program, 30…Guidance Device, 33P…Guidance Program, 34…Touch Panel Monitor, 36, 53, 54…Camera, 37, 55…Microphone, 38, 56…Speaker, 40…AI Server, 50…HMD, 52…Gaze Sensor, 61…Motion Sensor, 62…Operating Device, 90…Virtual Space, 91…Panoramic Image, 92…Virtual Camera, 93…Reference Line of Sight, 94…Field of View, 9 5, 96… Area, 97… Virtual space image, 110… Avatar generation means, 120… Request feasibility determination means, 130… Request control means, 140… AI dialogue means, 150… Request storage means, 160… Input format conversion means, 180… Voice conversion means, 210… Virtual space definition means, 220… Placement means, 230… Camera movement means, 240… Image generation means, 250… Image display means, 260… User input means, 270… User output means, 310… Image acquisition means, 320… Guidance input means, 330… Guidance output means, 340… Image display means, 350… Operation control means, 410… Message processing means, 910… Guidance counter object, 920… Concierge avatar, 930… Dialogue display area, 940… Question input area, 950… Input icon

Claims

1. On the computer, If an object associated with artificial intelligence and the user who uses the object meet certain conditions, A program that allows the user to make inquiries to the artificial intelligence through the object.

2. In the program described in claim 1, A program that causes the computer to permit the query if the object and the user are at a predetermined distance from each other.

3. In the program described in claim 2, A program that causes the computer to allow the query when the user is facing the object.

4. In the program described in claim 1, A program that causes the computer to permit the query in response to the user's image input to the object.

5. In the program described in claim 1, A program that causes the computer to query the artificial intelligence using information stored in the storage means when the state in which the object and the user do not satisfy a predetermined condition changes to a state in which they satisfy a predetermined condition.