program
The program optimizes AI interaction by controlling queries through virtual or real-space objects based on conditions, enhancing user convenience and reducing unnecessary charges.
Patent Information
- Application Number
- JP2024193916
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-11-05
AI Technical Summary
Existing technologies lack a user-friendly mechanism for interacting with artificial intelligence via virtual objects, leading to potential inefficiencies and unnecessary charges.
A program that allows interaction with AI through virtual or real-space objects only when a predetermined condition is met, such as proximity or orientation, thereby controlling query processes and incurring a fee only when valid queries are made.
Enhances user convenience by preventing unnecessary AI queries and charges, optimizing interaction with AI systems.
Smart Images

Figure 0007752744000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a program. [Background technology]
[0002] BACKGROUND ART Conventionally, there has been known a technique for controlling information processing in response to an input via a virtual object (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-042593 Summary of the Invention [Problem to be solved by the invention]
[0004] The present invention aims to improve user convenience. [Means for solving the problem]
[0005] In order to solve the above problem, the program of the present invention is a program that causes a computer to permit a query process to be made to the artificial intelligence by a user via an object when the user and the object satisfy a predetermined condition, the program being configured such that the artificial intelligence and the computer are connected via the Internet, and a predetermined fee is incurred when the query process is executed, and the permitted query process is based on an input from the user to the object. Te-sei done Question and transmitting the inquiry information to the artificial intelligence. [Effects of the Invention]
[0006] According to the present invention, user convenience is improved. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a diagram illustrating an overview of a system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a hardware configuration diagram of a server. [Figure 3] FIG. 1 is a hardware configuration diagram of an HMD set, which is an example of a user terminal. [Figure 4] FIG. 10 is a hardware configuration diagram of a tablet terminal, which is another example of a user terminal. [Figure 5] FIG. 1 is a hardware configuration diagram of a digital signage device that is an example of a guidance device. [Figure 6] FIG. 1 is a diagram conceptually illustrating one aspect of a virtual space. [Figure 7] 10 is a diagram showing a YZ cross section of a field of view in a virtual space as viewed from the X direction. [Figure 8] 10 is a diagram showing an XZ cross section of a field of view in a virtual space as viewed from the Y direction. [Figure 9] FIG. 2 is a functional block diagram of a server and a user terminal according to the first embodiment. [Figure 10] 5 is a flowchart showing processing by a server according to the first embodiment. [Figure 11] 10 is a flowchart showing processing of a first user terminal according to the first embodiment. [Figure 12] FIG. 2 is an image diagram of a virtual space according to the first embodiment. [Figure 13] FIG. 2 is an image diagram of an information counter in a virtual space according to the first embodiment. [Figure 14] FIG. 2 is an image diagram of an information counter in a virtual space according to the first embodiment. [Figure 15] FIG. 10 is a functional block diagram of a server and a guidance device according to a second embodiment. [Figure 16] 10 is a flowchart showing the process of the guidance device according to the second embodiment. [Figure 17] FIG. 10 is an external view of a guide device according to a second embodiment. [Figure 18]FIG. 10 is a display image diagram of a guidance device according to a second embodiment. [Figure 19] FIG. 10 is a display image diagram of a guidance device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In this specification, the term "the present embodiment" refers to both the first and second embodiments described below. The embodiment of the present invention described below is an example of how the present invention is embodied, and does not limit the scope of the present invention to the scope of the described embodiment. Therefore, the present invention can be implemented by making various modifications to the embodiment. First, an overview of the system 1 according to this embodiment will be described.
[0009] [System 1 Overview] FIG. 1 is a diagram illustrating an overview of a system 1 according to an embodiment of the present invention. The system according to the present invention has a function for controlling permission to communicate with an external device with which the system has a communication connection. More specifically, when the relative relationship between an object included in the system according to the present invention and having a communication connection function with an external device and a user of the system satisfies a predetermined condition, the system controls so as to permit a communication connection with the external device via the object. In other words, when the relative relationship between the object and the external device does not satisfy a predetermined condition, the system controls so as not to permit a communication connection with the external device via the object.
[0010] In this embodiment, when the system 1 "allows" information communication with an external device with which it has a communication connection, it includes, for example, enabling a configuration that operates when a device constituting the system 1 transmits or receives information to or from an external device. It also includes enabling a processing function that a device constituting the system 1 has for transmitting or receiving information to or from an external device.
[0011] Fig. 1 shows an example of the configuration of a system 1 according to this embodiment. Fig. 1(A) illustrates the system configuration of a system 1A, which is a first example of the system 1. Fig. 1(B) illustrates the system configuration of a system 1B, which is a second example of the system 1.
[0012] As shown in FIG. 1(A), system 1A mainly includes a server 10 and user terminals 20A, 20B, and 20C (hereinafter, these may be collectively referred to as "user terminals 20"). Although three user terminals 20 are shown in FIG. 1A, examples of user terminals 20 included in system 1A are not limited to this. Server 10 and user terminals 20 are connected to each other so as to be able to communicate with each other via a communication network 2. Specific examples of communication network 2 are not particularly limited, and may include, for example, the Internet, a mobile communication system (e.g., 4G, 5G, etc.), a wireless network such as Wi-Fi (registered trademark), or a combination of these.
[0013] 1(B), the system 1B mainly includes a server 10 and a guiding device 30. The server 10 and the guiding device 30 are connected to each other so as to be able to communicate with each other via a communication network 2. A specific example of the communication network 2 is not particularly limited, but is the same as the first example.
[0014] In addition, both system 1A and system 1B are connected to an AI server 40 so that they can communicate with each other via a communication network 2. The AI server 40 is an example of an information processing device equipped with a trained model (artificial intelligence) capable of interacting in natural language.
[0015] [System 1A Overview] The system 1A provides a function for responding to inquiries from users of a plurality of user terminals 20A to 20C in a virtual space provided by the server 10. More specifically, a user character (hereinafter sometimes referred to as a "user avatar") associated with a user using the user terminal 20 is placed in the virtual space, and an image of the virtual space viewed from a preset viewpoint is displayed on the user terminal 20.
[0016] System 1A provides a function for making inquiries to AI server 40, allowing the user to obtain a predetermined answer by using an object placed in the virtual space. Note that in system 1A, an "object for making an inquiry to AI server 40" is, for example, a concierge avatar (concierge object) with which a user avatar (user object) placed in the virtual space generated by user terminal 20 interacts.
[0017] That is, in the system 1A, a concierge avatar, which is one of the objects with which a user avatar (user object) generated and deployed in the user terminal 20 interacts, is an example of an object associated with the artificial intelligence. The concierge avatar is an example of an object having a query function for the artificial intelligence. A query made by a user using the user terminal 20 to the artificial intelligence is generated as a request to the artificial intelligence by performing predetermined information processing on information input by the user to the server 10 via the concierge avatar. For example, when a user asks a question to the artificial intelligence, the user asks the concierge avatar via the user avatar a question (which may be non-text information such as voice, or text information consisting of a question), and the server 10 generates a request to query the artificial intelligence from the input question. The request generated by the server 10 is then notified to the AI server 40 as an API request. In response to the API request, the AI server 40 notifies the server 10 of an API response including an answer to the question. The server 10 converts the answer included in the API response into a form recognizable by the user corresponding to the user avatar via the concierge avatar and outputs the converted answer to the user terminal 20.
[0018] The user terminal 20 moves a user avatar in the virtual space in conjunction with a person's operation of the operating device. The movements of the avatar include, for example, moving within the virtual space, moving various parts of the body, changing posture, changing facial expressions, speaking, and moving objects placed in the virtual space.
[0019] Furthermore, the user terminal 20 transmits avatar data indicating changes in the avatar's movements and status to the server 10 via the communication network 2. The server 10 transmits the avatar data received from the user terminal 20 to other user terminals 20 via the communication network 2. The user terminal 20 then updates the movements and status of the corresponding avatar in the virtual space based on the avatar data received from the server 10.
[0020] Furthermore, the system 1A provides a virtual space that corresponds to real space. That is, the virtual space according to this embodiment is a space that imitates real space. However, the virtual space provided by the system 1 does not need to completely match real space, and may be appropriately deformed. Real space refers to, for example, a town, commercial facility, shop, theme park, event venue, park, station, etc. that actually exists.
[0021] In system 1A, when the relative relationship between a character (concierge avatar) placed on an object equivalent to an information counter placed in virtual space and a user avatar satisfies a predetermined condition, server 10 allows the generation of a request corresponding to input information (e.g., inquiry information including a question) or the transmission of the generated request to AI server 40.
[0022] In other words, if the user avatar does not satisfy certain conditions, the system 1A (server 10) does not allow the concierge avatar to generate a request corresponding to the input inquiry information, or does not allow the generated request to be sent to the AI server 40. Here, "not allowing the generation of a request" includes, for example, accepting the input of inquiry information via the user avatar, but not executing the process of generating a request based on that inquiry information.
[0023] In the system 1A, the function of controlling whether or not requests to the AI server 40 are permitted as described above is realized by a computer program (hereinafter simply referred to as the "program") described below.
[0024] The server 10 analyzes the information received from the user terminal 20 and executes a process to determine whether or not the situation allows the generation of a request to the AI server 40 and the request device. The server 10 does not allow the generation of a request to the AI server 40 if, for example, the information input source is the user terminal 20 and the user avatar is not in a predetermined relative relationship with the concierge avatar in the virtual space.
[0025] "When a predetermined relative relationship is not established within the virtual space" refers to, for example, when the distance between the user avatar and the concierge avatar within the virtual space is not equal to or less than a predetermined distance, i.e., when the user avatar is not approaching the concierge avatar. Also, "when a predetermined relative relationship is not established within the virtual space" refers to when the user avatar is not facing the concierge avatar within the virtual space, i.e., when the user avatar is approaching the concierge avatar but is not facing it. Also, "when a predetermined relative relationship is not established within the virtual space" refers to when the distance between the user avatar and the concierge avatar within the virtual space does not remain equal to or less than a predetermined distance for a certain period of time. Similarly, when the user avatar faces the concierge avatar, this facing state does not continue for a certain period of time.
[0026] As described above, a user avatar and a concierge avatar are placed in the virtual space according to this embodiment. The user avatar and the concierge avatar are characters, such as people, animals, and plants, that can be operated by a person using the user terminal 20 in the virtual space. The user avatar operates in the virtual space in conjunction with the user's operation via the user terminal 20. In contrast, the appearance of the concierge avatar is determined by the server 10. As an example, the server 10 may select a concierge avatar from among pre-registered concierge avatar candidates. Furthermore, the concierge avatar may operate in a manner determined by the server 10, or may be stopped.
[0027] [System 1B Overview] System 1B provides a function for making inquiries to AI server 40, allowing a person to obtain a predetermined answer by using an object placed in real space. In system 1B, the "object for making an inquiry to AI server 40" is, for example, a guidance device 30 that displays a concierge avatar (concierge object) corresponding to information input by a person who wishes to make an inquiry to the artificial intelligence. The guidance device 30 also outputs a reply to the inquiry from the artificial intelligence.
[0028] That is, in system 1B, the guidance device 30 displaying a concierge avatar is an example of an object associated with artificial intelligence. The guidance device 30 corresponds to an object having a function of querying the artificial intelligence. A query made by a person using the guidance device 30 to the artificial intelligence is generated as a request to the artificial intelligence by performing predetermined information processing on information input by the person to the server 10 via the concierge avatar. For example, when a person asks an artificial intelligence a question, the person asks the concierge avatar. The question may be non-textual information such as voice, or textual information consisting of a question sentence. The server 10 generates a request to query the artificial intelligence from the question input to the concierge avatar. The server 10 then notifies the AI server 40 of the generated request as an API request. In response to the API request, the AI server 40 notifies the server 10 of an API response including an answer to the question. The server 10 converts the answer included in the API response into a form recognizable by a person via the concierge avatar and outputs it via the concierge avatar.
[0029] The guidance device 30 is a device that detects the movement, voice, etc. of people present in the real space. The guidance device 30 also acquires information operated by people and their voice, and transmits it to the server 10 via the communication network 2. For example, the type of guidance device 30 connected to the system 1 is not limited to one type, and multiple types may be combined.
[0030] As an example, the guidance device 30 is a stationary display device installed in commercial facilities, shops, theme parks, event venues, parks, stations, etc., and a specific example is a digital signage that presents responses to user inquiries, such as the location of a store or its features.
[0031] The system 1B does not allow the guidance device 30 to generate a request corresponding to the input inquiry information or to send the generated request to the AI server 40 if the person does not meet certain conditions.
[0032] In the system 1B, the function of controlling whether or not requests to the AI server 40 are permitted as described above is realized by a program described later.
[0033] The server 10 analyzes the information received from the guidance device 30 and executes a process to determine whether or not a situation exists in which it is possible to allow the generation of a request to the AI server 40. The server 10 does not allow the generation of a request to the AI server 40 when, for example, a person is not in a predetermined relative relationship with the guidance device 30 in real space.
[0034] "When there is no predetermined relative relationship in real space" refers to, for example, when the distance between the person and the guiding device 30 is not equal to or less than a predetermined distance, i.e., when the person is not approaching the guiding device 30. Also, this refers to when the person is not facing the guiding device 30, i.e., when the distance between the person and the guiding device 30 is close but the person is not in a position to make an inquiry to the guiding device 30. Also, this refers to when the voice input to the guiding device 30 does not reach a predetermined volume, i.e., when the person is approaching the guiding device 30 and is facing the guiding device 30 but the information input as an inquiry is not identifiable.
[0035] [Common items for System 1A and System 1B] In this specification, a request from the server 10 to the AI server 40 may be referred to as "question processing," and a response from the AI server 40 to the server 10 may be referred to as "answer processing."
[0036] The AI server 40 may be charged, for example, based on the volume of requests sent from the server 10 and the corresponding responses. Requests and responses are composed of character string data, and the volume of the requests and responses is determined by factors such as the number of characters comprising the character string data. Factors that determine the volume of requests and responses include the number of characters, the length of use time for voice reading, and the volume of data transferred when outputting images or videos. Therefore, regardless of the content of the inquiry information received by the server 10 from the user terminal 20 and the guidance device 30, a request is generated from the data input to the server 10, and when the generated request is sent from the server 10 to the AI server 40, a charge is incurred.
[0037] Both system 1A and system 1B according to this embodiment have a function for preventing the server 10 from generating a request corresponding to the input, or storing the generated request without transmitting it to the AI server 40, if the server 10 determines that the input from the user terminal 20 or the guiding device 30 does not correspond to an inquiry (question) to an artificial intelligence. These functions can prevent unnecessary requests from being sent to the AI server 40, thereby preventing unnecessary charges.
[0038] [Server 10 configuration] 2 is a hardware configuration diagram of server 10. Server 10 is realized by, for example, a general-purpose computer such as a workstation or a personal computer. As shown in FIG. 2, server 10 mainly includes a processor 11, memory 12, storage 13, an input / output interface 14, and a communication interface 15. Each component of server 10 is connected to a communication bus 19.
[0039] The processor 11 performs the processes described below by executing a series of instructions included in a server program 13P stored in the memory 12 or the storage 13. The processor 11 is realized as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor unit (MPU), a field-programmable gate array (FPGA), or other devices.
[0040] The memory 12 temporarily stores a server program 13P and data. The server program 13P is loaded from, for example, the storage 13. The data includes data input to the server 10 and data generated by the processor 11. For example, the memory 12 is realized as a RAM (Random Access Memory) or other volatile memory.
[0041] The storage 13 permanently stores the server program 13P and data. The storage 13 is realized, for example, as a ROM (Read-Only Memory), a hard disk drive, a flash memory, or other non-volatile storage device. The storage 13 may also be realized as a removable storage device such as a memory card. As yet another example, the storage 13 may be connected to the server 10 as an external storage device instead of being built into the server 10. With this configuration, for example, in a situation where multiple user terminals 20 are used, such as an amusement facility, it becomes possible to collectively update the server program 13P and data.
[0042] The input / output interface 14 is an interface for connecting external devices such as a monitor, input device (e.g., keyboard, pointing device), external storage device, speaker, camera, microphone, sensor, etc. to the server 10. The processor 11 communicates with the external devices through the input / output interface 14. The input / output interface 14 is realized using, for example, a Universal Serial Bus (USB), a Digital Visual Interface (DVI), a High-Definition Multimedia Interface (HDMI, registered trademark), or other terminals.
[0043] The communication interface 15 communicates with other devices (e.g., the guide device 30 and the user terminal 20) connected to the communication network 2. The communication interface 15 is realized as, for example, a wired communication interface such as a LAN (Local Area Network), or a wireless communication interface such as Wi-Fi (Wireless Fidelity), Bluetooth (registered trademark), or NFC (Near Field Communication).
[0044] [Configuration of user terminal 20] Fig. 3 is a hardware configuration diagram of an HMD (Head Mounted Display) set, which is an example of the user terminal 20. Fig. 4 is a hardware configuration diagram of a tablet terminal, which is another example of the user terminal 20. The user terminal 20 is realized as, for example, the HMD set shown in Fig. 3, the tablet terminal shown in Fig. 4, a smartphone, a feature phone, a laptop computer, a desktop computer, or a computer having similar functions and capable of displaying a virtual space 90, which will be described later, to the user.
[0045] 3, a user terminal 20 realized as an HMD set includes a computer 26 having a processor 21, a memory 22, a storage 23, an input / output interface 24, and a communication interface 25. The components of the computer 26 are connected to a communication bus 29. The basic configuration of the processor 21, memory 22, storage 23, input / output interface 24, communication interface 25, and communication bus 29 is the same as that of the processor 11, memory 12, storage 13, input / output interface 14, communication interface 15, and communication bus 19 shown in FIG. 2. The storage 23 also holds a terminal program 23P.
[0046] The user terminal 20 realized as an HMD set also includes an HMD 50, a motion sensor 61, and an operation device 62 as external devices of the computer 26. The HMD 50, the motion sensor 61, and the operation device 62 are connected to the processor 21 via the input / output interface 24.
[0047] The HMD 50 is worn on the user's head and provides the user with a virtual space. More specifically, the HMD 50 may include both a so-called head-mounted display equipped with a monitor and a head-mounted device to which a smartphone or other terminal equipped with a monitor can be attached. The HMD 50 mainly includes a monitor 51 (display device), a gaze sensor 52, cameras 53 and 54, a microphone 55, and a speaker 56.
[0048] As an example, the monitor 51 is realized as a non-transmissive display device. The monitor 51 is disposed on the main body of the HMD 50 so as to be positioned in front of both eyes of the user. The non-transmissive monitor 51 is realized as, for example, a liquid crystal monitor or an organic EL (Electro Luminescence) monitor.
[0049] As another example, the monitor 51 is realized as a see-through display device. In this case, the HMD 50 is not a closed type that covers the user's eyes, but an open type such as glasses. The monitor 51 may include a configuration that simultaneously displays a part of an image that constitutes a virtual space and the real space. As one example, the see-through monitor 51 may display an image of the real space captured by a camera mounted on the HMD 50. As another example, the see-through monitor 51 may be configured to have adjustable transmittance. The see-through monitor 51 may also be configured to have a high transmittance in part of its display area so that the real space can be directly viewed.
[0050] Furthermore, the monitor 51 may employ the following configuration to allow the person wearing the HMD 50 to view a three-dimensional image. As one example, the monitor 51 may include a sub-monitor for displaying an image for the right eye and a sub-monitor for displaying an image for the left eye. As another example, the monitor 51 may be configured to display an image for the right eye and an image for the left eye as a single unit. In this case, the monitor 51 includes a high-speed shutter. The high-speed shutter operates to alternately display an image for the right eye and an image for the left eye so that the image is recognized by only one of the eyes.
[0051] The gaze sensor 52 detects the direction in which the user's right and left eyes are looking. In other words, the gaze sensor 52 detects the user's gaze. The gaze sensor 52 is realized, for example, by a sensor with an eye tracking function. The gaze sensor 52 preferably includes a sensor for the right eye and a sensor for the left eye. The gaze sensor 52 irradiates the user's right and left eyes with infrared light and detects the rotation angle of each eyeball by receiving light reflected from the cornea and iris of the irradiated light. The gaze sensor 52 then identifies the user's gaze based on the detected rotation angles.
[0052] The camera 53 captures an image of the upper part of the face of the user wearing the HMD 50 (more specifically, the user's eyes, eyebrows, etc.). The camera 54 captures an image of the lower part of the face of the user wearing the HMD 50 (more specifically, the user's nose, mouth, etc.). For example, the camera 53 is attached to the side of the housing of the HMD 50 that faces the user, and the camera 54 is attached to the side opposite the side facing the user. Note that the HMD 50 may be provided with one camera that captures an image of the entire face of the user, instead of the two cameras 53 and 54.
[0053] The microphone 55 converts the user's speech into an audio signal (electrical signal) and outputs it to the computer 26. The speaker 56 converts the audio signal output from the computer 26 into sound and outputs it to the user. Note that the HMD 50 may include earphones instead of the speaker 56.
[0054] The motion sensor 61 has a position tracking function for detecting the movement of the HMD 50. As one example, the motion sensor 61 may read a plurality of infrared rays emitted by the HMD 50 to detect the position and tilt of the HMD 50 in real space. As another example, the motion sensor 61 may be realized by a camera. In this case, the motion sensor 61 analyzes image information of the HMD 50 output from the camera to detect the position and tilt of the HMD 50. As yet another example, the motion sensor 61 may be realized by an angular velocity sensor, a geomagnetic sensor, or an acceleration sensor.
[0055] The operation device 62 is connected to the computer 26 via a wired or wireless connection. The operation device 62 accepts input (operation) of commands to the computer 26 by a user. As one example, the operation device 62 may be a so-called controller that is held and operated by the person wearing the HMD 50. As another example, the operation device 62 may be configured to be attachable to the body or part of the clothing of the person wearing the HMD 50, and may detect the movement of the person using a motion sensor. However, specific examples of the operation device 62 are not limited to these, and may also be a keyboard, a pointing device, a touch panel, etc.
[0056] 4, the user terminal 20 realized as a tablet terminal mainly includes a processor 31, a memory 22, a storage 23, a communication interface 25, a monitor 51, cameras 53 and 54, a microphone 55, a speaker 56, a motion sensor 61, and an operation device 62. The components of the tablet terminal are connected to a communication bus 29. The basic configuration of the processor 21, memory 22, storage 23, communication interface 25, monitor 51, cameras 53 and 54, a microphone 55, a speaker 56, a motion sensor 61, and an operation device 62 is the same as that of an HMD set, so the configuration specific to the tablet terminal will be described below.
[0057] The monitor 51 is provided on the surface of a flat housing. The camera 53 is a so-called in-camera that is attached to the surface of the flat housing and captures an image of the face of a user viewing the monitor 51. The camera 54 is a so-called out-camera that is attached to the back surface of the flat housing (the surface opposite the monitor 51) and captures an image of the surroundings. The motion sensor 61 detects the motion of the housing (for example, rotation around three mutually perpendicular axes). An example of an operation device 62 suitable for a tablet terminal is a touch panel that is superimposed on the monitor 51 and receives various touch operations by the user (for example, tapping, sliding, flicking, pinching in, pinching out, etc.).
[0058] 5, a guidance device 30 realized as a digital signage device mainly includes a processor 31, a memory 32, a storage 33, a communication interface 35, a touch panel monitor 34, a camera 36, a microphone 37, and a speaker 38. The components of the digital signage device are connected to a communication bus 39.
[0059] The basic configuration of the processor 31, memory 32, storage 33, communication interface 35, and communication bus 39 is the same as that of the processor 11, memory 12, storage 13, communication interface 15, and communication bus 19 shown in Fig. 2. The storage 33 also holds a guidance program 33P.
[0060] The guidance device 30 realized as a digital signage device has, as its unique components, a touch panel monitor 34, a camera 36, a microphone 37, and a speaker 38. Note that instead of the touch panel monitor 34, a normal monitor may also be provided.
[0061] The touch panel monitor 34 has a configuration similar to that of the monitor 51 included in the user terminal 20 realized as a tablet terminal, with the operation device 62 superimposed on it, and operates as a touch panel that accepts various touch operations by the user (for example, tapping, sliding, flicking, pinching in, pinching out, etc.). The touch panel monitor 34 can display an image equivalent to a keyboard so that the user can input a question by text, and the user can input a character string of the question by touching the keyboard. The touch panel monitor 34 also displays video based on advertising information stored in the storage 33 of the guiding device 30.
[0062] Camera 36 is installed on the front of the housing of touch panel monitor 34, and captures an image of a user approaching guide device 30. When camera 36 captures an image of a user, if the area occupied by a single user in the image occupies a certain percentage of the area of the entire image, server 10 determines that the user is approaching guide device 30 to ask a question.
[0063] The microphone 37 functions as an input interface when a user using the guidance device 30 inputs a question by voice to the guidance device 30. When the user's voice input to the microphone 37 is transmitted to the server 10, it is converted into a character string and used to generate a request.
[0064] The speaker 38 outputs a voice that is changed from the answer character string included in the response returned from the AI server 40 when a request generated based on a question input by a user to the guidance device 30 is sent to the AI server 40 via the server 10. The speaker 38 also outputs voice and music based on advertising information stored in the storage 33 of the guidance device 30.
[0065] [Virtual Space 90 Overview] Fig. 6 is a diagram conceptually illustrating one aspect of virtual space 90. Fig. 7 is a diagram illustrating a YZ cross section of viewing area 94 in virtual space 90 as viewed from the X direction. Fig. 8 is a diagram illustrating an XZ cross section of viewing area 94 in virtual space 90 as viewed from the Y direction.
[0066] As shown in Fig. 6, virtual space 90 has a spherical structure that covers the entire 360-degree area around center C. To avoid complicating the explanation, Fig. 6 illustrates the celestial sphere in the upper half of virtual space 90. Meshes are defined in virtual space 90. The position of each mesh is defined in advance as coordinate values in an XYZ coordinate system, which is a global coordinate system defined in virtual space 90. Each partial image that makes up panoramic image 91 (still image, video, etc.) that can be deployed in virtual space 90 is associated with a corresponding mesh in virtual space 90.
[0067] For example, the coordinate system of the virtual space 90 is defined as an XYZ coordinate system with the center C as the origin. It is assumed that the XYZ coordinate system is parallel to the real coordinate system. The horizontal direction, vertical direction (up-down direction), and front-to-back direction in the XYZ coordinate system are defined as the X-axis, Y-axis, and Z-axis, respectively. Therefore, the X-axis (horizontal direction) of the XYZ coordinate system is parallel to the x-axis of the real coordinate system, the Y-axis (vertical direction) of the XYZ coordinate system is parallel to the y-axis of the real coordinate system, and the Z-axis (front-to-back direction) of the XYZ coordinate system is parallel to the z-axis of the real coordinate system.
[0068] A virtual camera 92 associated with the user terminal 20 is placed in the virtual space 90. The position of the virtual camera 92 in the virtual space 90 corresponds to the viewpoint of a person (a person wearing an HMD set) in the virtual space 90. The orientation of the virtual camera 92 corresponds to the line of sight (reference line of sight 93) of the person in the virtual space 90. Then, the processor 21 defines a field of view 94 (the angle of view of the virtual camera 92) in the virtual space 90 based on the position and orientation of the virtual camera 92.
[0069] As shown in Fig. 7, the field of view 94 includes an area 95 in the YZ cross section. The area 95 is a range of polar angle α centered on the reference line of sight 93 in a vertical cross section (YZ cross section) including the reference line of sight 93 in the virtual space 90. As shown in Fig. 8, the field of view 94 includes an area 96 in the XZ cross section. The area 96 is a range of azimuth angle β centered on the reference line of sight 93 in a horizontal cross section (XZ cross section) including the reference line of sight 93 in the virtual space 90.
[0070] Processor 21 generates (extracts) a partial image included in field of view 94 of panoramic image 91 deployed in virtual space 90 as virtual space image 97 captured by virtual camera 92. Processor 21 then displays generated virtual space image 97 on monitor 51. That is, field of view 94 corresponds to the field of view of a person wearing an HMD set in virtual space 90. Furthermore, field of view 94 moves in virtual space 90 in accordance with changes in the position and orientation of virtual camera 92, and virtual space image 97 displayed on monitor 51 is updated. That is, the person's field of view moves.
[0071] For example, the processor 21 moves the virtual camera 92 within the virtual space 90 in conjunction with a human operation received by the operating device 62. The processor 21 also changes the orientation of the virtual camera 92 (i.e., the reference line of sight 93) in conjunction with the movement of the user terminal 20 detected by the motion sensor 61 (e.g., rotation around three mutually orthogonal axes). Furthermore, the processor 21 displays on the monitor 51 a virtual space image 97 captured by the virtual camera 92 after the position and orientation have been changed.
[0072] [First embodiment] [Functional block diagram of server 10 and user terminal 20] The operation of the system 1A according to the first embodiment will be described with reference to Figures 9 to 14. Figure 9 is a functional block diagram of the system 1A. The functional blocks of the server 10 are such that a server program 13P loaded into memory 12 of the server 10 causes the server 10 to function as avatar generation means 110, request acceptance determination means 120, request control means 130, AI dialogue means 140, input format conversion means 160, output format conversion means 170, and request storage means 150.
[0073] In addition, the terminal program 23P loaded into the memory 22 of the user terminal 20 causes the user terminal 20 (computer 26) to function as a virtual space definition means 210, a placement means 220, a camera movement means 230, an image generation means 240, an image display means 250, a user input means 260, and a user output means 270.
[0074] The avatar generation means 110 generates a user avatar corresponding to each user terminal 20, and transmits (relays) avatar data received from one of the user terminals 20 via the communication network 2 to the other user terminals 20. The avatar generation means 110 also generates a concierge avatar and transmits (relays) the avatar data to be placed at an information counter provided in the virtual space 90 to all user terminals 20A to 20C. This causes all user terminals 20A to 20C to hold the same avatar data. As a result, the same avatar is placed at the same position in the virtual space 90 defined by all user terminals 20A to 20C. That is, the avatars are synchronized in all user terminals 20A to 20C.
[0075] The request permission determining means 120 determines whether to permit or deny the request based on the determination condition that the relative relationship between each user avatar acquired from the avatar generating means 110 and the concierge avatar satisfies a predetermined condition.
[0076] If the relative relationship between the user avatar and the concierge avatar satisfies a predetermined condition, the request grant / deny determination means 120 notifies the request control means 130 that it permits the operation of sending a request using input from the user terminal 20 to the AI server 40. On the other hand, if the relative relationship between the user avatar and the concierge avatar does not satisfy a predetermined condition, the request grant / deny determination means 120 notifies the request control means 130 that it does not permit (disallow) the operation of sending a request using input from the user terminal 20 to the AI server 40.
[0077] The request control means 130 instructs the AI dialogue means 140 to permit or deny a request to the AI server 40 based on the notification from the request permission determination means 120. If the notification from the request control means 130 is "request permitted", the request based on the input data to the AI dialogue means 140 is sent to the AI server 40. If the notification from the request control means 130 is "request not permitted", the request based on the input data to the AI dialogue means 140 is not sent to the AI server 40. If the notification from the request permission determination means 120 is "request not permitted", the request control means 130 stops the operation of the guiding device 30. Stop The operation of the guidance device 30 may be stopped by stopping the processor 31, stopping the supply of power to the guidance device 30, or the like.
[0078] Input data to the AI dialogue means 140 is based on "question information" input from the input format conversion means 160. The question information is composed of text data. When the notification from the request control means 130 is "request permitted," the AI dialogue means 140 generates a request based on the text data input from the input format conversion means 160 and executes processing to send the generated request to the AI server 40. When the notification from the request control means 130 is "request not permitted," the AI dialogue means 140 executes processing to not generate a request based on the text data input from the input format conversion means 160, or, even if a request is generated, not to send it to the AI server 40. Note that when the AI dialogue means 140 receives text data from the input format conversion means 160 when the request is not permitted, it generates a request and stores the generated request in the request storage means 150. When the request is changed from not permitted to permitted, the AI dialogue means 140 executes processing to send the request stored in the request storage means 150 to the AI server 40.
[0079] If the question input from the user input means 260 of the user terminal 20 is character data, the input format conversion means 160 inputs the input character data to the AI dialogue means 140. If the data input from the user input means 260 of the user terminal 20 is voice data, the input format conversion means 160 converts the input voice data into character data and inputs the converted character data to the AI dialogue means 140.
[0080] The output format conversion means 170 converts the response received by the AI dialogue means 140 from the message processing means 410 of the AI server 40 into voice data and outputs it to the user output means 270. When the voice data is output to the user output means 270, the answer included in the response is output as voice from the speaker 56.
[0081] If the request is not permitted, the request storage means 150 stores the request generated by the AI dialogue means 140 when input data is received from the input format conversion means 160 to the AI dialogue means 140.
[0082] [Function block of user terminal 20] The virtual space definition means 210 defines a virtual space 90 corresponding to the real space. More specifically, the virtual space definition means 210 expands virtual space data indicating the virtual space 90 in the memory 22. The virtual space data indicates, for example, a panoramic image 91, the shapes and positions of virtual objects such as buildings and plants placed in the virtual space 90, and the shape and position of an information counter (information office) placed in the virtual space 90. The virtual space data may be downloaded from the server 10 in advance and stored in the storage 23, or may be downloaded from the server 10 when the virtual space 90 is defined. The specific process of defining the virtual space 90 is already known, so a detailed description thereof will be omitted.
[0083] The placement means 220 places characters in the virtual space 90 defined by the virtual space definition means 210. Here, a case will be described where the terminal program 23P is executed on the user terminal 20A. The characters placed in the virtual space 90 by the placement means 220 include an avatar of the user of the user terminal 20A (first user terminal) (hereinafter referred to as "user's avatar"), avatars of users of other user terminals 20B and 20C (second user terminals) (hereinafter referred to as "other users' avatars"), and a concierge avatar.
[0084] First, the placement means 220 places the principal avatar in virtual space 90 based on avatar data pre-stored in storage 23. Then, the placement means 220 transmits the avatar data of the principal avatar to server 10 via communication network 2. Furthermore, when the placement means 220 receives a user's operation instructing the principal avatar to move via operation device 62, the placement means 220 moves the principal avatar in virtual space 90 in accordance with the operation. Then, the placement means 220 updates the avatar data to indicate the state of the principal avatar after the operation, and transmits the updated avatar data to server 10 via communication network 2.
[0085] Furthermore, the placement means 220 receives avatar data of the other person's avatar from the server 10 via the communication network 2. Then, the placement means 220 places the other person's avatar in the virtual space 90 based on the avatar data received from the server 10. Furthermore, the placement means 220 rearranges the other person's avatar every time the avatar data is updated, thereby causing the other person's avatar to operate within the virtual space 90.
[0086] It is assumed that the concierge avatar is always placed at a predetermined position in the virtual space 90. That is, the concierge avatar is placed at a predetermined position in the virtual space 90 in advance, similar to the objects placed in the virtual space 90 that do not move within the virtual space 90.
[0087] The camera moving means 230 moves the virtual camera 92 within the virtual space 90 based on the movement of the user terminal 20 detected by the motion sensor 61 or an instruction to move the avatar via the operation device 62. Moving the user terminal 20 and instructing the avatar to move via the operation device 62 are examples of movement operations. In a setting in which the relative position between the virtual camera 92 and the user's avatar does not change, the camera moving means 230 moves the virtual camera 92 to follow the user's avatar moved by the placement means 220. Alternatively, the camera moving means 230 may move the virtual camera 92 independently of the user's avatar, in accordance with an instruction via the motion sensor 61 or the operation device 62.
[0088] The image generation means 240 generates a virtual space image 97 by capturing an image of the virtual space 90 with the virtual camera 92. More specifically, the image generation means 240 extracts an image corresponding to the field of view 94 from the panoramic image 91 as the virtual space image 97, and stores virtual space image data representing the virtual space image 97 in the memory 22. The image generation means 240 also controls how characters are included in the virtual space image 97.
[0089] As one example, the image generating means 240 may generate the virtual space image 97 by changing the display mode of a concierge avatar included in the field of view 94 (i.e., the angle of view of the virtual camera 92) based on the relative distance from the user avatar. That is, when the user avatar moves to a position within a certain distance from the concierge avatar, the virtual space image 97 may be generated so as to include the concierge avatar and an input screen for a question to the concierge avatar. As another example, the image generating means 240 may output a message pre-registered in association with the concierge avatar based on a selection operation of the user avatar performed on the operation device 62. As another example, the image generating means 240 may display a question input by the user avatar in the virtual space image 97 based on an operation performed on the operation device 62. As yet another example, the image generating means 240 may switch the display mode of the avatar based on an operation performed on the operation device 62. The image generating means 240 may generate the virtual space image 97 so as to display the response notified to the user output means 270 from the server 10 in a manner that is recognizable by the user.
[0090] The image display means 250 displays the virtual space image 97 on the monitor 51. More specifically, the image display means 250 expands the virtual space image data generated by the image generation means 240 in the graphic memory of the monitor 51. Then, every time a character in the virtual space 90 is updated, the image generation means 240 generates a new virtual space image 97, and the image display means 250 displays the newly generated virtual space image 97 on the monitor 51. This allows the user of the user terminal 20 to view, through the monitor 51, an image showing the field of view of the character operating in the virtual space 90.
[0091] When a user character inputs a question to a concierge avatar in the virtual space 90, the user input means 260 transmits the input data to the input format conversion means 160.
[0092] The user output means 270 displays the question posed to the concierge avatar and, if the output format of the response to the question is text data, also displays the answer included in the response within the virtual space 90. If the output format of the response is audio data output from the output format conversion means 170, the user output means 270 outputs the audio data to the user as audio.
[0093] [AI Server 40 Function Blocks] Although the AI server 40 is not included in the system 1 according to this embodiment, it has a message processing means 410 that receives requests from the system 1, generates responses to the requests, and transmits the generated responses to the server 10. The message processing means 410 includes a machine-learned model that generates optimal answers to questions using natural language, and has the function of realizing mutual data communication with the AI dialogue means 140 of the server 10 via an API. When a question for the artificial intelligence is input from the user terminal 20 or the guidance device 30 and a request related to this question is transmitted to the message processing means 410, the message processing means 410 generates and outputs a response including an answer corresponding to this request.
[0094] Next, the operation of system 1A will be described. System 1A assumes that the avatar of the user of user terminal 20A is in virtual space 90. System 1A then provides an information guidance service in virtual space 90 to the user of user terminal 20A.
[0095] FIG. 10 is a flowchart showing the processing of the server 10 according to the first embodiment. FIG. 11 is a flowchart showing the processing of the user terminal 20A according to the first embodiment. FIG. 12 is an image diagram of an information counter placed in a virtual space according to the first embodiment. FIGS. 13 and 14 are image diagrams of a concierge avatar placed in a virtual space according to the first embodiment. The processing of FIG. 10 by the server 10 and the processing of FIG. 11 by the user terminal 20A are executed in parallel.
[0096] The server 10 repeatedly executes the process shown in Fig. 10 at predetermined time intervals. First, the server 10 (request permission determination means 120) determines whether or not input data from any of the user terminals 20A to 20C satisfies the request permission condition (S1001).
[0097] If the server 10 (request permission determination means 120) determines in step S1001, which is repeatedly executed at predetermined time intervals, that the request permission conditions are not met (S1001: No), it disallows the request from the AI dialogue means 140 to the AI server 40 (message processing means 410) (S1002). After that, the process ends and is executed again at the next execution time.
[0098] In step S1001, for example, from the state shown in FIG. 12, when the user moves through the virtual space 90 and approaches within a certain distance of the guidance counter object 910, and the server 10 (avatar generation means 110) detects that the user avatar has approached the concierge avatar, the server 10 (request acceptance determination means 120) determines that the request acceptance condition is met (S1001: Yes).
[0099] In this case, the server 10 (avatar generation means 110) generates avatar data as shown in Fig. 13(A) and sends it to the placement means 220 of the user terminal 20A. The avatar data sent here is used to display a concierge avatar 920, a dialogue display area 930, and a question input area 940 together in the user's field of view 94, as exemplified in Fig. 13(A).
[0100] 13(A), when the user avatar is facing the concierge avatar 920, a question can be input via the user avatar. Based on the condition derived from the relative relationship between the user avatar and the concierge avatar 920, the server 10 (request acceptance determination means 120) determines that the request acceptance condition is satisfied (S1001: Yes).
[0101] Next, the server 10 (request control means 130) permits a request from the AI dialogue means 140 to the AI server 40 (message processing means 410) (S1003). After that, the server 10 waits to execute the processing from step S1006 onwards until it receives input from the user terminals 20A to 20C (S1004: Yes) and until a certain period of time has passed without any input (S1004: No, S1005: No). Note that the processing from S1001 onwards is executed at any timing and frequency based on the user avatar corresponding to each of the user terminals 20A to 20C.
[0102] When the server 10 (avatar generation means 110) receives input data from one of the user terminals 20A to 20C via the communication network 2 (S1004: Yes), it determines in step S1005 whether a certain time has passed until the request input is completed (S1006: No). For example, as shown in FIG. 12, after the guidance counter object 910 is in the user's field of view in the virtual space 90, if there is voice input from the microphone 55 and the voice data is input to the server 10 (S1004: Yes), and before the request input is completed (S1006: No), if a certain time has passed (S1005: Yes), the server 10 (request control means 130) disallows the request from the AI dialogue means 140 to the AI server 40 (message processing means 410) (S1012). In this case, the server 10 (AI dialogue means 140) generates a request based on the voice data that has been partially input (S1013), but does not send it to the AI server 40, and stores the generated request in storage 13 by the request storage means 150 (S1014).
[0103] If it is determined in step S1006 that the input of the request is complete (S1006: Yes), the server 10 (AI dialogue means 140) reads out the request generated in storage 13 (S1007: Yes) and sends it to the AI server 40 (S1008). If the server 10 (AI dialogue means 140) does not have a request generated in storage 13 (S1007: No), it generates a request using the character string that has been input in the question input area 940 (S1011), as shown in Figure 13(A), and sends the request to the AI server 40 (message processing means 410) (S1008).
[0104] When a request is sent from the server 10 (AI dialogue means 140) to the AI server 40 (message processing means 410), a response to the request is generated by the message processing means 410 and sent to the server 10 (AI dialogue means 140) (S1008). As a result, as shown in Figure 13(B), in response to the response received by the server 10 (AI dialogue means 140), the user terminal 20 (user output means 270) updates the display in the dialogue display area 930 (S1009).
[0105] 14(A) are displayed and the request is permitted (S1005), if a valid question is not entered and a message prompting the user to enter a question is displayed in the dialogue display area 930 and a predetermined time has elapsed (S1006: Yes), as in the determination process of step S1006, the server 10 (request control means 130) controls the AI dialogue means 140 to disallow the request (S1012). Also, for example, if the request is disallowed with the concierge avatar 920 displayed (S1006: Yes, S1012), the displayed input icon 950 may be grayed out as shown in FIG. 14(B) so that the user can recognize that the request is disallowed.
[0106] Next, processing of the user terminal 20 in the system 1 according to the first embodiment will be described with reference to the flowchart of Fig. 11. In the first embodiment, as shown in Fig. 11, the user terminal 20A (virtual space definition means 210) defines the virtual space 90 by expanding the virtual space data stored in the storage 23 into the memory 12 (S1101). Next, the user terminal 20A (placement means 220 and camera movement means 230) places a user avatar and a virtual camera 92 in the virtual space 90 defined in step S1101, and transmits the avatar data to the server 10 via the communication network 2 (S1102).
[0107] Next, the user terminal 20A (placement means 220) determines whether the user avatar is approaching the guidance counter object 910 placed at a predetermined position in the virtual space 90 (S1103). This determination is based on the condition of whether the virtual distance between the user avatar and the guidance counter object in the virtual space 90 is closer than a predetermined distance. If the user avatar corresponding to the user terminal 20A moves in the virtual space and the distance to the guidance counter object 910 is closer than the predetermined distance (S1103: Yes), the guide counter object 910 and the concierge avatar 920 are placed in the virtual space 90, for example, as shown in FIG. 13(A) (S1104).
[0108] In step S1103, if the distance to the guidance counter object 910 is not closer than the predetermined distance (S1103: No), a virtual space 90 including only the user avatar is generated (S1110), and the generated virtual space 90 is displayed on the user terminal 20A (S1111).
[0109] Next, the user terminal 20A (image generation means 240 and image display means 250) executes a display process for the virtual space image 97 to reflect the information counter object 910 and concierge avatar 920 placed in the virtual space 90 in the virtual space image 97 (S1105).
[0110] Next, if a question is input to the user terminal 20A (S1106: Yes), the input data is sent to the server 10 (S1107). Next, if a response is received from the server 10 (S1108: Yes), a virtual space image including the answer included in the response is displayed (S1109), as shown in FIG. 13(B).
[0111] [Effects of the first embodiment] According to the first embodiment, when voice input is received from the microphone 55 while the user avatar is not approaching the guidance counter object 910 inside the virtual space 90, control can be performed so that a request to the AI server 40 is not made based on that voice input. In other words, if the input is determined not to be a question, it is possible to avoid sending an unnecessary request to the AI server 40 and receiving a response.
[0112] Furthermore, according to the first embodiment, when a user avatar and a concierge avatar 920 meet certain conditions, the transmission of a request to the AI server 40 is permitted, and in other cases, the request is not permitted. This makes it possible to control the request so that it is not sent if an input unrelated to the question is received when providing a service that uses the AI server 40 to respond to a question in an interactive format.
[0113] According to the first embodiment, since the AI server 40 generally uses a pay-per-use system that corresponds to the amount of requests and responses, the system 1 can reduce unnecessary charges.
[0114] Furthermore, according to the first embodiment, the display mode can be changed to indicate to the user whether communication with the AI server 40 is in a disallowed state, as shown in Fig. 14(B). This allows the user to know that when they try to ask a question, it will take time to process the state transition from a request-disallowed state to a request-allowed state in order to receive the answer. This allows the user to know in advance that there will be a time lag between inputting a question and receiving the answer, improving user convenience.
[0115] [Second embodiment] Next, the operation of the system 1B according to the second embodiment will be described with reference to Figures 9, 10, and 15 to 19. Note that a detailed description of commonalities with the first embodiment will be omitted, and the description will focus on processing unique to the second embodiment. Note that the first embodiment and the second embodiment can be combined in part or in whole.
[0116] [Functional block diagram of server 10 and guidance device 30] Fig. 15 is a functional block diagram of system 1B. As shown in Fig. 9, the functional blocks of server 10 have the same configuration as in the first embodiment, but the connection between the information input source and output destination is different. Therefore, a detailed description of server 10 will be omitted, and the functional configuration of server 10 related to system 1B will be described while explaining the functional blocks of guiding device 30.
[0117] [Function block of the guidance device 30] As shown in Figure 15, the guidance program 33P loaded into the memory 32 of the guidance device 30 causes the guidance device 30 to function as an image acquisition means 310, a guidance input means 320, a guidance output means 330, an image display means 340, and an operation control means 350.
[0118] The image acquisition means 310 acquires an image of a person using the guidance device 30 using the camera 36 provided in the guidance device 30 and transmits the image to the request acceptance determination means 120. The image transmitted by the image acquisition means 310 is image data including a person asking a question to the AI server 40 via the guidance device 30. The request acceptance determination means 120 analyzes the input image data and determines whether or not the person is approaching the guidance device 30 and is ready to input a question. For example, the request acceptance determination means 120 analyzes the input image data and determines whether or not the person is approaching the guidance device 30 within a predetermined distance, and if the person is approaching within the predetermined distance, sends a notification to the request control means 130 permitting the operation of transmitting a request using input from the guidance input means 320 to the AI server 40.
[0119] The guidance input means 320 has a function of, when a person touches the touch panel monitor 34 to input a question, transmitting the input character string to the server 10 (input format conversion means 160). Furthermore, when a person inputs a question by voice toward the guidance device 30, the guidance input means 320 transmits the voice collected by the microphone 37 provided in the guidance device 30 to the input format conversion means 160. When the data format input to the input format conversion means 160 is voice data, the input format conversion means 160 converts the voice data into character string data and outputs it.
[0120] The guidance input means 320 also outputs the audio collected by the microphone 37 to the request acceptance determination means 120. For example, the request acceptance determination means 120 determines whether the volume of the input audio exceeds a predetermined threshold, and if the predetermined volume is exceeded, sends a notification to the request control means 130 permitting the operation of sending a request using the input from the guidance input means 320 to the AI server 40. If the predetermined volume is not exceeded, the request control means 130 sends a notification not permitting the operation of sending a request using the input from the guidance input means 320 to the AI server 40.
[0121] The guidance output means 330 outputs the response received by the AI dialogue means 140 to the touch panel monitor 34 of the guidance device 30. When the guidance output means 330 receives voice data from the voice conversion means 180, it outputs the response from the speaker 38 as a voice based on the voice data.
[0122] The image display means 340 displays the data input by the guidance input means 320 and the data from the guidance output means 330 as images on the touch panel monitor 34 of the guiding device 30. The image display means 340 also displays a concierge avatar accepting questions on the touch panel monitor 34. The image display means 340 changes the display mode of the concierge avatar depending on the input state of the guidance input means 320 and the output state of the guidance output means 330. Specifically, if no data input from the guidance input means 320 is detected for a certain period of time, the image display means 340 displays the concierge avatar issuing a message prompting input. If no data input from the guidance input means 320 is detected for a certain period of time, the image display means 340 stops the operation of the concierge avatar and displays in a manner that is recognizable to the user that the request has been denied.
[0123] The operation control means 350 controls the operation of the guiding device 30 based on a notification from the server 10 (request acceptance determination means 120). For example, if the request acceptance determination means 120 notifies the request control means 130 of a notification disapproving a request to the AI server 40 based on input from the guidance input means 320 or the image acquisition means 310, the request control means 130 sends a notification to the operation control means 350 instructing it to stop operation. Upon receiving the notification instructing it to stop operation, the operation control means 350 stops the guiding device 30 from inputting questions to the server 10. Furthermore, upon receiving the notification instructing it to stop operation, the operation control means 350 puts the guiding device 30 into suspend mode or cuts off the operating power, thereby preventing the guiding device 30 from inputting questions to the server 10. This prevents requests to the AI server 40 from being generated due to inadvertent input from the guiding device 30 to the server 10.
[0124] In the server 10 according to this embodiment, if the data input from the guidance input means 320 of the guidance device 30 is character data, the input format conversion means 160 inputs the input character data to the AI dialogue means 140. If the data input from the guidance input means 320 of the guidance device 30 is voice data, the input format conversion means 160 converts the input voice data into character data and inputs the converted character data to the AI dialogue means 140.
[0125] Furthermore, the output format conversion means 170 of the server 10 converts the response received by the AI dialogue means 140 from the message processing means 410 of the AI server 40 into voice data and outputs it to the guidance output means 330. When the voice data is output to the guidance output means 330, the answer included in the response is output as voice from the speaker 38.
[0126] The functional blocks of the AI server 40 are the same as those in the first embodiment, and therefore will not be described here.
[0127] Next, the operation of the system 1B will be described. As in the first embodiment, the server 10 repeatedly executes the process shown in FIG. 10 at predetermined time intervals. First, the server 10 waits to execute the process from step S1004 onwards until it receives input from the guiding device 30 (S1001) or until a certain period of time has passed without any input (S1001: NO, S1012: NO). Note that the input data from the guiding device 30 is transmitted irregularly. In other words, the process from step S1004 onwards is executed at any timing and frequency.
[0128] The server 10 (request permission determination means 120) determines whether the input from the guiding device 30 satisfies the request permission condition (S1004). The subsequent processing is similar to that of the first embodiment, except that the information input source and output destination are the guiding device 30, not the user terminal 20. Note that in the second embodiment, the functions of the server 10 may be implemented in the guiding device 30, and the server 10 may be omitted.
[0129] In the second embodiment, a guidance device 30 used as digital signage is assumed, as exemplified in Fig. 16. The guidance device 30 is, for example, a device in which a touch panel monitor 34 is attached to a housing that can be set upright. A camera 36, a microphone 37, and a speaker 38 are attached to the housing.
[0130] The guidance device 30 according to the second embodiment displays a concierge avatar 920 on the touch panel monitor 34 (S1601). For example, if a person at a predetermined distance from the touch panel monitor 34 inputs a question by voice to the concierge avatar 920 illustrated in Fig. 16(A), and it is determined that a person is inputting a question to the touch panel monitor 34 (S1602: Yes), the guidance input means 320 transmits the voice collected by the microphone 37 to the server 10 (input format conversion means 160) (S1603).
[0131] In step 1602, the determination of whether or not a person is at a predetermined distance from the touch panel monitor 34 is made, for example, by determining whether or not the volume of the voice uttered by the person as an input to the touch panel monitor 34 exceeds a predetermined threshold. In this case, if the volume of the voice collected by the microphone 37 exceeds the predetermined volume, it is assumed that the person has uttered a question toward the touch panel monitor 34, and the process proceeds to step 1603.
[0132] Furthermore, the determination in step 1603 can be made not based on the volume of sound collected by the microphone 37, but also by, for example, estimating the approach of a person using the camera 36 provided in the touch panel monitor 34. That is, a person present within the angle of view captured by the camera 36 is detected, and a determination is made based on whether the proportion of the area occupied by the person within the angle of view exceeds a threshold. For example, if a person is approaching the touch panel monitor 34, the person present within the angle of view captured by the camera 36 is captured larger, thereby increasing the proportion of the area occupied within the angle of view. In this case, it is assumed that the person has posed a question to the touch panel monitor 34, and the process proceeds to step 1603. Note that if the camera 36 is equipped with a distance detection function, the distance detected by the camera 36 may be used to determine, as described above, that a person is approaching the touch panel monitor 34 and is about to pose a question.
[0133] Also, in step 1602, if the camera 36 captures a person and determines whether the person is making a particular facial expression, it may be determined that the person is about to input a question onto the touch panel monitor 34.
[0134] Also, in step 1602, if a person touches the touch panel monitor 34, it can be determined that the person is trying to input a question into the touch panel monitor 34, and therefore the fact that a touch operation has occurred can be transmitted to the server 10 as input data.
[0135] If it is determined in step 1602 that "a situation exists in which a question may be input," the server 10 determines in step S1001 that "input has occurred." Therefore, a request based on the input voice is then sent to the AI server 40 (message processing means 410), and the corresponding response is received by the server 10 (AI dialogue means 140), and then by the guidance output means 330 via the output format conversion means 170 (S1604: Yes).
[0136] The guiding device 30 (guidance output means 330) outputs the answer using the speaker 38 (S1605). Furthermore, the guiding device 30 outputs the answer in text using the image display means 340, as shown in Fig. 18(A), and changes the display of the concierge avatar 920 (S1606).
[0137] Furthermore, if there is no input from the user for a certain period of time (S1602: No, S1607: Yes), the guiding device 30 according to the second embodiment changes the display of the concierge avatar 920 so that the user can recognize that the request from the server 10 to the AI server 40 has been denied (S1608). For example, as illustrated in FIG. 18(B), the display is changed to a state that allows the user to recognize that the concierge avatar 920 is waiting for input and that no input has been received for a while.
[0138] Furthermore, when a request to the AI server 40 is permitted, the guiding device 30 displays a message so that the user can recognize that the device is accepting questions, as shown in Fig. 17(A). If no question is input within a certain period of time, the display changes to a state in which the concierge avatar 920 is waiting to respond to a question, as shown in Fig. 17(B).
[0139] [Effects of the second embodiment] According to the second embodiment, it is possible to use AI to provide answers to facility guides and the like using digital signage (guidance device 30) installed in commercial facilities, etc. In this case, when the microphone 37 of the guidance device 30 collects surrounding sounds, if the input does not correspond to a question, the device is controlled not to send a request to the AI server. This makes it possible to avoid sending unnecessary requests and receiving unnecessary responses due to input that does not correspond to a question.
[0140] Furthermore, according to the second embodiment, when a person satisfies a predetermined condition for the guidance device 30, the transmission of a request to the AI server 40 can be permitted, and in other cases the request can be disallowed. For example, the server 10 (request acceptability determination means 120) determines whether a person is facing the guidance device 30 based on an image captured by the camera 36. This allows control so that when a service that responds to questions in an interactive format using the AI server 40 is provided, if an input unrelated to the question is received, the request cannot be sent.
[0141] According to the second embodiment, since the AI server 40 generally uses a pay-per-use system that corresponds to the amount of requests and responses, the system 1 can reduce unnecessary charges.
[0142] Furthermore, according to the second embodiment, the display mode of the guidance device 30 can be changed to indicate whether communication with the AI server 40 is in a prohibited state, as illustrated in FIGS. 17 and 18, allowing a person viewing the display of the guidance device 30 to easily recognize the situation. For example, in the display mode of FIG. 18(B), when a user attempts to ask a question, the user can recognize in advance that a processing time is required for the state transition from the request-prohibited state to the request-permitted state in order to obtain an answer. Therefore, the user can recognize in advance that there will be a time lag between inputting a question and receiving an answer, improving user convenience.
[0143] Furthermore, the program according to the present invention is not limited to a single program, but may be a collection of multiple programs. Furthermore, the program according to the present invention is not limited to one executed by a single device, but may be shared and executed by multiple devices. Furthermore, the division of roles between the server 10 and the user terminal 20 is not limited to the example described above. That is, part of the processing of the server 10 may be executed by the user terminal 20, or part of the processing of the user terminal 20 may be executed by the server 10.
[0144] Furthermore, some or all of the means implemented by the program can be implemented by hardware such as an integrated circuit. Furthermore, the program may be provided recorded on a non-transitory recording medium that can be read by a computer. Examples of recording media include hard disks, SD cards, DVDs, and servers on the Internet.
[0145] [Note] The following is a summary of some of the present invention. [assignment] For example, an object of the present invention is to improve user convenience. [Solution] (1) To the computer, If an object associated with the artificial intelligence and a user who uses the object meet certain conditions, A program that allows the user to make inquiries to the artificial intelligence via the object. (2) In the program described in (1) above, The program allows the inquiry if the object and the user are located at a predetermined distance from each other. (3) In the program described in (2) above, The program allows the query if the user is facing the object. (4) In the program according to any one of (1) to (3), A program that causes the computer to allow the inquiry depending on an image of the user input into the object. (5) In the program described in (1) to (4), A program that causes the computer to make an inquiry to the artificial intelligence using information stored in a storage means when the object and the user change from a state in which they do not satisfy a predetermined condition to a state in which they do satisfy the condition. The solution of the above program may be applied to other categories such as systems, methods, media, and devices, as appropriate. [Action and effect] According to the above solution (1), when an object in a virtual space or an object installed in real space that has the function of making inquiries (questions) to artificial intelligence, i.e., a concierge avatar in a virtual space or a digital guidance display device installed in real space, and a user (user avatar) in the virtual space or a user (person) in real space who uses this, meets certain conditions, the user is allowed to make inquiries to the artificial intelligence, thereby improving the convenience for users who use artificial intelligence.
[0146] According to the above solution (2), in the case of an avatar in which the object is placed in a virtual space, when the distance in the virtual space between the user avatar and the concierge avatar is closer than a virtual distance set as a threshold, an inquiry (question) to the artificial intelligence is permitted. Also, in the case of a guidance display device in which the object is placed in the real world, when the distance between a person and the guidance display device is closer than a distance set as a threshold, an inquiry (question) to the artificial intelligence is permitted. This improves the convenience of users who use the artificial intelligence.
[0147] According to the above solution (3), in the case where the object is an avatar placed in a virtual space, when the user avatar is facing the concierge avatar and the distance in the virtual space is closer than the virtual distance, an inquiry (question) to the artificial intelligence is permitted. Also, in the case where the object is a guidance display device placed in the real world, when a person is facing the guidance device and the distance between the person and the guidance display device is closer than a threshold, an inquiry (question) to the artificial intelligence is permitted. This improves the convenience of users who use the artificial intelligence.
[0148] According to the above solution (4), if the image of a person captured by the guidance display device satisfies a predetermined condition, the person is permitted to make an inquiry (question) to the AI, thereby improving the convenience of the user who uses the AI.
[0149] According to the above solution (5), when a query to the AI is not permitted, the temporarily recorded information is read out and used to query (ask) the AI when the query to the AI is permitted, thereby improving the convenience of the user who uses the AI. [Explanation of symbols]
[0150] 1, 1A, 1B... System, 2... Communication network, 10... Server, 11, 21, 31... Processor, 12, 22, 21... Memory, 13, 23, 33... Storage, 13P... Server program, 14, 24... Input / output interface, 15, 25, 35... Communication interface, 19, 29, 39... Communication bus, 20, 20A, 20B, 20C... User terminal, 23P... Terminal program, 30... Guidance device, 33P... Guidance program, 34... Touch panel monitor, 36, 53, 54... Camera, 37, 55... Microphone, 38, 56... Speaker, 40... AI server, 50... HMD, 52... Gaze sensor, 61... Motion sensor, 62... Operation device, 90... Virtual space, 91... Panoramic image, 92... Virtual camera, 93... Reference line of sight, 94... Field of view area, 9 5, 96...area, 97...virtual space image, 110...avatar generation means, 120...request acceptance determination means, 130...request control means, 140...AI dialogue means, 150...request storage means, 160...input format conversion means, 180...voice conversion means, 210...virtual space definition means, 220...placement means, 230...camera movement means, 240...image generation means, 250...image display means, 260...user input means, 270...user output means, 310...image acquisition means, 320...guidance input means, 330...guidance output means, 340...image display means, 350...action control means, 410...message processing means, 910...guidance counter object, 920...concierge avatar, 930...interaction display area, 940...question input area, 950...input icon
Claims
1. On the computer, If an object associated with the artificial intelligence and a user who uses the object meet certain conditions, A program that allows the user to allow the artificial intelligence to process an inquiry via the object, The artificial intelligence and the computer are connected via the Internet, a predetermined fee is incurred when the inquiry process is executed, The permitted query processing includes a process of transmitting query information generated based on an input from the user for the object to the artificial intelligence. program.
2. 2. The program according to claim 1, A program that causes the computer to permit the inquiry if the object and the user are within a predetermined distance.
3. 3. The program according to claim 2, A program that causes the computer to allow the inquiry if the user is facing the object.
4. 2. The program according to claim 1, A program that causes the computer to allow the inquiry depending on an image of the user input into the object.
5. 2. The program according to claim 1, A program that causes the computer to make an inquiry to the artificial intelligence using information stored in a storage means when the object and the user change from a state in which they do not satisfy a predetermined condition to a state in which they do satisfy the condition.
Citation Information
Patent Citations
Robot, control method of the same, and program
JP2019219509A
Program, information processing device, and method
JP2020042593A
Dialogue system, dialogue control method, and program
JP2024112283A
JPP7511068B
System and method for cross-platform sharing of virtual assistants
US20190332400A1