Information processing apparatus

The information processing apparatus addresses the challenge of remembering tags for virtual objects in AR technology by displaying corresponding tag names for identified virtual objects, enhancing user recall and identification efficiency.

JP7693824B2Active Publication Date: 2025-06-17NTT DOCOMO INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023557896
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-04
Filing Date
2022-09-30
Publication Date
2025-06-17
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

In AR technology, as the number of pairs of virtual objects and tags increases, it becomes difficult for users to grasp the correspondence between each virtual object and each tag, leading to difficulties in remembering the tags themselves, which hinders the identification of virtual objects.

Method used

An information processing apparatus that displays virtual objects in a virtual space on a head-worn display device and includes a unit to identify virtual objects based on user instructions, with a name identification unit that retrieves and displays the corresponding tag names for identified virtual objects.

Benefits of technology

The apparatus enables users to easily remember and identify virtual objects by visually displaying the corresponding tag names, thereby simplifying the process of recalling tags for virtual objects in a virtual space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007693824000001
    Figure 0007693824000001
  • Figure 0007693824000002
    Figure 0007693824000002
  • Figure 0007693824000003
    Figure 0007693824000003
Patent Text Reader

Abstract

This information processing device is provided with: a display control unit which, on a display device mounted on the user's head, displays multiple virtual objects arranged in a virtual space; a virtual object specifying unit which specifies a first virtual object from multiple virtual objects on the basis of instruction information generated in response to a user operation; and a name specifying part which, in the case that a name corresponding to the first virtual object specified by the virtual object specifying unit is stored in a storage device, specifies said corresponding name as a first name. The display control unit displays the first name on the display device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus.

Background Art

[0002] In AR (Augmented Reality) technology, the real environment perceived by a user is extended by a computer. By using this technology, for example, it becomes possible to superimpose and display a virtual space on the real space visually recognized by the user through AR glasses worn on the head.

[0003] In AR technology, a tag may be associated with a virtual object arranged in a virtual space. For example, Patent Document 1 discloses a technique for performing face authentication processing on a captured image of a person captured by a head-mounted display and assigning, as tag information, the name of the person, which is the face authentication result, to the face image extracted from the captured image.

[0004] On the other hand, regarding the technology for photographing the real space, for example, Patent Document 2 discloses a technique for setting a tag indicating an object included in a photographed image. In the technique according to Patent Document 2, it is possible to search for a photographed image in which the object indicated by the tag is photographed by using the tag.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0006] In the AR technology, when setting tags for virtual objects and using the tags to identify the virtual objects corresponding to the tags, as the number of pairs of virtual objects and tags increases, it becomes difficult for the user to grasp the correspondence between each virtual object and each tag, and it also becomes difficult for the user to remember the tags themselves. As a result, when the user cannot recall the tag corresponding to the virtual object they want to identify, they cannot easily identify the virtual object.

[0007] Therefore, an object of the present invention is to provide an information processing apparatus that can easily remind a user of tags as names for identifying virtual objects arranged in a virtual space.

Means for Solving the Problem

[0008] An information processing apparatus according to a preferred aspect of the present invention includes a display control unit that displays a plurality of virtual objects arranged in a virtual space on a display device worn on the user's head, and a virtual object identification unit that identifies a first virtual object among the plurality of virtual objects based on instruction information generated according to the user's operation. When the name corresponding to the first virtual object identified by the virtual object identification unit is stored in the storage device, it includes a name identification unit that identifies the corresponding name as the first name. The display control unit is an information processing apparatus that displays the first name on the display device.

Effects of the Invention

[0009] According to the present invention, it becomes possible to easily remind a user of names for identifying virtual objects arranged in a virtual space.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 10C

Figure 11A

Figure 11B

Figure 12A

Figure 12B

Figure 13A

Figure 13B

Figure 14A

Figure 14B

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26A

Figure 26B

Figure 27

Figure 28

Figure 29

Mode for Carrying Out the Invention

[0011] 1: First Embodiment Hereinafter, with reference to FIGS. 1 to 19, the configuration of the information processing system 1 including the information processing apparatus according to the first embodiment of the present invention will be described.

[0012] 1.1: Configuration of the First Embodiment 1.1.1: Overall Configuration FIG. 1 is a diagram showing the overall configuration of an information processing system 1 according to a first embodiment of the present invention. The information processing system 1 is a system that provides a virtual space to a user U1 wearing an AR glass 20 described later by AR technology.

[0013] The information processing system 1 includes a terminal device 10, an AR glass 20, and a server 30. The terminal device 10 and the AR glass 20 are connected to each other so as to be communicable. Also, the terminal device 10 and the server 30 are connected to each other so as to be communicable via a communication network NET. In FIG. 1, as a set of the terminal device 10 and the AR glass 20, a total of three sets, namely, a set of the terminal device 10-1 and the AR glass 20-1, a set of the terminal device 10-2 and the AR glass 20-2, and a set of the terminal device 10-3 and the AR glass 20-3 are described. However, the number of the sets is merely an example, and the information processing system 1 can include an arbitrary number of sets of the terminal device 10 and the AR glass 20. Also, the terminal device 10 is an example of an information processing device.

[0014] The terminal device 10 is a device for causing a virtual object arranged in a virtual space to be displayed on an AR glass 20 worn on the head by a user. The virtual space is, for example, a spherical space. Also, the virtual object is, for example, a virtual object showing data such as a still image, a moving image, a 3DCG model, an HTML file, and a text file, and a virtual object showing an application. Here, examples of the text file include a memo, a source code, a diary, and a recipe. Also, examples of the application include a browser, an application for using SNS, and an application for generating a document file. Note that the terminal device 10 is preferably a portable terminal device such as a smartphone and a tablet.

[0015] The AR glasses 20 are a see-through wearable display worn on the user's head. The AR glasses 20 display virtual objects on display panels provided on each of the lenses for both eyes under the control of the terminal device 10. Note that the AR glasses 20 are an example of a display device.

[0016] The server 30 provides various data and cloud services to the terminal device 10 via the communication network NET.

[0017] 1.1.2: Configuration of AR Glasses FIG. 2 is a perspective view showing the appearance of the AR glasses 20. As shown in FIG. 2, the AR glasses 20 have temples 91 and 92, a bridge 93, body parts 94 and 95, and lenses 41L and 41R, similar to ordinary glasses. An imaging device 27 is provided on the bridge 93. The imaging device 27 images the outside world and outputs imaging data indicating the captured image. In addition, a sound collection device 24 for collecting sound is provided on each of the temples 91 and 92. The sound collection device 24 outputs voice data indicating the collected voice. Note that the position of the sound collection device 24 is not limited to the temples 91 and 92, and may be, for example, either the bridge 93 or the body parts 94 and 95.

[0018] Each of the lenses 41L and 41R includes a half mirror. In the body part 94, a liquid crystal panel or an organic EL panel for the left eye (hereinafter collectively referred to as a display panel) and an optical member for guiding the light emitted from the display panel for the left eye to the lens 41L are provided. The half mirror provided on the lens 41L transmits the light of the outside world and guides it to the left eye, and reflects the light guided by the optical member to make it incident on the left eye. In the body part 95, a display panel for the right eye and an optical member for guiding the light emitted from the display panel for the right eye to the lens 41R are provided. The half mirror provided on the lens 41R transmits the light of the outside world and guides it to the right eye, and reflects the light guided by the optical member to make it incident on the right eye.

[0019] The display 29 described below includes a lens 41L, a display panel for the left eye, and an optical member for the left eye, as well as a lens 41R, a display panel for the right eye, and an optical member for the right eye.

[0020] In the above configuration, the user can observe the image on the display panel in a state where it is superimposed on the state of the outside world. Also, in the AR glasses 20, among the binocular images with parallax, the left-eye image is displayed on the display panel for the left eye, and the right-eye image is displayed on the display panel for the right eye. By utilizing binocular parallax, the user U1 can perceive the displayed image as if it has depth and a three-dimensional effect.

[0021] Figs. 3 and 4 are schematic diagrams of the virtual space VS provided to the user U1 by using the AR glasses 20. As shown in Fig. 3, various virtual objects VO1 to VO5 indicating various contents such as a browser, cloud services, images, and videos are arranged in the virtual space VS. The user U1 can experience the virtual space VS as a private space in the public space by wearing the AR glasses 20 on which the virtual objects VO1 to VO5 arranged in the virtual space VS are displayed and moving around in the public space. Subsequently, the user U1 can act in the public space while receiving the benefits brought by the virtual objects VO1 to VO5 arranged in the virtual space VS.

[0022] Also, as shown in Fig. 4, it is also possible to share the virtual space VS among a plurality of users U1 to U3. By sharing the virtual space VS among a plurality of users U1 to U3, the plurality of users U1 to U3 can share one or more virtual objects VO and communicate with each other via the shared virtual object VO.

[0023] FIG. 5 is a block diagram showing a configuration example of the AR glasses 20. The AR glasses 20 include a processing device 21, a storage device 22, a line-of-sight detection device 23, a sound collection device 24, a GPS device 25, a motion detection device 26, an imaging device 27, a communication device 28, and a display 29. Each element of the AR glasses 20 is interconnected by one or more buses for communicating information.

[0024] The processing device 21 is a processor that controls the entire AR glasses 20 and is configured using, for example, one or more chips. The processing device 21 is configured using, for example, a central processing unit (CPU) including an interface with peripheral devices, an arithmetic unit, registers, etc. Note that part or all of the functions of the processing device 21 may be realized by hardware such as a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), an FPGA (Field Programmable Gate Array), etc. The processing device 21 executes various processes in parallel or sequentially.

[0025] The storage device 22 is a recording medium that can be read from and written to by the processing device 21 and stores a plurality of programs including a control program PR1 executed by the processing device 21.

[0026] The line-of-sight detection device 23 detects the line of sight of the user U1 and outputs line-of-sight data indicating the direction of the user U1's line of sight to the processing device 21 described later based on the detection result. The line-of-sight detection by the line-of-sight detection device 23 may use any method. For example, the line-of-sight data may be detected based on the position of the eye head and the position of the iris.

[0027] The sound collection device 24 collects sound and outputs sound data based on the collected sound to the processing device 21 described later.

[0028] The GPS device 25 receives radio waves from a plurality of satellites and generates position data from the received radio waves. The position data indicates the position of the AR glasses 20. The position data may be in any format as long as the position can be specified. The position data indicates, for example, the latitude and longitude of the AR glasses 20. As an example, the position data is obtained from the GPS device 25. However, the AR glasses 20 may obtain the position data by any method. The acquired position data is output to the processing device 21.

[0029] The motion detection device 26 detects the motion of the AR glasses 20 and outputs motion data to the processing device 21. Examples of the motion detection device 26 include inertial sensors such as an acceleration sensor that detects acceleration and a gyro sensor that detects angular acceleration. The acceleration sensor detects the acceleration of the orthogonal X-axis, Y-axis, and Z-axis. The gyro sensor detects the angular acceleration with the X-axis, Y-axis, and Z-axis as the central axes of rotation. The motion detection device 26 can generate attitude information indicating the attitude of the AR glasses 20 based on the output information of the gyro sensor. The motion data includes acceleration data indicating the acceleration of each of the three axes and angular acceleration data indicating the angular acceleration of each of the three axes.

[0030] The imaging device 27 outputs imaging data obtained by imaging the external world. The imaging device 27 includes, for example, a lens, an image sensor, an amplifier, and an AD converter. The light collected through the lens is converted into an imaging signal, which is an analog signal, by the image sensor. The amplifier amplifies the imaging signal and outputs it to the AD converter. The AD converter converts the amplified imaging signal, which is an analog signal, into imaging data, which is a digital signal. The converted imaging data is output to the processing device 21. The imaging data output to the processing device 21 is output to the terminal device 10 via the communication device 28. The terminal device 10 recognizes various gestures of the user U1 based on the imaging data and controls the terminal device 10 according to the recognized gestures. That is, the imaging device 27 functions as an input device for inputting the instructions of the user U1, such as a pointing device and a touch panel.

[0031] The communication device 28 is hardware as a transceiver device for communicating with other devices. The communication device 28 is also called, for example, a network device, a network controller, a network card, a communication module, etc. The communication device 28 may include a connector for a wired connection and an interface circuit corresponding to the connector. Further, the communication device 28 may include a wireless communication interface. Examples of the connector for a wired connection and the interface circuit include products compliant with wired LAN, IEEE1394, or USB. Examples of the wireless communication interface include products compliant with wireless LAN and Bluetooth (registered trademark), etc.

[0032] The display 29 is a device for displaying images. The display 29 displays various images under the control of the processing device 21. As described above, the display 29 includes the lens 41L, the display panel for the left eye, and the optical member for the left eye, and the lens 41R, the display panel for the right eye, and the optical member for the right eye. As the display panel, various display panels such as a liquid crystal display panel and an organic EL display panel are preferably used.

[0033] The processing device 21 functions as an acquisition unit 211 and a display control unit 212, for example, by reading and executing the control program PR1 from the storage device 22.

[0034] The acquisition unit 211 acquires a control signal from the terminal device 10. More specifically, the acquisition unit 211 acquires a control signal for controlling the display on the AR glasses 20, which is generated by a display control unit 114 (to be described later) provided in the terminal device 10.

[0035] In addition, the acquisition unit 211 acquires the gaze data input from the gaze detection device 23, the voice data input from the sound collection device 24, the position data input from the GPS device 25, the motion data input from the motion detection device 26, and the imaging data input from the imaging device 27. Then, the acquisition unit 211 outputs the acquired gaze data, voice data, position data, motion data, and imaging data to the communication device 28.

[0036] The display control unit 212 controls the display on the display 29 based on the control signal from the terminal device 10 acquired by the acquisition unit 211.

[0037] 1.1.3: Configuration of the Terminal Device FIG. 6 is a block diagram showing a configuration example of the terminal device 10. The terminal device 10 includes a processing device 11, a storage device 12, a communication device 13, a display 14, an input device 15, and an inertial sensor 16. Each element of the terminal device 10 is interconnected by one or a plurality of buses for communicating information. Note that the term "device" in this specification may be read as other terms such as a circuit, a device, or a unit.

[0038] The processing device 11 is a processor that controls the entire terminal device 10 and is configured using, for example, one or a plurality of chips. The processing device 11 is configured using, for example, a central processing unit (CPU) including an interface with peripheral devices, an arithmetic unit, and registers. Note that some or all of the functions of the processing device 11 may be realized by hardware such as a DSP, an ASIC, a PLD, or an FPGA. The processing device 11 executes various processes in parallel or sequentially.

[0039] The storage device 12 is a recording medium that can be read from and written to by the processing device 11 and stores a plurality of programs including a control program PR2 executed by the processing device 11, first information IF1, and second information IF2.

[0040] FIG. 7 is a diagram showing an example of the first information IF1. In the example shown in FIG. 7, the first information IF1 is tabular information. The first information IF1 associates identification information that uniquely identifies the virtual object VO, a tag TG corresponding to the virtual object VO, position information indicating the position of the virtual object VO in the celestial sphere-shaped virtual space VS, and image information of each virtual object VO. The identification information is hereinafter referred to as ID. The position information is three-dimensional coordinates in the virtual space VS. Note that when each virtual object VO does not have a tag TG corresponding to itself, the column of the tag TG is blank. That is, in the first information IF1, some or all of the plurality of virtual objects VO arranged in the virtual space VS are associated with the plurality of tags TG one-to-one. Note that the tag TG is an example of the name of the virtual object VO. Also, the plurality of virtual objects VO arranged in the virtual space VS are a plurality of virtual objects VO that can be visually recognized when the user U1 changes the posture. The virtual object VO can be arranged in the virtual space VS based on an instruction from the user U1. Also, the virtual object VO may be arranged in the virtual space VS when a predetermined condition is satisfied even without an instruction from the user U1. For example, among J virtual objects VO, K virtual objects VO are arranged in the virtual space VS. And among the K virtual objects VO arranged in the virtual space VS, tags TG are assigned to L virtual objects VO. J, K, and L are integers, and J≧K≧L. The first information IF1 is information regarding the K virtual objects VO. Note that the first information IF1 may be acquired from the server 30 via the communication device 13.

[0041] FIG. 8 is a diagram showing an example of the second information IF2. In the example shown in FIG. 8, the second information IF2 is tabular information. The second information IF2 associates the ID of the virtual object VO arranged in the virtual space VS with the attribute of the virtual object VO. The second information IF2 is information regarding K virtual objects VO. Note that the second information IF2 may be acquired from the server 30 via the communication device 13. Here, the “attribute” is an item obtained by classifying the content of each virtual object VO by features or properties. The virtual object VO includes data such as a still image, a moving image, a 3DCG model, an HTML file, and a text file, as well as an application. In this example, an attribute is assigned to each of the plurality of virtual objects VO arranged in the virtual space VS, but an attribute may be assigned to a part of the plurality of virtual objects VO. An attribute that is common to two or more virtual objects VO may be assigned. On the other hand, there is no tag TG that is common to two or more virtual objects VO. Therefore, by using attributes, virtual objects VO having the same features or properties can be grouped. In the example shown in FIG. 8, for example, by using the attribute “animal”, the virtual objects VO with IDs = 6, ID = 7, and ID = 8 are grouped. On the other hand, by using the tag TG, one virtual object VO can be specified. In the example shown in FIG. 7, for example, by using the tag TG “#HORSE”, the virtual object VO with ID = 8 is specified.

[0042] Returning to the description of FIG. 6, the communication device 13 is hardware as a transmission / reception device for communicating with other devices. The communication device 13 is also called, for example, a network device, a network controller, a network card, or a communication module. The communication device 13 may include a connector for a wired connection and an interface circuit corresponding to the connector. Further, the communication device 13 may include a wireless communication interface. Examples of the connector for a wired connection and the interface circuit include products compliant with wired LAN, IEEE 1394, or USB. Examples of the wireless communication interface include products compliant with wireless LAN and Bluetooth (registered trademark).

[0043] The display 14 is a device for displaying images and character information. The display 14 displays various images under the control of the processing device 11. For example, various display panels such as a liquid crystal display panel and an organic EL (Electro Luminescence) display panel are preferably used as the display 14.

[0044] The input device 15 receives operations from the user U1 wearing the AR glasses 20 on the head. For example, the input device 15 includes a pointing device such as a keyboard, a touch pad, a touch panel, or a mouse. Here, when the input device 15 includes a touch panel, it may also serve as the display 14.

[0045] The inertial sensor 16 is a sensor that detects inertial forces. The inertial sensor 16 includes, for example, one or more sensors among an acceleration sensor, an angular velocity sensor, and a gyro sensor. The processing device 11 detects the posture of the terminal device 10 based on the output information of the inertial sensor 16. Further, the processing device 11 receives the selection of the virtual object VO, the input of characters, and the input of instructions in the celestial sphere-shaped virtual space VS based on the posture of the terminal device 10. For example, when the user U1 operates the input device 15 with the central axis of the terminal device 10 directed toward a predetermined area in the virtual space VS, the virtual object VO arranged in the predetermined area is selected. The operation of the user U1 on the input device 15 is, for example, a double tap. In this way, the user U1 can select the virtual object VO without looking at the input device 15 of the terminal device 10 by operating the terminal device 10.

[0046] When a virtual keyboard is arranged in the virtual space VS, characters are input by operating the input device 15 with the central axis of the terminal device 10 directed toward the key that the user U1 wants to input. Also, for example, a predetermined instruction is input by moving the terminal device 10 left and right while the user U1 presses the input device 15.

[0047] As a result, the terminal device 10 functions as a portable controller that controls the virtual space VS.

[0048] The processing device 11 functions as an acquisition unit 111, an operation recognition unit 112, a specification unit 113, a display control unit 114, a determination unit 115, and an audio recognition unit 116 by reading and executing the control program PR2 from the storage device 12.

[0049] The acquisition unit 111 acquires instruction information corresponding to the operation of the user U1 wearing the AR glasses 20 on the head. The instruction information is information that designates a specific virtual object VO. Here, the operation of user U1 is, for example, the input by user U1 to the terminal device 10 using the input device 15. More specifically, the operation of user U1 may be the pressing of a specific part as the input device 15 provided in the terminal device 10. Alternatively, the operation of user U1 may be an operation using the terminal device 10 as a portable controller.

[0050] Or, the operation of user U1 may be the visual observation of user U1 with respect to the AR glasses 20. When the operation of user U1 is visual observation, the instruction information is the viewpoint of user U1 in the AR glasses 20. In this case, the instruction information is transmitted from the AR glasses 20 to the terminal device 10.

[0051] Alternatively, the operation of user U1 may be a gesture of user U1. As will be described later, the motion recognition unit 112 recognizes various gestures of user U1. The acquisition unit 111 may acquire instruction information corresponding to various gestures of user U1.

[0052] Also, the acquisition unit 111 acquires the first information IF1 and the second information IF2 from the server 30 by using the communication device 13. The acquisition unit 111 stores the acquired first information IF1 and second information IF2 in the storage device 12. Further, the acquisition unit 111 acquires the first information IF1 and the second information IF2 from the storage device 12.

[0053] The motion recognition unit 112 recognizes various gestures of user U1 based on the imaging data obtained from the AR glasses 20. More specifically, as described above, the imaging device 27 provided in the AR glasses 20 outputs the imaging data obtained by imaging the external world. When a part of the body of user U1 wearing the AR glasses 20 on the head is included in the imaging data, the motion recognition unit 112 recognizes various gestures of user U1 based on the imaging data acquired from the AR glasses 20.

[0054] As shown in FIG. 9, the specific part 113 includes a virtual object specifying part 113-1 and a name specifying part 113-2. Based on the instruction information acquired by the acquisition part 111, the virtual object specifying part 113-1 specifies one virtual object VO among a plurality of virtual objects VO arranged in the virtual space VS. Hereinafter, the one virtual object VO specified by the virtual object specifying part 113-1 is referred to as the first virtual object VO. Further, when the first information IF1 acquired by the acquisition part 111 includes one tag TG corresponding to the specified virtual object VO, the name specifying part 113-2 specifies the one tag TG. One tag TG is an example of the first name.

[0055] The name specifying part 113-2 specifies the tag TG corresponding to the first virtual object VO by referring to the first information IF1. In other words, when the tag TG corresponding to the first virtual object VO is stored in the storage device 12, the name specifying part 113-2 specifies the corresponding tag TG as the first name.

[0056] Explaining with the example of the first information IF1 shown in FIG. 7, assume that the virtual object specifying part 113-1 specifies the virtual object VO with ID = 7 based on the above instruction information. In this case, in the first information IF1, since the tag TG corresponding to the virtual object VO with ID = 7 is "#FOX", the name specifying part 113-2 specifies the tag TG of "#FOX". On the other hand, assume that the virtual object specifying part 113-1 specifies the virtual object VO with ID = 6 based on the above instruction information. In this case, since there is no tag TG corresponding to the virtual object VO with ID = 6 in the first information IF1, the name specifying part 113-2 does not specify the tag TG.

[0057] Figs. 10A to 10C are explanatory diagrams of how to use tag TG. As an example, as shown in Fig. 10A, it is assumed that in the virtual space VS perceived by user U1, virtual objects VO6 to VO8 each showing a deer, a fox, and a horse are arranged. Also, it is assumed that a character string "#FOX" is registered as tag TG7 in the virtual object VO7 of the fox. As an example, as shown in Fig. 10B, user U1 utters the character string indicated by the tag TG corresponding to each content while pressing a specific part as the input device 15 provided in the terminal device 10. Here, as an example, it is assumed that user U1 uttered the character string "#FOX" which is the tag TG7 corresponding to the fox. Due to the generation of tag TG7 by user U1, as shown in Fig. 10C, user U1 can call the virtual object VO7 corresponding to the uttered character string in the direction of the line of sight.

[0058] The display control unit 114 causes the AR glasses 20 as a display device to display a plurality of virtual objects VO arranged in the virtual space VS. Also, the display control unit 114 causes the AR glasses 20 to display the tag TG specified by the name identification unit 113-2. More specifically, the display control unit 114 generates image data to be displayed on the AR glasses 20, and transmits the generated image data to the AR glasses 20 via the communication device 13.

[0059] Figures 11A and 11B are explanatory diagrams of a first operation example of the display control unit 114. First, the acquisition unit 111 acquires, as a trigger, an operation signal generated by an operation on the input device 15 by the user U1. Then, as shown in FIG. 11A, the display control unit 114 divides the celestial sphere-shaped virtual space VS into a plurality of regions R1 to R17 by a plurality of straight lines corresponding to the latitudes and longitudes of the celestial sphere. Then, the display control unit 114 flattens each of the divided plurality of regions R1 to R17. Further, as shown in FIG. 11B, the display control unit 114 causes the display 29 of the AR glasses 20 to display a two-dimensional image SI obtained by flattening the plurality of flattened regions R1 to R17. Specifically, the display control unit 114 arranges, in the two-dimensional image SI, the regions R5 and R6 located to the right of the region R17 at the top of the sky as seen from the user U1 located at the center of the celestial sphere-shaped space, to the right of the region R17. Further, the display control unit 114 arranges, in the two-dimensional image SI, the regions R13 and R14 located to the left of the region R17 at the top of the sky as seen from the user U1 located at the center of the celestial sphere-shaped space, to the left of the region R17. Further, the display control unit 114 arranges, in the two-dimensional image SI, the regions R9 and R10 located in front of the region R17 at the top of the sky as seen from the user U1 located at the center of the celestial sphere-shaped space, below the region R17. Further, the display control unit 114 arranges, in the two-dimensional image SI, the regions R1 and R2 located behind the region R17 at the top of the sky as seen from the user U1 located at the center of the celestial sphere-shaped space, above the region R17. Further, the display control unit 114 arranges, in the two-dimensional image SI, the regions R7 and R8 located in the right front of the region R17 at the top of the sky as seen from the user U1 located at the center of the celestial sphere-shaped space, in the lower right of the region R17. Further, the display control unit 114 arranges, in the two-dimensional image SI, the regions R11 and R12 located in the left front of the region R17 at the top of the sky as seen from the user U1 located at the center of the celestial sphere-shaped space, in the lower left of the region R17. Further, the display control unit 114 arranges, in the two-dimensional image SI, the regions R3 and R4 located in the right rear of the region R17 at the top of the sky as seen from the user U1 located at the center of the celestial sphere-shaped space, in the upper right of the region R17.In addition, from the perspective of the user U1 located at the center of the spherical space, the display control unit 114 arranges the regions R15 and R16 located left and rearward centered on the region R17 at the top of the sky in the upper left of the region R17 in the two-dimensional image SI. Further, as shown in FIG. 11B, the vertical and horizontal directions of each of the regions R1 to R17 coincide with the vertical and horizontal directions as seen from the user U1 when the user U1 located at the center of the spherical space confronts the regions R1 to R17.

[0060] Then, on the display 29 of the AR glasses 20, the display control unit 114 causes the tag TG specified by the name identification unit 113-2 to be displayed adjacent to the region R that includes the virtual object VO specified by the virtual object identification unit 113-1. In the example shown in FIG. 11B, the display control unit 114 causes the tag TG7 specified by the name identification unit 113-2 to be displayed adjacent to the region R3 that includes the virtual object VO7 specified by the virtual object identification unit 113-1.

[0061] Note that as described above, the display control unit 114 uses an operation signal corresponding to an operation on the input device 15 by the user U1 as a trigger to flatten the plurality of regions R1 to R17. However, the operation signal that triggers the flattening is not limited to being generated in response to an operation on the input device 15. For example, as described above, the processing device 11 functions as the motion recognition unit 112. The above operation signal may be generated according to the gesture of the user U1 detected when the processing device 11 functions as the motion recognition unit 112. Alternatively, the above operation signal may be generated according to the posture of the terminal device 10.

[0062] Figs. 12A and 12B are explanatory diagrams of a second operation example of the display control unit 114. As shown in Fig. 12A, normally, the display control unit 114 causes a spherical virtual space VS in which the user U1 is located at the nadir to be displayed on a display 29 of the AR glasses 20 worn on the head of the user U1. Further, based on an operation signal generated by an operation of the input device 15 by the user U1, as shown in Fig. 12B, the display control unit 114 causes a three-dimensional image TI obtained by reducing the virtual space VS to be displayed on the display 29 of the AR glasses 20. Then, on the three-dimensional image TI, the display control unit 114 displays the virtual object VO specified by the virtual object specifying unit 113-1 and the tag TG corresponding to the virtual object VO adjacent to each other. In the example shown in Fig. 12B, the display control unit 114 displays the virtual object VO7 specified by the virtual object specifying unit 113-1 and the tag TG7 corresponding to the virtual object VO7 adjacent to each other on the three-dimensional image TI.

[0063] Note that, as described above, the display control unit 114 causes the three-dimensional image TI to be displayed on the display 29 based on an operation signal generated by an operation of the input device 15 by the user U1. However, the operation signal that triggers the display of the three-dimensional image TI is not limited to being generated in response to an operation on the input device 15. For example, similar to the above, the operation signal may be generated in response to a gesture of the user U1 detected when the processing device 11 functions as the motion recognition unit 112. Alternatively, the operation signal may be generated according to the posture of the terminal device 10.

[0064] Returning to the description of Fig. 6, the determination unit 115 determines whether or not the viewpoint of the user U1 is located within a specific virtual object VO displayed on the AR glasses 20 as a display device for a predetermined time or longer.

[0065] As described above, the AR glass 20 includes the line-of-sight detection device 23. Further, the line-of-sight detection device 23 detects line-of-sight data based on, for example, the position of the head and the position of the iris of the user U1. The determination unit 115 sets the coordinates of the point where the direction of the line of sight indicated by the line-of-sight data collides with the spherical surface of the celestial sphere as the virtual space VS as the position of the viewpoint of the user U1. As illustrated in FIG. 7, the first information IF1 includes the position information of each first virtual object VO. The first virtual object VO is arranged at the position indicated by the position information. The determination unit 115 determines whether or not the viewpoint of the user U1 is located within the area where the first virtual object VO is arranged for a predetermined time or longer.

[0066] FIGS. 13A and 13B are explanatory diagrams of a third operation example of the display control unit 114 particularly when cooperating with the determination unit 115. As shown in FIG. 13A, in the virtual space VS, when it is determined that the viewpoint of the user U1 is located within a specific virtual object VO for a predetermined time or longer, the display control unit 114, as shown in FIG. 13B, displays a tag TG corresponding to the virtual object VO in the vicinity of the virtual object VO. Here, the "vicinity" of the virtual object VO specifically refers to a range within a predetermined distance from the virtual object VO. In the examples shown in FIGS. 13A and 13B, when it is determined that the viewpoint of the user U1 is located within a specific virtual object VO7 for a predetermined time or longer, the display control unit 114 displays "#FOX", which is the tag TG7 corresponding to the virtual object VO7, in the vicinity of the virtual object VO7.

[0067] Returning to the explanation of FIG. 6, the voice recognition unit 116 recognizes the voice spoken by the user U1.

[0068] As described above, the AR glasses 20 are provided with a sound collection device 24. The voice spoken by the user U1 wearing the AR glasses 20 on the head is collected by the sound collection device 24 and converted into voice data. The voice data is output from the AR glasses 20 to the terminal device 10. The voice recognition unit 116 recognizes the content of the speech based on the voice data acquired from the AR glasses 20. More specifically, the voice recognition unit 116 converts the voice data into text data.

[0069] FIGS. 14A and 14B are explanatory diagrams of a fourth operation example of the display control unit 114, particularly when cooperating with the voice recognition unit 116. As shown in FIG. 14A, when the user U1 wearing the AR glasses 20 on the head speaks the attribute of the virtual object VO, the voice recognition unit 116 recognizes the voice data indicating the voice spoken by the user U1. When the recognition result by the voice recognition unit 116 indicates one of the attributes included in the second information IF2, the virtual object specifying unit 113-1 specifies one or more virtual objects VO corresponding to the one attribute based on the second information IF2.

[0070] In the example shown in FIG. 14A, the user U1 speaks the word "animal" as the above attribute. The voice recognition unit 116 recognizes the voice spoken by the user U1 as the character string "animal". The virtual objects VO corresponding to the recognition result of "animal" are the virtual object VO6 which is a 3D model of a deer, the virtual object VO7 which is a 3D model of a fox, and the virtual object VO8 which is a 3D model of a horse. Therefore, the virtual object specifying unit 113-1 specifies the virtual objects VO6 to VO8 as the virtual objects VO corresponding to the attribute of "animal" among the plurality of virtual objects VO.

[0071] Also, based on the first information IF1, the name identification unit 113-2 identifies one or more tags TG corresponding to each of some or all of the one or more virtual objects VO identified by the virtual object identification unit 113-1. In the example shown in FIG. 14A, examples of the one or more virtual objects VO identified by the virtual object identification unit 113-1 include a virtual object VO6 which is a 3D model of a deer, a virtual object VO7 which is a 3D model of a fox, and a virtual object VO8 which is a 3D model of a horse. Also, it is assumed that there is no tag TG6 corresponding to the virtual object VO6 which is a 3D model of a deer. It is assumed that the tag TG7 corresponding to the virtual object VO7 which is a 3D model of a fox is "#FOX". It is assumed that the tag TG8 corresponding to the virtual object VO8 which is a 3D model of a horse is "#HORSE". Therefore, some of the identified one or more virtual objects VO are the virtual object VO7 which is a 3D model of a fox and the virtual object VO8 which is a 3D model of a horse. Further, one or more tags TG corresponding to each of some of the identified one or more virtual objects VO are two tags TG, i.e., the tag TG7 of "#FOX" and the tag TG8 of "#HORSE". Accordingly, the name identification unit 113-2 identifies two tags TG, i.e., the tag TG7 of "#FOX" and the tag TG8 of "#HORSE". On the other hand, if, as the tag TG6 corresponding to the virtual object VO6 which is a 3D model of a deer, "#DEER" is assigned, the name identification unit 113-2 will identify one or more tags TG corresponding to each of all of the one or more virtual objects VO identified by the virtual object identification unit 113-1.

[0072] As shown in FIG. 14B, the display control unit 114 causes the AR glasses 20 as a display device to display one or more virtual objects VO identified based on the second information IF2 by the virtual object identification unit 113-1. Also, when there is a tag TG corresponding to the virtual object VO, the display control unit 114 causes the tag TG to be displayed in association with the corresponding virtual object VO.

[0073] In the example shown in FIG. 14B, icons IC6 to IC8 each indicating a virtual object VO6 to VO8 identified by the virtual object identification unit 113-1 are displayed in the pop-up P1. Further, in the pop-up P1, a tag TG7 of "#FOX" is attached to the icon IC7, and a tag TG8 of "#HORSE" is attached to the icon IC8.

[0074] Also, when the recognition result by the voice recognition unit 116 matches the tag TG included in the first information IF1, the display control unit 114 changes the display content regarding the first virtual object VO corresponding to the tag TG. As an example, as described with reference to FIGS. 10A to 10C, the display control unit 114 may move the virtual object VO corresponding to the tag TG in the direction of the user U1's line of sight. Alternatively, when the virtual object VO corresponding to the tag TG is an application, the display control unit 114 may display a screen for selecting whether to start the application.

[0075] Although not shown in FIG. 6, the terminal device 10 may include a GPS device similar to the GPS device 25 provided in the AR glasses 20. In this case, the AR glasses 20 may not include the GPS device 25.

[0076] 1.1.4: Server Configuration FIG. 15 is a block diagram showing a configuration example of the server 30. The server 30 includes a processing device 31, a storage device 32, a communication device 33, a display 34, and an input device 35. Each element of the server 30 is interconnected by one or more buses for communicating information.

[0077] The processing device 31 is a processor that controls the entire server 30 and is configured using, for example, one or more chips. The processing device 31 is configured using, for example, a central processing unit (CPU) that includes an interface with peripheral devices, an arithmetic unit, registers, and the like. Note that part or all of the functions of the processing device 31 may be realized by hardware such as a DSP, an ASIC, a PLD, or an FPGA. The processing device 31 executes various processes in parallel or sequentially.

[0078] The storage device 32 is a recording medium that can be read from and written to by the processing device 31, and stores a plurality of programs including a control program PR3 executed by the processing device 31, first information IF1, and second information IF2.

[0079] The communication device 33 is hardware as a transmission / reception device for communicating with other devices. The communication device 33 is also called, for example, a network device, a network controller, a network card, or a communication module. The communication device 33 includes a connector for a wired connection and may include an interface circuit corresponding to the connector. Further, the communication device 33 may include a wireless communication interface. Examples of the connector and the interface circuit for the wired connection include products compliant with a wired LAN, IEEE1394, and USB. Examples of the wireless communication interface include products compliant with a wireless LAN and Bluetooth (registered trademark).

[0080] The display 34 is a device that displays image and character information. The display 34 displays various images under the control of the processing device 31. For example, various display panels such as a liquid crystal display panel and an organic EL display panel are suitably used as the display 34.

[0081] The input device 35 is a device that accepts operations by the administrator of the information processing system 1. For example, the input device 35 includes a pointing device such as a keyboard, a touch pad, a touch panel, or a mouse. Here, when the input device 35 includes a touch panel, it may also serve as the display 34. In particular, the administrator of the information processing system 1 can input and edit the first information IF1 and the second information IF2 by using the input device 35.

[0082] The processing device 31 functions as an output unit 311 and an acquisition unit 312, for example, by reading and executing the control program PR3 from the storage device 32.

[0083] The output unit 311 outputs the first information IF1 and the second information IF2 stored in the storage device 32 to the terminal device 10 by using the communication device 33. Also, the output unit 311 outputs to the terminal device 10 the data necessary for the terminal device 10 to provide the virtual space VS to the user U1 wearing the AR glasses 20 on the head. The data includes data related to the virtual object VO itself and data related to an application (not shown) for using the cloud service.

[0084] The acquisition unit 312 acquires various data from the terminal device 10 by using the communication device 33. The data includes, for example, data indicating the operation content for the virtual object VO input to the terminal device 10 by the user U1 wearing the AR glasses 20 on the head. Also, when the user U1 uses the above cloud service, the data includes the input data to the above application.

[0085] 2: Operations of the First Embodiment FIGS. 16 to 19 are flowcharts showing the operations of the information processing system 1 according to the first embodiment. Hereinafter, the operations of the information processing system 1 will be described with reference to FIGS. 16 to 19.

[0086] 1.2.1: First Operation FIG. 16 is a flowchart for explaining the first operation of the information processing system 1.

[0087] In step S1, the processing device 11 functions as an acquisition unit 111 to acquire an operation signal. As illustrated with reference to FIGS. 11A and 11B, the operation signal is a trigger for causing a two-dimensional image SI obtained by flattening a celestial sphere-shaped virtual space VS to be displayed on a display 29 provided in the AR glasses 20.

[0088] In step S2, the processing device 11 functions as a display control unit 114 to cause the two-dimensional image SI obtained by flattening the celestial sphere-shaped virtual space VS to be displayed on the display 29 provided in the AR glasses 20.

[0089] In step S3, the processing device 11 functions as an acquisition unit 111 to acquire instruction information. The instruction information is information for specifying a first virtual object VO among a plurality of virtual objects VO arranged in the virtual space VS.

[0090] In step S4, the processing device 11 functions as a virtual object specifying unit 113-1 to specify the first virtual object VO among the plurality of virtual objects VO based on the instruction information. Further, as a result of the processing device 11 functioning as the virtual object specifying unit 113-1 to specify the first virtual object VO, an ID corresponding to the first virtual object VO is output.

[0091] In step S5, the processing device 11 functions as the name identification unit 113-2, and refers to the first information IF1 stored in the storage device 12. Then, the processing device 11 functions as the name identification unit 113-2, and determines whether a tag TG corresponding to the first virtual object VO identified by functioning as the virtual object identification unit 113-1 is included in the first information IF1. Specifically, the processing device 11 functions as the name identification unit 113-2, and determines whether the tag TG corresponding to the ID output in step S4 is included in the first information IF1. When the tag TG is included in the first information IF1, that is, when the determination result in step S5 is YES, the processing device 11 functions as the name identification unit 113-2, identifies the tag TG as the first name, and then executes the process of step S6. When the tag TG is not included in the first information IF1, that is, when the determination result in step S5 is NO, the processing device 11 ends all processes.

[0092] In step S6, the processing device 11 functions as the display control unit 114, and causes the tag TG identified by the name identification unit 113-2 to be displayed on the display 29 provided in the AR glass 20. More specifically, the processing device 11 functions as the display control unit 114, and causes the identified tag TG to be displayed in the two-dimensional image SI displayed on the display 29. Note that in step S5, when the tag TG is not identified, the processing device 11 omits the process of step S6.

[0093] 1.2.2: Second operation FIG. 17 is a flowchart for explaining the second operation of the information processing system 1.

[0094] In step S11, the processing device 11 functions as the acquisition unit 111, and acquires an operation signal. As illustrated with reference to FIGS. 12A and 12B, the operation signal is a trigger for causing the display 29 provided in the AR glass 20 to display a three-dimensional image TI obtained by reducing the celestial sphere-shaped virtual space VS.

[0095] In step S12, the processing device 11 functions as a display control unit 114 to cause the display 29 provided in the AR glasses 20 to display a three-dimensional image TI obtained by reducing the celestial sphere-shaped virtual space VS.

[0096] In step S13, the processing device 11 functions as an acquisition unit 111 to acquire instruction information. The instruction information is information for specifying a first virtual object VO among a plurality of virtual objects VO arranged in the virtual space VS.

[0097] In step S14, the processing device 11 functions as a virtual object specifying unit 113-1 to specify the first virtual object VO among the plurality of virtual objects VO based on the instruction information. Further, as a result of the processing device 11 functioning as the virtual object specifying unit 113-1 to specify the first virtual object VO, an ID corresponding to the first virtual object VO is output.

[0098] In step S15, the processing device 11 functions as a name specifying unit 113-2 to refer to first information IF1 stored in the storage device 12. Then, the processing device 11 functions as the name specifying unit 113-2 to determine whether a tag TG corresponding to the first virtual object VO specified by functioning as the virtual object specifying unit 113-1 is included in the first information IF1. Specifically, the processing device 11 functions as the name specifying unit 113-2 to determine whether the tag TG corresponding to the ID output in step S14 is included in the first information IF1. When the tag TG is included in the first information IF1, that is, when the determination result in step S15 is YES, the processing device 11 functions as the name specifying unit 113-2 to specify the tag TG as the first name and then execute the processing in step S16. When the tag TG is not included in the first information IF1, that is, when the determination result in step S15 is NO, the processing device 11 ends all processing.

[0099] In step S16, the processing device 11 functions as a display control unit 114 to cause the tag TG specified by the name identification unit 113-2 to be displayed on the display 29 provided in the AR glass 20. More specifically, the processing device 11 functions as a display control unit 114 to cause the specified tag TG to be displayed in the three-dimensional image TI displayed on the display 29. Note that if the tag TG is not specified in step S15, the processing device 11 omits the processing of step S16.

[0100] 1.2.3: The Third Operation FIG. 18 is a flowchart for explaining the third operation of the information processing system 1.

[0101] In step S21, the processing device 11 functions as an acquisition unit 111 to acquire gaze data related to the gaze of the user U1 in the AR glass 20. More specifically, the processing device 21 of the AR glass 20 functions as an acquisition unit 211 to output the acquired gaze data to the communication device 28. The communication device 28 outputs the gaze data acquired from the processing device 21 to the terminal device 10. The processing device 11 of the terminal device 10 functions as an acquisition unit 111 to acquire the gaze data from the AR glass 20 using the communication device 13.

[0102] In step S22, the processing device 11 functions as a determination unit 115 to determine whether the viewpoint of the user U1 is located within the first virtual object VO displayed on the AR glasses 20 for a predetermined time or more. More specifically, the processing device 11 functions as a determination unit 115 to obtain the position of the viewpoint of the user U1 as instruction information based on the gaze data obtained in step S21. Then, the processing device 11, as the determination unit 115, determines whether the viewpoint of the user U1 has been located within a specific virtual object VO for a predetermined time or more. If the determination result is true, that is, if the determination result in step S22 is "YES", the processing device 11 executes the process of step S23. If the determination result is false, that is, if the determination result in step S22 is "NO", the processing device 11 executes the process of step S21.

[0103] In step S23, the processing device 11 functions as a virtual object identification unit 113-1 to identify the first virtual object VO among the plurality of virtual objects VO based on the determination result in step S22. More specifically, the processing device 11 functions as a virtual object identification unit 113-1 to identify, in step S22, the virtual object VO in which the viewpoint of the user U1 has been determined to be located for a predetermined time or more as the first virtual object VO. Further, as a result of the processing device 11 functioning as the virtual object identification unit 113-1 to identify the first virtual object VO, the ID corresponding to the first virtual object VO is output.

[0104] In step S24, the processing device 11 functions as a name identification unit 113-2 to refer to the first information IF1 stored in the storage device 12. Then, the processing device 11 functions as a name identification unit 113-2 to determine whether a tag TG corresponding to the first virtual object VO identified by functioning as a virtual object identification unit 113-1 is included in the first information IF1. Specifically, the processing device 11 functions as a name identification unit 113-2 to determine whether the tag TG corresponding to the ID output in step S23 is included in the first information IF1. When the tag TG is included in the first information IF1, that is, when the determination result in step S24 is YES, the processing device 11 functions as a name identification unit 113-2 to identify the tag TG as the first name and then execute the process in step S25. When the tag TG is not included in the first information IF1, that is, when the determination result in step S24 is NO, the processing device 11 ends all processes.

[0105] In step S25, the processing device 11 functions as a display control unit 114 to display the tag TG identified as the first name on the display 29 provided in the AR glass 20. More specifically, the processing device 11 functions as a display control unit 114 to display the identified tag TG near the first virtual object VO on the display 29. Here, the "vicinity" of the first virtual object VO specifically refers to a range within a predetermined distance from the first virtual object VO. In step S25, when the tag TG is not identified, the processing device 11 omits the process in step S26.

[0106] 1.2.4: The Fourth Operation FIG. 19 is a flowchart for explaining the fourth operation of the information processing system 1.

[0107] In step S31, the processing device 11 functions as a voice recognition unit 116 to recognize the voice uttered by the user U1. More specifically, the processing device 21 of the AR glasses 20 functions as an acquisition unit 211 to acquire voice data indicating the voice of the user U1 from the sound collection device 24. Further, the processing device 21 of the AR glasses 20 functions as an acquisition unit 211 to output the acquired voice data to the communication device 28. The communication device 28 outputs the voice data acquired from the processing device 21 to the terminal device 10. The processing device 11 of the terminal device 10 functions as an acquisition unit 111 to acquire the voice data from the AR glasses 20 using the communication device 13. Further, the processing device 11 of the terminal device 10 functions as a voice recognition unit 116 to perform voice recognition on the voice data. The character string as the voice recognition result corresponds to the instruction information in the above first operation to third operation. In this operation example, it is assumed that the character string as the voice recognition result is a character string indicating an attribute.

[0108] In step S32, the processing device 11 functions as a virtual object identification unit 113-1 to identify the virtual object VO. More specifically, the processing device 11 functions as a virtual object identification unit 113-1 to refer to the second information IF2 stored in the storage device 12. Further, the processing device 11 functions as a virtual object identification unit 113-1 to identify one or more virtual objects VO whose attributes corresponding to the character string as the voice recognition result are included in the second information IF2. Further, as a result of the processing device 11 functioning as a virtual object identification unit 113-1 to identify one or more virtual objects VO, the ID corresponding to the one or more virtual objects VO is output.

[0109] In step S33, the processing device 11 functions as a name identification unit 113-2 to refer to the first information IF1 stored in the storage device 12. Then, the processing device 11 functions as a name identification unit 113-2 and determines whether a tag TG corresponding to the first virtual object VO identified by functioning as a virtual object identification unit 113-1 is included in the first information IF1. Specifically, the processing device 11 functions as a name identification unit 113-2 and determines whether the tag TG corresponding to the ID output in step S32 is included in the first information IF1. When the tag TG is included in the first information IF1, that is, when the determination result in step S33 is YES, the processing device 11 functions as a name identification unit 113-2, identifies the tag TG as the first name, and then executes the process of step S34. When the tag TG is not included in the first information IF1, that is, when the determination result in step S33 is NO, the processing device 11 ends all processes.

[0110] In step S34, the processing device 11 functions as a display control unit 114 to cause the tag TG identified as the first name to be displayed on the display 29 provided in the AR glasses 20. As an example, the processing device 11 functions as a display control unit 114 to cause the tag TG identified to be displayed on the display 29 in a form appended to the virtual object VO within the pop-up P1. Note that in step S33, when the tag TG is not identified, the processing device 11 omits the process of step S34.

[0111] 1.3: Effects Achieved by the First Embodiment According to the above description, the terminal device 10 as an information processing device includes a display control unit 114, a virtual object specifying unit 113-1, and a call name specifying unit 113-2. The display control unit 114 causes an AR glass 20 as a display device worn on the head of the user U1 to display a plurality of virtual objects VO arranged in the virtual space VS. The virtual object specifying unit 113-1 specifies a first virtual object VO among the plurality of virtual objects VO based on instruction information generated according to the operation of the user U1. When a tag TG corresponding to the first virtual object VO specified by the virtual object specifying unit 113-1 is stored in the storage device 12, the call name specifying unit 113-2 specifies the corresponding tag TG as a first call name. The display control unit 114 causes the AR glass 20 as a display device to display the tag TG as the first call name.

[0112] By using the above configuration, the terminal device 10 as an information processing device can easily remind the user U1 wearing the AR glass 20 on the head of a tag TG as a first call name for specifying a first virtual object VO arranged in the virtual space VS. Specifically, it may be difficult for the user U1 to grasp the correspondence between each virtual object VO and each tag TG and to remember the tag TG. In such a case, for the AR glass 20 as a display device worn on the head of the user U1, the tag TG corresponding to the virtual object VO specified based on the instruction information generated according to the operation of the user U1 is displayed in an associated manner with the virtual object VO. The user U1 can remember the tag TG corresponding to the virtual object VO by visually recognizing the displayed tag TG on the AR glass 20.

[0113] In addition, the display control unit 114 causes the AR glass 20 as a display device to display a two-dimensional image SI obtained by flattening the virtual space VS. Further, the display control unit 114 displays the tag TG as the first call name in association with the first virtual object VO within the two-dimensional image SI.

[0114] By using the terminal device 10 as the information processing device with the above configuration, the user U1 can recall the tag TG corresponding to the virtual object VO that the user wants to call in the virtual space VS after grasping the position of the virtual object VO.

[0115] In addition, the display control unit 114 causes the AR glass 20 as the display device to display a three-dimensional image TI obtained by shrinking the virtual space VS. Further, the display control unit 114 causes the tag TG as the first call name to be displayed in association with the first virtual object VO in the three-dimensional image TI.

[0116] By using the terminal device 10 as the information processing device with the above configuration, the user U1 can recall the tag TG corresponding to the virtual object VO that the user wants to call in the virtual space VS after grasping the position of the virtual object VO.

[0117] In addition, the operation of the user U1 described above is visual on the AR glass 20 as the display device. The above instruction information indicates the viewpoint of the user U1 on the AR glass 20 as the display device. The terminal device 10 as the information processing device further includes a determination unit 115. The determination unit 115 determines whether the viewpoint of the user U1 is located within the first virtual object VO displayed on the AR glass 20 as the display device for a predetermined time or more. When the determination unit 115 determines that the viewpoint of the user U1 is located within the first virtual object VO for a predetermined time or more, the display control unit 114 causes the tag TG as the first call name corresponding to the first virtual object VO to be displayed on the AR glass 20 as the display device.

[0118] By using the terminal device 10 as the information processing device with the above configuration, the user U1 can recall the tag TG as the first call name corresponding to the first virtual object VO without the need for operations other than the operation related to the line of sight with respect to the AR glass 20 as the display device.

[0119] In addition, when the recognition result of the voice uttered by the user U1 matches the tag TG as the first call name, the display control unit 114 causes the display content regarding the first virtual object VO to be changed.

[0120] By using the terminal device 10 as the information processing device with the above configuration, the user U1 can, for example, move the first virtual object VO corresponding to the tag TG that matches the recognition result of the voice in the direction of the line of sight of the user U1. Alternatively, for example, when the first virtual object VO corresponding to the tag TG that matches the recognition result of the voice is an application, the user U1 can start the application.

[0121] Also, according to the above description, in the terminal device 10 as the information processing device, the virtual object specifying unit 113-1 specifies one or more virtual objects VO based on the voice uttered by the user U1 and the recognition result of the voice representing at least one of the attributes of the plurality of virtual objects VO. The call name specifying unit 113-2 specifies a tag TG as the call name corresponding to each of some or all of the one or more virtual objects VO specified by the virtual object specifying unit 113-1. The display control unit 114 causes the tag TG as the corresponding call name specified by the call name specifying unit 113-2 to be displayed on the AR glasses 20 as the display device in association with each of some or all of the virtual objects VO.

[0122] By using the terminal device 10 as the information processing device with the above configuration, the user U1 narrows down the plurality of virtual objects VO arranged in the virtual space VS to one or more virtual objects VO corresponding to the same attribute. When the tag TG corresponding to the virtual object VO that the user U1 wants to call is included in the tags TG corresponding to the narrowed-down virtual objects VO, it becomes easier for the user U1 to recall the tag TG.

[0123] 2: Second Embodiment Hereinafter, with reference to FIGS. 20 to 29, the configuration of the information processing system 1A including the information processing apparatus according to the second embodiment of the present invention will be described. In the following description, for the sake of simplicity of explanation, the same reference numerals are used for the same components as those in the first embodiment, and the description of their functions may be omitted. Further, in the following description, for the sake of simplicity of explanation, mainly, the differences between the second embodiment and the first embodiment will be described.

[0124] 2.1: Configuration of the First Embodiment 2.1.1: Overall Configuration FIG. 20 is a diagram showing the overall configuration of the information processing system 1A according to the second embodiment of the present invention. The information processing system 1A is different from the information processing system 1 according to the first embodiment in that it includes a terminal device 10A instead of the terminal device 10.

[0125] 2.1.2: Configuration of the Terminal Device FIG. 21 is a block diagram showing a configuration example of the terminal device 10A. The terminal device 10A is different from the terminal device 10 according to the first embodiment in that it includes a processing device 11A instead of the processing device 11 and a storage device 12A instead of the storage device 12.

[0126] Unlike the storage device 12 according to the first embodiment, the storage device 12A is different in that it is not essential to store the second information IF2 and in that it stores the learning model LM1.

[0127] The learning model LM1 is a learning model used by the name identification unit 113-2A described later. Specifically, the learning model LM1 is a learning model for calculating the similarity between a first word and a second word. As an example, the learning model LM1 numerically vectorizes the meaning of words and calculates the similarity between the first word and the second word based on how much the directions of the vectors of the first word and the second word are the same. However, the above method of calculating the similarity is only an example and is not limited thereto. As long as the learning model LM1 can calculate the similarity between the first word and the second word, other methods may be used.

[0128] The learning model LM1 is generated by learning teacher data in the learning phase. The teacher data used to generate the learning model LM1 has a plurality of sets of a first word and a second word and a numerical value indicating similarity.

[0129] Also, the learning model LM1 is generated outside the terminal device 10A. In particular, it is preferable that the learning model LM1 is generated in the server 30. In this case, the terminal device 10A acquires the learning model LM1 from the server 30 via the communication network NET.

[0130] The processing device 11A functions as an acquisition unit 111, an operation recognition unit 112, a specification unit 113A, a display control unit 114A, a voice recognition unit 116, and an update unit 117 by reading and executing the control program PR4 from the storage device 12A. Note that the acquisition unit 111, the operation recognition unit 112, and the voice recognition unit 116 are the same as the acquisition unit 111, the operation recognition unit 112, and the voice recognition unit 116 as functions of the processing device 11 according to the first embodiment, and thus the description thereof is omitted.

[0131] FIG. 22 is a functional block diagram showing the configuration of the specification unit 113A. The specification unit 113A is different from the specification unit 113 according to the first embodiment in that it includes a name specification unit 113-2A instead of the name specification unit 113-2.

[0132] As a first operation, the name specification unit 113-2A performs the same operation as the name specification unit 113-2. Specifically, when the tag TG corresponding to the first virtual object VO specified by the virtual object specification unit 113-1 is included in the first information IF1 stored in the storage device 12, the name specification unit 113-2A specifies the corresponding tag TG as the first name.

[0133] Further, as a second operation, the name identification unit 113-2A identifies a plurality of tags TG corresponding to some or all of the plurality of virtual objects VO arranged in the virtual space VS as a first name. More specifically, the name identification unit 113-2A identifies the tag TG of the virtual object VO in which the tag TG is included in the first information IF1 among the plurality of virtual objects VO arranged in the virtual space VS.

[0134] Furthermore, as a third operation, when the recognition result of the first voice spoken by the user U1 is a second name that does not match any of the plurality of tags TG included in the first information IF1, the name identification unit 113-2A identifies the tag TG that is most similar to the second name among the plurality of tags TG included in the first information IF1 as the first name.

[0135] More specifically, the name identification unit 113-2A inputs the character string as the recognition result of the first voice spoken by the user U1 and one of the plurality of tags TG included in the first information IF1 into the above learning model LM1. The learning model LM1 outputs the similarity between the character string as the recognition result of the first voice and the above one tag TG. The name identification unit 113-2A performs the same operation for all the tags TG described in the first information IF1 and obtains the similarity between the character string as the recognition result of the first voice and all the tags TG. Further, the name identification unit 113-2A identifies the tag TG with the highest similarity value among all the tags TG.

[0136] Returning to the explanation in FIG. 21, the display control unit 114A causes the AR glasses 20 as a display device to display the tag TG identified by the name identification unit 113-2A.

[0137] FIG. 23 is an explanatory diagram of a first operation example of the display control unit 114A. As shown in FIG. 23, although the user U1 uttered "DOG" as the tag TG corresponding to the virtual object VO to be called, assume that the character string "DOG", which is the recognition result of the first uttered voice, is a second call name not included in the first information IF1. In this case, the call name specifying unit 113-2A specifies the tag TG of "HORSE" as the tag TG most similar to the second call name of "DOG" among the plurality of tags TG included in the first information IF1. Then, the display control unit 114A causes a pop-up P2 to be displayed in the virtual space VS. Further, the display control unit 114A causes a message for confirming whether the tag TG that the user U1 originally intended to utter is not "HORSE" to be displayed to the user U1 within the pop-up P2.

[0138] FIG. 24 is an explanatory diagram of a second operation example of the display control unit 114A. As shown in FIG. 24, although the user U1 uttered some character string as the tag TG corresponding to the virtual object VO to be called, assume that the character string, which is the recognition result of the first uttered voice, is a second call name not included in the first information IF1. In this case, the call name specifying unit 113-2A specifies a plurality of tags TG corresponding to some or all of the plurality of virtual objects VO arranged in the virtual space VS based on the first information IF1. Then, the display control unit 114A causes a pop-up P3 to be displayed in the virtual space VS. Further, the display control unit 114A causes the icons of the plurality of virtual objects VO corresponding to the plurality of specified tags TG, that is, reduced displays, and the plurality of specified tags TG to be associated and displayed in a list within the pop-up P3.

[0139] FIG. 25 is an explanatory diagram of a third operation example of the display control unit 114A. As shown in FIG. 25, assume that a user U1 uttered some string as a tag TG corresponding to a virtual object VO that the user U1 wants to call, and the string, which is the recognition result of the first uttered voice, is a second call name not included in the first information IF1. In this case, the call name specifying unit 113-2A specifies a plurality of tags TG corresponding to some or all of the plurality of virtual objects VO arranged in the virtual space VS based on the first information IF1. That is, the call name specifying unit 113-2A specifies all the tags TG included in the first information IF1. Then, the display control unit 114A causes the plurality of tags TG to be displayed in the vicinity of some or all of the plurality of virtual objects VO corresponding to the plurality of tags TG in the virtual space VS. Here, the “vicinity” of some or all of the plurality of virtual objects VO specifically refers to a range within a predetermined distance from each virtual object VO.

[0140] Assume that after the control unit 114A executes the operations shown in the above first operation example, the user U1 uttered the tag TG identified by the name identification unit 113-2A. Alternatively, assume that after the display control unit 114A executes the operations shown in the above second operation example or third operation example, the user U1 uttered any one of the one or more tags TG identified by the name identification unit 113-2A. In the operations shown in these first to third operation examples, although the user U1 uttered some string as the tag TG for the virtual object VO to be called, if the string was not included in the first information IF1, the uttered voice is the first voice. Thereafter, when the user U1 utters any one of the one or more tags TG identified by the name identification unit 113-2A, the uttered voice is the second voice. When the recognition result of the second voice uttered by the user U1 matches the tag TG identified by the name identification unit 113-2A, the display control unit 114A causes the display related to the virtual object VO corresponding to the matched tag TG to be changed. Specifically, the display control unit 114A may move the virtual object VO corresponding to the matched tag TG in the direction of the user U1's line of sight. Alternatively, when the virtual object VO corresponding to the matched tag TG is an application, the display control unit 114 may display a screen for selecting whether to start the application.

[0141] Returning to the explanation of FIG. 21, as described above, when the recognition result of the first voice spoken by the user U1 is not included in the first information IF1 and the number of times the recognition result of the second voice matches a specific tag TG included in the first information IF1 reaches a predetermined number of times, the update unit 117 associates the virtual object VO corresponding to the specific tag TG with the tag TG as the second name of the recognition result of the first voice. More specifically, as described in the explanations of the above first operation example to the third operation example, assume that the user U1 spoke some string as the tag TG for the virtual object VO to be called, but the recognition result of the first voice by the speech was not included in the first information IF1. Next, assume that the user U1 spoke a string as any one of the one or more tags TG identified by the name identification unit 113-2A, and the recognition result of the second voice by the speech was included in the first information IF1. Thus, assume that after the recognition result of the first voice was not included in the first information IF1, the fact that the recognition result of the second voice was included in the first information IF1 reached a predetermined number of times. In this case, the update unit 117 sets the string as the recognition result of the first voice as a new tag TG, and then associates the virtual object VO corresponding to the tag TG as the recognition result of the second voice with the new tag TG.

[0142] Here, the storage device 12A stores the virtual object VO, the recognition result of the first voice, and the number of times the recognition result of the second voice was included in the first information IF1 in a tabular format after the first voice was spoken. By referring to this table by the processing device 11A, the update unit 117 determines whether or not the number of times the recognition result of the first voice is not included in the first information IF1 and the recognition result of the second voice is included in the first information IF1 has reached a predetermined number of times. Then, when the number of times reaches a predetermined number of times, the update unit 117 sets the string as the recognition result of the first voice as a new tag TG, and then associates the virtual object VO corresponding to the tag TG as the recognition result of the second voice with the new tag TG.

[0143] Figures 26A and 26B are explanatory diagrams showing the functions of the update unit 117. As shown in Figure 26A, assume that the tag TG corresponding to the virtual object VO with ID = 1 is "#WEB", and the tag TG corresponding to the virtual object VO with ID = 8 is "#HORSE". Also, in the first operation example to the third operation example described by referring to Figures 23 to 25, assume that the recognition result of the first voice is "DOG" as the second call name, and the recognition result of the second voice is "HORSE" included in the first information IF1. In this case, as shown in Figure 26B, the update unit 117 associates the virtual object VO with ID = 8 corresponding to the tag TG of "#HORSE" with the tag TG of "#DOG". That is, the tag TG corresponding to the virtual object VO with ID = 8 becomes two tags TG of "#HORSE" and "#DOG".

[0144] 2.2: Operations of the Second Embodiment Figures 27 to 29 are flowcharts showing the operations of the information processing system 1A according to the second embodiment. Hereinafter, the operations of the information processing system 1A will be described by referring to Figures 27 to 29.

[0145] 2.2.1: The First Operation Figure 27 is a flowchart for explaining the first operation of the information processing system 1A.

[0146] In step S41, the processing device 11A functions as a voice recognition unit 116 to recognize the voice uttered by the user U1. More specifically, the processing device 21 of the AR glasses 20 functions as an acquisition unit 211 to acquire voice data indicating the voice of the user U1 from the sound collection device 24. Also, the processing device 21 of the AR glasses 20 functions as an acquisition unit 211 to output the acquired voice data to the communication device 28. The communication device 28 outputs the voice data acquired from the processing device 21 to the terminal device 10A. The processing device 11A of the terminal device 10A functions as an acquisition unit 111 to acquire the voice data from the AR glasses 20 using the communication device 13. Also, the processing device 11A of the terminal device 10A functions as a voice recognition unit 116 to perform voice recognition on the voice data. The character string as the voice recognition result corresponds to the instruction information in the first embodiment.

[0147] In step S42, the processing device 11A functions as a call name identification unit 113-2A to determine whether the recognition result of the voice uttered by the user U1 corresponds to any of a plurality of tags TG corresponding to a plurality of virtual objects VO included in the first information IF1. If the determination result is true, that is, if the determination result in step S42 is "YES", the processing device 11A executes the process of step S45. In this case, the recognition result of the voice uttered by the user U1 is an example of the above first call name. If the determination result is false, that is, if the determination result in step S42 is "NO", the processing device 11A executes the process of step S43. In this case, the recognition result of the voice uttered by the user U1 is an example of the above second call name.

[0148] In step S43, the processing device 11A functions as a call name identification unit 113-2A to identify, as the first call name, the tag TG that is most similar to the second call name as the recognition result of the voice of the user U1 among the plurality of tags TG included in the first information IF1, that is, the plurality of first call names.

[0149] In step S44, the processing device 11A functions as the display control unit 114A to display the most similar tag TG identified in step S43 in the virtual space VS. For example, the processing device 11A functions as the display control unit 114A to display the pop-up P2 in the virtual space VS. Further, the processing device 11A functions as the display control unit 114A to display, within the pop-up P2, a message for the user U1 to confirm the tag TG that the user U1 was originally about to utter. Thereafter, the processing device 11A executes the process of step S41.

[0150] In step S45, the processing device 11A functions as the display control unit 114A to change the display of the virtual object VO corresponding to the tag TG as the first call name, which is the recognition result in step S42.

[0151] Note that when the recognition result of the first voice uttered by the user U1 is not included in the first information IF1 and the number of times the recognition result of the second voice matches a specific tag TG included in the first information IF1 reaches a predetermined number of times, after step S45, the update unit 117 may associate the virtual object VO corresponding to the specific tag TG with the tag TG as the second call name, which is the recognition result of the first voice.

[0152] 2.2.2: Second operation FIG. 28 is a flowchart for explaining the second operation of the information processing system 1A.

[0153] In step S51, the processing device 11A functions as the voice recognition unit 116 to recognize the voice uttered by the user U1. Note that the details of the operation are the same as those in step S41 of the first operation, so the description thereof is omitted.

[0154] In step S52, by functioning as the name identification unit 113-2A, the processing device 11A determines whether the recognition result of the voice uttered by the user U1 corresponds to any of the plurality of tags TG included in the first information IF1. If the determination result is true, that is, if the determination result in step S52 is "YES", the processing device 11A executes the process of step S55. In this case, the recognition result of the voice uttered by the user U1 is an example of the above-mentioned first name. If the determination result is false, that is, if the determination result in step S52 is "NO", the processing device 11A executes the process of step S53. In this case, the recognition result of the voice uttered by the user U1 is an example of the above-mentioned second name.

[0155] In step S53, by functioning as the name identification unit 113-2A, the processing device 11A identifies a plurality of tags TG corresponding to some or all of the plurality of virtual objects VO arranged in the virtual space VS based on the first information IF1.

[0156] In step S54, by functioning as the display control unit 114A, the processing device 11A causes a pop-up P3 to be displayed in the virtual space VS. Further, by functioning as the display control unit 114A, the processing device 11A associates and displays in a list, within the pop-up P3, the icons of the plurality of virtual objects VO corresponding to the plurality of tags TG identified in step S53, that is, the reduced displays, with the identified plurality of tags TG. Thereafter, the processing device 11A executes the process of step S51.

[0157] In step S55, by functioning as the display control unit 114A, the processing device 11A causes the display of the virtual object VO corresponding to the tag TG as the first name, which is the recognition result in step S52, to be changed.

[0158] Note that, similar to the first operation shown in FIG. 27, when the recognition result of the first voice spoken by the user U1 is not included in the first information IF1 and the number of times the recognition result of the second voice matches a specific tag TG included in the first information IF1 reaches a predetermined number, after step S55, the update unit 117 may associate the virtual object VO corresponding to the specific tag TG with the tag TG as the second name of the recognition result of the first voice.

[0159] 2.2.3: Third Operation FIG. 29 is a flowchart for explaining the third operation of the information processing system 1A.

[0160] In step S61, the processing device 11A functions as the voice recognition unit 116 to recognize the voice spoken by the user U1. Note that since the details of the operation are the same as those in step S41 in the first operation and step S51 in the second operation, the description thereof is omitted.

[0161] In step S62, the processing device 11A functions as the name identification unit 113-2A to determine whether the recognition result of the voice spoken by the user U1 corresponds to any of the plurality of tags TG included in the first information IF1. If the determination result is true, that is, if the determination result in step S62 is "YES", the processing device 11A executes the process of step S65. In this case, the recognition result of the voice spoken by the user U1 is an example of the first name described above. If the determination result is false, that is, if the determination result in step S62 is "NO", the processing device 11A executes the process of step S63. In this case, the recognition result of the voice spoken by the user U1 is an example of the second name described above.

[0162] In step S63, the processing device 11A functions as the name identification unit 113-2A to identify a plurality of tags TG corresponding to some or all of the plurality of virtual objects VO arranged in the virtual space VS based on the first information IF1.

[0163] In step S64, by functioning as the display control unit 114A, the processing device 11A causes the plurality of tags TG to be displayed near a part or all of the plurality of virtual objects VO corresponding to the plurality of tags TG specified in step S63. Here, the "vicinity" of a part or all of the plurality of virtual objects VO specifically refers to a range within a predetermined distance from each virtual object VO. Thereafter, the processing device 11A executes the process of step S61.

[0164] In step S65, by functioning as the display control unit 114A, the processing device 11A causes the display of the virtual object VO corresponding to the tag TG as the first call name, which is the recognition result in step S62, to be changed.

[0165] Note that, similar to the first operation shown in FIG. 27 and the second operation shown in FIG. 28, when the recognition result of the first voice spoken by the user U1 is not included in the first information IF1 and the number of times the recognition result of the second voice matches a specific tag TG included in the first information IF1 reaches a predetermined number, after step S65, the update unit 117 may associate the virtual object VO corresponding to the specific tag TG with the tag TG as the second call name, which is the recognition result of the first voice.

[0166] 2.3: Effects of the Second Embodiment According to the above description, the terminal device 10A as an information processing device includes a display control unit 114A and a call name specifying unit 113-2A. The display control unit 114A causes a plurality of virtual objects VO arranged in the virtual space VS to be displayed on the AR glasses 20 as a display device worn on the head of the user U1. The call name specifying unit 113-2A specifies the first call name that is most similar to the second call name among the plurality of first call names when the recognition result of the first voice spoken by the user U1 is a second call name that does not match any of the plurality of first call names corresponding to the plurality of virtual objects VO. The display control unit 114A causes the first call name specified by the call name specifying unit 113-2A to be displayed on the AR glasses 20 as a display device.

[0167] When the terminal device 10A as an information processing device uses the above configuration, even if the speech recognition result of the speech of the user U1 does not correspond to the tag TG included in the first information IF1, it is possible to recall the tag TG included in the first information IF1. In particular, the user U1 can recall one tag TG that is most similar to the recognition result of the speech he himself uttered.

[0168] Also, after the first call name specified by the call name specifying unit 113-2A is displayed on the AR glasses 20 as a display device, when the recognition result of the second speech uttered by the user U1 matches the specified first call name, the display control unit 114A causes the display related to the virtual object VO corresponding to the specified first call name among the plurality of virtual objects VO to be changed. Further, the terminal device 10A as an information processing device further includes an update unit 117 that associates the virtual object VO with a second call name when the number of times the recognition result of the first speech becomes the second call name reaches a predetermined number of times.

[0169] When the terminal device 10A as an information processing device uses the above configuration, when the number of times the user U1 attempts to call a certain virtual object VO using a second call name not included in the first information IF1 reaches a predetermined number of times, it becomes possible to associate the second call name with the virtual object VO.

[0170] According to the above description, the terminal device 10A as an information processing device includes a display control unit 114A and a name identification unit 113-2A. The display control unit 114A causes an AR glass 20 as a display device worn on the head of the user U1 to display a plurality of virtual objects VO arranged in the virtual space VS. The name identification unit 113-2A identifies a plurality of first names corresponding to some or all of the plurality of virtual objects VO. When the recognition result of the voice spoken by the user U1 is a second name that does not match any of the plurality of first names, the display control unit 114A associates each of the plurality of first names identified by the name identification unit 113-2A with the corresponding virtual object VO among some or all of the plurality of virtual objects VO and displays them.

[0171] When the terminal device 10A as an information processing device uses the above configuration, even when the voice recognition result of the user U1's speech does not correspond to the tag TG included in the first information IF1, the user U1 can recall the tag TG included in the first information IF1. In particular, the user U1 can visually recognize all of the tags TG included in the first information IF1 within the virtual space VS.

[0172] Also, after the plurality of first names identified by the name identification unit 113-2A are displayed on the AR glass 20 as a display device, when the recognition result of the second voice spoken by the user U1 matches any of the identified plurality of first names, the display control unit 114A changes the display related to the virtual object VO corresponding to the matching first name among the plurality of virtual objects VO. Further, the terminal device 10A as an information processing device further includes an update unit 117 that associates the virtual object VO with the second name when the number of times the recognition result of the first voice becomes the second name reaches a predetermined number of times.

[0173] By using the above configuration, when the number of times the user U1 attempts to call a certain virtual object VO using a second name that is not included in the first information IF1 reaches a predetermined number of times, the terminal device 10A as an information processing device can associate the second name with the virtual object VO.

[0174] 3: Modification Example The present disclosure is not limited to the embodiments illustrated above. Specific modification modes are illustrated below. Two or more modes arbitrarily selected from the following illustrations may be combined.

[0175] 3.1: Modification Example 1 The terminal device 10 according to the first embodiment includes a voice recognition unit 116 as a function of the processing device 11. Similarly, the terminal device 10A according to the second embodiment includes a voice recognition unit 116 as a function of the processing device 11A. However, the terminal devices 10 and 10A do not necessarily have to include the voice recognition unit 116. Specifically, the voice recognition unit 116 may be an external device of the terminal devices 10 and 10A and may be connected to the terminal devices 10 and 10A in a communicable manner. In this case, the voice recognition device corresponding to the voice recognition unit 116 may exist on the cloud and be communicably connected to the terminal devices 10 and 10A via the communication network NET.

[0176] 3.2: Modification Example 2 The terminal device 10 according to the first embodiment includes an acquisition unit 111 as a function of the processing device 11. The acquisition unit 111 acquires the first information IF1 and the second information IF2 from the storage device 12. Similarly, the terminal device 10A according to the second embodiment includes an acquisition unit 111 as a function of the processing device 11A. The acquisition unit 111 acquires the first information IF1 from the storage device 12A. However, the acquisition sources of the first information IF1 and the second information IF2 of the acquisition unit 111 do not have to be the storage devices 12 or 12A. Specifically, the acquisition unit 111 may directly acquire the first information IF1 and the second information IF2 from the server 30.

[0177] 3.3: Modification Example 3 The terminal device 10 according to the first embodiment includes an operation recognition unit 112 as a function of the processing device 11. Similarly, the terminal device 10A according to the second embodiment includes an operation recognition unit 112 as a function of the processing device 11A. The operation recognition unit 112 recognizes the gestures of the user U1. However, the method for recognizing the gestures of the user U1 is not limited to the above method. For example, the AR glasses 20 may recognize the gestures of the user U1 by including an operation recognition unit similar to the operation recognition unit 112.

[0178] 3.4: Modification Example 4 The terminal device 10 according to the first embodiment includes a name identification unit 113-2 as a function of the processing device 11. Similarly, the terminal device 10A according to the second embodiment includes a name identification unit 113-2A as a function of the processing device 11A. The name identification units 113-2 and 113-2A identify a tag TG as a name corresponding to the virtual object VO identified by the virtual object identification unit 113-1. On the other hand, the name identification units 113-2 and 113-2A do not identify the tag TG when there is no tag TG corresponding to the virtual object VO identified by the virtual object identification unit 113-1. In this case, the processing devices 11 and 11A may be provided with a function of setting a new tag TG for the virtual object VO for which the corresponding tag TG does not exist.

[0179] 3.5: Modification Example 5 In the information processing system 1 according to the first embodiment, the terminal device 10 and the AR glasses 20 are realized as separate entities. Similarly, in the information processing system 1A according to the second embodiment, the terminal device 10A and the AR glasses 20 are realized as separate entities. However, the method for realizing the terminal device 10 or 10A and the AR glasses 20 in the embodiments of the present invention is not limited to this. For example, the AR glasses 20 may have the same functions as the terminal device 10 or 10A, so that the terminal device 10 or 10A and the AR glasses 20 may be realized in a single housing.

[0180] 3.6: Modification Example 6 The terminal device 10A according to the second embodiment includes an update unit 117 as a function of the processing device 11A. When the recognition result of the first voice spoken by the user U1 is not included in the first information IF1 and the number of times the recognition result of the second voice matches a specific tag TG included in the first information IF1 reaches a predetermined number of times, the update unit 117 associates the virtual object VO corresponding to the specific tag TG with the tag TG as the second name called in the recognition result of the first voice. That is, the update unit 117 associates a plurality of tags TG with one virtual object VO. However, the operation of the update unit 117 is not limited to this. For example, instead of associating a plurality of tags TG with one virtual object VO, the update unit 117 may associate one tag TG with the one virtual object VO and set a set of the one virtual object VO and one tag TG for each type of tag TG.

[0181] 4: Others (1) In the above-described embodiments, the storage devices 12, 12A, 22, and 32 are exemplified by ROM and RAM, but flexible disks, magneto-optical disks (e.g., compact disks, digital versatile disks, Blu-ray (registered trademark) disks), smart cards, flash memory devices (e.g., cards, sticks, key drives), CD-ROMs (Compact Disc-ROMs), registers, removable disks, hard disks, floppy (registered trademark) disks, magnetic strips, databases, servers, and other appropriate storage media. Also, the program may be transmitted from a network via a telecommunication line. Also, the program may be transmitted from the communication network NET via a telecommunication line.

[0182] (2) In the above-described embodiments, the information, signals, etc. described may be represented using any of a variety of different technologies. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0183] (3) In the above-described embodiments, the input / output information, etc. may be stored in a specific location (e.g., memory), or may be managed using a management table. The input / output information, etc. may be overwritten, updated, or appended. The output information, etc. may be deleted. The input information, etc. may be transmitted to other devices.

[0184] (4) In the above-described embodiments, the determination may be made based on a value represented by 1 bit (0 or 1), may be made based on a boolean value (Boolean: true or false), or may be made by comparing numerical values (e.g., comparison with a predetermined value).

[0185] (5) The processing procedures, sequences, flowcharts, etc. exemplified in the above-described embodiments may be rearranged as long as there is no contradiction. For example, regarding the methods described in the present disclosure, the elements of various steps are presented using an exemplary order and are not limited to the specific order presented.

[0186] (6) Each function exemplified in FIGS. 1 to 29 is realized by any combination of at least one of hardware and software. Also, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one physically or logically combined device, or may be realized using two or more physically or logically separated devices directly or indirectly (e.g., using wired, wireless, etc.) connected, and these multiple devices. The functional block may be realized by combining software with the above one device or the above multiple devices.

[0187] (7) The programs exemplified in the above-described embodiments should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether called by the names of software, firmware, middleware, microcode, hardware description languages, or other names.

[0188] Also, software, instructions, information, etc. may be transmitted and received via a transmission medium. For example, when software is transmitted from a website, server, or other remote source using at least one of wired technologies (such as coaxial cables, optical fiber cables, twisted pairs, digital subscriber lines (DSL), etc.) and wireless technologies (such as infrared rays, microwaves, etc.), at least one of these wired and wireless technologies is included within the definition of the transmission medium.

[0189] (8) In each of the foregoing embodiments, the terms "system" and "network" are used interchangeably.

[0190] (9) The information, parameters, etc. described in the present disclosure may be represented using absolute values, relative values from a predetermined value, or corresponding other information.

[0191] (10) In the above-described embodiments, the terminal device 10, the terminal device 10A, and the server 30 may include cases where they are mobile stations (MS: Mobile Station). A mobile station may be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terms. Also, in the present disclosure, terms such as "mobile station", "user terminal", "user equipment (UE)", "terminal", etc. may be used interchangeably.

[0192] (11) In the above-described embodiments, the terms "connected" and "coupled", or any variations thereof, mean any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be a physical coupling or connection, a logical coupling or connection, or a combination thereof. For example, "connected" may be read as "accessed". As used in the present disclosure, two elements can be considered to be "connected" or "coupled" to each other using at least one of one or more wires, cables, and printed electrical connections, and also, as some non-limiting and non-exhaustive examples, electromagnetic energy having wavelengths in the radio frequency region, microwave region, and optical (both visible and invisible) region.

[0193] (12) In the above-described embodiments, the description "based on" does not mean "only based on" unless otherwise specified. In other words, the description "based on" means both "only based on" and "at least based on".

[0194] (13) As used in this disclosure, the terms "determining" may encompass a wide variety of operations. "Determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up (e.g., searching in a table, database, or another data structure), ascertaining, and considering something as having been "determined". Also, "determining" may include considering something as having been "determined" after receiving (e.g., receiving information), transmitting (e.g., transmitting information), inputting, outputting, accessing (e.g., accessing data in memory), etc. Further, "determining" may include considering something as having been "determined" after resolving, selecting, choosing, establishing, comparing, etc. That is, "determining" may include considering something as having been "determined" after performing some operation. Also, "determining" may be replaced with "assuming", "expecting", "considering", etc.

[0195] (14) In the above-described embodiments, when the terms "include", "including", and their variants are used, these terms are intended to be inclusive, similar to the term "comprising". Further, the term "or" used in this disclosure is not intended to be an exclusive disjunction.

[0196] In the present disclosure, for example, when articles are added by translation as in the case of a, an, and the in English, the present disclosure may include that the nouns following these articles are in the plural form.

[0197] (16) In the present disclosure, the term "A and B are different" may mean that "A and B are different from each other". Note that the term may also mean that "A and B are different from C respectively". Terms such as "separate", "coupled", etc. may also be interpreted in the same way as "different".

[0198] (17) Each aspect / embodiment described in the present disclosure may be used alone, in combination, or switched for use during execution. Also, the notification of predetermined information (for example, the notification of "being X") is not limited to the explicitly performed notification, and may be performed implicitly (for example, without performing the notification of the predetermined information).

[0199] As described above in detail about the present disclosure, it is obvious to those skilled in the art that the present disclosure is not limited to the embodiments described in the present disclosure. The present disclosure can be implemented as modified and changed aspects without departing from the spirit and scope of the present disclosure determined by the description of the claims. Therefore, the description of the present disclosure is for the purpose of illustrative explanation and has no restrictive meaning for the present disclosure.

Explanation of Reference Signs

[0200] 1, 1A... Information processing system, 10, 10A... Terminal device, 11, 11A... Processing device, 12, 12A... Memory device, 13... Communication device, 14... Display, 15... Input device, 20... AR glasses, 21... Processing device, 22... Memory device, 23... Line-of-sight acquisition device, 24... Sound collection device, 25... GPS device, 26... Motion detection device, 27... Imaging device, 28... Communication device, 29... Display, 30... Server, 31... Processing device, 32... Memory device, 33... Communication device, 34... Display, 35... Input device, 41L, 41R... Lenses, 91, 92... Temples, 93... Bridge, 94, 95... Trunks, 111... Acquisition unit, 112... Motion recognition unit, 113, 113A... Identification unit, 113-1... Virtual object identification unit, 113-2, 113-2A... Name identification unit, 114, 114A... Display control unit, 115... Judgment unit, 116... Voice recognition unit, 117... Update unit, 211... Acquisition unit, 212... Display control unit, 311... Output unit, 312... Acquisition unit, IF1... First information, IF2... Second information, LM1... Learning model, P1, P2, P3... Pop-up, PR1, PR2, PR3, PR4... Control program, R... Region, TG... Tag, U1, U2, U3... Users, VO... Virtual object

Claims

1. A display control unit that causes a display device worn on a user's head to display a plurality of virtual objects arranged in a virtual space; A virtual object specifying unit that specifies a first virtual object among the plurality of virtual objects based on instruction information generated according to the operation of the user; If a name corresponding to the first virtual object specified by the virtual object specifying unit is stored in the storage device, a name specifying unit that specifies the corresponding name as a first name; comprising The operation of the user is visual inspection of the display device, the instruction information indicates the user's viewpoint in the display device, The display device further includes a determination unit that determines whether or not the viewpoint is located within the first virtual object for a predetermined time or longer; When it is determined by the determination unit that the viewpoint is located within the first virtual object for the predetermined time or longer, the display control unit causes the display device to display the first name corresponding to the first virtual object. An information processing apparatus.

2. The display control unit causes the display device to display a two-dimensional image obtained by flattening the virtual space, and displays the first name in association with the first virtual object in the two-dimensional image. The information processing apparatus according to claim 1.

3. The display control unit causes the display device to display a three-dimensional image obtained by shrinking the virtual space, and displays the first name in association with the first virtual object in the three-dimensional image. The information processing apparatus according to claim 1.

4. When the recognition result of the voice spoken by the user matches the first name, the display control unit changes the display content related to the first virtual object. The information processing apparatus according to any one of claims 1 to 3.

5. A display control unit that causes a display device worn on a user's head to display a plurality of virtual objects arranged in a virtual space, When the recognition result of the first voice spoken by the user is a second name that does not match any of the plurality of first names corresponding to the plurality of virtual objects, among the plurality of first names, a name specifying unit that specifies the first name that is most similar to the second name Comprising The display control unit causes the display device to display the first name specified by the name specifying unit. An information processing device.

6. After the display control unit displays the first name specified by the name specifying unit on the display device, when the recognition result of the second voice spoken by the user matches the specified first name, the display control unit causes the display of the virtual object corresponding to the specified first name among the plurality of virtual objects to be changed. When the number of times the recognition result of the first voice becomes the second name reaches a predetermined number of times, the information processing device further includes an update unit that associates the virtual object with the second name. The information processing device according to claim 5.

7. A display control unit that causes a display device worn on a user's head to display a plurality of virtual objects arranged in a virtual space, A name specifying unit that specifies a plurality of first names corresponding to some or all of the plurality of virtual objects Comprising When the recognition result of the voice spoken by the user is a second name that does not match any of the plurality of first names, the display control unit causes each of the plurality of first names specified by the name specifying unit to be displayed in association with the corresponding virtual object among some or all of the plurality of virtual objects. An information processing device.

8. After a plurality of first names specified by the name specifying unit are displayed on the display device, when the recognition result of the second voice spoken by the user matches any of the plurality of specified first names, the display control unit causes the display of the virtual object corresponding to the matched first name among the plurality of virtual objects to be changed. When the number of times the recognition result of the first voice becomes the second name reaches a predetermined number of times, the information processing apparatus further includes an update unit that associates the virtual object with the second name. The information processing apparatus according to claim 7.

Citation Information

Patent Citations

  • Voice recognition device

    JP2012093422A

  • Information providing method, program, and information providing apparatus

    JP2019012536A

  • Recognition device, recognition method, and recognition program

    JP2020016784A

  • Information processing device

    JP6908953B1

  • Video display device and method

    WO2020110270A1