Wearable terminal device, program, and image processing method

The wearable terminal device enhances mixed reality experiences by capturing and interacting with virtual images through gestures, addressing the limitations of existing interfaces for VR and MR devices.

JP7724341B2Active Publication Date: 2025-08-15KYOCERA CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024138504
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2025-08-15
Estimated Expiration
2041-06-25

AI Technical Summary

Technical Problem

Existing wearable terminal devices for VR and MR lack intuitive and user-friendly interfaces for capturing and interacting with virtual images in real space, limiting the effectiveness of mixed reality experiences.

Method used

The wearable terminal device includes a camera to capture the user's visual field, a display unit, and processors to combine real and virtual images, allowing users to interact with virtual images by gestures and display them at predetermined positions in real space, with expandable or contractible captured images.

Benefits of technology

Enables intuitive user interaction and enhanced mixed reality experiences by allowing users to capture and manipulate virtual images within their real environment, improving usability and immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007724341000001
    Figure 0007724341000001
  • Figure 0007724341000002
    Figure 0007724341000002
  • Figure 0007724341000003
    Figure 0007724341000003
Patent Text Reader

Abstract

SOLUTION: A wearable terminal device worn and used by a user comprises a camera that captures a space as a user's visible area and at least one processor. The at least one processor identifies a part of the visible area in the space captured by the camera as a capture area based on a first gesture operation by the user, and stores the capture image corresponding to the capture area in a storage unit.EFFECT: A portion of a visible area desired by the user can be stored as a capture image in the storage unit, thus enhancing convenience for the user.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a wearable terminal device, a program, and an image processing method. [Background technology]

[0002] Conventionally, VR (virtual reality), MR (mixed reality), and AR (augmented reality) are known as technologies that allow a user to experience virtual images and / or virtual spaces using a wearable terminal device worn on the user's head. A wearable terminal device has a display unit that covers the user's field of vision when worn by the user. By displaying virtual images and / or virtual spaces on this display unit according to the user's position and orientation, a visual effect is achieved that makes it seem as if the virtual images and / or virtual spaces are actually present (for example, Patent Document 1 and Patent Document 2).

[0003] MR is a technology that allows users to experience mixed reality, which is a fusion of real space and virtual images, by displaying a virtual image that appears to exist in a specific location in real space while allowing the user to view the real space. VR is a technology that allows users to view a virtual space instead of the real space in MR, allowing the user to experience the feeling of being in the virtual space.

[0004] The virtual image displayed in VR and MR has a predetermined display position in the space where the user is located, and is displayed on the display unit and viewed by the user when the display position is within the user's visual field. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] US Patent Application Publication No. 2019 / 0087021 [Patent Document 2] US Patent Application Publication No. 2019 / 0340822 Summary of the Invention

[0006] The wearable terminal device of the present disclosure includes a camera that captures an image of a space as a user's visual field, a display unit that is visible to the user, and at least one processor, wherein the at least one processor displays a virtual image on the display unit so that the virtual image is visible in the space, and controls the camera to capture the image based on a predetermined operation by the user. a composite image obtained by combining the photographed image of the space and the virtual image, A captured image including a portion of the viewable area in the space and at least a portion of the virtual image is displayed. Extract Store in the memory The captured image is visually recognized as being placed at a predetermined position in the space, and is displayed on the display unit as a captured virtual image different from the virtual image. The captured virtual image is expanded or contracted in accordance with a predetermined gesture operation by a user on the displayed captured virtual image. When the captured virtual image is expanded, an image of the expanded portion is extracted from the composite image. .

[0007] The program of the present disclosure is executed by a computer capable of controlling a wearable terminal device including a camera that captures an image of a space as a user's visual field, a display unit that is visible to the user, and at least one processor. The program causes the computer to perform a process of displaying a virtual image on the display unit so that the virtual image is visible in the space, and a process of displaying a virtual image by the camera based on a predetermined operation of the user. a composite image obtained by combining the photographed image of the space and the virtual image, A captured image including a portion of the viewable area in the space and at least a portion of the virtual image is displayed. Extract A process of storing the data in a storage unit; a process of displaying the captured image on the display unit as a captured virtual image that is visually recognized as being placed at a predetermined position in the space and that is different from the virtual image; a process of expanding or contracting the captured virtual image in response to a predetermined gesture operation by a user on the displayed captured virtual image, and, when the captured virtual image is expanded, extracting an image of the expanded portion from the composite image; Execute the following.

[0008] The image processing method of the present disclosure is an image processing method executed by a computer capable of controlling a wearable terminal device including a camera that captures an image of a space as a user's visual field, a display unit that is visible to the user, and at least one processor. In the image processing method, a virtual image is displayed on the display unit so that the virtual image is visible in the space, and an image is captured by the camera based on a predetermined operation of the user. a composite image obtained by combining the photographed image of the space and the virtual image, A captured image including a portion of the viewable area in the space and at least a portion of the virtual image is displayed. Extract Store in the memory The captured image is visually recognized as being placed at a predetermined position in the space, and is displayed on the display unit as a captured virtual image different from the virtual image. The captured virtual image is expanded or contracted in accordance with a predetermined gesture operation by a user on the displayed captured virtual image. When the captured virtual image is expanded, an image of the expanded portion is extracted from the composite image. . [Brief explanation of the drawings]

[0009] [Figure 1]FIG. 1 is a schematic perspective view showing a configuration of a wearable terminal device. [Figure 2] 10A and 10B are diagrams illustrating an example of a visual recognition area and a virtual image visually recognized by a user wearing a wearable terminal device. [Figure 3] FIG. 1 is a diagram illustrating a visual recognition area in space. [Figure 4] FIG. 2 is a block diagram showing a main functional configuration of the wearable terminal device. [Figure 5] FIG. 10 is a diagram illustrating an example of a first gesture operation for specifying a capture region. [Figure 6] FIG. 10 is a diagram illustrating an example of a first gesture operation for specifying a capture region. [Figure 7] FIG. 1 is a diagram showing an image of the entire space captured by a camera. [Figure 8] 10A and 10B are diagrams illustrating a third gesture operation for displaying a menu virtual image. [Figure 9] FIG. 10 is a diagram showing an example of a viewable area in a state where a captured virtual image is displayed. [Figure 10] FIG. 10 is a diagram illustrating an example of a method for displaying a captured virtual image. [Figure 11] FIG. 10 is a diagram illustrating an example of a method for displaying a captured virtual image. [Figure 12] 10A and 10B are diagrams showing a captured virtual image with a highlighted outer frame and a method for moving the captured virtual image. [Figure 13] FIG. 1 illustrates a method for extending a captured virtual image. [Figure 14] 10A and 10B are diagrams illustrating the operation of removing a virtual image region from a captured virtual image. [Figure 15] 10A and 10B illustrate the operation of removing a spatial image region from a captured virtual image. [Figure 16] FIG. 10 is a diagram illustrating an operation of duplicating a virtual image included in a captured virtual image. [Figure 17] FIG. 10 is a diagram illustrating an operation of duplicating a virtual image included in a captured virtual image. [Figure 18]FIG. 10 is a diagram showing a movement method designation button for designating a movement method of a captured virtual image. [Figure 19] 10A and 10B are diagrams illustrating an operation of moving a virtual image within the frame of a captured virtual image. [Figure 20] 10A and 10B are diagrams illustrating an operation of moving a virtual image outside the frame of a captured virtual image. [Figure 21] 10A and 10B are diagrams illustrating an operation of extracting and displaying person information from a captured virtual image. [Figure 22] FIG. 10 is a diagram showing another example of the extracted information virtual image. [Figure 23] 10A and 10B are diagrams illustrating an operation of extracting and displaying information about an article from a captured virtual image. [Figure 24] 10A and 10B are diagrams illustrating an operation of extracting and displaying location information from a captured virtual image. [Figure 25] 10A and 10B are diagrams illustrating another example of an operation of extracting and displaying information from a captured virtual image. [Figure 26] 10A and 10B are diagrams illustrating an operation of extracting and displaying text information from a captured virtual image. [Figure 27] 10A to 10C are diagrams illustrating operations when various processes are executed on extracted character information. [Figure 28] 10A and 10B are diagrams illustrating an operation of extracting and displaying two-dimensional code information from a captured virtual image. [Figure 29] 10A and 10B are diagrams illustrating an operation of displaying position information when a capture image is generated. [Figure 30] 10A and 10B are diagrams illustrating an operation of displaying user information when a capture image is generated. [Figure 31] FIG. 10 is a diagram showing a display operation when capture is prohibited. [Figure 32] 10A and 10B are diagrams illustrating a display operation when a captured image is stored in association with audio. [Figure 33] FIG. 10 is a diagram showing a captured virtual image in a case where audio is associated with the captured image. [Figure 34] FIG. 10 shows a captured virtual image displayed on the surface of another object. [Figure 35] FIG. 10 is a diagram showing a captured virtual image that moves together with a display object. [Figure 36] FIG. 10 is a diagram showing a captured virtual image rotating together with a display object. [Figure 37] 10 is a flowchart showing a control procedure for a captured virtual image display process. [Figure 38] 10 is a flowchart showing a control procedure for a captured virtual image display process. [Figure 39] FIG. 10 is a schematic diagram illustrating the configuration of a display system according to a second embodiment. [Figure 40] FIG. 2 is a block diagram showing a main functional configuration of an external device. [Figure 41] FIG. 10 is a diagram illustrating a viewable area of a wearable terminal device during screen sharing. [Figure 42] FIG. 10 is a diagram showing a dialog image displayed on an instructor screen of an external device during screen sharing. [Figure 43] 10A and 10B are diagrams illustrating a virtual dialog image displayed in the viewable area of a wearable terminal device during screen sharing. [Figure 44] FIG. 10 is a schematic diagram illustrating the configuration of a display system according to a third embodiment. [Figure 45] FIG. 2 is a block diagram showing a main functional configuration of the information processing device. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments will be described with reference to the drawings. However, for the sake of convenience, the drawings referred to below show only the main components necessary for explaining the embodiments in a simplified form. Therefore, the wearable terminal device 10, external device 20, and information processing device 80 of the present disclosure may include any components not shown in the drawings referred to.

[0011] [First embodiment] As shown in FIG. 1, the wearable terminal device 10 includes a main body 10a, a visor 141 (display member) attached to the main body 10a, and the like.

[0012] The main body 10a is an annular member whose circumference is adjustable. Various devices such as a depth sensor 153 and a camera 154 are built into the main body 10a. When the main body 10a is worn on the head, the user's field of vision is covered by a visor 141.

[0013] The visor 141 is optically transparent. The user can view the real space through the visor 141. An image such as a virtual image is projected and displayed on a display surface of the visor 141 facing the user's eyes from a laser scanner 142 (see FIG. 4) built into the main body 10a. The user views the virtual image through light reflected from the display surface. At this time, the user also views the real space through the visor 141, providing a visual effect as if the virtual image were present in the real space.

[0014] As shown in FIG. 2, when a virtual image 30 (first virtual image) is displayed, the user views the virtual image 30 facing a predetermined direction at a predetermined position in a space 40. In this embodiment, the space 40 is a real space viewed by the user through a visor 141. The virtual image 30 is projected onto the optically transparent visor 141, and is therefore viewed as a translucent image superimposed on the real space. While FIG. 2 illustrates a planar window screen as an example of the virtual image 30, the present invention is not limited thereto. The virtual image 30 may be, for example, an object such as an arrow, or various types of three-dimensional images (three-dimensional virtual objects). When the virtual image 30 is a window screen, the virtual image 30 has a front surface (first surface) and a back surface (second surface), and necessary information is displayed on the front surface, and normally, no information is displayed on the back surface.

[0015] The wearable terminal device 10 detects the user's visual field 41 based on the position and orientation of the user in the space 40 (in other words, the position and orientation of the wearable terminal device 10). As shown in FIG. 3 , the visual field 41 is a region in the space 40 that is located in front of the user U wearing the wearable terminal device 10. For example, the visual field 41 is a region within a predetermined angular range in the left-right and up-down directions from the front of the user U. In this case, when a three-dimensional object corresponding to the shape of the visual field 41 is cut by a plane perpendicular to the front direction of the user U, the shape of the cut surface is rectangular. Note that the shape of the visual field 41 may be determined so that the shape of the cut surface is other than rectangular (for example, circular or elliptical). The shape of the visual field 41 (for example, the angular range in the left-right and up-down directions from the front) can be identified by, for example, the following method.

[0016] In the wearable terminal device 10, adjustment of the field of view (hereinafter referred to as calibration) is performed according to a predetermined procedure at a predetermined timing, such as when the device is first started up. This calibration specifies the range that can be seen by the user, and thereafter, the virtual image 30 is displayed within this range. The shape of the visible range specified by this calibration can be set as the shape of the visible area 41.

[0017] Furthermore, calibration is not limited to being performed according to the above-described predetermined procedure, and calibration may be performed automatically during normal operation of the wearable terminal device 10. For example, if the user does not react to a display that should elicit a reaction, the range in which the display is being performed may be considered to be outside the user's field of view, and the field of view (and the shape of the visible area 41) may be adjusted. Alternatively, a test display may be performed at a position that is determined to be outside the field of view, and if the user reacts to the display, the range in which the display is being performed may be considered to be within the user's field of view, and the field of view (and the shape of the visible area 41) may be adjusted.

[0018] The shape of visible area 41 may be predetermined and fixed at the time of shipment, etc., without being based on the results of adjusting the field of view. For example, the shape of visible area 41 may be determined to be the maximum displayable range in terms of the optical design of display unit 14.

[0019] The virtual image 30 is generated in a state where the display position and orientation in the space 40 are determined in response to a predetermined operation by the user. The wearable terminal device 10 projects and displays, on the visor 141, the virtual image 30 whose display position is determined to be within the visible area 41, among the generated virtual images 30. In FIG. 2, the visible area 41 is indicated by a dashed line.

[0020] The display position and orientation of the virtual image 30 on the visor 141 are updated in real time in response to changes in the user's viewing area 41. That is, the display position and orientation of the virtual image 30 change in response to changes in the viewing area 41 so that the user recognizes that "the virtual image 30 is located in the space 40 at a set position and orientation." For example, when the user moves from the front side of the virtual image 30 to the back side, the shape (angle) of the displayed virtual image 30 gradually changes in response to this movement. Furthermore, when the user turns toward the back side of the virtual image 30 after going around to the back side of the virtual image 30, the back side of the virtual image 30 is displayed so that the back side of the virtual image 30 is visible. Furthermore, in response to changes in the viewing area 41, any virtual image 30 whose display position is outside the viewing area 41 is no longer displayed, and if any virtual image 30 whose display position is within the viewing area 41 is present, that virtual image 30 is newly displayed.

[0021] 2, when a user holds out their hand (or finger) in front of them, the direction in which the hand is extended is detected by the wearable terminal device 10, and a virtual line 411 extending in that direction and a pointer 412 are displayed on the display surface of the visor 141 for the user to see. The pointer 412 is displayed at the intersection of the virtual line 411 and the virtual image 30. If the virtual line 411 does not intersect with the virtual image 30, the pointer 412 may be displayed at the intersection of the virtual line 411 and a wall surface or the like of the space 40. If the distance between the user's hand and the virtual image 30 is within a predetermined reference distance, the display of the virtual line 411 may be omitted, and the pointer 412 may be displayed directly at a position corresponding to the position of the user's fingertip.

[0022] By changing the direction in which the user extends his / her hand, the direction of the virtual line 411 and the position of the pointer 412 can be adjusted. By performing a predetermined gesture while adjusting the pointer 412 so that it is positioned over a predetermined operation object (for example, the function bar 31, the window shape change button 32, the close button 33, etc.) included in the virtual image 30, the wearable terminal device 10 detects the gesture and allows the user to perform a predetermined operation on the operation object. For example, by performing a gesture to select the operation object (for example, a gesture of pinching the fingertips) while the pointer 412 is positioned over the close button 33, the virtual image 30 can be closed (deleted). Furthermore, by performing a selection gesture while the pointer 412 is positioned over the function bar 31 and then performing a gesture of moving the hand back and forth or left and right while the selection state is maintained, the virtual image 30 can be moved in the depth direction and left and right directions. Operations on the virtual image 30 are not limited to these.

[0023] In this way, the wearable terminal device 10 of this embodiment realizes a visual effect as if the virtual image 30 exists in real space, and can accept user operations on the virtual image 30 and reflect them in the display of the virtual image 30. In other words, the wearable terminal device 10 of this embodiment provides MR.

[0024] Next, the functional configuration of the wearable terminal device 10 will be described with reference to FIG. The wearable terminal device 10 includes a CPU 11 (Central Processing Unit), a RAM 12 (Random Access Memory), a storage unit 13, a display unit 14, a sensor unit 15, a communication unit 16, a microphone 17, and a speaker 18, and these units are connected via a bus 19. Of the components shown in FIG. 4, all units except for the visor 141 of the display unit 14 are built into the main body unit 10a and operate using power supplied from a battery also built into the main body unit 10a.

[0025] The CPU 11 is a processor that performs various calculation processes and controls the overall operation of each part of the wearable terminal device 10. The CPU 11 performs various control operations by reading and executing a program 131 stored in the storage unit 13. By executing the program 131, the CPU 11 performs, for example, a visual recognition area detection process and a display control process. Of these, the visual recognition area detection process is a process of detecting the visual recognition area 41 of the user within the space 40. Furthermore, the display control process is a process of displaying, on the display unit 14, virtual images 30 whose positions in the space 40 are determined to be within the visual recognition area 41.

[0026] 4 shows a single CPU 11, but is not limited to this. Two or more processors such as CPUs may be provided, and the processing executed by the CPU 11 of this embodiment may be shared and executed by these two or more processors.

[0027] The RAM 12 provides a working memory space for the CPU 11 and stores temporary data.

[0028] The storage unit 13 is a non-transitory recording medium readable by the CPU 11 as a computer. The storage unit 13 stores a program 131 executed by the CPU 11, various setting data, and the like. The program 131 is stored in the storage unit 13 in the form of computer-readable program code. The storage unit 13 may be, for example, a non-volatile storage device such as an SSD (Solid State Drive) equipped with a flash memory.

[0029] The data stored in the storage unit 13 includes virtual image data 132 related to the virtual image 30. The virtual image data 132 includes data related to the display content of the virtual image 30 (e.g., image data), data on the display position, and data on the orientation.

[0030] The display unit 14 includes a visor 141, a laser scanner 142, and an optical system that guides light output from the laser scanner 142 onto the display surface of the visor 141. The laser scanner 142 irradiates the optical system with pulsed laser light, which is on / off controlled for each pixel, in a predetermined direction while scanning the optical system in accordance with a control signal from the CPU 11. The laser light incident on the optical system forms a display screen consisting of a two-dimensional pixel matrix on the display surface of the visor 141. The type of the laser scanner 142 is not particularly limited, and for example, a type in which a mirror is operated by a MEMS (microelectromechanical system) to scan the laser light can be used. The laser scanner 142 has three light-emitting elements that emit, for example, RGB laser light. The display unit 14 can display color by projecting light from these light-emitting elements onto the visor 141.

[0031] The sensor unit 15 includes an acceleration sensor 151, an angular velocity sensor 152, a depth sensor 153, a camera 154, and an eye tracker 155. The sensor unit 15 may further include sensors not shown in FIG.

[0032] The acceleration sensor 151 detects acceleration and outputs the detection result to the CPU 11. From the detection result by the acceleration sensor 151, translational movement of the wearable terminal device 10 in three orthogonal axial directions can be detected.

[0033] The angular velocity sensor 152 (gyro sensor) detects the angular velocity and outputs the detection result to the CPU 11. From the detection result by the angular velocity sensor 152, the rotational movement of the wearable terminal device 10 can be detected.

[0034] The depth sensor 153 is an infrared camera that detects the distance to a subject using a ToF (Time of Flight) method, and outputs the distance detection result to the CPU 11. The depth sensor 153 is provided on the front surface of the main body 10a so that it can capture an image of the visible area 41. By repeatedly performing measurements using the depth sensor 153 each time the position and orientation of the user changes in the space 40 and combining the results, it is possible to perform three-dimensional mapping of the entire space 40 (i.e., to obtain the three-dimensional structure).

[0035] The camera 154 captures an image of the space 40 using a group of RGB image sensors, acquires color image data as the image capture result, and outputs the acquired color image data to the CPU 11. The camera 154 is provided on the front surface of the main body 10a so as to capture an image of the space 40 as the visible area 41. The image of the space 40 captured by the camera 154 is used to detect the position and orientation of the wearable terminal device 10, and is also transmitted from the communication unit 16 to an external device and used to display the visible area 41 of the user of the wearable terminal device 10 on the external device. The image of the space 40 captured by the camera 154 is also used as the image of the visible area 41 when the visible area 41 is stored as a captured image, as will be described later.

[0036] The eye tracker 155 detects the user's line of sight and outputs the detection result to the CPU 11. The method for detecting the line of sight is not particularly limited, but for example, a method can be used in which an eye tracking camera captures an image of the reflection point of near-infrared light on the user's eye, and the image captured by the camera 154 is analyzed to identify the object the user is looking at. A part of the configuration of the eye tracker 155 may be provided on the periphery of the visor 141, for example.

[0037] The communication unit 16 is a communication module including an antenna, a modulation / demodulation circuit, a signal processing circuit, etc. The communication unit 16 transmits and receives data to and from an external device via wireless communication in accordance with a predetermined communication protocol. The communication unit 16 can also communicate audio data with the external device. That is, the communication unit 16 transmits audio data collected by the microphone 17 to the external device and receives audio data transmitted from the external device to output audio from the speaker 18.

[0038] The microphone 17 converts sounds such as the user's voice into electrical signals and outputs them to the CPU 11.

[0039] The speaker 18 converts the input audio data into mechanical vibrations and outputs the vibrations as sound.

[0040] In the wearable terminal device 10 configured as above, the CPU 11 performs the following control operations.

[0041] The CPU 11 performs three-dimensional mapping of the space 40 based on distance data to the subject input from the depth sensor 153. The CPU 11 repeatedly performs this three-dimensional mapping each time the user's position and orientation change, updating the results each time. The CPU 11 also performs three-dimensional mapping for each continuous space 40. Therefore, when the user moves between multiple rooms separated by walls or the like, the CPU 11 recognizes each room as a single space 40 and performs three-dimensional mapping for each room separately.

[0042] The CPU 11 detects the user's visual recognition area 41 within the space 40. Specifically, the CPU 11 identifies the position and orientation of the user (wearable terminal device 10) in the space 40 based on the detection results of the acceleration sensor 151, angular velocity sensor 152, depth sensor 153, camera 154, and eye tracker 155, and the accumulated results of three-dimensional mapping. Then, the CPU 11 detects (identifies) the visual recognition area 41 based on the identified position and orientation and the predetermined shape of the visual recognition area 41. The CPU 11 also continuously detects the user's position and orientation in real time, and updates the visual recognition area 41 in conjunction with changes in the user's position and orientation. Note that the detection of the visual recognition area 41 may be performed using some of the detection results of the acceleration sensor 151, angular velocity sensor 152, depth sensor 153, camera 154, and eye tracker 155.

[0043] In response to a user operation, the CPU 11 generates virtual image data 132 related to the virtual image 30. That is, when the CPU 11 detects a predetermined operation (gesture) instructing the generation of the virtual image 30, the CPU 11 specifies the display content (e.g., image data), display position, and orientation of the virtual image, and generates virtual image data 132 including data representing the results of these specifications.

[0044] CPU 11 causes display unit 14 to display virtual image 30, the display position of which is determined within visible area 41. CPU 11 identifies virtual image 30 based on display position information included in virtual image data 132, and generates image data for a display screen to be displayed on display unit 14 based on the positional relationship between visible area 41 and the display position of virtual image 30 at that time. CPU 11 causes laser scanner 142 to perform a scanning operation based on this image data, and forms a display screen including the virtual image on the display surface of visor 141. In other words, CPU 11 displays virtual image 30 on the display surface of visor 141 so that virtual image 30 is visible in space 40 visible through visor 141. By continuously performing this display control process, CPU 11 updates the display content on display unit 14 in real time in accordance with the user's movement (changes in visible area 41). If the wearable terminal device 10 is set to retain the virtual image data 132 even when the power is turned off, the next time the wearable terminal device 10 is started up, the existing virtual image data 132 is read, and if there is a virtual image 30 within the visible area 41, it is displayed on the display unit 14.

[0045] Note that the virtual image data 132 may be generated based on instruction data acquired from an external device via the communication unit 16, and the virtual image 30 may be displayed based on the virtual image data 132. Alternatively, the virtual image data 132 itself may be acquired from an external device via the communication unit 16, and the virtual image 30 may be displayed based on the virtual image data 132. For example, an image from the camera 154 of the wearable terminal device 10 may be displayed on an external device operated by a remote instructor, and an instruction to display the virtual image 30 may be received from the external device, and the instructed virtual image 30 may be displayed on the display unit 14 of the wearable terminal device 10. This makes it possible, for example, to display a virtual image 30 indicating the content of a task near a task target, and to instruct the user of the wearable terminal device 10 to perform the task.

[0046] CPU 11 detects the position and orientation of the user's hand (and / or finger) based on the images captured by depth sensor 153 and camera 154, and displays a virtual line 411 extending in the detected direction and a pointer 412 on display unit 14. CPU 11 also detects a gesture of the user's hand (and / or finger) based on the images captured by depth sensor 153 and camera 154, and executes processing according to the content of the detected gesture and the position of pointer 412 at that time.

[0047] Next, the operation of the wearable terminal device 10 will be described, focusing on the operation of capturing the visible area 41.

[0048] The wearable terminal device 10 is equipped with a camera 154, and can capture the user's visual field 41 at that time by having the camera 154 photograph the space 40 at a timing according to the user's operation and storing the photographed image. However, the simple method of simply storing an image of the entire visual field 41 photographed by the camera 154 does not necessarily allow the user to use the captured image for the intended purpose, and is therefore inconvenient. For this reason, there has been a demand for an improved user interface that takes user convenience into account for the function of capturing the visual field 41 in the wearable terminal device 10.

[0049] In contrast, the wearable terminal device 10 of the present disclosure is equipped with various functions related to capturing the visible area 41. Below, operations related to these functions and the processing executed by the CPU 11 to realize these operations will be described.

[0050] As shown in FIGS. 5 and 6 , the CPU 11 of the wearable terminal device 10 of the present disclosure specifies a part of the visual recognition area 41 in the space 40 captured by the camera 154 as a capture area R (see FIGS. 6 and 7 ) based on the user's first gesture operation, and stores a capture image C corresponding to the specified capture area R in the storage unit 13. This allows a part of the visual recognition area 41 desired by the user to be stored in the storage unit 13 as the capture image C, thereby improving user convenience. The capture image C includes an image of a part of the visual recognition area 41 that corresponds to the capture area R. Furthermore, if a virtual image 30 is displayed in the visual recognition area 41 and the virtual image 30 is included in the capture area R, the virtual image 30 is also reflected in the capture image C. Therefore, the capture image C is an image that is a cutout of a part of the field of view that the user is viewing through the visor 141 when the capture area R is specified.

[0051] The first gesture operation for specifying the capture area R may be a gesture operation in which the user moves the hand so that the trajectory of a predetermined part of the user's hand surrounds a part of the viewable area 41, as shown in FIG. 5 . In this case, the area surrounded by the trajectory of the predetermined part of the user's hand can be specified as the capture area R. The first gesture operation may be, for example, a gesture operation in which at least one finger is held up and the capture area R is surrounded by the trajectory of the fingertip. By holding up the finger, the first gesture operation can be distinguished from other gesture operations and false detections can be reduced. The first gesture operation may also be an operation using both hands. Alternatively, the area surrounded by the trajectory of the pointer 412, instead of the user's hand or finger, may be specified as the capture area R. In FIG. 5 , a range of the viewable area 41 that includes the person 44 and a part of the virtual image 30, is specified as the capture area R.

[0052] As shown in FIG. 6 , the first gesture operation for specifying the capture region R may be a gesture operation for moving a capture frame r of a preset size to a desired position in the visible region 41 and confirming the position. The operation for moving the capture frame r may be an operation for moving the user's hand (or pointer) while aligning the position of the user's finger (or pointer 412) with the capture frame r. The operation for confirming the position of the capture frame r may also be an operation for tapping the air with the user's hand (finger). Here, the tapping operation may be an operation of moving the hand (finger) away from the user and then moving it closer, repeated twice. The size of the capture frame r may also be changed by a gesture operation for moving the hand while pinching a part of the capture frame r with the fingers (or selecting it with the pointer 412). Alternatively, two or more capture frames r of different sizes may be displayed, and the user may select one of the capture frames r. Once the position of the capture frame r is confirmed, the area surrounded by the capture frame r is specified as the capture region R.

[0053] When the capture area R is identified, the CPU 11 stores in the storage unit 13 a captured image D (see FIG. 7) of the entire space 40 captured by the camera 154 at that time. The CPU 11 also stores in the storage unit 13 a composite image E (corresponding to the viewable area 41 shown in the upper diagrams of FIGS. 5 and 6) obtained by combining the captured image D with the virtual image 30 displayed on the display unit 14. The CPU 11 also extracts a partial image corresponding to the capture area R from the composite image E and stores it in the storage unit 13 as a captured image C. In this way, when at least a portion of the virtual image 30 serving as the first virtual image is included in the capture area, the CPU 11 stores in the storage unit 13 a captured image C including at least a portion of the virtual image 30. This makes it possible to store the user's field of view, including the virtual image 30, as the captured image C.

[0054] When performing the first gesture operation, the user first performs a predetermined operation to transition the wearable terminal device 10 to a state in which the first gesture operation can be accepted. For example, as shown in FIG. 8, the user performs a predetermined third gesture operation to cause a menu virtual image 61 (third virtual image) to be displayed on the display unit 14. Then, in the menu virtual image 61, the user selects a capture operation start button 611 or 612 to start accepting the first gesture operation. The third gesture operation to display the menu virtual image 61 may be a gesture operation in which both hands are moved to a predetermined positional relationship. For example, as shown in the upper left of FIG. 8, the third gesture operation may be a gesture operation in which the fingers of the right hand point at the wrist of the left hand. When the CPU 11 detects this third gesture operation, the CPU 11 causes the display unit 14 to display the menu virtual image 61. The third gesture operation may also be a gesture operation in which the hand (or the pointer 412) is moved in a predetermined direction from outside the visible area 41 to within the visible area 41, as shown in the upper right of FIG. 8. When the CPU 11 detects this third gesture operation, it slides in the menu virtual image 61 (from the left side in the example of FIG. 8) so as to follow the hand movement, and displays it on the display unit 14. In this way, in response to the user's third gesture operation, the CPU 11 displays the menu virtual image 61 as a third virtual image for starting to accept the first gesture operation on the display unit 14. This makes it possible to easily start the capture operation.

[0055] When a gesture operation to select the capture operation start button 611 of the menu virtual image 61 is performed, the CPU 11 accepts a first gesture operation to surround the capture area R with a fingertip or the like, as shown in Fig. 5. Furthermore, when a gesture operation to select the capture operation start button 612 is performed, the CPU 11 accepts a first gesture operation to specify the capture area R by a capture frame r, as shown in Fig. 6. In this way, the CPU 11 specifies the capture area R using different methods depending on the type of first gesture operation, and specifies the type of first gesture operation to be accepted depending on the operation on the menu virtual image 61 as the third virtual image. This allows the user to select a desired first gesture operation to specify the capture area R.

[0056] It should be noted that the state may be changed to one in which the first gesture operation is accepted without going through the menu virtual image 61. For example, the state may be changed to one in which the first gesture operation is accepted in response to the operation of tapping the air with the user's hand (finger) as described above. Also, an icon for starting capture may be provided on the function bar 31 of the virtual image 30, and the state may be changed to one in which the first gesture operation is accepted in response to the operation of selecting the icon.

[0057] As shown in FIG. 9 , the capture image C stored in the storage unit 13 can be displayed on the display unit 14 as a capture virtual image 50 (second virtual image). When the CPU 11 stores the capture image C in the storage unit 13, the CPU 11 may cause the display unit 14 to display the capture image C as the capture virtual image 50 (second virtual image). This allows the capture virtual image 50 to be automatically displayed without the user performing a particular gesture operation. Furthermore, the CPU 11 may cause the display unit 14 to display the capture image C as the capture virtual image 50 (second virtual image) in response to the user's second gesture operation. This allows the user to display the capture virtual image 50 at a desired position at an intended timing.

[0058] 9, when virtual image 30 is displayed as the first virtual image before displaying captured virtual image 50 as the second virtual image, CPU 11 may display captured virtual image 50 at a position closer to the user in space 40 than virtual image 30. This allows captured virtual image 50 to be displayed in a manner that allows the entire captured virtual image 50 to be viewed.

[0059] When displaying the capture virtual image 50 in response to the second gesture operation, the second gesture operation is not particularly limited as long as it corresponds to a hand or finger movement. The second gesture operation may be, for example, a specific movement of the hand (left hand and / or right hand), a specific movement of the fingers, opening and closing of the fingers, or a combination thereof. The second gesture operation may also be a movement of the pointer 412. A capture virtual image 50 may be displayed in response to a plurality of different types of second gesture operations. When the second gesture operation is performed, the CPU 11 displays the capture virtual image 50 at a relative position to the user that is predetermined in accordance with the type of the second gesture operation. This relative position may be, for example, near the palm or fingers, or at the user's line of sight. FIG. 10 illustrates an example in which the capture virtual image 50 is displayed near the fingertips of the left hand in response to a second gesture operation in which the fingers of the right hand point at the wrist of the left hand. In this way, the CPU 11 displays the capture virtual image 50 as the second virtual image at a predetermined relative position to the user in response to the second gesture operation, and the relative position is set in advance for each type of second gesture operation. This allows the capture virtual image 50 to be displayed at a desired position with a simple operation.

[0060] The method of displaying the capture virtual image 50 is not limited to the above, and may be, for example, a method of selecting one of one or more capture images C stored in the storage unit 13 to be displayed as the capture virtual image 50, as shown in FIG. 11 . In the example of FIG. 11 , as shown in the upper diagram, by performing a gesture operation of moving a hand (or a pointer 412) from the outside (right side) of the viewable area 41 into the viewable area 41, a list area L of capture images C slides in and is displayed from the right end of the viewable area 41 so as to follow the movement of the hand. Next, as shown in the lower diagram of FIG. 11 , by performing a gesture operation of placing a finger on one of the capture images C in the list area L (or selecting it with the pointer 412) and dragging it into the viewable area 41, the capture image C can be displayed on the display unit 14 as the capture virtual image 50. If the list area L contains only one capture image C, the list area L may be erased in response to the drag operation. If the list area L includes two or more capture images C, the list area L may be displayed even after the drag operation, allowing the user to continue dragging other capture images C. The image representing the list area L may also be determined based on a predetermined condition. For example, the CPU 11 may display a specific image associated with a user ID used when logging in to the wearable terminal device 10 as the image representing the list area L, based on the user ID.

[0061] 12, the CPU 11 may display the outer frame of the captured virtual image 50 as the second virtual image in a predetermined highlighted manner. This makes it easier to view the captured virtual image 50. One example of the highlighted manner is to make the outer frame of the captured virtual image 50 thicker than the outer frame of the virtual image 30, as shown in FIG. 12, but is not limited to this.

[0062] 12, the CPU 11 moves the captured virtual image 50 as the displayed second virtual image in response to a fourth gesture operation by the user on the captured virtual image 50. This allows the user to arbitrarily change the display position of the captured virtual image 50. The fourth gesture operation is not particularly limited, but may be an operation of moving a hand while selecting a predetermined portion (for example, the upper function bar) or an arbitrary portion of the captured virtual image 50.

[0063] As shown in FIG. 13 , the user can expand the display area of the capture virtual image 50 by performing a predetermined fifth gesture operation. The fifth gesture operation may be, for example, an operation of moving a hand while pinching a part of the outer frame of the capture virtual image 50 with fingers (or while selecting it with the pointer 412). The range of the capture area R reflected in the capture virtual image 50 is expanded in response to the expansion of the capture virtual image 50. The expanded capture virtual image 50 includes the initial display area 50a before expansion and the expanded expanded area 50b. The image of the expanded area 50b may be extracted from the composite image E of the entire viewable area 41 that was stored when the capture image C was generated. In other words, a part of the composite image E corresponding to the expanded capture area R is extracted to generate a new capture image C, and the expanded capture virtual image 50 is displayed based on the new capture image C. It should be noted that the capture virtual image 50 can also be reduced in response to the fifth gesture operation. In this way, the CPU 11 expands or reduces the capture virtual image 50 in response to the user's fifth gesture operation on the captured virtual image 50 as the displayed second virtual image, and expands or reduces the range of the capture region R reflected in the captured virtual image 50 in response to the expansion or contraction. This makes it possible to change the capture range even after the captured virtual image 50 is displayed.

[0064] As shown in the upper diagram of FIG. 14 , when a captured virtual image 50 includes a virtual image area 51 corresponding to the virtual image 30 and a spatial image area 52 corresponding to the background space 40, the virtual image area 51 can be selectively deleted as shown in the lower diagram. Hereinafter, for convenience, the virtual image area 51 in the captured virtual image 50 may be referred to as the “virtual image 30 included in the captured virtual image 50.” The virtual image area 51 is deleted in response to a sixth gesture operation by the user. The sixth gesture operation is not particularly limited, and may be, for example, a double-tap or long-press operation on the virtual image area 51. After the virtual image area 51 is deleted from the captured virtual image 50, the background space 40 is displayed in the area where the virtual image area 51 was displayed before the deletion. The image of the space 40 is extracted from a captured image D of the space 40 (see FIG. 7 ) that was stored when the captured image C was generated. In this way, in response to the user's sixth gesture operation on the captured virtual image 50 serving as the displayed second virtual image, the CPU 11 deletes the virtual image 30 serving as the first virtual image included in the captured virtual image 50 and displays an image of the space 40 corresponding to the deleted area in the deleted area. This allows the virtual image 30 to be deleted from the captured virtual image 50 even after the captured virtual image 50 has been displayed. In this case, the CPU 11 may display the virtual image 30 included in the captured virtual image 50 in a predetermined highlighted manner. This makes it easier to visually recognize the virtual image area 51 to be deleted. An example of the highlighted manner is, but is not limited to, thickening the outer frame of the virtual image area 51 as shown in the upper diagram of FIG. 14 .

[0065] As shown in the upper diagram of FIG. 15 , when a captured virtual image 50 includes a virtual image area 51 and a spatial image area 52, the spatial image area 52 can be selectively deleted from the captured virtual image 50 as shown in the lower diagram. The deletion of the virtual image area 51 is performed in response to a seventh gesture operation by the user. The seventh gesture operation is not particularly limited, and may be, for example, a double-tap or long-press operation on the spatial image area 52. The captured virtual image 50 after the spatial image area 52 has been deleted may include only the virtual image area 51. The image of the virtual image area 51 may be extracted from the captured image C before the deletion. Alternatively, the image of the virtual image area 51 may be extracted from the virtual image data 132 stored in the storage unit 13. That is, an ID for identifying the virtual image 30 may be stored when the captured image C is stored, and the image data of the virtual image area 51 may be identified and acquired from the virtual image data 132 by referring to the ID. In this way, the CPU 11 deletes the portion of the capture virtual image 50 other than the virtual image 30 in response to the user's seventh gesture operation on the captured virtual image 50 as the displayed second virtual image. This allows the background space 40 to be deleted from the captured virtual image 50 even after the capture virtual image 50 has been displayed. In this case, the CPU 11 may display the portion of the capture virtual image 50 other than the virtual image 30 (spatial image area 52) in a predetermined highlighted manner. This makes it easier to visually recognize the spatial image area 52 to be deleted. An example of the highlighted manner is to thicken the outer frame of the spatial image area 52 as shown in the upper diagram of FIG. 15 , but is not limited to this.

[0066] 16, in response to a tenth gesture operation by the user on the captured virtual image 50 as the displayed second virtual image, the CPU 11 may copy at least a part of the captured virtual image 50 to display a virtual image 34 as the fifth virtual image. This allows a part of the captured virtual image 50 to be handled as a separate virtual image 34, thereby improving user convenience.

[0067] From another perspective, as shown in the upper diagram of FIG. 16 , when a portion of the virtual image 30 serving as the first virtual image is included in the captured virtual image 50 serving as the second virtual image, the CPU 11 may duplicate the portion of the virtual image 30 to display a virtual image 34 serving as the fourth virtual image in response to an eighth gesture operation by the user on the displayed virtual image 30. This allows a portion of the virtual image 30 included in the captured virtual image 50 to be extracted and treated as a separate virtual image 34, thereby improving user convenience. Here, the virtual image area 51 in the original captured virtual image 50 may be displayed in a predetermined suppressed manner after duplication. This makes it possible to clearly indicate the duplicated object. The suppressed manner is not particularly limited, and may be, for example, a manner in which the image is made semi-transparent. The image of the virtual image 34 (a portion of the virtual image 30) to be duplicated may be extracted from the virtual image data 132 stored in the storage unit 13. That is, the ID of the virtual image 30 may be stored when the captured image C is stored, and a portion of the image data of the virtual image 30 may be obtained from the virtual image data 132 by referring to the ID. The tenth and eighth gesture operations are not particularly limited, but may be, for example, an operation of double-tapping the spatial image area 52 or an operation of pressing and holding the spatial image area 52.

[0068] As shown in the upper diagram of FIG. 17 , when an eighth gesture operation is performed on a virtual image area 51 of a captured virtual image 50 that corresponds to a portion of the virtual image 30, the entire virtual image 30 may be restored and displayed as a virtual image 35, as shown in the lower diagram of FIG. 17 . Therefore, the content of the virtual image 35 will be the same as the content of the virtual image 30. In this case, the image of the virtual image 35 to be duplicated may be extracted from the virtual image data 132 stored in the storage unit 13. That is, the ID of the virtual image 30 may be stored when the captured image C is stored, and the image data of the entire virtual image 30 may be obtained from the virtual image data 132 by referring to the ID. Note that, when the duplicated virtual image 35 is displayed, a portion of the virtual image 30 (the virtual image area 51) that was inside the captured virtual image 50 may be deleted. This produces a visual effect in which the virtual image 30 appears to have moved from within the frame of the captured virtual image 50 to outside the frame.

[0069] In addition, when copying part or all of the virtual images 30 included in the captured virtual image 50 as shown in Figures 16 and 17, only virtual images 30 that are permitted to be copied may be copied, and virtual images 30 that are not permitted to be copied may not be copied.

[0070] As shown in FIGS. 18 to 20 , when a virtual image area 51 (at least a part of the virtual image 30) is included in a captured virtual image 50, the virtual image area 51 may be able to be moved within and / or outside the frame of the captured virtual image 50. To perform such movement, the user performs a predetermined gesture operation (for example, an operation of pressing and holding inside the captured virtual image 50) to display movement method designation buttons 91 to 93 for designating a method of moving the virtual image area 51, as shown in FIG. 18 . The movement method designation button 91 is a button for moving the virtual image area 51 within the frame of the captured virtual image 50, the movement method designation button 92 is a button for moving the virtual image area 51 within and / or outside the frame of the captured virtual image 50, and the movement method designation button 93 is a button for moving the virtual image area 51 outside the frame of the captured virtual image 50.

[0071] When the user performs a ninth gesture operation of selecting the movement method specification button 91, the CPU 11 accepts the start of an operation of moving the virtual image area 51 (at least a part of the virtual image 30) within the frame of the captured virtual image 50. In this state, as shown in FIG. 19 , when the user performs a gesture operation of dragging the virtual image area 51 within the frame of the captured virtual image 50, the virtual image area 51 moves within the frame of the captured virtual image 50 in accordance with the dragging. Here, the outer shape of the captured virtual image 50 does not change, and therefore the size of the virtual image area 51 displayed within the captured virtual image 50 (the display range of the virtual image 30) expands or contracts as the virtual image area 51 moves. In accordance with the movement of the virtual image area 51, a shadow 51 a may be displayed in the area where the virtual image area 51 was located before the movement.

[0072] When the user performs a ninth gesture operation of selecting the movement method specification button 93, the CPU 11 accepts the start of an operation of moving the virtual image area 51 (a part of the virtual image 30) outside the frame of the captured virtual image 50. In this state, as shown in FIG. 20 , when the user performs a gesture operation of dragging the virtual image area 51 outside the frame of the captured virtual image 50, a virtual image 35 that is a copy of the entire virtual image 30 is displayed outside the frame of the captured virtual image 50. Alternatively, as in the virtual image 34 of FIG. 16 , a portion of the captured virtual image 50 that corresponds to the virtual image area 51 may be moved directly outside the frame. Furthermore, the virtual image area 51 that was within the frame of the captured virtual image 50 may be deleted in accordance with the display of the virtual image 35 (or virtual image 34).

[0073] When the user performs a ninth gesture operation of selecting the movement method designation button 92, the CPU 11 accepts the start of both movement of the virtual image area 51 within the frame of the captured virtual image 50 and movement of the virtual image area 51 outside the frame of the captured virtual image 50. That is, when the user performs a gesture operation of dragging the virtual image area 51 within the frame of the captured virtual image 50, the CPU 11 moves the virtual image area 51 within the frame of the captured virtual image 50 as shown in FIG. 19 , and when the user performs a gesture operation of dragging the virtual image area 51 outside the frame of the captured virtual image 50, the CPU 11 displays the virtual image 35 (or the virtual image 34) outside the frame of the captured virtual image 50 as shown in FIG.

[0074] Thus, in response to the user's ninth gesture operation on captured virtual image 50 as the displayed second virtual image, CPU 11 accepts the start of an operation to move virtual image 30 as the first virtual image included in captured virtual image 50 by one of a plurality of movement methods, the plurality of movement methods including a method of expanding or reducing the display range of virtual image 30 within captured virtual image 50, and a method of moving virtual image 30 outside captured virtual image 50 and displaying it as virtual image 34. This allows virtual image 30 within captured virtual image 50 to be moved in a desired manner.

[0075] The wearable terminal device 10 of the present disclosure can extract and display information from a captured virtual image 50. That is, when an extraction target from which information can be extracted is included in the captured virtual image 50 as a second virtual image, the CPU 11 extracts information from the extraction target and displays it on the display unit 14. The information extraction target may be at least one of a person, an object, a place, text, and code information. This allows easy access to information that can be extracted from the person, object, place, text, code information, and the like included in the captured virtual image 50. Various aspects of information extraction will be described below with reference to FIGS. 21 to 30.

[0076] As shown in FIG. 21 , if the captured virtual image 50 includes an image of a person 44 and information can be extracted from the image of the person 44, an extracted information virtual image 62 (sixth virtual image) including the extracted information is displayed. The extracted information virtual image 62 may be displayed when a predetermined user operation is performed while the captured virtual image 50 is displayed, or may be displayed automatically along with the display of the captured virtual image 50. The extracted information virtual image 62 displays information such as a facial photograph 621 of the person 44, their name, affiliation, and contact ID. If the facial photograph 621 cannot be acquired, a message indicating that it cannot be acquired may be displayed. In response to a gesture operation to select the facial photograph 621, an operation to access the person 44 (e.g., a call operation) may be initiated. The contact ID is a code or number used to contact the person 44, and may be, for example, a phone number. The method for extracting information about person 44 from captured virtual image 50 is not particularly limited, but for example, a method may be used in which the characteristics of the image of person 44 are analyzed, a database in which information about multiple people is registered in advance is referenced to identify a person who matches the results of the characteristic analysis, and information about that person is obtained from the database.

[0077] Furthermore, the extracted information virtual image 62 may display an icon 622 (sign) of an application program (hereinafter referred to as an app) to be executed to contact the person 44. The app is predetermined according to the type of information extraction target (here, a person). When a gesture operation to select the icon 622 is performed, an app corresponding to the icon 622 (here, a phone book app) is executed. In this manner, when the captured virtual image 50 as the second virtual image includes an extraction target from which information can be extracted, the CPU 11 may display the icon 622 as a sign to launch an application predetermined according to the type of the extraction target on the display unit 14. This makes it possible to easily launch an appropriate app according to the type of extraction target. Note that the display of the icon 622 may be omitted, and an app for contacting the person 44 may be launched and a call operation or the like may be initiated in response to a gesture operation to select the face photo 621 or a contact ID (such as a phone number).

[0078] In FIG. 21, various information and icons 622 are displayed on the front (first surface) of the extracted information virtual image 62, and no information is displayed on the back (second surface). However, this is not limited to this example, and, for example, as shown in FIG. 22, icons 622 may be displayed on the back surface of the extracted information virtual image 62. That is, the CPU 11 displays the extracted information virtual image 62 as a sixth virtual image including information extracted from the extraction target on the display unit 14, and displays the icons 622 as signs on at least one of the first surface and the second surface opposite the first surface of the extracted information virtual image 62. This allows the icons 622 to be displayed in an easily accessible position depending on the display position and orientation of the extracted information virtual image 62.

[0079] As shown in FIG. 23 , if the captured virtual image 50 includes an image of an item 45 (here, a mask), and if information can be extracted from the image of the item 45, an extracted information virtual image 63 (sixth virtual image) containing the extracted information is displayed. The extracted information virtual image 63 displays an image 631 of the item 45 obtained from the Web (e.g., an e-commerce site), as well as information about the item 45, such as its product name, manufacturer, contact ID, and price. If the image 631 cannot be obtained, a message indicating this may be displayed. In response to a gesture operation to select the image 631, a website (e.g., an e-commerce site) where information about the item 45 can be accessed may be displayed. The contact ID may be a code or number used to obtain information about or purchase the item 45, and may be, for example, a telephone number or URL. The method for extracting information about the item 45 from the captured virtual image 50 is not particularly limited. For example, the method may involve identifying the item 45 using a method similar to the method for identifying the person 44 described above, and then acquiring information about the item 45 from a database.

[0080] Furthermore, the extracted information virtual image 63 may display an icon 632 (sign) of an app to be executed to access information about the item 45. The app is predetermined depending on the type of information to be extracted (here, the item). When a gesture operation to select the icon 632 is performed, the app corresponding to the icon 632 (here, a browser app) is executed. Note that the display of the icon 632 may be omitted, and an app (browser app, etc.) to access information about the item 45 may be launched in response to a gesture operation to select the image 631 or a contact ID (URL, etc.) of the item 45, and an e-commerce website where the item 45 can be purchased may be displayed, for example.

[0081] As shown in FIG. 24 , if location information can be extracted from an image of the space 40 in the background of the captured virtual image 50, an extracted information virtual image 64 (sixth virtual image) containing the extracted information is displayed. The extracted information virtual image 64 displays an image 641 representing the location, the name of the location, building information if the location is a building, and a contact ID. For example, if the location is a building, the image 641 may be the logo of the company that owns the building. If the image 641 cannot be acquired, a message indicating that it cannot be acquired may be displayed. In response to a gesture operation to select the image 641, a map of the location may be displayed. The displayed map may be a standard map if the location is outside a building, or an indoor map if the location is inside a building. The contact ID may be a code or number used to access location information, such as a phone number or URL. The method for extracting location information from the captured virtual image 50 is not particularly limited. For example, a method of identifying a location using a method similar to the method for identifying the person 44 described above and acquiring location information from a database may be used.

[0082] Furthermore, the extracted information virtual image 64 may display an icon 642 (sign) of an app to be executed to access location information. The app is predetermined depending on the type of information to be extracted (location in this case). When a gesture operation to select the icon 642 is performed, the app corresponding to the icon 642 (a map app in this case) is executed. Note that the display of the icon 642 may be omitted, and an app (such as a map app) to access location information may be launched in response to a gesture operation to select the location image 641 or a contact ID (such as a URL), and a map showing the location of the location may be displayed, for example.

[0083] The operation when information can be extracted from the captured virtual image 50 is not limited to the above. For example, as shown on the left side of Fig. 25, when person information can be extracted from the captured virtual image 50, an icon 622 of an app (e.g., a calling app) that is pre-associated with the "person" from which information is to be extracted may be automatically displayed. Furthermore, the app may be executed in response to a gesture operation to select the icon 622, and a virtual image 623 of the app may be displayed in a state where, for example, a call with the extracted person has started.

[0084] 25, when information about an item can be extracted from the captured virtual image 50, an icon 632 of an app (for example, a browser app) that is pre-associated with the "item" from which information is to be extracted may be automatically displayed. In addition, the app may be executed in response to a gesture operation to select the icon 632, and a virtual image 633 of the app may be displayed, for example, including a website from which information about the extracted item can be accessed.

[0085] 25, when location information can be extracted from the captured virtual image 50, an icon 642 of an app (for example, a map app) that is pre-associated with the "location" from which information is to be extracted may be automatically displayed. In addition, the app may be executed in response to a gesture operation to select the icon 642, and a virtual image 643 of the app including, for example, a map showing the position of the extracted location may be displayed.

[0086] In FIG. 25, the display of the icons 622, 632, and 642 may be skipped and the virtual images 623, 633, and 643 of the applications may be displayed directly.

[0087] As shown in FIG. 26, if the captured virtual image 50 includes an image of the character 46 and information can be extracted from the image of the character 46, an extracted information virtual image 65 (sixth virtual image) including the extracted information is displayed. The extracted information virtual image 65 displays the content of the character 46 extracted by OCR (Optical Character Recognition) or the like, and operation buttons 651 to 653 for executing various processes on the extracted character. The size of the character 46 displayed in the extracted information virtual image 65 may be larger than the size of the character 46 in the captured virtual image 50. Furthermore, the size of the character 46 displayed in the extracted information virtual image 65 may be set in advance.

[0088] As shown in Fig. 27, when a gesture operation to select one of operation buttons 651 to 653 is performed, a process corresponding to operation button 651 to 653 is executed for extracted character 46. When a gesture operation to select operation button 651 displaying "Copy" is performed, icon 661 of an application (for example, an editor application) that is preset as an application for editing text data is automatically displayed, as shown on the left side of Fig. 27. Furthermore, the application is executed in response to the gesture operation to select icon 661, and a virtual image 67 of the application is displayed in a state in which character 46 can be edited.

[0089] Furthermore, when a gesture operation is performed to select the operation button 652 labeled "Translation: Auto → EG," an icon 662 of an application (e.g., an editor application) that is preset as an application for editing text data is automatically displayed, as shown in the center of FIG. 27. The icon 662 may be the same as the icon 661. When a gesture operation to select the icon 662 is performed, the extracted characters 46 are translated in accordance with a predetermined translation setting. Here, the setting is such that the language of the extracted characters 46 is automatically determined and translated into English. Alternatively, a translation setting related to the translation destination language, etc., may be set in accordance with an eleventh gesture operation by the user on a setting button (not shown). In this case, the characters 46 are translated in accordance with the translation setting set by the user. Then, the application corresponding to the icon 662 is executed, and a virtual image 68 of the application is displayed in a state in which the translated characters can be edited. In this way, the CPU 11 translates the information extracted from the characters 46 in accordance with the predetermined translation setting or in accordance with the translation setting in accordance with the eleventh gesture operation by the user, and displays the translated information on the display unit 14. This allows the extracted characters 46 to be easily translated and displayed.

[0090] Furthermore, when a gesture operation is performed to select the operation button 653 labeled "Specify App," an icon 663 of the app (here, a browser app) designated by the user is automatically displayed, as shown on the right side of FIG. 27. The user may designate an app in advance, or the user may select an app each time from a selection of apps displayed in response to the selection of the operation button 653. The app is executed in response to the gesture operation to select the icon 663, and a virtual image 69 of the app that performs processing related to the information of the extracted characters 46 is displayed. In the example of FIG. 27, a virtual image 69 of the browser app is displayed with the search results for the extracted characters 46 displayed.

[0091] In FIG. 27, the display of the icons 661 to 663 may be skipped and the virtual images 67 to 69 of the applications may be displayed directly.

[0092] As shown in FIG. 28, if the captured virtual image 50 includes an image of a two-dimensional code 47 (code information) and information can be extracted from the image of the two-dimensional code 47, an extracted information virtual image 71 (sixth virtual image) containing the extracted information is displayed. The extracted information virtual image 71 displays information obtained by decoding the two-dimensional code 47 and information related to the information. In the example of FIG. 71, information about "K Corporation" is extracted from the two-dimensional code 47, and the company's address, telephone number, and URL are displayed. In response to a gesture operation to select the telephone number or URL, an app (such as a calling app or a browser app) for accessing "K Corporation" may be launched. Note that the code information from which information can be extracted is not limited to the two-dimensional code 47 and may be a barcode or a code consisting of letters, symbols, etc.

[0093] The captured virtual image 50 may include location information of the wearable terminal device 10 when the captured image C was generated. In this case, the CPU 11 acquires location information of the wearable terminal device 10 when the first gesture operation described above is performed, and stores the acquired location information in the storage unit 13 in association with the captured image C. This allows the location where the captured image C was generated to be easily referenced at any time after capture. The location information may be acquired from a positioning satellite of a GNSS (Global Navigation Satellite System) such as a GPS (Global Positioning System), or may be various types of local location information. The local location information may be acquired from signals transmitted from, for example, a wireless LAN access point, a beacon station, or a local 5G base station. Furthermore, when the captured virtual image 50 is displayed based on the captured image C, a virtual image 72 including information related to the location when the captured image C was generated may be displayed, as shown in FIG. 29. Furthermore, instead of (or in addition to) the text information shown in FIG. 29, a map indicating the capture location may be displayed. In addition to the location information, information on the date and time when the capture image C was generated may be acquired and displayed on the virtual image 72.

[0094] The captured virtual image 50 may include information identifying the user who was using the wearable terminal device 10 when the captured image C was generated. In this case, the CPU 11 identifies the user when the first gesture operation was performed and stores user information related to the identified user in the storage unit 13 in association with the captured image C. This makes it possible to easily refer to the operator who generated the captured image C at any time after capture. The method for identifying the user is not particularly limited, and may be obtained from login information of the user who generated the captured image C, for example. Furthermore, when the captured virtual image 50 is displayed based on the captured image C, a virtual image 73 including information related to the user and the date and time when the captured image C was generated may be displayed, as shown in FIG. 30 . Note that, as information related to the user, an icon capable of displaying an image of the user's face may be superimposed on the captured virtual image 50 or displayed near the captured virtual image 50.

[0095] Because the wearable terminal device 10 can be used for various purposes in various locations, it may not be appropriate to generate and store a captured image C depending on the location when the capture operation is performed and the capture target included in the visible area 41. In such cases, the wearable terminal device 10 may have a function to prevent the captured image C from being stored in the storage unit 13. Hereinafter, prohibiting the generation of a captured image C and storing it in the storage unit 13 will also be referred to as "prohibiting capture," and allowing the generation of a captured image C and storing it in the storage unit 13 will also be referred to as "allowing capture."

[0096] For example, as shown in FIG. 31 , the CPU 11 may determine whether a predetermined capture-prohibited object is included in the capture area R, and if it determines that the capture area R includes a capture-prohibited object, may not store the capture image C in the storage unit 13. This may limit the generation of the capture image C including the capture-prohibited object and its storage in the storage unit 13. In the example of FIG. 31 , the capture area R specified by the user's first gesture operation includes a person 44 who has been set as a capture-prohibited object in advance. In this case, even if the capture area R is specified, the capture image C is not stored in the storage unit 13. Note that the capture-prohibited object is not limited to a person, but may also be an object such as a work of art, such as a painting, or real estate, such as a building. The capture-prohibited object may also be an object protected by copyright. The method for determining whether an object is a capture-prohibited object is not particularly limited. For example, a method may be used in which the image characteristics of the capture-prohibited object are stored in the storage unit 13 or an external server, and the image included in the capture area R is compared with the image characteristics of the capture-prohibited object.

[0097] When capture is prohibited, for example, the display mode of the capture operation start buttons 611 and 612 may be changed in the menu virtual image 61 shown in the lower diagram of Fig. 8 to disable the operation. Alternatively, the capture operation start buttons 611 and 612 may be hidden.

[0098] When a capture area R includes a prohibited object, if an authority who lifts the prohibition on capture permits capture, the prohibition may be lifted and a captured image C may be stored in the storage unit 13. For example, as shown in the lower diagram of FIG. 31 , in response to a first gesture operation that identifies the capture area R, a dialog virtual image 74 is displayed that asks the user whether or not to request permission for capture. When a gesture operation is performed on the OK button 741 of the dialog virtual image 74, a signal requesting lifting of the prohibition on capture is transmitted to an external device used by the authority. When a permission signal permitting lifting of the prohibition on capture is received from the external device, the captured image C is stored in the storage unit 13. The subsequent display operation of the captured virtual image 50 is the same as that described above.

[0099] Furthermore, the authority level of the user operating the wearable terminal device 10 may be referenced based on the user ID used when logging in to the wearable terminal device 10, and whether or not to prohibit capture may be determined according to the authority level. The authority level may be determined based on, for example, position, department, or qualification.

[0100] Furthermore, whether or not capture is prohibited may be determined based on the current location of the wearable terminal device 10. In this case, the CPU 11 acquires location information of the wearable terminal device 10 when the first gesture operation is performed, and if the location indicated by the location information satisfies a predetermined prohibited location condition, the CPU 11 does not store the captured image C in the storage unit 13. This makes it possible to prohibit capture when the wearable terminal device 10 is located in a specific location. The location information may be acquired from a GNSS positioning satellite such as GPS, or may be various types of local location information. The local location information may be acquired from signals transmitted from, for example, a wireless LAN access point, a beacon station, or a local 5G base station. If the acquired location information is within a predetermined prohibited area, it is determined that the prohibited location condition is satisfied.

[0101] Furthermore, the CPU 11 may not store the captured image C in the storage unit 13 when the communication unit 16 receives a specific signal. This allows capturing to be prohibited at any timing by transmitting a specific signal to the wearable terminal device 10. The specific signal may be any signal that is preset as a signal for prohibiting capturing. The prohibition of capturing may be lifted when a predetermined time has elapsed after receiving the specific signal. Furthermore, even when a specific signal is received, capturing may be permitted if the visible area 41 does not include a virtual object such as the virtual image 30 and includes only the background space 40.

[0102] Furthermore, the CPU 11 may not store the capture image C in the storage unit 13 when a predetermined connection condition related to the connection state of the communication unit 16 to the communication network is satisfied. This enables capture prohibition control according to the connection state to the communication network. The connection condition related to the capture prohibition control may be determined arbitrarily. For example, when connected to a public communication network, it may be determined that the connection condition is satisfied and capture may be prohibited, whereas when connected to a private communication network (such as local 5G or wireless LAN), it may be determined that the connection condition is not satisfied and capture may be permitted. As another example, when the wearable terminal device 10 is online, it may be determined that the connection condition is satisfied and capture may be prohibited, whereas when offline, it may be determined that the connection condition is not satisfied and capture may be permitted.

[0103] Furthermore, if at least a portion of a virtual image (first virtual image) is included in the capture area R and the virtual image satisfies a predetermined prohibition condition, the CPU 11 may prevent the capture image C from being stored in the storage unit 13. This makes it possible to prevent a virtual image that is not appropriate as a capture target from being stored as the capture image C. For example, if the virtual image is a screen of a specific application and a specific capture prohibition flag is set in the application, the virtual image of the application can be prevented from being stored as the capture image C.

[0104] As described above, in each of the cases where capture is prohibited based on the current location, when a specific signal is received, when a connection condition related to the connection state to the communication network is satisfied, and when a virtual image included in the capture area R satisfies the prohibition condition, the prohibition of capture may be lifted if an authorized person permits it. In this case, as shown in the lower diagram of FIG. 31 , a dialog virtual image 74 may be displayed to inquire the user whether or not to request permission for capture. Furthermore, the authority level of the user operating the wearable terminal device 10 may be referenced, and whether or not to prohibit capture may be determined depending on the authority level.

[0105] The capture image C may be stored in association with audio data. That is, when storing the capture image C in the storage unit 13, the CPU 11 may acquire audio data and store the audio data in association with the capture image C in the storage unit 13. This allows audio information to be added to the capture image C. For example, when a capture area R is identified in response to a first gesture operation as shown in the upper diagram of FIG. 32 , a dialog virtual image 75 inquiring of the user as to whether or not to record audio may be displayed as shown in the lower diagram of FIG. 32 . When a gesture operation is performed to operate the OK button 751 of the dialog virtual image 75, audio data is acquired (recorded) by the microphone 17 and stored in the storage unit 13 in association with the capture image C in the capture area R. Note that the timing of recording is not limited thereto; for example, the audio may be recorded when the capture virtual image 50 is displayed. When a capture virtual image 50 including a capture image C associated with audio data is displayed, a play button 53 for playing the audio data may also be displayed as shown in FIG. 33 . In response to a gesture operation to select the play button 53, the sound of the audio data associated with the captured image C is output from the speaker 18. The play button 53 may be displayed behind the captured virtual image 50.

[0106] In the above description, the captured virtual image 50 is displayed alone in the space 40. However, the present invention is not limited to this, and the captured virtual image 50 may be displayed so as to satisfy a predetermined positional relationship with another object (display target). That is, the CPU 11 may display the captured virtual image 50 as a second virtual image on the surface of a predetermined display target located in the space 40, or at a position that satisfies a predetermined positional relationship with the display target. The display target may also be an object, a person, or any virtual image other than the captured virtual image 50 in the space 40. This allows the captured virtual image 50 to be displayed in association with other objects, people, virtual images (virtual objects, etc.), etc.

[0107] 34, a captured virtual image 50 may be displayed at a position corresponding to the outer surface of a spherical object 48 located within the space 40, in accordance with the shape of the outer surface. Alternatively, the contour of the surface of an object or virtual image located within the space 40 may be identified, and a captured virtual image 50 with an outer shape that matches the contour may be displayed on the surface of the object or virtual image. The display position of the captured virtual image 50 is not limited to the surface of the object, but may be a position near the object.

[0108] Furthermore, the CPU 11 may move the display position of the captured virtual image 50 as the second virtual image in accordance with the movement of the display object within the space 40. This makes it possible to dynamically express the relationship between the captured virtual image 50 and other objects, people, virtual images, or the like. For example, as shown in FIG. 35 , when the captured virtual image 50 is displayed above the head of a person 44 as the display object, the captured virtual image 50 may be moved in accordance with the movement of the person 44. Furthermore, as shown in FIG. 35 , the CPU 11 may display the captured virtual image 50 so that the size of the captured virtual image 50 increases as the distance between the user and the display object decreases. This allows the size of the captured virtual image 50 to express a sense of perspective, making it appear as if the captured virtual image 50 is moving more naturally within the space 40.

[0109] Furthermore, CPU 11 may change the orientation of captured virtual image 50 as the second virtual image so as to follow a change in the orientation of the display object. For example, in the upper diagram of FIG. 36 , captured virtual image 50 is displayed directly above arrow-shaped three-dimensional virtual object 54 (virtual image) as the display object, with its surface facing the direction of the arrow. Here, when the orientation of virtual object 54 is changed as shown in the lower diagram, the orientation of captured virtual image 50 is changed to match the orientation of virtual object 54. Here, only the states before and after the orientation change are shown, but when captured virtual image 50 rotates from the state shown in the upper diagram of FIG. 36 to the state shown in the lower diagram, captured virtual image 50 may be rotated to follow the rotation.

[0110] Next, a capture virtual image display process for performing various operations related to the display of the above-mentioned capture virtual image 50 will be described with reference to the flowcharts of Figures 37 and 38. Here, a representative process for performing a representative operation is illustrated. As explained for each operation above, the operations and processes related to the display of the capture virtual image 50 are not limited to these.

[0111] As shown in FIG. 37, when the capture virtual image display process is started, the CPU 11 displays the menu virtual image 61 on the display unit 14 in response to the user's third gesture operation (step S101). The CPU 11 determines whether or not a method for specifying the capture area R has been specified (step S102). For example, the CPU 11 determines that a method for specifying the capture area R has been specified when an operation for selecting the capture operation start button 611 or 612 has been performed in the menu virtual image 61 of FIG. 8. If it is determined that a method for specifying the capture area R has not been specified ("NO" in step S102), the CPU 11 executes step S102 again. If it is determined that a method for specifying the capture area R has been specified ("YES" in step S102), the CPU 11 accepts the specification of the capture area R using the specified specification method (step S103) and determines whether or not the capture area R has been specified (step S104). If it is determined that the capture region R has not been designated ("NO" in step S104), the CPU 11 executes step S104 again.

[0112] If it is determined that the capture area R has been specified ("YES" in step S104), the CPU 11 determines whether or not capture is prohibited (step S105). As exemplified above, capture is prohibited when the capture area R includes an object that is prohibited from being captured, when capture is prohibited based on the current location, when capture is prohibited in response to the receipt of a specific signal, when capture is prohibited in response to the satisfaction of a connection condition related to the connection state to the communication network, and when capture is prohibited in response to the virtual image included in the capture area R satisfying the prohibition condition.

[0113] If it is determined that the situation is one in which capture is prohibited ("YES" in step S105), the CPU 11, in response to a user instruction, transmits a signal to an external device operated by an authorized person requesting that the capture prohibition be lifted, and determines whether or not an authorization signal permitting the lifting of the capture prohibition has been received from the external device (step S106). If the authorization signal has not been received within a predetermined period of time ("NO" in step S106), the CPU 11 terminates the capture virtual image display process. If the authorization signal has been received within the predetermined period of time ("YES" in step S106), or if it is determined in step S105 that the situation is not one in which capture is prohibited ("NO" in step S105), the CPU 11 generates a capture image C of the capture area R and stores it in the storage unit 13 (step S107).

[0114] The CPU 11 determines whether or not a second gesture operation for displaying the capture image C as the capture virtual image 50 has been performed (step S108), and if it is determined that the second gesture operation has not been performed ("NO" in step S108), executes step S108 again. If it is determined that the second gesture operation has been performed ("YES" in step S108), the CPU 11 displays the capture virtual image 50 including the capture image C on the display unit 14 (step S109).

[0115] The CPU 11 determines whether the captured virtual image 50 includes an information extraction target (step S110). If it is determined that the information extraction target is included ("YES" in step S110), the CPU 11 displays an extracted information virtual image including the extracted information and an icon of a predetermined application on the display unit 14 (step S111). The CPU 11 determines whether a gesture operation for selecting an icon has been performed (step S112). If it is determined that the gesture operation has been performed ("YES" in step S112), the CPU 11 executes the application corresponding to the icon and displays a virtual image of the application on the display unit 14 (step S113). When step S113 is completed, if it is determined that the information extraction target is not included in step S110 ("NO" in step S110), or if it is determined that a gesture operation for selecting an icon has not been performed ("NO" in step S112), the CPU 11 terminates the captured virtual image display process.

[0116] Second Embodiment Next, the configuration of a display system 1 according to a second embodiment will be described. As shown in Fig. 39, the display system 1 according to the second embodiment differs from the first embodiment in that it includes a wearable terminal device 10 and a plurality of external devices 20. Below, differences from the first embodiment will be described, and commonalities will not be described.

[0117] As shown in FIG. 39, the wearable terminal device 10 and multiple external devices 20 included in the display system 1 are communicatively connected via a network N. The network N can be, for example, the Internet, but is not limited to this. The display system 1 may include multiple wearable terminal devices 10. The display system 1 may also include one external device 20. For example, in this embodiment, a user performing a predetermined task wears the wearable terminal device 10. A remote instructor who gives instructions to the user wearing the wearable terminal device 10 from a remote location via the wearable terminal device 10 operates the external device 20.

[0118] As shown in Figure 40, the external device 20 includes a CPU 21, a RAM 22, a memory unit 23, an operation display unit 24, a communication unit 25, a microphone 26, a speaker 27, etc., and each of these units is connected by a bus 28.

[0119] The CPU 21 is a processor that performs various types of arithmetic processing and controls the overall operation of each part of the external device 20. The CPU 21 reads and executes a program 231 stored in the storage unit 23 to perform various control operations.

[0120] The RAM 22 provides a working memory space for the CPU 21 and stores temporary data.

[0121] The storage unit 23 is a non-transitory recording medium readable by the CPU 21 as a computer. The storage unit 23 stores a program 231 executed by the CPU 21, various setting data, and the like. The program 231 is stored in the storage unit 23 in the form of computer-readable program code. The storage unit 23 may be, for example, a non-volatile storage device such as an SSD equipped with a flash memory or an HDD (Hard Disk Drive).

[0122] The operation display unit 24 includes a display device such as a liquid crystal display, and input devices such as a mouse and a keyboard. The operation display unit 24 displays various information such as the operation status and processing results of the display system 1 on the display device. The display includes, for example, an instructor screen 42 (see FIG. 42 ) that includes an image of the visible area 41 captured by the camera 154 of the wearable terminal device 10. The operation display unit 24 also converts a user's input operation on the input device into an operation signal and outputs the operation signal to the CPU 21.

[0123] The communication unit 25 transmits and receives data to and from the wearable terminal 10 in accordance with a predetermined communication protocol. The communication unit 25 can also perform voice data communication with the wearable terminal 10. That is, the communication unit 25 transmits voice data collected by the microphone 26 to the wearable terminal 10, and receives voice data transmitted from the wearable terminal 10 to output voice from the speaker 27. The communication unit 25 may be capable of communicating with devices other than the wearable terminal 10.

[0124] The microphone 26 converts sounds such as the voice of a remote instructor into electrical signals and outputs them to the CPU 21.

[0125] The speaker 27 converts input audio data into mechanical vibrations and outputs the vibrations as sound.

[0126] In the display system 1 of this embodiment, bidirectional data communication is performed between the wearable terminal device 10 and one or more external devices 20, allowing various data to be shared and collaborative work to be performed. For example, data of an image captured by the camera 154 of the wearable terminal device 10 and data of the displayed virtual image 30 are transmitted to the external device 20 and displayed as an instructor screen 42 on the operation display unit 24, allowing the remote instructor to recognize in real time what the user of the wearable terminal device 10 is viewing through the visor 141. Furthermore, voice calls can be performed by transmitting voices collected by the microphone 17 of the wearable terminal device 10 and the microphone 26 of the external device 20 via bidirectional voice data communication. Therefore, the period during which voice data communication is being performed by the wearable terminal device 10 and the external device 20 includes the period during which a voice call is being performed between the user of the wearable terminal device 10 and the remote instructor. The remote instructor can give instructions and support to the user of the wearable terminal device 10 through voice communication while viewing the real-time camera image on the instructor screen 42.

[0127] When the above-described capture virtual image 50 is displayed on the display unit 14 of the wearable terminal device 10, the display of the capture virtual image 50 can be reflected as is on the instructor screen 42 of the external device 20. However, there are cases where it is desirable not to display the capture virtual image 50 on the instructor screen 42 of the external device 20, for example, when the capture virtual image 50 contains confidential information. Therefore, in the display system 1 of the present disclosure, it is possible to set in advance whether or not to reflect the display of the capture virtual image 50 on the instructor screen 42. Data related to this setting may be stored in the storage unit 13 of the wearable terminal device 10 or in the storage unit 23 of the external device 20. Furthermore, even when the setting is such that the capture virtual image 50 is not reflected on the instructor screen 42 of the external device 20, the display of the capture virtual image 50 may be reflected on the instructor screen 42 if the user of the wearable terminal device 10 gives permission.

[0128] For example, assume that user A is logged in to the wearable terminal device 10, user B is logged in to the external device 20, and screen sharing is occurring between the wearable terminal device 10 and the external device 20. Also assume that the display of the captured virtual image 50 in the visible area 41 of the wearable terminal device 10 is set not to be reflected on the instructor screen 42 of the external device 20. In this state, when the captured virtual image 50 is displayed on the wearable terminal device 10 as shown in FIG. 41 , the instructor screen 42 of the external device 20 shows an area corresponding to the captured virtual image 50, but the content of the image is not displayed, as shown in FIG.

[0129] 42, a dialog image 76 may be displayed on the instructor screen 42 to inquire of the remote instructor whether or not to request permission to display the captured virtual image 50. When an operation to select the OK button 761 on the dialog image 76 is performed, a request signal requesting permission to display the captured virtual image 50 is transmitted to the wearable terminal device 10.

[0130] 43, in response to receiving the request signal, the wearable terminal device 10 displays on the display unit 14 a dialog virtual image 77 inquiring whether or not to allow display of a captured image. When a gesture operation is performed to select the OK button 771 of the dialog virtual image 77, a permission signal that allows display of the captured virtual image 50 is transmitted to the external device 20. When the permission signal is received, the content of the captured virtual image 50 is displayed on the instructor screen 42 of the external device 20.

[0131] Third Embodiment Next, the configuration of a display system 1 according to a third embodiment will be described. The third embodiment differs from the first embodiment in that an external information processing device 80 executes some of the processing that was executed by the CPU 11 of the wearable terminal device 10 in the first embodiment. Below, differences from the first embodiment will be described, and commonalities will not be described. The third embodiment may be combined with the second embodiment.

[0132] As shown in FIG. 44, the display system 1 includes a wearable terminal device 10 and an information processing device 80 (server) communicatively connected to the wearable terminal device 10. At least a part of the communication path between the wearable terminal device 10 and the information processing device 80 may be wireless communication. The hardware configuration of the wearable terminal device 10 may be the same as that of the first embodiment, but a processor for performing the same processing as that executed by the information processing device 80 may be omitted. Furthermore, when this embodiment is combined with the second embodiment, the information processing device 80 may be connected to a network N.

[0133] As shown in FIG. 45, an information processing device 80 includes a CPU 81, a RAM 82, a storage unit 83, an operation display unit 84, a communication unit 85, and the like, and these units are connected to each other via a bus 86.

[0134] The CPU 81 is a processor that performs various types of arithmetic processing and controls the overall operation of each unit of the information processing device 80. The CPU 81 reads and executes a program 831 stored in the storage unit 83 to perform various control operations.

[0135] The RAM 82 provides a working memory space for the CPU 81 and stores temporary data.

[0136] The storage unit 83 is a non-transitory recording medium readable by the CPU 81 as a computer. The storage unit 83 stores a program 831 executed by the CPU 81, various setting data, and the like. The program 831 is stored in the storage unit 83 in the form of computer-readable program code. The storage unit 83 may be, for example, an SSD equipped with a flash memory, or a non-volatile storage device such as an HDD.

[0137] The operation display unit 84 includes a display device such as a liquid crystal display, and input devices such as a mouse and a keyboard. The operation display unit 84 displays various information such as the operation status and processing results of the display system 1 on the display device. Here, the operation status of the display system 1 may include images captured in real time by the camera 154 of the wearable terminal device 10. The operation display unit 84 also converts user input operations on the input device into operation signals and outputs the operation signals to the CPU 21.

[0138] The communication unit 85 communicates with the wearable terminal device 10 to transmit and receive data. For example, the communication unit 85 receives data including some or all of the detection results by the sensor unit 15 of the wearable terminal device 10, and information related to user operations (gestures) detected by the wearable terminal device 10. The communication unit 85 may also be capable of communicating with devices other than the wearable terminal device 10.

[0139] In the display system 1 configured as above, the CPU 81 of the information processing device 80 executes at least part of the processing executed by the CPU 11 of the wearable terminal device 10 in the first embodiment. For example, the CPU 81 may perform three-dimensional mapping of the space 40 based on the detection results of the depth sensor 153. The CPU 81 may also detect the user's visible area 41 within the space 40 based on the detection results of each component of the sensor unit 15. The CPU 81 may also generate virtual image data 132 related to the virtual image 30 in response to the user's operation of the wearable terminal device 10. The CPU 81 may also detect the position and orientation of the user's hand (and / or fingers) based on images captured by the depth sensor 153 and the camera 154. The CPU 81 may also execute processing related to the generation of the captured image C and the display of the captured virtual image 50.

[0140] The processing results by the CPU 81 are transmitted to the wearable terminal device 10 via the communication unit 85. The CPU 11 of the wearable terminal device 10 operates each unit of the wearable terminal device 10 (for example, the display unit 14) based on the received processing results. The CPU 81 may also transmit a control signal to the wearable terminal device 10 to control the display of the display unit 14 of the wearable terminal device 10.

[0141] In this way, by executing at least part of the processing in the information processing device 80, it is possible to simplify the device configuration of the wearable terminal device 10 and reduce manufacturing costs. Furthermore, by using a higher performance information processing device 80, it is possible to increase the speed and precision of various processes related to MR. This makes it possible to increase the precision of 3D mapping of the space 40, improve the display quality of the display unit 14, and increase the response speed of the display unit 14 to the user's actions.

[0142] 〔others〕 The above embodiment is an example, and various modifications are possible.

[0143] For example, in each of the above embodiments, the optically transparent visor 141 is used to allow the user to view the real space, but this is not limited to this. For example, the light-blocking visor 141 may be used to allow the user to view the image of the space 40 captured by the camera 154. That is, the CPU 11 may display on the display unit 14 the image of the space 40 captured by the camera 154 and the virtual image 30 superimposed on the image of the space 40. With such a configuration, MR in which the virtual image 30 is blended with the real space can also be realized.

[0144] Furthermore, by using a pre-generated image of a virtual space instead of an image of a real space captured by camera 154, it is possible to realize VR that makes the user feel as if they are in the virtual space. In this VR, the user's visual recognition area 41 is also specified, and the portion of the virtual space that is inside visual recognition area 41 and virtual image 30 whose display position is determined inside visual recognition area 41 are displayed. Therefore, the background of captured image C in this case is the virtual space.

[0145] The wearable terminal device 10 is not limited to the one having the annular main body 10a illustrated in Fig. 1, and may have any structure as long as it has a display unit that is visible to the user when worn. For example, it may have a structure that covers the entire head like a helmet. It may also have a frame that is hung on the ears like glasses, with various devices built into the frame.

[0146] The various virtual images do not necessarily have to be stationary in the space 40, but may move within the space 40 along a predetermined trajectory.

[0147] Although the example in which a user's gesture is detected and accepted as an input operation has been described, the present invention is not limited to this. For example, an input operation may be accepted by a controller that the user holds in their hand or wears on their body.

[0148] Although an example of a voice call between the wearable terminal device 10 and the external device 20 has been described, this is not limiting and a video call may also be possible. In this case, a webcam for capturing an image of the remote operator may be provided in the external device 20, and image data captured by the webcam may be transmitted to the wearable terminal device 10 and displayed on the display unit 14.

[0149] In addition, the specific details of the configurations and controls shown in the above embodiments can be appropriately changed without departing from the spirit of the present disclosure. Furthermore, the configurations and controls shown in the above embodiments can be appropriately combined without departing from the spirit of the present disclosure. [Industrial Applicability]

[0150] The present disclosure can be used in a wearable terminal device, a program, and an image processing method. [Explanation of symbols]

[0151] 1 Display System 10 Wearable terminal device 10a Main body 11 CPU (processor) 12 RAM 13 Storage section 131 Programs 132 Virtual Image Data 14 Display section 141 Visor (display component) 142 Laser Scanner 15 Sensor section 151 Accelerometer 152 Angular velocity sensor 153 Depth Sensor 154 Camera 155 Eye Tracker 16 Communications Department 17. Mike 18 speakers 19 Bus 20 External equipment 21 CPU 23 Memory section 231 Programs 24 Operation display section 30 Virtual Image (First Virtual Image) 31 Feature Bar 32 Window shape change button 33 Close button 34, 35 Virtual images (4th virtual image, 5th virtual image) 40 space 41 Visibility Zone 411 Virtual Line 412 Pointer 42 Instructor screen 44 People 45 Goods 46 characters 47 Two-dimensional code (code information) 48 Object 50 Capture Virtual Image (Second Virtual Image) 50a Initial display area 50b Extension Area 51 Virtual Image Area 52 Spatial Image Region 53 Play button 54 Virtual Objects 61 Menu Virtual Image (Third Virtual Image) 62-65, 71 Extracted information virtual image (6th virtual image) 622, 632, 642, 661~663 Icons (signs) 67~69, 72, 73, 623, 633, 643 Virtual images 74, 75, 77 Dialogue Virtual Image 76 Dialogue Images 80 Information processing equipment 81 CPU 83 Memory section 831 Programs 84 Operation display section C Captured image D. Captured image E. Composite image L List area N Network R Capture Area r Capture frame U User

Claims

1. a camera that captures a space as a user's visual field; a display unit visible to the user; at least one processor; The at least one processor: displaying the virtual image on the display unit so that the virtual image is visible in the space; Based on a predetermined operation of the user, extracting a captured image including a portion of the visible area in the space and at least a portion of the virtual image from a composite image obtained by combining the captured image of the space by the camera and the virtual image, and storing the captured image in a storage unit; displaying the captured image on the display unit as a captured virtual image that is visually recognized as being disposed at a predetermined position in the space and that is different from the virtual image; The captured virtual image is expanded or contracted in response to a predetermined gesture operation by a user on the displayed captured virtual image, and when the captured virtual image is expanded, an image of the expanded portion is extracted from the composite image. Wearable terminal device.

2. The wearable terminal device according to claim 1 , wherein the predetermined operation by the user includes a first gesture operation.

3. The wearable terminal device according to claim 1 , wherein the at least one processor causes the display unit to display the captured virtual image when the captured image is stored in the storage unit.

4. The wearable terminal device according to claim 1 , wherein the at least one processor causes the display unit to display the captured virtual image in response to a second gesture operation by a user.

5. the at least one processor, in response to the second gesture operation, causes the captured virtual image to be displayed at a predetermined position relative to the user; The wearable terminal device according to claim 4 , wherein the relative position is set in advance for each type of the second gesture operation.

6. The wearable terminal device according to claim 1 , wherein the at least one processor displays the captured virtual image at a position closer to the user in the space than the virtual image.

7. The wearable terminal device according to claim 2 , wherein the at least one processor causes the display unit to display a virtual menu image for starting to accept the predetermined operation in response to a third gesture operation by the user.

8. The at least one processor: specifying a capture area corresponding to the capture image by a method different from each other depending on a type of the first gesture operation; The wearable terminal device according to claim 7 , wherein the type of the first gesture operation to be accepted is identified in accordance with an operation on the menu virtual image.

9. The wearable terminal device according to any one of claims 1 to 8, wherein the at least one processor displays an outer frame of the captured virtual image in a predetermined highlighted manner.

10. The wearable terminal device according to claim 1 , wherein the at least one processor moves the captured virtual image in response to a fourth gesture operation by the user on the displayed captured virtual image.

11. The wearable terminal device according to any one of claims 1 to 10, wherein the at least one processor deletes the virtual image included in the captured virtual image in response to a sixth gesture operation by the user on the displayed captured virtual image, and displays an image of the space corresponding to the deleted area in the deleted area.

12. The wearable terminal device according to claim 11 , wherein the at least one processor displays the virtual image included in the captured virtual image in a predetermined emphasized manner.

13. a computer capable of controlling a wearable terminal device including a camera that captures an image of a space as a user's visual recognition area, a display unit that is visible to the user, and at least one processor; a process of displaying the virtual image on the display unit so that the virtual image is visually recognized in the space; Based on a predetermined operation of the user, A process of extracting a captured image including a part of the visible area in the space and at least a part of the virtual image from a composite image obtained by combining the captured image of the space by the camera and the virtual image, and storing the captured image in a storage unit; a process of displaying the captured image on the display unit as a captured virtual image that is visually recognized as being disposed at a predetermined position in the space and that is different from the virtual image; a process of expanding or contracting the captured virtual image in response to a predetermined gesture operation by a user on the displayed captured virtual image, and extracting an image of the expanded portion from the composite image when the captured virtual image is expanded; A program that executes the following.

14. An image processing method executed by a computer capable of controlling a wearable terminal device including a camera that captures an image of a space as a user's visual recognition area, a display unit that is visible to the user, and at least one processor, displaying the virtual image on the display unit so that the virtual image is visible in the space; Based on a predetermined operation of the user, extracting a captured image including a portion of the visible area in the space and at least a portion of the virtual image from a composite image obtained by combining the captured image of the space by the camera and the virtual image, and storing the captured image in a storage unit; displaying the captured image on the display unit as a captured virtual image that is visually recognized as being disposed at a predetermined position in the space and that is different from the virtual image; The captured virtual image is expanded or contracted in response to a predetermined gesture operation by a user on the displayed captured virtual image, and when the captured virtual image is expanded, an image of the expanded portion is extracted from the composite image. Image processing methods.

Citation Information

Patent Citations

  • Display device, method for controlling display device, and program

    JP2016161734A

  • Human body gesture-based region and volume selection for hmd

    JP2016514298A

  • Passive optical and inertial tracking in slim form-factor

    US20190087021A1

  • Representation of user position, movement, and gaze in mixed reality space

    US20190340822A1

  • Picture-Taking Within Virtual Reality

    US20190377416A1