Image display system

The video display system addresses misalignment issues in head-mounted displays by calculating and presenting relative positions and directions within 360-degree videos, improving user experience in virtual reality applications.

JP7702675B2Active Publication Date: 2025-07-04PANASONIC HOLDINGS CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023514693
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-16
Filing Date
2022-04-18
Publication Date
2025-07-04
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

Existing head-mounted display systems struggle to provide appropriate video content to users during virtual reality experiences, particularly in scenarios where user movement and guide instructions lead to misalignment between the intended and perceived directions of attention within a 360-degree video.

Method used

A video display system that includes a photographing unit for capturing wide-angle videos, data acquisition for cue information, metadata configuration, and a display state estimation unit to calculate relative positions and directions, enabling accurate presentation of video content based on user orientation and guide instructions.

Benefits of technology

The system effectively presents appropriate video content by aligning user attention with guide instructions, reducing confusion and enhancing the virtual reality experience by ensuring the user correctly identifies the intended direction of attention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702675000001
    Figure 0007702675000001
  • Figure 0007702675000002
    Figure 0007702675000002
  • Figure 0007702675000003
    Figure 0007702675000003
Patent Text Reader

Abstract

This video display system comprises: an observation device having an imaging unit that generates a wide angle view video, a data acquisition unit that acquires data concerning the position and / or direction of an attention object to be closely watched by a user of a display device in the wide angle view video and queue information for providing notification of a change in state of the observation system; and a VR device having a reception unit that receives the wide angle view video, the data, and the queue information, and a difference calculation unit for calculating, on the basis of a difference between the position and / or direction of the display device in the wide angle view video and the position and / or direction of the attention object on metadata, at least one of a relative position and a relative direction which are respectively the position and the direction, of the attention object, relative to the position and / or direction of the display device in the wide angle view video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a video display system, an observation device, an information processing device, an information processing method, and a program.

Background Art

[0002] In recent years, development of so-called head-mounted displays, which are head-mounted display devices, has been actively carried out. For example, Patent Document 1 discloses a head-mounted display capable of presenting (that is, displaying) a video of content and a video of the outside world. In the head-mounted display disclosed in Patent Document 1, by adjusting the brightness of at least one of the video of content and the video of the outside world, the discomfort given to the user when switching between the video of content and the video of the outside world is reduced.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] By the way, as an application that takes advantage of the high immersion of a display device such as a head-mounted display, there is an application such as pseudo-experiencing an experience at a certain location by watching a video from a remote location. At this time, it is required that an appropriate video be provided to the display device.

[0005] The present disclosure has been made in view of the above, and an object thereof is to provide a video display system and the like capable of displaying an appropriate video.

Means for Solving the Problems

[0006] In order to achieve the above object, one aspect of the video display system according to the present disclosure is a video display system for displaying a display video by a display device, including a photographing unit that generates a wide-angle video, data regarding at least one of the position and direction of a fixation target that causes a user of the display device to fixate within the wide-angle video, and a data acquisition unit that acquires cue information for notifying a change in the state of the observation system, a metadata configuration unit that sets the data from the data acquisition unit as metadata together with other information, a transmission unit that transmits the wide-angle video together with the metadata, an observation device including these components, a reception unit that receives the wide-angle video, the data, and the cue information, a display state estimation unit that estimates at least one of the position and direction of the display device within the wide-angle video of the display device, a difference calculation unit that calculates at least one of a relative position, which is the position of the fixation target relative to at least one of the position and direction of the display device within the wide-angle video of the display device, and a relative direction, which is the direction of the fixation target relative to the same, based on the difference between at least one of the estimated position and direction of the display device within the wide-angle video of the display device and at least one of the position and direction of the fixation target on the metadata, a presentation unit that presents information regarding at least one of the calculated relative position and relative direction, an instruction based on the cue information, and the state of the observation system to the user of the display device, a video generation unit that generates the display video including information regarding at least one of the position and direction of the display device within the wide-angle video of the display device estimated by the display state estimation unit from the received wide-angle video, an instruction based on the cue information, and a partial image corresponding to a visual field portion according to the state of the observation system, and a VR device including the display device that displays the display video.

[0007] Also, one aspect of the information processing apparatus according to the present disclosure is an information processing apparatus used in a video display system for causing a display device to display at least a part of the display video within a wide viewing angle video, the information processing apparatus receiving metadata based on data obtained by receiving an input, the metadata being data regarding at least one of the position and direction of a fixation target that causes a user of the display device to fixate on the wide viewing angle video, and a difference calculation unit that calculates and outputs at least one of a relative position that is the position of the fixation target relative to at least one of the position and direction within the wide viewing angle video of the display device, and a relative direction that is the direction of the fixation target relative, based on the difference from at least one of the position and direction of the fixation target on the metadata.

[0008] Also, one aspect of the information processing method according to the present disclosure is an information processing method for causing a display device to display at least a part of the display video within a wide viewing angle video, the method receiving metadata based on data obtained by receiving an input, the metadata being data regarding at least one of the position and direction of a fixation target that causes a user of the display device to fixate on the wide viewing angle video, and calculating and outputting at least one of a relative position that is the position of the fixation target relative to the orientation of the display device, and a relative direction that is the direction of the fixation target relative, based on the difference between at least one of the estimated position and direction within the wide viewing angle video of the display device and at least one of the position and direction of the fixation target on the metadata.

[0009] These general or specific aspects may be implemented by a system, an apparatus, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be implemented by any combination of a system, an apparatus, an integrated circuit, a computer program, and a recording medium.

Advantages of the Invention

[0010] According to the present disclosure, a video display system or the like capable of displaying an appropriate video is provided.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48

Figure 49

DETAILED DESCRIPTION OF THE INVENTION

[0012] (Knowledge underlying the disclosure) In recent years, display devices have been developed that can be worn on the user's head to place a display unit in front of the eyes, enabling the user to visually recognize the seemingly displayed image on a large screen. Such display devices are called head-mounted displays (HMDs) and have the characteristic that images can be optically visually recognized as a large screen. In addition, in some HMDs, the user can feel the video being watched stereoscopically by displaying videos that generate visual differences corresponding to each of the user's right and left eyes. And due to the improvement of communication quality in recent years, videos taken by observation devices placed in remote locations can be viewed almost in real time with a delay of about several milliseconds to several tens of milliseconds, making it possible to have an experience as if being on the spot without visiting the local area. Utilizing this technology, virtual tourism experiences such as sightseeing trips, exhibition tours, inspections, factory tours, and museum / zoo / aquarium tours (hereinafter also referred to as pseudo-tourism or VR (Virtual Reality) tourism) have also come to be realized.

[0013] In such VR tourism, cameras that can capture 360-degree (full longitude) videos (so-called omnidirectional cameras) are used as observation devices. The videos taken by the observation devices are wide-angle videos and form a three-dimensional image space. The user side of the display device can display the videos (that is, cut out and display the images within the visual field range in an arbitrary direction of the images constituting the three-dimensional image space). For example, if the display device is equipped with a function that can detect the direction the user is facing, a part of the video corresponding to the user's orientation in the three-dimensional image space can be cut out and displayed, so it is possible to provide a viewing experience that meets the needs of many users from a single camera video.

[0014] Here, when a user is watching a part of the video in an arbitrary direction and is instructed by a guide guiding within this video to look at a predetermined object of attention, the user may not understand where in the three-dimensional image space this instruction is to look. For example, when the user is looking in the 3 o'clock direction and a guide existing in the 12 o'clock direction gives a voice instruction such as "Please look at your left hand" to make the user look in the 9 o'clock direction, the user will look in the 12 o'clock direction instead of the 9 o'clock direction. Thus, when the guide gives an instruction to make the user look at the object of attention on the premise that the user is looking straight ahead, the user may not understand the object of attention intended by the guide.

[0015] Therefore, in the present disclosure, in order to suppress such a situation where the object of attention cannot be understood, it is an object to provide a video display system capable of presenting the direction corresponding to this object of attention to the user. In the present disclosure, it will be described that a 360-degree wide-angle video is captured by an observation device, but the wide-angle video may be a video captured in an arbitrary angle range such as 270 degrees or more, 180 degrees or more, or 150 degrees or more. Such a wide-angle video only needs to have a field of view angle wider than at least the field of view angle of the video displayed by the user on the display device side. Also, in the present disclosure, a video display system assuming the occurrence of movement of the video in the horizontal plane will be described, but it is also applicable to the movement of the video occurring on an intersection plane intersecting the horizontal plane including a vertical component.

[0016] Hereinafter, a conventional video display system and the like will be described in more detail with reference to the drawings. FIG. 1 is a diagram for explaining a conventional example. As shown in FIG. 1, conventionally, a service called VR tourism (first-person experience) has been provided. In VR tourism, when the local VR space is appropriately reproduced, a tourism experience as if one were at that place is possible. Examples of services using 360° camera shooting include FirstAirlines (https: / / firstairlines.jp / index.html) and Travel Assistant (https: / / www.tokyotravelpartners.jp / kaigotabisuke-2 / ). Examples of services using 3D CG (computer graphics) include Google Earth VR and Boulevard (https: / / www.blvrd.com / ).

[0017] FIG. 2 is a diagram for explaining a conventional example. As shown in FIG. 2, in addition to VR tourism, a service (also referred to as a third-person experience) is also provided in which a video shot on-site is displayed on a display device such as a television and the video is viewed from a third-person perspective. In the third-person experience, there is a service specialized for users guided by experts, and it has features such as the possibility of monetization according to an individual's hobbies.

[0018] FIG. 3 is a diagram for explaining a conventional example. As shown in FIG. 3, when developing VR tourism, as a basic configuration, a VR system main body, a controller 311, a computer or a smartphone 313, a network, a cloud 314, an observation system 315, etc. are required. The VR system main body used to be heavy and only of the HMD type that covers the face considerably, but with the small-sized glasses-type VR glasses, it has become easier to use for a long time and is becoming more widely used. The VR system main body includes an All in One type that includes functions necessary for the VR system main body and a tethered type that entrusts some functions to a computer or a smartphone. The controller is used for selecting a menu, moving in the VR space, etc. The computer or smartphone may be only for communication function or may constitute a part of the VR system.

[0019] The network and the cloud 314 connect the observation system and the VR system, and in some cases, implement some functions of the observation system or the VR system on a computer system on the cloud. For the observation system, a 360° camera with a wireless function or a 360° camera wirelessly or wiredly connected to a smartphone or a computer, a 180° camera, or a wide-angle camera is used. Through these devices, etc., the user 312 can visually recognize the guide and the building or scenery of the tourist destination within the VR space.

[0020] Note that in the description, VR tourism using a 360° camera is taken as an example, but any VR glasses that allow participants to change their viewpoints, such as those using a 180° camera, may be used. Also, although an example of photographing and guiding an actual scenery may be described, instead of the actual scenery, a virtual space composed of computer graphics can be used with a virtual camera, and the guide can also enter the virtual space using VR glasses, etc., and realize tourism by playing back an image within the virtual space. Therefore, the present invention can also be applied to such applications. Typical examples of the above include areas and spaces where ordinary travelers cannot easily go, such as moon travel, and VR travel to spaces.

[0021] FIG. 4 is a diagram for explaining a conventional example. In FIG. 4, in the case of 360° camera shooting, the VR tourism service (without guide: upper part (hereinafter referred to as Conventional Example 1), with guide: middle part (hereinafter referred to as Conventional Example 2)) and an example of a three-person experience, the Zoom (registered trademark) tourism service conventional example (lower part (hereinafter referred to as Conventional Example 3)) are shown in a schematic configuration. Hereinafter, in the present invention, the term "voice" or "voice data" and "voice information" shall include not only conversations but also audio signals including music and, in some cases, ultrasonic waves outside the audible band. In the VR tourism service, on the observation system (tourist destination) side, pre-recorded videos are sent, or the VR system side operates a 360° camera, a robot, or a drone so that the VR system side can view VR videos. As shown in the middle part, it is also possible that there is a guide or a camera operator on the observation system side, and the VR videos such as 360° cameras can be enjoyed as VR in the VR system. Also, as shown in the lower part, in the three-person experience, a 2D video is sent from the observation system side in 2D by a multi-person remote conversation service with voice and video such as Zoom, and the video of the tourist destination can be viewed and enjoyed remotely.

[0022] FIG. 5 is a diagram for explaining a conventional example. The overall system configuration of Conventional Example 2 will be described. Conventional Example 1 is different from Conventional Example 2 in that pre-recorded VR videos are used or operations are performed from the VR system side, and the difference will also be explained. The observation system of Conventional Example 2 is composed of a camera for VR shooting, for example, a 360° camera, and a communication device for sending the captured information to a remote location.

[0023] A 360° camera for VR shooting synthesizes (stitches) the images of a plurality of cameras shooting in different directions into one video, maps it onto a plane, for example, by the equirectangular projection (ERP) method, appropriately compresses it as an ERP image, and transmits it from a communication device to a remote VR system together with audio data captured by a microphone. The 360° camera may be mounted on a robot, a drone, etc. The 360° camera or a robot or a drone equipped with it is operated by a photographer or a guide. However, in Conventional Example 1, there are cases where it is operated on the VR system side, or cases where pre-recorded images, etc. are received on the VR system side. Thus, the three-dimensional image space means that the images constituting the space are not only those that allow the user to experience depth, but also the resulting displayed image is a planar image and includes a plurality of planar images arranged on a virtual three-dimensional surface.

[0024] On the VR system side, contrary to the observation system, the received planar video (ERP image) is converted into a spherical video, and a part is cut out according to the orientation and position of the observer, etc., and displayed on a VR display device. In Conventional Example 3, since the received video is 2D, it is displayed as 2D, and mostly a 2D display device such as a tablet, a smartphone, or a TV is used. The above is the same for the case of receiving pre-recorded video in Conventional Example 1.

[0025] When operating on the VR system side, when the observation system operates in conjunction with the orientation and position on the VR system side, there are cases where the observation system operates by key operations such as a mouse, a tablet, a joystick, a keyboard, etc., or by selecting a menu or an icon on the screen. In such cases, appropriate control data is sent from the VR system side to the observation system, and it is necessary for the state of the observation system, that is, the orientation and position, etc., to be sent to the VR system.

[0026] FIG. 6 is a diagram for explaining a conventional example. Using the comparison between the 360° video and the normal video shown in FIG. 6, the resolution when viewing the 360° video on a VR system will be described. When viewing a 4K video of 360° on a VR device with a field of view (FOV) of 100 degrees, the resolution of the video cut out for VR display is only 1067×600 (about twice that of SD video). Since a VR system using a panel with a resolution of 2K×2K for one eye is displayed on a square panel and is stretched further by a factor of 2 in the vertical direction, the resulting video has a very low resolution.

[0027] For an 8K video, the VR display resolution is 2133×1200, and in terms of data volume, it is 1.23 times the area of Full HD (1920×1080). However, since it is stretched by a factor of 2 in the vertical direction, the resulting video is of about the Full HD level. For 11K shooting (10560×5940), the VR resolution is 2933×1650, which is comparable to a VR system.

[0028] In order to provide a VR tourism experience with high resolution and high sense of presence, shooting at a minimum of 8K and preferably 11K is required. Shooting at 8K and 11K requires large equipment, high video transfer rates, and large capacities. Therefore, both shooting and distribution become expensive.

[0029] Therefore, it is essential to avoid VR sickness, make it easy to understand and use, and enable many users to use it, thereby reducing the unit cost per user. Also, the effective use of VR recording content becomes important for the business to succeed.

[0030] FIG. 7 is a diagram for explaining a conventional example. The main functional configuration examples of Conventional Examples 1 and 2 will be described by function. The observation systems 751 of Conventional Examples 1 and 2 include a VR imaging means 762 (VR imaging camera) for performing VR imaging, a VR video processing means 758 for processing the video captured by the VR imaging means 762 into an image suitable for transmission, a VR video compression means 756 for compressing the VR video processed by the VR video processing means 758 into a data rate and video signal format suitable for transmission, an audio input means 763 composed of a microphone for inputting a guide or surrounding audio, an audio compression means 760 for converting the audio signal input by the audio input means 763 into a data rate and audio signal format suitable for transmission, a graphics generation means 759 for generating auxiliary information as graphics, a multiplexing means 757 for converting the video signal, audio signal, and graphics information compressed by the VR video compression means 756, the graphics generation means 759, and the audio compression means 760 into a signal suitable for transmission, a communication means 754 for sending the multiplexed communication observation signal to a plurality of VR systems 701 and receiving the communication audio signal from the plurality of VR systems 701, a separation means 755 for extracting the compressed audio signal from the communication audio signal received by the communication means 754, an audio decoding means 761 for extracting the audio signal from the compressed audio signal from the separation means 755, and an audio output means 764 for outputting the audio signal decoded by the audio decoding means 761 as sound.

[0031] In this example, it is assumed that the VR video processing means 758, the VR video compression means 756, and the graphics generation means 759 are realized within the GPU, and the audio compression means 760, the multiplexing means 757, the separation means 755, and the audio decoding means 761 are realized within the CPU. However, it is not necessarily limited to this. In a simpler configuration, the CPU and the GPU may be realized as one processor, but their functional configurations and operations are the same.

[0032] The VR shooting means 762 is, for example, a 360° camera, but is composed of a plurality of cameras that shoot in different directions. In the VR video processing means, the outputs of the plurality of cameras are synthesized (stitched) into one video, and this is mapped onto a plane by, for example, the equirectangular projection (ERP) method and output as an ERP image.

[0033] Conversely, the VR systems 701 of the conventional examples 1 and 2 include a communication means 716 that receives the communication observation signal sent from the observation system 751 or transmits the voice input in the VR system 701 to the observation system 751 as communication voice information, a separation means 715 that separates and outputs a compressed VR video (ERP image), graphics information, and compressed voice information from the communication observation signal from the communication means 716, a VR video decoding means 710 that decodes the compressed VR video (ERP image) from the separation means 715, converts the ERP image from the VR video decoding means 710 into a spherical video, cuts out a part according to the control information from the VR control means 707, and makes it a video that can be displayed on the VR display means 704, and a VR display control means 708 that outputs a VR video to be displayed on the VR display means 704 together with the graphics information of the graphics generation means 712 that converts the graphics information to be displayed from the graphics information output from the separation means 715. The VR display means 704 outputs the VR video from the VR display control means 708 for viewing with both eyes. The outputs of the rotation detection means 703 that detects the inclination of the VR display means 704 in the front, rear, left, and right directions or the direction of the white of the eyes and the position detection means 702 that detects the position of the VR display means 704 in the left, right, front, rear, and height directions are sent to the VR control means 707, and the video displayed on the VR display means 704 by the output of the VR control means 707 and the voice output by the voice reproduction means 709 by the voice reproduction control means are appropriately controlled. The compressed voice information separated by the separation means 715 is decoded by the voice decoding means 713 and sent to the voice reproduction control means 709 as voice information. In the voice reproduction control means 709, according to the control information from the VR control means 707, the balance in the left, right, front, rear, and height directions, and in some cases, frequency characteristics, delay processing, or synthesis of an alarm as the VR system 701, etc. are performed. Also in the graphics generation means 712, graphics for displaying the system menu, warnings, etc. of the VR system 701 are generated and overlaid on the VR image and displayed on the VR display means 704.The VR system 701 is provided with voice input means 706 for inputting the voice of the user of the VR system 701. The voice information from the voice input means 706 is compressed by the voice compression means 714 and sent as compressed voice information to the multiplexing means 717, where it is sent as communication voice information from the communication means 716 to the observation system 751.

[0034] FIG. 8 is a diagram for explaining a conventional example. As a typical realization example of the observation system of the conventional example 2, a realization example of a 360° camera 801 will be described.

[0035] A typical example of the 360° camera 801 combines two imaging systems, namely, an ultra-wide-angle lens 854, a shutter 853, and an image sensor 852, to capture a 360° video in the front, back, up, and down directions. Since there may be cases where two or more imaging systems are combined to capture higher-quality videos, in this example, the VR shooting camera 804 is illustrated as having two or more imaging systems. The imaging system may be configured by combining independent cameras. In that case, generally, there is a high-speed digital video I / F after the ADC 851 of the video, and thus it is connected to a high-speed digital video input connected to a video system bus connected to the GPU (Graphics Processing Unit) 803 or the CPU (Central Processing Unit) 802. Here, it will be described as an integrated unit.

[0036] The main components of the 360° camera 801 include the VR shooting camera 804 composed of the plurality of imaging systems described above, the GPU 803 mainly for processing video data and graphics, the CPU 802 for general data processing, processing related to input / output, and controlling the entire 360° camera 801, the EEPROM (Electrical Erasable Programmable ROM) 813 for storing programs to operate the CPU 802 and GPU 803, the RAM 814 used for storing data for the operation of the CPU 802 and GPU 803, the SD card (registered trademark) 821 which is a removable memory for storing video, audio, or programs, the wireless communication element 820 for performing wireless communication via WiFi (registered trademark) or Bluetooth (registered trademark) for data exchange with the outside and receiving operations from the outside, the buttons and display elements 808 for operations and display, the battery 807 and the power control element 812, the audio input section composed of a plurality of microphones (microphone group 819) or microphone terminals 825, the microphone amplifier 818, and the ADC 817, the audio output section composed of the speaker 826 or headphone terminals 824, the amplifier 823, and the DAC 822, the video system bus mainly connecting the VR shooting camera 804 and the CPU 802 and used for reading digital video data, the memory bus connecting the aforementioned EEPROM 813, RAM 814, SD card 821 with the GPU 803 and CPU 802 for data exchange with the memory, the system bus to which the aforementioned CPU 802, GPU 803, wireless communication element 820, audio input section, and audio output section are connected for control and data exchange, the I / O bus to which the aforementioned buttons and display elements 808, power control element 812, and although not shown, the audio input section, audio output section, and VR shooting camera 804, etc. are connected for control and low-speed data exchange, and several bus conversion units 815 and 816 for connecting each bus. The motion / position detection unit 860 is further connected to the I / O bus. For some processes, whether they are performed by the GPU 803 or the CPU 802 may be different in this example, and the bus configuration may also be different from this example, but there is no difference in the functional configuration and operation described later.

[0037] Each VR imaging camera 804 includes a lens 854 for capturing a wide-angle video, an imaging device 852 that converts the light collected by the lens 854 into an electrical signal, a shutter 853 that is located between the lens 854 and the imaging device 852 and blocks light, and an aperture (not shown here) that is located at the same position as the shutter 853 and controls the intensity of the light from the lens 854. It is composed of an ADC 851 that converts the analog electrical signal from the imaging device 852 into a digital video signal. Although not shown, each is controlled by the CPU 802 through an I / O bus, and the state is notified to the CPU 802.

[0038] Buttons include a power switch 806 for turning the power on / off, a shooting start / end button 811 for operating the start / end of shooting, a shooting mode selection button 809 (which may not be provided) for changing the shooting mode, and a zoom button 810 for moving the lens 854 and digitally controlling the angle of view for zooming in and out.

[0039] The power control element 812 may be integrated with the battery 807. It stabilizes the voltage, manages the battery capacity, etc., and supplies power to all components (not shown). Furthermore, it supplies power to the HMD / VR glasses through USB or AV output.

[0040] Each function realized by the GPU 803 is realized by dedicated hardware and programs such as image processing. Generally, the functions realized by the CPU 802 are realized by general-purpose hardware and programs. As an example, the GPU 803 is used to realize a VR video processing unit 842, a VR video compression unit 841, a graphics generation unit 843, etc. Also, as an example, the CPU 802 is used to realize a memory control unit 835, a multiplexing unit 832, an audio compression unit 833, an audio decoding unit 834, and a separation unit 831.

[0041] FIG. 9 is a diagram for explaining a conventional example. Based on FIG. 9, an example of realizing a VR system 901 will be described as a typical realization example of the observation system of Conventional Example 2. In this realization example, it is assumed that the VR system 901 is composed of a computer or smartphone 951 and an HMD or VR glasses 902 connected thereto. There is also an example of realizing it with only the HMD or VR glasses 902. In that case, it can be considered that the functions of both CPUs and GPUs are combined, and the peripheral functions are also integrated.

[0042] The main components of the computer / smartphone 951 in the VR system 901 include a high-speed communication element 970 such as WiFi or Ethernet (registered trademark) for connecting to the observation system, a GPU 954 mainly for processing video data and graphics, a CPU 965 for general data processing and overall control of the computer / smartphone 951, a non-volatile memory 962 such as a hard disk or flash memory for storing programs for operating the CPU 965 and GPU 954, a RAM 961 for storing data for the operation of the CPU 965 and GPU 954, a power switch 963 and a power control element 964 for supplying power to each part, an AV output 952 for outputting video and audio signals to the HMD / VR glasses 902, an I / F such as a USB 953 for controlling the HMD / VR glasses 902 and acquiring data therefrom, a memory bus for connecting the RAM 961 and non-volatile memory 962 and for the CPU 965 and GPU 954 to access, a system bus for the CPU 965 and GPU 954 to access the AV output 952, USB 953, and communication element 970, a bus connection (bus conversion unit 960) for connecting the system bus and the memory bus, and although not shown here, a display device, an input device for operation, and other general-purpose I / Fs, etc.

[0043] For some processes, whether they are performed by the GPU 954 or the CPU 965 may be different from this example, and the bus configuration may also be different from this example, but there is no difference in the functional configuration and operations described later. As an example, the GPU 954 is used to implement the motion / position detection processing unit 955, the VR control unit 956, the VR display control unit 957, the VR video decoding unit 958, and the graphics generation unit 959, etc. Also, as an example, the CPU 965 is used to implement the audio decoding unit 966, the audio playback control unit 967, the multiplexing unit 968, and the demultiplexing unit 969.

[0044] Also, the AV output 952 and the USB 953 can be replaced with a high-speed bidirectional I / F, such as an I / F like USB Type-C (registered trademark). In that case, the HMD / VR glasses 902 side will also be connected with the same I / F or connected with a converter for I / F conversion. Generally, when sending video via the USB 953, appropriate video compression is performed to compress the data volume, so appropriate video compression is performed by the CPU 965 or the GPU 954, and the video is sent to the HMD / VR glasses 902 via the USB 953.

[0045] The main components of the HMD / VR glasses 902 in the VR system 901 include an audio input unit consisting of a microphone 906 for inputting audio, a microphone amplifier 917, and an ADC 918; an audio output unit consisting of a speaker 907 or a headphone terminal 908, an amplifier 919, and a DAC 920; a VR display unit consisting of two sets of lenses 904 and a display element 905 for the user to view VR images; a motion / position sensor 903 consisting of a motion / position detection unit and an orientation detection unit composed of a gyro sensor, a camera, or an ultrasonic microphone, etc.; a wireless communication element 927 such as Bluetooth for communicating with a controller (not shown); a volume button 909 for controlling the output volume from the audio output unit; a power switch 921 for turning on / off the power of the HMD / VR glasses; a power control element 924 for power control; the aforementioned EEPROM 913, RAM 914, an SD card, a memory bus connecting the GPU 910 and the CPU 915 to perform data exchange with the memory; the aforementioned CPU 915, GPU 910, wireless communication element 927, an AV input 925 for receiving video signals and audio signals from a computer / smartphone 951; an I / F such as a USB 926 for receiving control signals from the computer / smartphone 951 and sending video, audio signals, and motion / position data; a CPU 915 mainly for controlling the overall HMD / VR glasses 902, such as performing control of audio compression (realized by the audio compression unit 916), switches, power, etc.; a GPU 910 mainly for performing video display processing (realized by the video display processing unit 912) for adjusting the video to the VR display unit and correcting and shaping the motion / position information sent to the computer / smartphone 951 from the information of the motion / position sensor 903; an EEPROM 913 for storing programs and data for operating the CPU 915 and the GPU 910; a RAM 914 for storing data during the operation of the CPU 915 and the GPU 910; a memory bus for connecting the CPU 915, GPU 910, RAM 914, and EEPROM 913; a system bus to which the CPU 915, GPU 910, USB 926, audio input unit, audio output unit, and wireless communication element 927 are connected to perform control and data exchange; the aforementioned buttons and power control element 924, motion / position sensor 903, and although not shown, the audio input unit,It is composed of an I / O bus that performs control and low-speed data exchange, including an audio output unit and a VR shooting camera, and several bus conversion units 922 that connect the respective buses. For some processes, whether they are performed by the GPU 910 or the CPU 910 may be different from this example, and the bus configuration may also be different from this example, but there is no difference in the functional configuration and operation described later.

[0046] Since the video data from the AV input 925 has a large data volume and is high-speed, it is illustrated as being directly taken in by the GPU 910 when the system bus does not have sufficient speed.

[0047] Note that the video information captured by the camera of the motion / position sensor 903 may be sent to the display element as information for the user to check the periphery of the HMD / VR glasses 902, or may be sent to the computer / smartphone 951 through the USB 926 for the user to monitor whether there is a dangerous situation.

[0048] The power control element 924 receives power supply from the USB 926 or the AV input 925, performs voltage stabilization, battery capacity management, etc., and supplies power to all components not shown. In some cases, the battery 923 may be provided internally or externally and connected to the power control element 924.

[0049] The states of the buttons and the cursor of the controller not shown are acquired by the CPU 915 through the wireless communication element 927 and are used for button operations, movement, and application operations in the VR space. The position and orientation of the controller are detected by a camera or an ultrasonic sensor in the motion / position detection unit. After appropriate processing is performed by the motion / position sensor, it is used for control by the CPU 915 and is also sent to the computer / smartphone 951 via the USB 926, and is used for the drawing of graphics and image processing executed by the program executed by the CPU 915 or the GPU 910. Since the basic operations are not directly related to the present invention, they are omitted.

[0050] FIG. 10 is a diagram for explaining a conventional example. An implementation example of an integrated VR system 1001 having a function for VR in a computer / smartphone and HMD / VR glasses will be described.

[0051] As can be seen in FIG. 10, the functions of the computer / smartphone and the HMD / VR glasses are integrated, and the respective functions of the CPU and GPU are realized by one CPU and GPU.

[0052] The communication element 1033 is typically WiFi for performing wireless communication, and since it does not have a power cable, it has a battery 1026. It has an interface with a general-purpose computer such as USB1034 for charging the battery 1026 and initial settings.

[0053] Since the integrated VR system 1001 does not require an AV output, AV input, or USB to connect the computer / smartphone and the HMD / VR glasses, high-quality and delay-free transmission of AV information and efficient control are possible. However, by integrating them, due to limitations in size, it may not be possible to use a high-performance CPU 1027 or GPU 1006 due to power, heat, and space limitations, and the VR function may be limited.

[0054] However, not being connected by a cable increases the degree of freedom and can expand the range of applications.

[0055] Also, by realizing part of the function on a computer in the cloud, etc., the lack of performance can be compensated for and a high-functional application can be realized.

[0056] The integrated VR system 1001, similar to the configuration described in FIGS. 8 and 9, further includes a lens 1002, a display element 1011, a microphone 1003, a microphone amplifier 1007, an ADC 1009, a speaker 1004, a headphone terminal 1005, an amplifier 1008, a DAC 1010, a RAM 1019, an EEPROM 1020, a bus conversion 1021, a motion position sensor 1022, a power switch 1023, a volume button 1024, and a power control element 1025. Also, video display processing 1012, motion / position detection processing 1013, VR control 1014, VR display control 1015, motion / position detection 1016, VR video decoding 1017, and graphics generation 1018 are realized using the GPU 1006. Further, audio compression 1028, audio decoding 1029, audio playback control 1030, multiplexing 1031, and demultiplexing 1032 are realized using the CPU 1027.

[0057] FIG. 11 is a diagram for explaining a conventional example. Based on FIG. 11, a more detailed configuration of a VR video processing unit 1103 that processes video captured by a VR shooting camera 1151 of the observation systems of Conventional Examples 1 and 2 will be described.

[0058] As described above, the VR shooting camera has a plurality of cameras cm for shooting 360° omnidirectional video, typically cameras cm with ultra-wide-angle lenses, and rectangular individual videos with the same pixels captured by each camera cm are input to a VR video processing unit 1103 realized by a program or a dedicated circuit in the GPU 1101.

[0059] In the VR video processing unit 1103, first, the input plurality of videos are evaluated for the shooting direction of each camera cm and the captured video, and the videos captured by each camera cm are input to a stitching processing unit 1105 that performs a process of synthesizing and connecting them so as to form a continuous spherical video. The spherical video data output from the stitching processing unit 1105 is mapped onto a plane by a VR video mapping unit 1104, for example, by the equirectangular projection (ERP) method, and output from the VR video processing unit 1103 as an ERP image and passed to the next VR video compression unit 1102.

[0060] Note that, although the connection between the video bus and the cameras is illustrated such that each camera is connected to the bus, within the VR shooting camera 1151, the videos shot by each camera may be combined into one signal and sent to the video bus in a time-division manner, and then input to the VR video processing unit 1103. In a simple configuration, since there are two cameras cm, instead of using a bus, the GPU 1101 may receive the outputs of the two cameras respectively, and the VR video processing unit 1103 may receive and process the videos shot in parallel.

[0061] FIG. 12 is a diagram for explaining a conventional example. Based on FIG. 12, a more detailed configuration of the VR display control unit 1204 of the VR systems of Conventional Examples 1 and 2 will be described.

[0062] As described above, the VR display control unit 1204 is implemented by a program or a dedicated circuit in the GPU 1201 of a computer / smartphone, and is composed of a mapping unit 1206 and a display VR video conversion unit 1205.

[0063] The operation is as follows. The communication element 1261 receives the communication data sent from the observation system, the compression video is separated by the separation unit 1232 of the CPU 1231, the GPU 1201 receives the video via the memory bus, and is decoded by the VR video decoding unit 1207 to become a planar video (ERP image). The planar video is converted into a 360° spherical video by the mapping unit 1206 of the VR display control unit 1204, and in the subsequent display VR video conversion 1205, the portion to be displayed by the VR display means 1202 is cut out based on the control information output by the VR control unit 1203.

[0064] Specifically, the center of the ERP image is used as the entire surface, and the origin of the 360° spherical video. The initial video of the VR video displayed on the VR display means 1202 is centered on the origin and, according to the capabilities of the VR display means 1202, the video for the right eye is slightly shifted to the right and the video for the left eye is slightly shifted to the left. In the height direction, the videos are cut out using the initial set values and are displayed on the display elements for the right eye and the left eye. From here, the cut-out position changes according to the rotation of the VR system to the left and right or looking up and down.

[0065] Generally, the video from a 360° camera does not change with the movement of the VR system. However, in the case of video generated by CG, the position changes due to the movement of the VR system or operations with a controller.

[0066] The initial value of the extraction from a 360° spherical video may be from the previous extraction position. Generally, however, a function to return to the initial position is provided.

[0067] FIG. 13 is a diagram for explaining a conventional example. An operation example of Conventional Example 2 will be described based on FIG. 13.

[0068] In the observation system, the audio input unit (microphone array, microphone terminal, microphone amplifier, ADC) inputs audio (S1325), and the audio compression unit performs audio compression (S1326).

[0069] At the same time, a plurality of cameras (lens, shutter, imaging device, ADC) of the VR shooting camera shoot a moving image (S1321). The stitching processing unit of the VR video processing unit stitches it into a spherical video with camera 1 at the center as the center (S1322). The VR video mapping unit generates an ERP image by an orthographic cylindrical projection method or the like (S1323), and the VR video compression unit appropriately compresses it (S1324).

[0070] The compressed ERP image and audio information are multiplexed by the multiplexing unit (S1327) into a transmissible format, and are sent (transmitted) to the VR system by the wireless communication element (S1328).

[0071] Over time, in some cases, it moves to a new direction and position (S1329), and the sending is repeated from the audio input and shooting with a plurality of VR shooting cameras.

[0072] Here, the graphics information may be superimposed on the video before video compression or may be multiplexed together with the video and audio as the graphics information, but this is omitted.

[0073] In a VR system, in a computer / smartphone, information sent from an observation system is received by a communication element (S1301) and sent to a separation unit. The separation unit separates the sent compressed video information and compressed audio information (S1302). The compressed audio information separated by the separation unit is sent to an audio decoder and decoded (S1303) to become uncompressed audio information. The audio information from the audio decoder is sent to an audio playback control unit, and audio processing is performed by the audio playback control unit based on the position / orientation information of the VR observation system sent via the system bus from the VR control unit of the GPU (S1304). The audio information on which audio processing has been performed is sent via the system bus, either through AV output or via USB, to the audio output unit (DAC, amplifier, speaker, and headphone terminal) of the HMD / VR glasses and output as audio (S1305). As audio processing, balance control of volume in the left and right or in space, change of frequency characteristics, delay, movement in space, similar processing for only a specific sound source, addition of sound effects, etc. are performed.

[0074] The compressed video signal is sent from the video data from the separation unit of the CPU of the computer / smartphone to the VR video decoder of the GPU via the memory bus, and is decoded in the VR video decoder (S1307) and input to the VR display control unit as an ERP image. In the VR display control unit, the ERP video is mapped to a 360° spherical video by a mapping unit (S1308), and an appropriate part is cut out from the 360° spherical video based on the position / orientation information of the VR system from the VR control unit in the display VR video conversion unit (S1309), and is displayed as a VR video by the VR display unit (display element, lens) (S1310).

[0075] Receiving from the observation system, video display, and audio output are repeated.

[0076] Note that regarding graphics here, there are cases where graphics are separated simultaneously with video / audio separation and superimposed on the VR video by the VR display control unit, or cases where they are generated within the VR system and superimposed on the VR video, etc., and the explanation is omitted here.

[0077] (Summary of the Disclosure) The summary of the present disclosure is as follows.

[0078] A video display system according to an aspect of the present disclosure is a video display system for displaying a display video by a display device, including a photographing unit that generates a wide viewing angle video, data regarding at least one of the position and direction of a fixation target that causes a user of the display device to fixate within the wide viewing angle video, and a data acquisition unit that acquires cue information for notifying a change in the state of an observation system; a metadata configuration unit that sets the data from the data acquisition unit together with other information as metadata; a transmission unit that transmits the wide viewing angle video together with the metadata; an observation device including these components; a reception unit that receives the wide viewing angle video, the data, and the cue information; a display state estimation unit that estimates at least one of the position and direction of the display device within the wide viewing angle video; a difference calculation unit that calculates at least one of a relative position, which is the position of the fixation target relative to at least one of the position and direction of the display device within the wide viewing angle video, and a relative direction, which is the direction of the fixation target relative to the display device, based on the difference between at least one of the estimated position and direction of the display device within the wide viewing angle video and at least one of the position and direction of the fixation target on the metadata; a presentation unit that presents information on at least one of the calculated relative position and relative direction, an instruction based on the cue information, and the state of the observation system to the user of the display device; a video generation unit that generates a display video including information on at least one of the position and direction of the display device within the wide viewing angle video estimated by the display state estimation unit from the received wide viewing angle video, and a partial image corresponding to a visual field portion according to an instruction based on the cue information and the state of the observation system; and a display device that displays the display video. A VR device including these components is provided.

[0079] Such a video display system calculates at least one of a relative position, which is a relative position of a target of attention with respect to at least one of the position and direction of the target of attention that causes a user of the display device to pay attention by using metadata, and a relative direction, which is a relative direction of the target of attention. Then, since at least one of the relative position and the relative direction is presented to the user, it is possible to suppress the problem that the user loses sight of the position of the target of attention when the user is movable. Therefore, according to the video display system, it is possible to display an appropriate video from the viewpoint of suppressing the problem that the user loses sight of the position of the target of attention when the user is movable.

[0080] Further, for example, the video display system may further include a camera that captures a video or an image generation unit that generates an image by calculation, and the wide-angle video may be a video captured by the camera or an image calculated by the image generation unit.

[0081] According to this, it is possible to display an appropriate video for the wide-angle video composed of the video captured by the camera or the image calculated by the image generation unit.

[0082] Further, for example, the presentation unit generates and outputs graphics indicating at least one of the calculated relative position and relative direction and information based on cue information, and superimposes the output graphics on some images to cause the video generation unit to present at least one of the relative position and relative direction.

[0083] According to this, at least one of the relative position and the relative direction can be presented to the user by the graphics.

[0084] Further, for example, the data reception unit receives an input of data regarding the direction of the target of attention, the display state estimation unit estimates the direction within the wide-angle video of the display device, and the graphics may display an arrow indicating the relative direction on the display video.

[0085] According to this, at least one of the relative position and the relative direction can be presented to the user by graphics that display an arrow indicating the relative movement direction on the display video.

[0086] Also, for example, the data reception unit may receive input of data regarding the direction of the object of fixation, the display state estimation unit may estimate the direction within the wide viewing angle video of the display device, and the graphics may display a mask that is an image for covering at least a part other than the relative direction side on the display video.

[0087] According to this, the relative direction can be presented to the user by graphics that display a mask that is an image for covering at least a part other than the relative direction side on the display video.

[0088] Also, for example, the data reception unit may receive input of data regarding the position of the object of fixation, the display state estimation unit may estimate the position within the wide viewing angle video of the display device, and the graphics may display a map indicating the relative position on the display video.

[0089] According to this, the relative movement direction can be presented to the user by graphics that display a map indicating the relative position on the display video.

[0090] Also, for example, further, an input interface for use in inputting data may be provided, and the data acquisition unit may acquire the data input via the input interface.

[0091] According to this, metadata can be configured from the data input via the input interface.

[0092] Also, for example, further, an input interface for designating at least one of the start and end timings of the movement within the wide viewing angle video of the user may be provided, and the data acquisition unit may acquire at least one of the start and end timings of the movement input via the input interface.

[0093] According to this, at least one of the start and end timings of the movement input via the input interface can be acquired.

[0094] Also, for example, the image constituting the wide-angle view image is an image output by an imaging unit that images the real space, and the input interface is an instruction marker held by an operator of the input interface in the real space, and the movement of the instruction marker indicates at least one of the position and direction of the fixation target. It may have an image analysis unit that receives at least one of the position and direction of the fixation target indicated by the instruction marker by analyzing an image including the instruction marker output by the imaging unit.

[0095] According to this, by indicating at least one of the position and direction of the fixation target by the movement of the instruction marker, at least one of the relative position and the relative direction can be calculated.

[0096] Also, for example, it may include at least some of the functions provided by the observation device and the VR device, be connected to the observation device and the VR device via a network, and include an information processing device that undertakes part of the processing of the observation device or the VR device.

[0097] According to this, a video display system can be realized by the observation device, the VR device, and the information processing device.

[0098] Also, for example, the recording information processing device includes a receiving unit that receives wide-angle view video, data, and queue information from the observation device as metadata, at least one of the position and direction of the fixation target on the metadata, and information for presenting information according to the queue information to the user of the display device. A presentation unit that generates information, and from the received wide-angle view video, adds the information generated by the presentation unit to a partial image corresponding to the visual field portion corresponding to at least one of the position and direction in the wide-angle view video of the display device estimated by the display state estimation unit to generate a display video. It may include a video generation unit, and a transmission unit that transmits the wide-angle view video, a partial image corresponding to the visual field portion corresponding to at least one of the relative position and the relative direction, and the metadata.

[0099] According to this, an image display system can be realized by the observation device, VR device, and information processing device configured as described above.

[0100] Also, for example, the information processing device may include: a receiving unit that receives a wide-angle image, data, and cue information from the observation device as metadata; a presentation unit that generates information for presenting at least one of the position and direction of a gaze target on the metadata and information according to the cue information to the user of the display device; a metadata configuration unit that generates metadata from the information generated by the presentation unit; and a transmission unit that transmits the metadata generated by the metadata configuration unit, the wide-angle image received by the receiving unit, and other information to the VR device.

[0101] According to this, an image display system can be realized by the observation device, VR device, and information processing device configured as described above.

[0102] Also, for example, the information processing device may include: a receiving unit that receives a wide-angle image, data, and cue information from the observation device as metadata and receives data regarding the orientation of the display device from the display device; a difference calculation unit that calculates a relative movement direction, which is the movement direction of the imaging unit relative to the orientation of the display device, based on the difference between the orientation of the display device and movement information regarding the movement of the imaging unit and the cue information; a presentation unit that generates and outputs graphics indicating the calculated relative movement direction, and superimposes the graphics on a part of the image corresponding to the visual field portion corresponding to the estimated orientation of the display device among the wide-angle images, thereby presenting information according to the relative movement direction and the cue information to the user of the display device; an image generation unit that corrects the graphics based on data regarding the orientation of the display device and generates a display image by superimposing the corrected graphics on the wide-angle image; and a transmission unit that transmits the display image and other information.

[0103] According to this, an image display system can be realized by the observation device, VR device, and information processing device configured as described above.

[0104] Further, for example, the information processing apparatus may be provided on a cloud connected to a wide area network and connected to the observation apparatus and the VR apparatus via the wide area network.

[0105] According to this, a video display system can be realized by the observation apparatus, the VR apparatus, and the information processing apparatus that is connected to the observation apparatus and the VR apparatus via a wide area network and provided on a cloud.

[0106] Further, for example, the queue information may be information indicating that at least one of the moving direction of the observation apparatus or the position and direction of the fixation target for causing the user of the display apparatus to fixate changes.

[0107] According to this, information indicating that at least one of the moving direction of the observation apparatus or the position and direction of the fixation target for causing the user of the display apparatus to fixate changes can be used as the queue information.

[0108] Further, the information processing apparatus according to one aspect of the present disclosure is an information processing apparatus used in a video display system for causing a display apparatus to display at least a part of the display video within a wide viewing angle video, and is data regarding at least one of the position and direction of a fixation target for causing the user of the display apparatus to fixate in the wide viewing angle video, and includes a receiving unit that receives metadata based on data obtained by receiving an input, and a difference calculation unit that calculates and outputs at least one of a relative position that is a relative position of the fixation target with respect to at least one of the position and direction within the wide viewing angle video of the display apparatus, and a relative direction that is a relative direction of the fixation target, based on a difference from at least one of the position and direction of the fixation target on the metadata.

[0109] Further, for example, graphics indicating at least one of the calculated relative position and relative direction, which are superimposed on a part of the images constituting the wide-angle video corresponding to the visual field portion according to at least one of the estimated position and direction within the wide-angle video of the display device, may further include a presentation unit that generates and outputs graphics for presenting at least one of the relative position and relative direction to the user of the display device.

[0110] According to these, by using them in the video display system described above, the same effects as those of the video display system described above can be achieved.

[0111] Also, an information processing method according to an aspect of the present disclosure is an information processing method for causing a display device to display at least a part of the display video within a wide-angle video, and is data regarding at least one of the position and direction of a fixation target that causes a user of the display device to fixate on the wide-angle video, and receives metadata based on the data obtained by receiving an input, and based on the difference between at least one of the estimated position and direction within the wide-angle video of the display device and at least one of the position and direction of the fixation target on the metadata, calculates and outputs at least one of a relative position that is the relative position of the fixation target with respect to the orientation of the display device and a relative direction that is the relative direction of the fixation target.

[0112] Such an information processing method can achieve the same effects as those of the video display system described above.

[0113] Also, a program according to an aspect of the present disclosure is a program for causing a computer to execute the information processing method described above.

[0114] Such a program can achieve the same effects as those of the video display system described above by using a computer.

[0115] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0116] Note that all the embodiments described below show comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement positions of the components, connection forms, steps, order of steps, etc. shown in the following embodiments are merely examples and are not intended to limit the scope of the claims. Among the components in the following embodiments, the components not described in the independent claims are described as optional components.

[0117] Note that each figure is not necessarily drawn precisely. In each figure, substantially the same configurations are denoted by the same reference numerals, and overlapping descriptions are omitted or simplified.

[0118] Also, in this specification, terms indicating the relationship between elements such as parallel, terms indicating the shape of elements such as rectangular, as well as numerical values and numerical ranges are not expressions representing only strict meanings, but are expressions meaning that they also include substantially equivalent ranges, for example, differences such as an error of about several percent.

[0119] (Embodiment) [Configuration] First, the outline of the video display system in the embodiment will be described with reference to FIGS. 14 and 15. FIG. 14 is a diagram showing the schematic configuration of the video display system according to the embodiment. FIG. 15 is a diagram showing an example of the video displayed in the video display system according to the embodiment.

[0120] As shown in FIG. 14, the video display system 500 of the present embodiment is realized by an observation device 300, a server device 200 connected via a network 150, and a display device 100 connected via the network 150.

[0121] The observation device 300 is a device housed inside the image holding device. More specifically, the observation device 300 is configured to hold the image obtained by shooting as information of a wide-angle video and supply it to the display device 100 so that a part of the video can be visually recognized on the display device 100. The observation device 300 is a so-called omnidirectional camera that can shoot videos of 360 degrees around. The observation device 300 may be, for example, a shooting device 300a that is held by hand for shooting, or an observation device 300b that is fixed with a tripod or the like. In the case of the shooting device 300a held by hand, it is easy to shoot while moving around. Hereinafter, these types will be collectively referred to as the observation device 300 without particularly distinguishing them. The observation device 300 has optical elements such as a fish-eye lens and can shoot a wide-angle region, for example, 180 degrees, with one sensor array. Then, using a plurality of combinations of optical elements and sensor arrays arranged to complement different wide-angle regions, a 360-degree wide-angle video can be shot. Note that, for the images shot by each of the plurality of sensor arrays, a process (stitching) of identifying and overlapping the corresponding elements with each other is performed. As a result, for example, one image that can mutually convert a plane such as an orthographic cylindrical view and a spherical surface is generated. By continuously generating such images in the time domain, a video (moving image) that changes in the time domain is generated. Note that the inside of the spherical video is also referred to as a 3D video space or a three-dimensional image space.

[0122] Also, in the present embodiment, two three-dimensional image spaces with a shift corresponding to the human parallax are generated. Such two three-dimensional image spaces may be generated from one three-dimensional image space by simulation or the like, or may be generated by two cameras with a parallax shift. In the present embodiment, a VR video in which the user can view an arbitrary direction of the three-dimensional image space can be displayed from inside the three-dimensional image space.

[0123] The network 150 is a communication network for connecting the observation device 300, the server device 200, and the display device 100 to be communicable with each other. Here, a communication network such as the Internet is used as the network 150, but it is not limited to this. Also, the connection between the observation device 300 and the network 150, the connection between the server device 200 and the network 150, and the connection between the display device 100 and the network 150 may be performed by wireless communication or by wired communication, respectively.

[0124] The server device 200 is a device for performing information processing and the like, and is realized by using, for example, a processor and a memory. The server device 200 may be realized by an edge computer or may be realized by a cloud computer. Also, one server device 200 may be provided for one video display system 500, or one server device 200 may be provided for a plurality of video display systems 500. That is, the server device 200 may perform various processes in a plurality of video display systems 500 in parallel. Note that the server device 200 is not an essential component in the video display system 500.

[0125] For example, by distributing and arranging each functional unit of the server device 200, which will be described later, to each of the observation device 300 and the display device 100, it is also possible to realize a video display system including only the observation device 300 and the display device 100. In particular, if the display device 100 is realized by an information processing terminal such as a smartphone having a display panel, the functional unit of the server device 200 can be easily realized by using a processor or the like of the information processing terminal. Alternatively, by giving the functions of the observation device 300 and the display device 100 to the server device 200, a part of the functions of the observation device 300 or the display device 100 can be reduced, and an existing observation device or display device can be diverted. That is, by aggregating various functions in the server device 200, it becomes possible to easily realize a video display system. Each functional unit of the server device 200 will be described later with reference to FIG. 16 and the like.

[0126] The display device 100 is a glass-type HMD that supports two separated lens barrels by locking the temple parts extending from each of the left and right sides to the auricle, thereby holding the two lens barrels at positions corresponding to the user's right and left eyes respectively. A display panel is built into each lens barrel of the display device 100. For example, as shown in FIG. 15, an image with parallax is projected toward each of the user's left and right eyes. In FIG. 15, (L) shows an image of one frame in the left-eye video, and (R) shows an image of the same one frame in the right-eye video. Note that the display device 100 does not have to be a terminal dedicated to such video display. It is also possible to implement the display device of the present disclosure using a display panel provided in a smartphone, a tablet terminal, a PC, or the like.

[0127] Hereinafter, with reference to FIGS. 16 to 18, the detailed configuration of the video display system 500 of the present embodiment will be described. FIG. 16 is a block diagram showing the functional configuration of the video display system according to the embodiment. As shown in FIG. 16 and as described in FIG. 14, the video display system 500 includes a display device 100, a server device 200, and an observation device 300.

[0128] The display device 100 includes a display unit 101 and a display state estimation unit 102. The display unit 101 is a functional unit that outputs an optical signal according to image information using a backlight, a liquid crystal panel, an organic EL, a micro LED, or the like. The display unit 101 controls the output optical signal so that an image is formed on the retina of the user's eye via optical elements such as a lens and an optical panel. As a result, the user can visually recognize the image formed on the retina. The display unit 101 can cause the user to visually recognize continuous images, that is, video, by continuously outputting the above-described images in the time domain. In this way, the display unit 101 displays video for the user of the display device 100.

[0129] The display state estimation unit 102 is a functional unit for estimating at which position and in which direction within the three-dimensional image space the user is viewing an image using the display device 100. It can also be said that the display state estimation unit 102 estimates at least one of the position and direction of the display device 100 within the three-dimensional image space. The display state estimation unit 102 is realized by various sensors such as an acceleration sensor and a gyro sensor built in an appropriate position of the display device 100. The display state estimation unit 102 estimates the position of the display device 100 within the three-dimensional image space by estimating how much and in which direction the position has changed with respect to a reference position preset in the display device 100. Also, the display state estimation unit 102 estimates the direction of the display device 100 within the three-dimensional image space by estimating how much and in which direction the posture has changed with respect to a reference direction preset in the display device 100. As described above, since the display device 100 is supported by the user's head (pinna and root of the nose), it moves together with the user's head.

[0130] By estimating the position and direction of the display device 100, a visual field portion corresponding to the position and direction can be cut out from the wide-angle image and displayed. That is, based on the position and direction of the display device 100 estimated by the display state estimation unit 102, the visual field area in the direction the user's head is facing within the three-dimensional image space can be displayed as the visual field area to be viewed. Note that the direction of the display device 100 estimated here is the direction along the normal direction of the display panel of the display device 100. Since the display panel is arranged to face the user's eyes, the user's eyes are usually located in the normal direction of the display panel. For this reason, the direction of the display device 100 coincides with the direction connecting the user's eyes and the display panel.

[0131] However, there may be a case where the direction of the display device 100 and the user's line of sight direction deviate due to the user's eye movement. In this case, if the display device 100 is equipped with a sensor (eye tracker) for detecting the user's line of sight, the detected user's line of sight may be used as the direction of the display device 100. That is, the eye tracker is another example of the display state estimation unit.

[0132] In addition to the above, the display device 100 is equipped with a power supply, various input switches, a circuit for driving the display panel, wired and wireless communication modules for input and output, an audio signal processing circuit such as a signal converter and an amplifier, and a microphone and a speaker for audio input and output. Details of these configurations will be described later.

[0133] The server device 200 includes a receiving unit 201, a difference calculation unit 202, a presentation unit 203, and a video generation unit 204. The receiving unit 201 is a processing unit that receives (acquires) various signals from an observation device 300 described later. The receiving unit 201 receives a wide-angle video captured by the observation device 300. In addition, the receiving unit 201 receives metadata acquired by the observation device 300. Further, the receiving unit 201 receives information regarding the position and orientation of the display device 100 estimated in the display device 100.

[0134] The difference calculation unit 202 is a processing unit that calculates a relative position, which is the position of the attention target 301 relative to the position of the display device 100, and a relative direction, which is the direction of the attention target 301 relative to the direction of the display device 100, based on the differences between the position and orientation of the display device 100 and the position and orientation of the attention target 301 included in the metadata. Details of the operation of the difference calculation unit 202 will be described later.

[0135] The prompting unit 203 is a processing unit that prompts the user of the display device 100 with the relative position and relative direction calculated by the difference calculation unit 202. Here, an example will be described in which the prompting unit 203 causes the video generation unit 204 to perform the above prompting by including content indicating the relative movement direction in the display video generated by the video generation unit 204. However, the prompting of the relative movement direction is not limited to the example of including it in the above display video. For example, it may be prompted as audio from a predetermined arrival direction corresponding to at least one of the relative position and relative direction within the three-dimensional sound field, or it may be prompted by vibrating a device such as a vibration device held by the user's both hands on the side corresponding to at least one of the relative position and relative direction. The detailed operation of the prompting unit 203 will be described later together with the detailed operation of the difference calculation unit 202.

[0136] The video generation unit 204 cuts out a part of the video corresponding to the visual field portion according to the position and direction of the display device 100 estimated by the display state estimation unit 102 from the received wide-angle video, and further generates a display video including content indicating at least one of the calculated relative position and relative direction if necessary. The detailed operation of the video generation unit 204 will be described later together with the detailed operations of the difference calculation unit 202 and the prompting unit 203. In addition, the server device 200 has a communication module for transmitting the generated display video to the display device 100.

[0137] The observation device 300 includes a storage unit 301, an input interface 302, a position input unit 303, a data acquisition unit 304, a metadata acquisition unit 305, and a transmission unit 306. The observation device 300 also has a photographing unit (not shown) which is a functional part related to image photographing and is integrally configured with other functional components of the observation device 300. Note that the photographing unit may be separated from other functional components of the photographing device 300 by wired or wireless communication. The photographing unit includes an optical element, a sensor array, an image processing circuit, etc. The photographing unit outputs, for example, the luminance value of the light of each pixel received on the sensor array via the optical element as 2D luminance value data. The image processing circuit performs post-processing such as noise removal of the luminance value data, and also performs processing for generating a three-dimensional image space from 2D image data such as stitching. In the present embodiment, an example of displaying a video in the three-dimensional image space formed from an actual image photographed by the photographing unit using a display device will be described. However, the three-dimensional image space may be a virtual image formed by a technology such as computer graphics. Therefore, the photographing unit is not an essential configuration.

[0138] The storage unit 301 is a storage device that stores image information (images constituting the three-dimensional image space) of the three-dimensional image space generated by the photographing unit. The storage unit 301 is realized using a semiconductor memory or the like.

[0139] The input interface 302 is a functional unit used when an input is made by a guide who guides VR sightseeing, etc. within a three-dimensional image space. For example, the input interface 302 includes a stick that can be tilted in each of 360-degree directions corresponding to the moving direction of the photographing unit 301, and a physical sensor that detects the tilting direction. When inputting the direction of the object of interest, the guide can input the direction of the object of interest to the system by tilting the stick in that direction. As another example of the input interface, there is an indicating marker with a fluorescent marker or the like attached to the tip of a pointer or the like held by the guide, and an image analysis unit that receives the direction of the object of interest indicated by the indicating marker by analyzing an image including the indicating marker output by the photographing unit. An input interface including the above may be used. Note that the input interface 302 is not an essential configuration, and the present embodiment can be realized if either one of the position input unit 303 described later is provided.

[0140] The position input unit 303 is a functional unit for inputting the position of the object of interest. The position input unit 303 is realized, for example, by executing a dedicated application on an information terminal such as a smartphone possessed by the guide. Map information of the entire area corresponding to the three-dimensional image space is displayed on the screen of the information terminal, and by selecting a predetermined position on this map information, the selected position is input to the system as the position of the object of interest. In this way, the position input unit 303 is an example of an input interface for inputting the position of the object of interest.

[0141] The data acquisition unit 304 is a functional unit that acquires data related to the position and direction of the object of fixation from the input interface 302, the position detection unit 303, and the like. The data acquisition unit 304 is connected to at least one of the input interface 302 and the position detection unit 303, and acquires a physical quantity corresponding to the data related to the position and direction of the object of fixation from these functional units. In this way, the data acquisition unit 304 is an example of a data reception unit that receives the input of data related to the direction of the object of fixation that the user of the display device 100 in the present embodiment focuses on.

[0142] The metadata acquisition unit 305 is a functional unit that acquires the metadata by converting the data related to the position and direction of the object of fixation acquired by the data acquisition unit 304 into metadata for adding to the captured video data. The acquired metadata may include various data used within the video display system 500 in addition to the data related to the position and direction of the object of fixation. That is, the metadata acquisition unit 305 is an example of a metadata configuration unit that constitutes metadata capable of reading a plurality of data from one piece of information by collecting a plurality of data into one.

[0143] The transmission unit 306 is a communication module that transmits the captured video (wide-angle video) stored in the storage unit 301 and the acquired metadata. The transmission unit 306 communicates with the reception unit 201 of the server device 200 to transmit the stored video and the acquired metadata and cause them to be received by the reception unit.

[0144] FIG. 17 is a more detailed block diagram showing the functional configuration of the observation device according to the embodiment. Further, FIG. 18 is a more detailed block diagram showing the functional configuration of the display device according to the embodiment. FIGS. 17 and 18 show the functional configuration around the observation device 300 and the display device 100 in more detail. Some of the functions shown in these figures may be realized by the configuration of the server device 200.

[0145] The queue information input means 51 corresponds to the input interface 302 and the position input unit 303, and inputs the position and direction of the object of interest by means of switches, tablets, smartphones, etc. physically operated by the operator or guide of the observation device 300. Also, the queue information input means 51 designates a target moving from a plurality of targets by queue data.

[0146] The queue information input means 51 may also obtain queue information from the video obtained from the VR video processing means 67 or the audio information obtained from the audio input means 71. The audio input means 71 is another example of the input interface. Note that the VR video processing means 67 is connected to the VR shooting means 69 corresponding to the shooting unit.

[0147] The queue information obtained from the queue information input means 51 is sent to the position / azimuth detection / memory means 53, processed together with the position and azimuth of the observation device 300, and in some cases, its state is stored and shaped into appropriate data, and sent as metadata to the multiplexing means 61. After multiplexing with video, audio, and graphics, it is transmitted by the communication means 55 to the display device 100 via the server device 200. In addition to the above, the observation device 300 includes a separation means 57, a VR video compression means 59, an audio compression means 63, an audio decoding means 65, and an audio output means 73.

[0148] In the display device 100, the communication means 39 receives the communication information from the observation device 300, separates the metadata by the separation means 37, and sends it to the position / azimuth / queue determination means 31. In the position / azimuth / queue determination means 31, the queue data is taken out from the metadata, performs predetermined processing, and sends it to the graphics generation means 33 for displaying the queue information as a graphic, superimposes and displays it on the VR video by the VR display means 15, sends it to the VR control means 21, appropriately processes the VR video by the VR display control means 23 together with the position / azimuth state of the display device 100, and displays it by the VR display means 15, or generates a guiding voice for guidance or appropriately processes the reproduced audio by the audio reproduction control means 25, etc. are performed.

[0149] As a specific example, when data such as "move to target A" is shown in the queue data, depending on the position of the display device 100, the relative position to target A is different. When the queue information is shown graphically, an arrow is displayed in an appropriate direction. More specifically, when target A is on the left side, a leftward arrow is displayed. When controlling the video, for example, control such as only the left side being cleared (the image quality on the right side deteriorates) is performed. When controlling with sound, an announcement such as "Please face the left side" is played. In this way, appropriate processing is performed by comparing the content of the queue data with the position and orientation of the display device 100. The display device 100 further includes a position detection means 11, a rotation detection means 13, an audio playback means 17, an audio input means 19, a VR control means 21, an audio compression means 27, an audio decoding means 35, and a multiplexing means 41. By each including the components shown in FIGS. 17 and 18 above in one or more combinations, the components shown in FIG. 3 are realized.

[0150] [Operation] Next, the operation of the video display system 500 configured as described above will be described with reference to FIGS. 19 to 32. FIG. 19 is a flowchart showing the operation of the video display system according to the embodiment.

[0151] When the operation of the video display system 500 is started, video is captured by the imaging unit, and the image is stored in the storage unit 301. At the same time, by the operation of the input interface 302, the position input unit 303, the data acquisition unit 304, and the metadata acquisition unit 305, metadata including data on the position and direction of the fixation target is acquired. The metadata is received by the server device 200 together with the video captured and stored via the transmission unit 306 and the reception unit 201 (S101).

[0152] In addition, the display state estimation unit 102 of the display device 100 continuously estimates the position and orientation of the display device 100. The display device 100 transmits the orientation of the display device 100 estimated by the display state estimation unit 102 to the server device 200. As a result, the server device 200 receives the estimated position and orientation of the display device 100 (S102). Note that the order of step S101 and step S102 may be swapped. The server device 200 determines whether there has been an input for instructing the position and orientation of the fixation target based on whether the data regarding the position and orientation of the fixation target is included (S103). If it is determined that there has been an input for instructing the position and orientation of the fixation target (Yes in S103), the server device 200 proceeds to an operation for presenting the relative position and relative orientation to the user of the display device 100. Specifically, the difference calculation unit 202 calculates the relative position of the fixation target relative to the position where the user is looking (i.e., corresponding to the position of the display device 100) as the relative position based on the position and orientation of the display device 100 and the data regarding the position and orientation of the fixation target on the metadata. In addition, the difference calculation unit 202 calculates the relative orientation of the fixation target relative to the direction where the user is looking (i.e., corresponding to the orientation of the display device 100) as the relative orientation based on the position and orientation of the display device 100 and the data regarding the position and orientation of the fixation target on the metadata (S104).

[0153] After calculating the relative position and relative direction, the presentation unit 203 generates graphics corresponding to this relative position and relative direction (S105). Then, the video generation unit 204 cuts out at least a part of the image corresponding to the visual field portion according to the orientation of the display device 100 from the three-dimensional image space (S106), and superimposes the graphics generated by the presentation unit 203 on the cut-out part of the video to generate a display video (S107). FIG. 20 is a conceptual diagram for explaining the generation of the display video according to the embodiment. In FIG. 20, (a) shows a part of the image cut out from the three-dimensional image space, (b) shows the graphics 99 generated by the presentation unit 203, and (c) shows the display video generated by superimposition. By superimposing the graphics 99 on a part of the image, an arrow 99a indicating at least one of the relative position and relative direction is displayed in the display video.

[0154] Also, FIG. 21 is a conceptual diagram for explaining the generation of the display video according to the embodiment. In FIG. 21, (a) shows the image captured by the imaging unit, and (b) shows the display video that the user is viewing. In FIG. 21, a guide is shown in (a), but this guide does not appear in the video. Therefore, in (b), a video without the guide is shown. Here, as shown in (a) of FIG. 21, the guide instructs the user facing the guide to gaze at the object of attention in the front direction as stated "What you can see straight ahead is". In (b) of FIG. 21, a schematic diagram representing the direction in which the user is looking is shown below the figure of the video that the user is viewing.

[0155] Here, the upper side of the paper surface is the front direction in the three-dimensional image space and is recognized as the direction in which the user is looking as guided by the guide. That is, the user shown in FIG. 21 is in a state of looking straight ahead. Here, as an input interface, the keyword "front" uttered by the guide is acquired by the voice input means 71 or the like, and this is acquired as data regarding the direction of the object of fixation. Then, as shown in FIG. 21(b), the display video is generated with the arrow 99a superimposed. Since the direction of this arrow coincides with the "front" as the direction of the object of fixation and the "front" as the direction of the display device 100, it is simply the arrow 99a indicating the "front" direction. Also, as shown in FIG. 21(b), the voice is reproduced. Here, since the "front" as the direction of the object of fixation and the "front" as the direction of the display device 100 coincide, the voice "What you can see in front" is reproduced together with the display video as it is.

[0156] FIG. 22 is a conceptual diagram for explaining the generation of the display video according to the embodiment. In FIG. 22, (a) shows the image captured by the imaging unit, and (b) shows the display video that the user is looking at. In FIG. 22, the guide is shown in (a), but this guide does not appear in the video. Therefore, in (b), the video without the guide is shown. Here, as shown in FIG. 22(a), the guide is instructing the user facing the guide to fixate on the object of fixation in the right hand direction as in "What you can see in the right hand". In FIG. 22(b), a schematic diagram representing the direction in which the user is looking is shown below the figure of the video that the user is visually recognizing.

[0157] Here, the upper side of the paper surface is the front direction in the three-dimensional image space and is recognized as the direction in which the user is looking as guided by the guide. That is, the user shown in FIG. 22 is in a state of looking straight ahead. Here, as an input interface, a keyword "right hand" uttered by the guide is acquired by voice input means 71 or the like and is acquired as data regarding the direction of the object of fixation. Then, as shown in FIG. 22(b), an arrow 99a is superimposed and a display video is generated. Since the direction of this arrow does not match between "right hand" as the direction of the object of fixation and "front" as the direction of the display device 100, the difference between "right hand" and "front" is calculated and the arrow 99a indicating the "right hand" direction is obtained. Also, as shown in FIG. 22(b), voice is reproduced. Here, since "right hand" as the direction of the object of fixation and "front" as the direction of the display device 100 do not match, the difference between "front" and "right hand" is calculated and the voice "What appears on the right" is reproduced together with the display video.

[0158] FIG. 23 is a conceptual diagram for explaining the generation of a display video according to the embodiment. In FIG. 23, (a) shows an image captured by the imaging unit, and (b) shows the display video that the user is looking at. In FIG. 23, the guide is shown in (a), and this guide may or may not appear on the video. However, in (b), the video of the visual field part without the guide is shown. Here, as shown in FIG. 23(a), the guide instructs the user facing the guide to fixate on the object of fixation in the right hand direction as in "What appears on the right". In FIG. 23(b), a schematic diagram representing the direction in which the user is looking is shown below the figure of the video that the user is viewing.

[0159] Here, the upper side of the paper surface is the front direction in the three-dimensional image space and is recognized as the direction in which the user is looking as guided by the guide. That is, the user shown in FIG. 23 is in a state of looking in the right hand direction. Here, as an input interface, the keyword "right hand" uttered by the guide is acquired by the voice input means 71 or the like and is acquired as data regarding the direction of the object of fixation. However, "right hand" here is a direction closer to the front side than the right hand direction in which the user is looking. Such a detailed direction of the object of fixation is input using the input interface 302 or the like. Then, as shown in FIG. 23(b), the display video is generated with the arrow 99a superimposed. Since the direction of this arrow does not match between "right hand closer to the front" as the direction of the object of fixation and "right hand" as the direction of the display device 100, the difference between "right hand closer to the front" and "right hand" is calculated, and the arrow 99a indicating the "left hand" direction is obtained.

[0160] FIG. 24 is a conceptual diagram for explaining the generation of the display video according to the embodiment. In FIG. 24, (a) shows an image captured by the imaging unit, and (b) shows the display video that the user is looking at. In FIG. 24, the guide is shown in (a), but this guide does not appear on the video. Therefore, in (b), a video without the guide is shown. Here, as shown in FIG. 24(a), the guide instructs the user facing the guide to fixate on the object of fixation located in the right hand direction as in "What you can see on the right hand". In FIG. 24(b), a schematic diagram representing the direction in which the user is looking is shown below the figure of the video that the user is visually recognizing.

[0161] Here, the upper side of the paper surface is the front direction in the three-dimensional image space and is recognized as the direction in which the user is looking as guided by the guide. That is, the user shown in FIG. 24 is in a state of looking in the front direction. Here, as an input interface, the position of the object of fixation input by the guide is acquired by the position input unit 303 or the like and is acquired as data regarding the position of the object of fixation. Then, as shown in FIG. 24(b), the map 99b is superimposed and the display video is generated. In this map, an arrow indicating the position corresponding to the position of the object of fixation in the map is attached. Since the position of the object of fixation does not coincide with the position of the display device 100, the map 99b with an arrow indicating the vicinity of the 4 o'clock direction is obtained. In the map 99b, the position of the user (that is, the position of the display device 100) is the central part.

[0162] FIG. 25 is a conceptual diagram for explaining the generation of the display video according to the embodiment. In FIG. 25, (a) shows the image captured by the imaging unit, and (b) shows the display video that the user is looking at. In FIG. 25, the guide is shown in (a), but this guide may or may not appear on the video. However, in (b), the video of the visual field part without the guide is shown. Here, as shown in FIG. 25(a), the guide instructs the user facing the guide to fixate on the object of fixation at the position in the right hand direction as in "What can be seen on the right hand". In FIG. 25(b), a schematic diagram showing the direction in which the user is looking is shown below the figure of the video that the user is visually recognizing.

[0163] Here, the upper side of the paper surface is the front direction in the three-dimensional image space and is recognized as the direction in which the user is looking as guided by the guide. That is, the user shown in FIG. 25 is in a state of looking in the right hand direction. Here, as an input interface, the position of the object of fixation input by the guide is acquired by the position input unit 303 or the like, and this is acquired as data regarding the position of the object of fixation. Then, as shown in FIG. 25(b), a map 99b is superimposed and a display video is generated. In this map, an arrow indicating the position corresponding to the position of the object of fixation in the map is attached. Since the position of the object of fixation does not match the position of the display device 100, the map 99b with an arrow indicating around the 2 o'clock direction is obtained.

[0164] FIG. 26 is a conceptual diagram for explaining the generation of the display video according to the embodiment. Since FIG. 26 shows another example of the operation in the same situation as FIG. 21, the description regarding the situation is omitted. As shown in FIG. 26(b), a display video in which the arrow 99a is not superimposed is generated. Since the "front" as the direction of the object of fixation and the "front" as the direction of the display device 100 match, an arrow 99a simply indicating the "front" direction would be displayed. However, since it is redundant to present an arrow indicating the front direction to a user looking at the front, the arrow 99a is not displayed here. On the other hand, as shown in FIG. 26(b), voice is reproduced. Here, since the "front" as the direction of the object of fixation and the "front" as the direction of the display device 100 match, the voice "What appears in front" is reproduced together with the display video as it is.

[0165] FIG. 27 is a conceptual diagram for explaining the generation of a display video according to the embodiment. Since FIG. 27 shows another example of the operation in the same situation as FIG. 22, the description of the situation is omitted. As shown in FIG. 27(b), instead of the arrow 99a, a mask 99c is superimposed to generate a display video. Since the "right hand" as the direction of the object of fixation does not match the "front" as the direction of the display device 100, the difference between the "right hand" and the "front" is calculated, and the mask 99c is used to guide the user to visually recognize the "right hand" direction. In this way, by generating and superimposing the mask 99c that covers a part of the side opposite to the relative direction side as the graphics 99, it is configured to give a large change on the display video. The arrow 99a is easy to visually understand its direction, but the change on the video may be difficult to understand. In this example, the above-mentioned drawbacks can be compensated for. That is, the user makes a line-of-sight movement to visually recognize the remaining video from the area covered by the mask 99c, and since the direction of the line-of-sight movement corresponds to the relative direction, there is an advantage that the change on the video is easy to understand and the relative direction is easy to naturally recognize. Here, the "covering" also includes being covered by a semi-transparent image in which the area to be covered can be seen partially through.

[0166] Also, as shown in FIG. 22(b), audio is reproduced. Here, since the "right hand" as the direction of the object of fixation does not match the "front" as the direction of the display device 100, the difference between the "front" and the "right hand" is calculated, and the audio "What you can see on the right" is reproduced together with the display video.

[0167] FIG. 28 is a conceptual diagram for explaining the generation of a display video according to the embodiment. Since FIG. 28 shows another example of the operation in the same situation as FIGS. 22 and 27, the description of the situation is omitted. As shown in FIG. 28(b), instead of the arrow 99a and the mask 99c, a downsampling filter 99d is superimposed to generate a display video. The downsampling filter 99d roughens the image of the portion where the filter is superimposed, like so-called mosaic processing. Then, similar to the mask 99c, from the region where visual recognition becomes difficult due to downsampling, the user makes a gaze movement to visually recognize a clearer video portion, and since the direction of the gaze movement corresponds to the relative direction, there is an advantage that the relative direction can be naturally and easily recognized.

[0168] FIG. 28(c) shows a situation where the user is visually recognizing the right hand direction. Since this right hand direction corresponds to the direction of the object of fixation, no downsampling filter 99d is superimposed on this region and the image remains clear. Further, FIG. 28(c) shows a situation where the user is visually recognizing the right hand direction (that is, the back direction with respect to the original front). Since a similar downsampling filter 99d is superimposed on the back direction to generate a display video, there is an advantage that the right hand direction to be visually recognized is easier for the user to understand. The explanations for FIGS. 28(c) and 28(d) are also valid for the example using the mask 99c (the example of FIG. 27).

[0169] FIG. 29 is a conceptual diagram for explaining the generation of a display video according to the embodiment. Since FIG. 29 shows another example of the operation in the same situation as FIGS. 21 and 29, the description of the situation is omitted. As shown in FIG. 29(b), a display video without the arrow 99a superimposed is generated. Since the "front" as the direction of the object of fixation and the "front" as the direction of the display device 100 coincide, an arrow 99a simply indicating the "front" direction would be displayed. However, since it is redundant to present an arrow indicating the front direction to a user looking at the front, the arrow 99a is not displayed here. On the other hand, as shown in FIG. 29(b), audio is reproduced. Here, since the "front" as the direction of the object of fixation and the "front" as the direction of the display device 100 coincide, the audio "What appears in front" is reproduced together with the display video as it is. However, the audio reproduced here is stereophonic audio that is perceived by the user as sound coming from the user's "front".

[0170] FIG. 30 is a conceptual diagram for explaining the generation of a display video according to the embodiment. Since FIG. 30 shows another example of the operation in the same situation as FIGS. 22, 27, and 28, the description of the situation is omitted. As shown in FIG. 30(b), a display video without the arrow 99a superimposed is generated. On the other hand, as shown in FIG. 30(b), audio is reproduced. Here, since the "right hand" as the direction of the object of fixation and the "front" as the direction of the display device 100 do not coincide, the difference between "front" and "right hand" is calculated, and the audio "What appears on the right hand" is reproduced together with the display video. And the audio reproduced here is stereophonic audio that is perceived by the user as sound coming from the user's "right hand", similar to the example of FIG. 29.

[0171] FIG. 31 is a conceptual diagram for explaining the generation of a display video according to the embodiment. In FIG. 31, (a) shows an image captured by a photographing unit, and (b) shows a display video that the user is viewing. In FIG. 31, a guide is shown in (a), but this guide may or may not appear in the video. However, in (b), a video of a visual field portion without the guide is shown. Here, as shown in (a) of FIG. 31, the guide is instructing the user facing the guide to gaze at a gazing target in the right hand direction as in "What can be seen in the right hand". In (b) of FIG. 31, a schematic diagram representing the direction in which the user is looking is shown below the figure of the video that the user is viewing.

[0172] Here, the upper side of the paper surface is the front direction in the three-dimensional image space and is the direction recognized as the direction in which the user is looking by the guide. That is, the user shown in FIG. 31 is in a state of looking in the back direction. Here, as an input interface, a keyword "right hand" uttered by the guide is acquired by voice input means 71 or the like, and this is acquired as data regarding the direction of the gazing target. Then, as shown in (b) of FIG. 31, the voice is reproduced. Here, since "right hand" as the direction of the gazing target does not match "back" as the direction of the display device 100, the difference between "right hand" and "back" is calculated, and a voice "What can be seen in the left hand" is reproduced together with the display video.

[0173] FIG. 32 is a conceptual diagram for explaining the generation of a display video according to the embodiment. In FIG. 32, (a) shows a map of each point included in the three-dimensional image space, and points A, B, C, and G on the map, and (b) to (d) show the display video that the user is viewing. In FIG. 32, the point A shown in (a) corresponds to the position of the user (the position of the display device 100) in (b), the point B shown in (a) corresponds to the position of the user (the position of the display device 100) in (c), the point D shown in (a) corresponds to the position of the user (the position of the display device 100) in (d), and the point G shown in (a) corresponds to the position of the object of fixation. In FIGS. 32(b) to (d), a schematic diagram representing the direction in which the user is looking is shown below the figure of the video that each user is viewing.

[0174] Here, the upper side of the paper surface is the front direction in the three-dimensional image space, and is the direction recognized as the direction in which the user is looking by the guide. That is, the user shown in FIG. 32(b) is in a state of looking in the front direction, the user shown in FIG. 32(c) is in a state of looking in the front direction, and the user shown in FIG. 32(d) is in a state of looking in the right hand direction. Here, as an input interface, a keyword such as "Please gather at ○○ (the name of the place G, etc.)" uttered by the guide is acquired by the voice input means 71 or the like, and this is acquired as data regarding the position of the object of fixation. Then, as shown in FIGS. 32(b) to (d), the voice is reproduced. Here, since neither the position of the object of fixation nor the position of the display device 100 coincides, the respective relative positions are calculated, and the voice "Please gather at ○○" is reproduced together with the display video. And the voice reproduced here is stereophonic sound that is perceived by the user as sound coming from the direction of the relative position, similar to the examples of FIGS. 29 and 30.

[0175] In this way, at least one of the relative position and the relative direction can be grasped by the user, so that at least one of the position and the direction of the object of fixation can be easily recognized. Thus, in the video display system 500, it is possible to display an appropriate video on the display device 100.

[0176] [Embodiment] The following will be described in more detail based on the embodiments of the implementation form. In this embodiment, mainly two use cases of VR tourism will be described. The two use cases are two cases for 360° camera shooting and for 3D CG space. In the former case, there is a shooter with a 360° camera at a remote location, who sends 360° video + metadata, and the viewer generates CG based on information such as the metadata on a VR device and synthesizes it with the sent 360° video for viewing. In this case, it is further subdivided into applications such as VR tourism (real-time tourism), VR tourism (real-time + recording), VR factory tour, on-site inspection: the guide (requiring instructions to the cameraman) guides, and VR real exhibition inspection: the guide (designating a privileged leader for each group) guides and conducts a guided tour.

[0177] On the other hand, in the latter case, it is assumed that multiple people participate in a 3D space and view it on a VR device with a function where one person acts as a guide (privileged user) to lead and explain to other participants. In this case, it is further subdivided into applications such as guided VR tourism (in a CG space): excluding the case where each person can freely move in a 3D space such as a museum, guided VR exhibition (VR market) inspection, VR group education (for example, 30 students. When students are studying in a VR-CG space (360° space), the teacher can give instructions, so it is expected to have advantages over an actual classroom), VR presentation: for example, when multiple people are viewing an architectural CG space, the presenter can control the actions of the audience.

[0178] In 360° video, when shooting and transmitting 360° video in real time, if the location is the same and only the viewing directions of different users are different, when the guide is taking a selfie and wants to insert a cue when looking at the direction in which they are reflected, the cue is transmitted in some way. On the viewer side, when the cue arrives, the center of the viewpoint can be forcibly moved there (changing the video cut-out position). At this time, in order to avoid VR sickness, processes such as displaying movement and switching are performed (for example, if forced to move smoothly, one may get sick). Alternatively, when the cue arrives, an arrow or target may be displayed on the viewer to prompt the user to turn their body in that direction.

[0179] On the other hand, when another person is shooting, if the guide gives an instruction or the cameraman decides, a specific direction of the camera is pointed in that direction, a cue is inserted and transmitted. The viewer side may be the same as above. As for the implementation method of the cue, a button for inserting a cue may be attached to the 360° camera, the application of image recognition technology, for example, when the guide presents a specific marker or a specific action (gesture, etc.), any one of the camera, server, or viewer recognizes it and confirms the cue (image marker cue), the use of inaudible sound, for example, using inaudible sound like a dog whistle for transmission to make a cue (specific sound source cue), or the use of voice recognition technology, for example, registering a specific keyword (activation word of an AI speaker) and having the viewer recognize it to make a cue (voice cue), etc. can be mentioned.

[0180] In 3DCG, assuming a case where, for example, Heiankyo is reproduced in CG and a guided tour is being conducted inside it (similar to the case where multiple people are viewing inside a large building in a VR meeting system), it is assumed that the positions of each person are also scattered. In this case, it is necessary to specify the location and the direction of the line of sight and forcibly move the current position to the specified location.

[0181] When a CG avatar corresponding to a guide instructs to look in a specific direction at a specific location for explanation or the like, the user is forced to face a specific direction, induced by displaying a spatial guide, or moved (warped) to a position where the current position is significantly different from the designated location and the designated object cannot be seen. However, since a sudden warp may cause surprise (which may lead to dizziness), some warning is displayed on the viewer, and a different warp button (escape button) is displayed, etc., and the user is made to take an action before moving (also aligning the direction during movement). As for the processing when a CG avatar instructs a gathering, it corresponds to the third case above, and the user is warped to the location where the guide is and made to face the direction of the guide.

[0182] As important elements in the present invention, there are two: (i) a method for detecting, recording, and transmitting direction and orientation, and (ii) a method for generating, transmitting, and controlling (forcing, alerting, etc.) cues such as turning participants in a specific direction and gathering them. (i) further has two sub-elements: detection of direction and orientation, and transmission of direction and orientation. Also, (ii) further has three sub-elements: generation of cues, transmission of cues, and a method for controlling the VR space according to the cues.

[0183] For cue generation, it is assumed that switches, menus are used, the display of a specific object is used, such as a guiding flag of a guide, a specific light pulse is used, specific keywords are used, such as directions like "front", "right" used by a tourist guide, the name of a building, actions like "gathering", or keywords are used when giving instructions like "Everyone, turn to the right" with a specific keyword attached to the head, or inaudible sound pulses for humans, such as those like an audio watermark are used.

[0184] For cue transmission, content similar to that for direction transmission is assumed.

[0185] For the method of controlling the VR space according to the queue, the following controls are assumed to be performed depending on the type of queue. That is, aligning the VR space with the marker position, moving the location of the VR space, controlling the appearance of the VR video (resolution, contrast, frame rate, viewing angle, etc.), switching between live and recorded videos, and so on.

[0186] FIG. 33 is a diagram showing an example of the functional configuration of the video display system according to the embodiment. FIG. 34 is a diagram showing an example of the functional configuration of the observation system according to the embodiment. FIG. 35 is a diagram showing an example of the functional configuration of the VR system according to the embodiment.

[0187] An example of realizing a 360° camera among the realization examples of the observation system (observation device) of the embodiment of the present invention will be described with reference to FIGS. 33 and 34.

[0188] The observation system 3351 (360° camera 3401) of the embodiment of the present invention is substantially the same as the realization example of the 360° camera of the conventional example 2, and the differences will be described. In the 360° camera of the embodiment of the present invention, an input button and a position / azimuth detection / memory unit 3402 are added as a queue information input unit (queue information input means 3352) as a program of the CPU 802, and the queue information input from the input button (data input unit) and the metadata based on the position / azimuth detected by the position / azimuth detection / memory unit (metadata conversion unit) 3402 are generated in the position / azimuth detection / memory unit, sent to the multiplexing unit 832, multiplexed, and sent to the VR system 3301 via the wireless communication element (transmission unit) 820.

[0189] The queue information is used to select a destination by selecting a previously specified direction, such as the right side or the left side, or a plurality of previously specified locations by number or menu using the input button, and to give the timing to start moving.

[0190] More specifically, there are designations of right, left, front, and back by buttons, designations of start / end by the azimuth change start / end button 811, designations of targets by number keys, designations by touch panels, and so on.

[0191] In some cases, the queue information is generated by analyzing the captured video not only by the input button but also by the position / orientation analysis units 3403 and 3404 indicated by the dashed lines, comparing it with a pre-specified image to detect the destination and the position / orientation of the moving direction, analyzing the gestures and body movements of the guide to detect the direction of the moving destination and the timing of starting to move, providing light-emitting elements such as buttons and LEDs on an object like the pointer held by the guide, detecting the direction and position of the moving destination by detecting the pulse emission pattern of the light-emitting element due to the operation of the guide, or selecting one from a plurality of pre-determined destinations and detecting the timing of starting to move. Alternatively, it is generated by analyzing the voice from the microphone, identifying the destination from the words uttered by the guide, detecting the moving direction and the timing of starting to move, and converting it into appropriate metadata by the position / orientation detection function executed by the CPU 802 and sending it to the VR system 3301.

[0192] An implementation example of the VR system 3301 (display device, HMD / VR glasses and computer / smartphone 3501) according to an embodiment of the present invention will be described with reference to FIGS. 33 and 35.

[0193] The difference between the VR system of the second conventional example and the VR system 3301 according to an embodiment of the present invention will be mainly described.

[0194] In the embodiment, a position / orientation / queue determination unit (viewpoint determination unit) 3502 and 3503 are added as programs to the CPU 965 and GPU 954 of the computer / smartphone 3501 of the VR system of the second conventional example.

[0195] In the position / orientation / queue determination unit 3503 of the CPU 965, metadata including queue data is received from the observation system 3351 via a communication element (reception unit), together with the position / orientation of the target object obtained from the observation system 3351 or a guide. When changing the voice according to the queue data, a guide voice is generated by the voice playback unit realized by the program of the CPU 965, or the reproduced voice is appropriately processed. When changing VR video or graphics according to the queue data, the position / orientation / queue determination unit 3503 of the CPU 965 sends the metadata to the GPU via the system bus.

[0196] In the GPU 954, the metadata received by the position / orientation / queue determination unit 3502 is processed, and information is sent to the graphics generation unit (graphics generation unit) 959 to display a figure based on the queue data. The graphics data from the graphics generation unit 959 is superimposed and displayed on the VR video by the VR display control unit 957. Alternatively, information based on the queue data is sent from the position / orientation / queue determination unit 3502 to the VR control unit 956. The position / orientation state of the VR system detected by the motion / position detection processing unit (detection unit) 955 from the information of the motion / position sensor 903 and the queue data are used by the VR display control unit (display control unit) 957 to appropriately process the VR video. Video data is sent from the AV output 952 to the AV input 925, and is displayed as a VR video on the display element (display unit) 905 by the video display processing unit 912. The above processing of voice, graphics, and VR video may be realized independently without other processing, or multiple processes may be realized, and the processes may be selected during the operation of the VR system or the observation system.

[0197] In addition, the position and orientation detection process of the observation system 3351 may also be realized by a computer system between the observation system such as the cloud and the VR system. In this case, either no metadata is sent from the observation system 3351, or the data input by the operator is sent as metadata. For example, with the position and orientation detection means in the cloud, the position, orientation, or movement of the observation system, guide, or target object is detected from the video, audio, or metadata sent from the observation system and sent to the VR system as metadata. This enables the existing 360° camera to also exhibit the effects of this embodiment.

[0198] Regarding some processes, whether they are performed by the GPU 954 or the CPU 965 may be different from this example, and the bus configuration may also be different from this example, but there is no difference in the functional configuration and operations described later.

[0199] Regarding the integrated VR system, this embodiment is almost the same as the conventional example. By realizing the functions with one CPU 965 and one GPU 954 respectively, a small integrated VR system can be realized.

[0200] Referring to FIG. 33 again, another configuration of the embodiment will be described.

[0201] In FIG. 33, the embodiment describes the integration of the two systems as the flow of data and control rather than the actual connection, with the observation system in FIG. 34 and the VR system in FIG. 35 as functional blocks.

[0202] In the observation system 3351, the VR shooting camera 804 in FIG. 34 corresponds to the VR shooting means 762 in FIG. 33. Similarly, the VR video processing unit corresponds to the VR video processing means 758, the VR video compression unit corresponds to the VR compression means 756, the microphone group, microphone terminal, microphone amplifier, and ADC correspond to the audio input means 763, the audio compression unit corresponds to the audio compression means 760, the input button corresponds to the queue information input means 3352, the motion / position detection unit, the position / orientation detection / memory unit, and the two position / orientation analysis units of the GPU and CPU correspond to the position / orientation detection / memory means 3353, the multiplexing unit corresponds to the multiplexing means 2652, the wireless communication element corresponds to the communication means 754, the separation unit corresponds to the separation means 755, the audio decoding unit corresponds to the audio decoding means 761, and the DAC, amplifier, headphone element, and speaker correspond to the audio output means 764. Since the video bus, memory bus, system bus, I / O bus, bus conversion, RAM, EEPROM, SD card, power switch, power control element, battery, display element, shooting mode selection button, zoom button, and shooting start / end button are not directly related to the operation of the present invention, their illustrations are omitted.

[0203] In the VR system 3301, the communication element in Fig. 35 corresponds to the communication means 716 in Fig. 33. Similarly, the separation unit corresponds to the separation means 715, the audio decoding unit corresponds to the audio decoding means 713, the audio playback control unit corresponds to the audio playback control means 709, the DAC, amplifier, speaker, and headphone jack correspond to the audio playback means 705, the VR video decoding unit corresponds to the VR video decoding means 710, the graphics generation unit corresponds to the graphics generation means 712, the position / orientation / queue determination units in the CPU and GPU respectively correspond to the position / orientation / queue determination means 3302, the motion / position sensors and the motion / position detection unit correspond to the position detection means and the rotation detection means, the motion / position detection processing unit and the VR control unit correspond to the VR control means 707, the VR display control unit corresponds to the VR display control means 708, the video display processing unit, the display element, and the lens correspond to the VR video display means 704, the microphone, microphone amplifier, and ADC correspond to the audio input means 706, the audio compression unit corresponds to the audio compression means 714, and the multiplexing unit corresponds to the multiplexing means 717. The video bus, memory bus, system bus, I / O bus, bus conversion, RAM, EEPROM, non-volatile memory, power switch, power control element, battery, volume button, AV output, AV input, and USB are not directly related to the operation of the present invention or are described as one system, so their illustrations are omitted. The wireless communication element is necessary for communication with the controller, but since the controller is omitted in Fig. 33, its illustration is omitted.

[0204] In an embodiment of the present invention, a queue information input means 3352 is provided. The queue information input means 3352 inputs the start and end timings of movement, the position or direction of movement, or the position and direction of the destination (target) by means of switches, tablets, smartphones, etc. physically operated by the operator or guide of the observation system 3351. Also, a target for moving from a plurality of targets is specified based on the queue data.

[0205] Alternatively, as indicated by the dashed line, the queue information input means 3352 may obtain queue information from the video obtained from the VR video processing means 758 or the audio information obtained from the audio input means.

[0206] The queue information obtained from the queue information input means 3352 is sent to the position and orientation detection and storage means 3353, where it is processed together with the position and orientation of the observation system 3351. In some cases, its state is stored and shaped into appropriate data, which is sent to the multiplexing means 2652 as metadata. After multiplexing together with video, audio, and graphics, it is sent to the VR system 3301 by the communication means 754.

[0207] In the VR system 3301, the communication means 716 receives the communication information from the observation system 3351. The separation means 715 separates the metadata and sends it to the position, orientation, and queue determination means 3302.

[0208] In the position, orientation, and queue determination means 3302, the queue data is extracted from the metadata, and determined processing is performed. It is sent to the graphics generation means 712 to display the queue information as a graphic, and is superimposed and displayed on the VR video by the VR display means 704. Alternatively, it is sent to the VR control means 707, and the VR video is appropriately processed by the VR display control means 708 together with the position and orientation state of the VR system 3301, and is displayed by the VR display means 704. Or, the voice playback control means 709 generates a guidance voice or appropriately processes the playback voice.

[0209] As a specific example, when the queue data indicates "move to target A", depending on the position of the VR system, the position of target A is different. When the queue information is shown in graphics, an arrow is displayed in an appropriate direction. Specifically, when target A is on the left side, a left - facing arrow is displayed. When controlling the video, for example, only the left side becomes clear. When controlling with voice, an announcement such as "Please face the left side" is played. In this way, appropriate processing is performed by comparing the content of the queue data with the position and orientation of the VR system.

[0210] Figures 36 and 37 are diagrams showing a configuration example of the metadata according to the embodiment. The configuration example of the metadata of this embodiment will be described.

[0211] The type of metadata contains a predefined code or character string indicating that it is the metadata of the present invention. The version number is a number for when the metadata structure is changed. For example, in the evaluation stage, it is used like a major version and a minor version such as 0.81 (0081), during the proof-of-concept experiment it is 0.92 (0092), and at the time of release it is 1.0 (0100), etc., and it is used with the idea of guaranteeing compatibility between the same major versions.

[0212] When the function code is 0, it indicates that the metadata information is invalid, and in other cases, it indicates the type of information in the metadata. For example, 0001 indicates that it is a format for describing the reference position, camera, positions and moving directions and speeds of the guide and target. 0002 indicates graphics data, 0003 indicates information of the VR system, 0011 is one with queue data sent from the observation system attached to 0001, 0021 is one with queue data and defines a moving target, etc.

[0213] In addition, the metadata contains a plurality of parameters related to the queue data, such as the type and size of the queue data. As parameters related to the queue data, as shown in FIG. 37, for example, 8 types from 0 to 7 are prepared for the type of queue data, and a numerical value specifying which queue data it is is input. Also, in this example, since a plurality of targets can be selectively specified, parameters for specifying the targets are included. Specifically, different numerical values are set for each of the plurality of targets, and the targets can be specified and selected by numerical values. Also, as one of the parameters, it is possible to input a parameter simply for specifying a direction. Here, the direction can be specified in units of 1°.

[0214] The reference position is the data of the position that serves as the reference for the position data, and it is defined in advance including units such as being represented by, for example, X (east-west distance), Y (north-south distance), Z (height-direction distance) or longitude, latitude and altitude for the entire system. When the reference position is 0, it indicates that the position at the time of resetting the entire system is used as the reference. Also, it is determined in advance whether the positions of the camera and the guide are absolute coordinates or relative coordinates from the reference position.

[0215] The moving direction and speed indicate the moving situation of the observation system or the guide. When there is queue data, it indicates how it will move in the future.

[0216] The number of targets indicates the destinations for tourism in the case of VR tourism. When the number of targets is 0, it indicates that there are no targets.

[0217] The verification code is a code for verifying whether the metadata data is incorrect during transmission, and for example, CRC etc. is used.

[0218] Regarding the order, content, and values of each item of the metadata, those having the same functions even if different from this configuration example may be used.

[0219] Figure 38 is a diagram showing another configuration example of the metadata according to the embodiment. In this example, the metadata in a state where a target is defined is shown as compared with the example in Figure 36.

[0220] Figure 39 is a diagram showing an example of the operation flow of the video display system according to the embodiment. The operation of the embodiment of the present invention will be described.

[0221] In the observation system, in the queue information input step, the input from an input button or the like is regarded as queue information by the queue information input means (S3932), and it is confirmed whether there is valid queue information by the position / orientation detection and storage means (S3933). If there is valid queue information (Yes in S3933), metadata is generated by the position / orientation detection and storage means for the input queue information (S3934). Next, the queue information is multiplexed with video, audio, and graphics information by the multiplexing means (S3927) and sent to the VR system by the communication means (S3928). If there is no valid queue information (No in S3933), no processing is performed (S3935).

[0222] Alternatively, queue information is extracted from the input voice information, VR video voice input, and input of VR video from the photographing means by the queue input means (S3930), converted into metadata by the position / orientation detection and storage means (S3931), multiplexed with video, audio, and graphics information by the multiplexing means (S2927), and sent to the VR system by the communication means (S3928).

[0223] In the VR system, metadata is separated by the separation means from the information received by the communication means (S3901, S3902), and the metadata is analyzed by the position / orientation / queue determination means in the metadata analysis step (S3906). If there is queue information, according to the queue information, the queue information is sent to the graphics generation means, the VR control means, or the audio control means. The graphics generation means generates graphics based on the queue information. Alternatively, the VR control means controls the VR video according to the queue information (S3907, S3908, S3909, S3910). Furthermore, addition or control of audio information based on the queue information by the audio control means may be performed (S3903, S3904, S3905). Which of the above processes is performed depends on the settings of the VR system or the entire system.

[0224] Also, for the steps not described above, the explanations here are omitted by referring to the explanations in the similar steps in FIG. 13. Specifically, step S3921 corresponds to step S1321, step S3922 corresponds to step S1322, step S3923 corresponds to step S1323, step S3924 corresponds to step S1324, step S3925 corresponds to step S1325, and step S3926 corresponds to step S1326.

[0225] FIG. 40 is a diagram for explaining the result of the operation of the video display system in the embodiment. In the example of FIG. 40, the case where arrow display graphics are superimposed is shown. Specifically, in the queue data, when data indicating "facing right" is shown, depending on the position of the VR system, there may be a target on the left side of the user of the VR system. In that case, when indicating queue information with graphics, a left-facing arrow is displayed. Which way the arrow points depends on whether the position and direction of the target are well-known with high accuracy in the observation system. When that information is input and sent to the VR system as metadata, and compared with the orientation of the user of the VR system, if the target is in front with a predetermined error, the arrow indicates the front of the user of the VR system. If it is more than the predetermined error to the left of the user, an arrow indicating the left side is shown, and if it is more than the predetermined error to the right, an arrow indicating the right is shown. If the target is behind the user within a separately defined error range, a backward-facing arrow is shown.

[0226] Even when the position and direction of the target are ambiguous in the observation system, if approximate information is input, it is converted into metadata and sent to the VR system, and the same processing as above is performed. The position and orientation determination means that displays the arrow receives the position and orientation of the observation system, guide, or target object as metadata, sends that information to the graphics generation means to display it as a graphic, and superimposes and displays it on the VR video by the VR display means.

[0227] For example, in the leftmost example in the figure, the state where an arrow indicating the front is displayed is shown. Next, in the example on the upper left, the state where an arrow is displayed as indicating the right side is shown. Next, in the middle example, when the user of the VR system is facing right, an arrow indicating the left side, when facing left, an arrow pointing backward, and when facing backward, an arrow pointing to the left rear (in the figure, an arrow indicating the left side is displayed when facing right) are shown. When the user changes the direction, the direction of the arrow changes accordingly as described above. Next, in the example on the upper right, the state in the case of MAP display is shown. In the case of MAP display, an arrow is displayed as indicating the right side, or the destination is indicated by a star mark or the like. Next, in the rightmost example, the case of MAP display is shown. Also in the case of MAP display, the state where the arrow and the MAP are appropriately rotated and displayed according to the direction is shown.

[0228] Figure 41 is a diagram for explaining the result of the operation of the video display system in the embodiment. The example of Figure 41 shows the case of processing VR video. Specifically, in the case where data indicating "facing right" is shown in the queue data, depending on the position of the VR system, there may be a target on the left side or behind the user of the VR system, and when controlling the video, for example, the control positions of the mask and the resolution change. Note that the determination of the direction is the same as in the case of the arrow.

[0229] For example, in the leftmost example in the figure, when the front is the target, no special processing is performed and the video is simply displayed as it is. Next, in the example on the upper left, areas other than the right side are masked to prompt facing right. Next, in the middle example, the resolution of areas other than the right side is reduced to prompt facing right. Next, in the example on the upper right, when facing right, the mask (example on the upper left) and the resolution (middle example) return to the original display. Next, in the rightmost example, when facing right from the beginning, nothing is displayed, or processing such as masking areas other than the central part or reducing the resolution is performed. When facing left, a small part of the right side may be masked, and when facing backward, the left side may be masked.

[0230] FIG. 42 is a diagram for explaining the result of the operation of the video display system in the embodiment. In the example of FIG. 42, the case where an audio guide is played is shown. Specifically, in the case of cue data indicating "turn right", depending on the position of the VR system, there may be a target on the left side of the user of the VR system. When controlling by voice, an announcement such as "please turn to the left side" will be played. Note that the determination of the direction is the same as in the case of the arrow.

[0231] For example, in the example at the left end of the figure, when explaining the building in the front, the state where the audio is played so that the guide's voice can be heard from the front is shown. Next, in the middle example, when explaining the building on the right hand, the guide's voice can be heard from the right, and when the back is the target, the voice is played so that it can be heard from the back, and when the left is the target, the voice is played so that it can be heard from the left (in the figure, an example where the guide's voice can be heard from the right when explaining the building on the right hand) is shown. Next, in the example at the right end, when the user of the VR system is facing right or back, since the right hand is actually the left hand, the guide's voice can be heard from the left of the user of the VR system. When there is confusion, replace "right" in the voice with "left", or indicate the direction with an arrow in combination with the graphics, etc.

[0232] FIG. 43 is a diagram for explaining the result of the operation of the video display system in the embodiment. In the example of FIG. 43, the case where a plurality of users are dispersed within the same image space is shown. Specifically, when there are a plurality of users of the VR system and they are in different places from the location of the guide within the VR space, the VR video seen by the users of the VR system will be the one that was pre-shot. Alternatively, there are a plurality of observation systems. At this time, in order for the guide to gather the participants to the location of the guide, when sending cue data for gathering by indicating the voice, button, or position of the map, depending on the position and orientation of the users of the VR system, the arrow or VR image is processed, or it is displayed or played for each user of the VR system by voice, vibration of the controller, etc.

[0233] In the figure, VR system A faces north, and since the front is A Shrine, the arrow is displayed pointing forward. VR system B faces north, and since the left hand is A Shrine, the arrow is displayed pointing left. VR system C faces east, and since the right hand is A Shrine, the arrow is displayed pointing right. Additionally, as shown as an alternative example of VR system C, MAP display may be performed.

[0234] When moving, it may be possible to use a so-called warp function with a controller to move. As a similar function, it may also be possible to move by selecting the function of "move to the position of the guide" from the menu.

[0235] This function is also effective when multiple VR systems are dispersed in a wide VR space such as VR meetings, VR events, and VR sports viewing.

[0236] Figure 44 is a diagram for explaining an example of the use of the video display system in the embodiment. As shown in Figure 44, the video display system can be used in an example of a VR exhibition. For example, when a guide or host announces a gathering or the like at Booth 2-3, participants in multiple VR systems can move from their respective positions.

[0237] Figure 45 is a diagram for explaining an example of the use of the video display system in the embodiment. As shown in Figure 45, the video display system can be used in an example of a VR mall. For example, when a guide or host announces a gathering or the like at Atrium A8, participants in multiple VR systems can move from their respective positions.

[0238] In this way, instead of sending VR images from the observation system, it can be applied to VR tourism, VR meetings (gathering from multiple sessions to the overall session), VR sports viewing (gathering from multiple viewing locations to the main location), VR events (gathering from multiple different event locations to the main venue), etc. in a VR space composed of CG.

[0239] In this case, functions and operations implemented in the observation system are realized by a VR system operated by a guide or the organizer.

[0240] Also, when moving in groups to different venues, not only for a gathering, the same method can be used. In this case, instead of the organizer or the guide issuing a queue, the members of the group can issue the queue.

[0241] FIG. 46 is a diagram for explaining another example of the movement method of the video display system in the embodiment. As shown in FIG. 46, when a movement queue is issued from a guide or the organizer, etc., the method of starting the movement differs depending on the situation of the VR screen being viewed by the user of the VR system. For example, when the movement direction is indicated by an arrow, a change in the screen, or a voice, the function of the controller (e.g., the warp function) is used to move. This method is not a problem for short distances but is laborious for long distances.

[0242] When an arrow is displayed, it may be possible to move to the target location by selecting the arrow. Similarly, in the case of MAP display, it may be possible to move to the corresponding location by selecting an arbitrary point within the MAP. Further, characters (such as A, B, C, etc.) corresponding to the target may be displayed, and it may be possible to move to the corresponding pre-set location by selecting any of these characters. When the screen is masked, it may be possible to move to the designated location by selecting the unmasked part.

[0243] FIG. 47 is a diagram for explaining a configuration example of realizing the video display system according to the embodiment using the cloud. In the configuration shown in FIG. 47, by having a function on the cloud to control graphics, VR videos, sound, and the vibration of the controller according to the position, orientation, and queue information of the observation system 4761 and the position and orientation of the VR system, and to provide appropriate information to the user of the VR system, the effects of the present invention can be exhibited even in a simple VR system.

[0244] In the configuration shown in Fig. 47, by having the position / orientation / queue detection and storage means 4740 in the cloud, the position and orientation of the observation system 4761 on the cloud are read from the metadata separated by the separation means 4742 from the data sent from the observation system, or the position and orientation of the observation system are read from the VR video sent from the observation system 4761, and graphics such as arrows are generated by the graphics generation means 4736. The VR control means 4707 determines the position and orientation of the observation system and the VR system position and orientation, and the VR display control means 4708 synthesizes the VR video and the graphics, or processes the VR video, and also changes the sound localization by the sound reproduction control means 4739, changes the content of the sound, etc., so that a display and sound output suitable for the position and orientation of the VR system 4701 can be achieved. Also, although not shown here, by appropriately controlling the controller of the VR system, it is possible to notify the user of the VR system of the direction and position by vibration or the like.

[0245] Note that the functions provided on the cloud are not limited to the configuration shown in Fig. 47, and the functions provided on the cloud can be selected so that the overall functions and operations are substantially the same according to the configuration and functions of the connected observation system or VR system. As an example, in the observation system, when the position and orientation of the observation system are not detected, but the position and orientation of the observation system are detected on the cloud and sent to the VR system superimposed on the video as graphics, there are limitations in changing the graphics according to the position and orientation of the VR system, but no special functions are required for the VR system. Also, in a configuration where the VR system is provided with position / orientation control means and graphics generation means for correcting the graphics according to the position and orientation of the VR system, it is also possible to change the graphics according to the position and orientation of the VR system.

[0246] Note that for configurations not described above, the descriptions here are omitted by referring to the descriptions of configurations with the same names in FIG. 33. Each of the position detection means 4702, rotation detection means 4703, VR display means 4704, audio reproduction means 4705, audio input means 4706, VR control means 4707, VR display control means 4708, audio decoding means 4709, audio compression means 4710, VR video decoding means 4711, separation means 4712, multiplexing means 4713, and communication means 4714 included in the VR system 4701, and the separation means 4732, VR video compression means 4733, multiplexing means 4734, communication means 4735, graphics generation means 4736, VR display control means 4737, VR video decompression means 4738, audio reproduction control means 4739, position / orientation / queue detection storage means 4740, communication means 4741, and separation means 4742 included in the computer system 4731, and the queue information input means 4762, multiplexing means 4763, communication means 4764, separation means 4765, VR video compression means 4766, audio compression means 4767, audio decoding means 4768, VR video processing means 4769, VR shooting means 4770, audio input means 4771, and audio output means 4772 included in the observation system 4761 respectively correspond to each of the position detection means 702, rotation detection means 703, VR display means 704, audio reproduction means 705, audio input means 706, VR control means 707, VR display control means 708, audio reproduction control means 709, VR video decoding means 710, graphics generation means 712, audio decoding means 713, audio compression means 714, separation means 715, communication means 716, multiplexing means 717, position / orientation / queue determination means 3302, queue information input means 3352, position / orientation detection / storage means 3353, communication means 754, separation means 755, VR video compression means 756, multiplexing means 757, VR video processing means 758, graphics generation means 759, audio compression means 760, audio decoding means 761, VR shooting means 762, audio input means 763, and audio output means 764 in a one-to-one, many-to-one, one-to-many, or many-to-many manner.

[0247] FIG. 48 is a diagram for explaining a configuration example of realizing the video display system according to the embodiment using a cloud. As shown in FIG. 48, the position / orientation / queue detection storage means 4740 of the observation system may be realized by a computer system between the observation system 4761 such as a cloud and the VR system 4701. In this case, either metadata indicating the direction is not sent from the observation system, or data of queue information input by the operator or the like is sent as metadata. For example, in the position / orientation / queue detection storage means 4840 in the cloud, the position, orientation, or movement of the observation system, guide, or target object is detected from the video, audio, or metadata sent from the observation system 4861 and sent to the VR system as metadata. Thereby, the effects of this embodiment can be exhibited even with an existing 360° camera.

[0248] Furthermore, the position / orientation determination means 4915 on the VR system side and the control of VR video and audio by the same may also be realized by a computer system between the VR system such as a cloud and the observation system. In this case, the same processing can be performed at one location, and effects such as easily giving the same effect to a plurality of VR systems simultaneously and being able to give the effects of the present invention to an existing system can be expected. However, in order to reflect the direction and position of the VR system, it is necessary to send the position and direction of the VR system from the VR system to the cloud side, and it is necessary to provide a processing unit corresponding to each VR system on the cloud side.

[0249] The configuration of FIG. 48 is an example in the case where the position and orientation of the VR system are not sent to the cloud side. In this case, it becomes difficult to display an arrow, change the audio, etc. according to the position and orientation of the VR system. However, with the VR display control means, it is possible to perform processes such as changing the resolution of the VR video, masking, and changing the sound localization according to the output of the position / orientation / queue detection storage means.

[0250] In addition, for the configurations not described above, the descriptions here are omitted by referring to the descriptions of the configurations with the same names in FIG. 33. Each of the position detection means 4802, rotation detection means 4803, VR display means 4804, audio reproduction means 4805, audio input means 4806, VR control means 4807, VR display control means 4808, audio decoding means 4809, audio compression means 4810, VR video decoding means 4811, separation means 4812, multiplexing means 4813, communication means 4814, and audio reproduction control means 4817 included in the VR system 4801, and the VR video compression means 4733, multiplexing means 4834, communication means 4835, graphics generation means 4836, VR display control means 4837, VR video decompression means 4838, audio reproduction control means 4839, position / orientation / queue detection and storage means 4840, communication means 4841, and separation means 4842 included in the computer system 4831, and the queue information input means 4862, multiplexing means 4863, communication means 4864, separation means 4865, VR video compression means 4866, audio compression means 4867, audio decoding means 4868, VR video processing means 4869, VR shooting means 4870, audio input means 4871, and audio output means 4872 included in the observation system 4861 respectively correspond to each of the position detection means 702, rotation detection means 703, VR display means 704, audio reproduction means 705, audio input means 706, VR control means 707, VR display control means 708, audio reproduction control means 709, VR video decoding means 710, graphics generation means 712, audio decoding means 713, audio compression means 714, separation means 715, communication means 716, multiplexing means 717, position / orientation / queue determination means 3302, queue information input means 3352, position / orientation detection and storage means 3353, communication means 754, separation means 755, VR video compression means 756, multiplexing means 757, VR video processing means 758, graphics generation means 759, audio compression means 760, audio decoding means 761, VR shooting means 762, audio input means 763, and audio output means 764 in a one-to-one, many-to-one, one-to-many, or many-to-many manner.

[0251] FIG. 49 is a diagram for explaining a configuration example of realizing the video display system according to the embodiment using the cloud. In the configuration of FIG. 48, it was difficult to display an arrow or change the voice according to the position and orientation of the VR system. However, in the configuration shown in FIG. 49, by providing the VR system with position / orientation determination means 4902 and 4903, the position and orientation of the observation system on the cloud are read from the metadata separated by the separation means 4912 from the data sent from the observation system, and graphics such as arrows are generated by the graphics generation means 4916 accordingly. This is converted into metadata together with the position / orientation information and the like sent from the observation system by the metadata conversion means 4937, multiplexed by the multiplexing means, and sent to the VR system.

[0252] In the VR system, graphics are generated from the metadata separated by the separation means, and the position and orientation of the observation system and the position and direction of the VR system obtained from the position detection means and the rotation detection means are determined by the VR control means, and the VR video and the graphics are appropriately synthesized by the VR display control means, or the VR video is processed. Also, by changing the sound localization or changing the content of the voice by the voice reproduction control means, it is possible to output a display and voice suitable for the position and orientation of the VR system. Although not shown here, it is possible to appropriately control the controller of the VR system and notify the user of the VR system of the direction and position by vibration or the like.

[0253] Note that the position / orientation information of the VR system detected by the position detection means and the rotation detection means of the VR system is used as metadata, multiplexed with other information in the multiplexing means, and sent to the computer system on the cloud by the communication means. This function is generally provided in a general VR system.

[0254] In addition, for configurations not described above, the explanations here are omitted by referring to the explanations in the configurations with the same names in FIG. 33. Each of the position detection means 4902, rotation detection means 4903, VR display means 4904, audio reproduction means 4905, audio input means 4906, VR control means 4907, VR display control means 4908, audio decoding means 4909, audio compression means 4910, VR video decoding means 4911, separation means 4912, multiplexing means 4913, communication means 4914, position / orientation determination means 4915, graphics generation means 4916, and audio reproduction control means 4917 included in the VR system 4901, as well as the multiplexing means 4934, communication means 4935, graphics generation means 4936, VR display control means 4937, position / orientation / queue detection and storage means 4940, communication means 4941, and separation means 4942 included in the computer system 4931, and each of the queue information input means 4962, multiplexing means 4963, communication means 4964, separation means 4965, VR video compression means 4966, audio compression means 4967, audio decoding means 4968, VR video processing means 4969, VR shooting means 4970, audio input means 4971, and audio output means 4972 included in the observation system 4961 corresponds one-to-one, many-to-one, one-to-many, or many-to-many to each of the position detection means 702, rotation detection means 703, VR display means 704, audio reproduction means 705, audio input means 706, VR control means 707, VR display control means 708, audio reproduction control means 709, VR video decoding means 710, graphics generation means 712, audio decoding means 713, audio compression means 714, separation means 715, communication means 716, multiplexing means 717, position / orientation / queue determination means 3302, queue information input means 3352, position / orientation detection and storage means 3353, communication means 754, separation means 755, VR video compression means 756, multiplexing means 757, VR video processing means 758, graphics generation means 759, audio compression means 760, audio decoding means 761, VR shooting means 762, audio input means 763, and audio output means 764.

[0255] (Other Embodiments) As described above, the embodiments etc. have been explained, but the present disclosure is not limited to the above embodiments etc.

[0256] In addition, although the components constituting the video display system have been exemplified in the above-described embodiments and the like, the functions of the components included in the video display system may be distributed in any manner among a plurality of parts constituting the video display system.

[0257] In addition, in the above-described embodiment, each component may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0258] In addition, each component may be realized by hardware. For example, each component may be a circuit (or an integrated circuit). These circuits may constitute one circuit as a whole, or may be separate circuits respectively. Further, these circuits may each be a general-purpose circuit or a dedicated circuit.

[0259] Furthermore, the general or specific aspects of the present disclosure may be realized by a system, an apparatus, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM. Further, the present disclosure may be realized by any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium.

[0260] In addition, forms obtained by applying various modifications that can be conceived by those skilled in the art to the embodiments and the like, or forms realized by arbitrarily combining the components and functions in the embodiments and the like without departing from the gist of the present disclosure are also included in the present disclosure.

Industrial Applicability

[0261] The present disclosure is useful in applications for displaying appropriate video on a display device.

Description of Reference Numerals

[0262] 11 Position detection means 13 Rotation detection means 15 VR display means 17 Audio playback means 19, 71 Audio input means 21 VR control means 23 VR display control means 25 Audio playback control means 27, 63 Audio compression means 29 VR video decoding means 31 Position / orientation / queue determination means 33 Graphics generation means 35, 65 Audio decoding means 37, 57 Separation means 39, 55 Communication means 41, 61 Multiplexing means 51 Queue information input means 53 Position / orientation detection / memory means 59 VR video compression means 67 VR video processing means 69 VR shooting means 73 Audio output means 99 Graphics 99a Arrow 99b Map 99c Mask 99d Downsampling filter 100 Display device 101 Display unit 102 Display state detection unit 150 Network 200 Server device 201 Receiver 202 Difference calculation unit 203 Presentation unit 204 Video generation unit 300 Observation device 300a, 300b Imaging device 301 Memory unit 302 Input interface 303 Position input unit 304 Data acquisition unit 305 Metadata acquisition unit 306 Transmitter 500 Image Display System

Claims

1. A video display system for displaying a display video by a display device, comprising: a photographing unit that generates a wide-angle video; a data acquisition unit that acquires data regarding at least one of the position and direction of a fixation target for causing a user of the display device to fixate within the wide-angle video, and queue information for notifying a change in the state of an observation system; a metadata configuration unit that sets the data from the data acquisition unit as metadata together with other information; and a transmission unit that transmits the wide-angle video together with the metadata; an observation device having the above components; a reception unit that receives the wide-angle video, the data, and the queue information; a display state estimation unit that estimates at least one of the position and direction of the display device within the wide-angle video; a difference calculation unit that calculates at least one of a relative position, which is the position of the fixation target relative to at least one of the position and direction of the display device within the wide-angle video of the display device, and a relative direction, which is the direction of the fixation target relative to the at least one of the position and direction of the display device within the wide-angle video of the display device, based on a difference between at least one of the estimated position and direction of the display device within the wide-angle video of the display device and at least one of the position and direction of the fixation target on the metadata; a presentation unit that presents information of at least one of the calculated relative position and relative direction, an instruction based on the queue information, and the state of the observation system to the user of the display device; a video generation unit that generates the display video including information of at least one of the position and direction of the display device within the wide-angle video of the display device estimated by the display state estimation unit from the received wide-angle video, an instruction based on the queue information, and a partial image corresponding to a visual field portion according to the state of the observation system; and a VR device having the display device that displays the display video. A video display system.

2. Further comprising a camera that photographs a video, or an image generation unit that generates an image by calculation, wherein the wide-angle video is a video photographed by the camera or an image calculated by the image generation unit. The video display system according to claim 1.

3. The presentation unit generates and outputs graphics indicating information based on at least one of the calculated relative position and relative direction and information based on the queue information. By superimposing the output graphics on a part of the image, the video generation unit is caused to present at least one of the relative position and the relative direction. The video display system according to claim 1.

4. The data acquisition unit receives input of data regarding the direction of the object of fixation. The display state estimation unit estimates the direction within the wide-angle field-of-view video of the display device. The graphics cause an arrow indicating the relative direction to be displayed on the display video. The video display system according to claim 3.

5. The data acquisition unit receives input of data regarding the direction of the object of fixation. The display state estimation unit estimates the direction within the wide-angle field-of-view video of the display device. The graphics cause a mask, which is an image for covering at least a part other than the relative direction side on the display video, to be displayed. The video display system according to claim 3.

6. The data acquisition unit receives input of data regarding the position of the object of fixation. The display state estimation unit estimates the position within the wide-angle field-of-view video of the display device. The graphics cause a map indicating the relative position to be displayed on the display video. The video display system according to claim 3.

7. Furthermore, it includes an input interface for use in inputting the data, and the data acquisition unit acquires the data input via the input interface. The video display system according to claim 1.

8. Furthermore, it includes an input interface for designating at least one of the start and end timings of the user's movement within the wide-angle field-of-view video, and the data acquisition unit acquires at least one of the start and end timings of the movement input via the input interface. The video display system according to claim 3.

9. The image constituting the wide-angle field-of-view video is an image output by an imaging unit that captures the real space, and the input interface includes an instruction marker held by an operator of the input interface in the real space, which indicates at least one of the position and direction of the object of fixation by the movement of the instruction marker, and an image analysis unit that receives at least one of the position and direction of the object of fixation indicated by the instruction marker by analyzing an image including the instruction marker output by the imaging unit. The video display system according to claim 7.

10. Comprising at least a part of the functions provided in the observation device and the VR device, connected to the observation device and the VR device via a network, and comprising an information processing device that undertakes a part of the processing of the observation device or the VR device The video display system according to claim 1.

11. The information processing device A receiving unit that receives the wide-angle video, the data, and the cue information from the observation device as the metadata; A presentation unit that generates information for presenting at least one of the position and direction of the fixation target on the metadata and information according to the cue information to the user of the display device; A video generation unit that adds the information generated by the presentation unit to a partial image corresponding to a visual field portion corresponding to at least one of the position and direction in the wide-angle video of the display device estimated by the display state estimation unit from the received wide-angle video to generate the display video; And a transmission unit that transmits the wide-angle video, a partial image corresponding to a visual field portion corresponding to at least one of the relative position and the relative direction, and the metadata. The video display system according to claim 10.

12. The information processing device A receiving unit that receives the wide-angle video, the data, and the cue information from the observation device as the metadata; A presentation unit that generates information for presenting at least one of the position and direction of the fixation target on the metadata and information according to the cue information to the user of the display device; A metadata configuration unit that generates the metadata from the information generated by the presentation unit; And a transmission unit that transmits the metadata generated by the metadata configuration unit, the wide-angle video received by the receiving unit, and other information to the VR device. The video display system according to claim 10.

13. The information processing device Receives the wide-angle video, the data, and the cue information from the observation device as the metadata, and receives data regarding the orientation of the display device from the display device A receiving unit; A difference calculation unit that calculates a relative movement direction, which is the movement direction of the imaging unit relative to the orientation of the display device, based on the difference between the orientation of the display device and movement information regarding the movement of the imaging unit and the cue information Graphics indicating the calculated relative movement direction, which is superimposed on a part of the video corresponding to the field-of-view portion according to the estimated orientation of the display device among the wide-angle videos, so as to present information according to the relative movement direction and the cue information to the user of the display device. A presentation unit that generates and outputs graphics; A video generation unit that corrects the graphics based on data regarding the orientation of the display device and superimposes the graphics on the wide-angle video to generate the display video; A transmission unit that transmits the display video and other information, and is provided on a cloud connected to a wide area network and is connected to the observation device and the VR device via the wide area network. The video display system according to claim 10; The video display system according to claim 10;

14. The information processing device is The cue information is Information indicating that at least one of the movement direction of the observation device or the position and direction of the fixation target that causes the user of the display device to fixate changes. The video display system according to any one of claims 1 to 14;

15. The information processing device is provided on a cloud connected to a wide area network and is connected to the observation device and the VR device via the wide area network. The video display system according to claim 10;

15. The cue information is Information indicating that at least one of the movement direction of the observation device or the position and direction of the fixation target that causes the user of the display device to fixate changes. The video display system according to any one of claims 1 to 14; The video display system according to any one of claims 1 to 14;

Citation Information

Patent Citations

  • Head-mounted display and brightness adjustment method

    JP2016090773A

  • Information processing unit and information processing system

    JP2019197939A

  • Method and apparatus for controlling the aiming direction of separated cameras

    JP2019503612A

  • SYSTEM AND METHOD FOR PRESENTING CONTENT - Patent application

    JP2019521547A

  • Image processing device, image processing method, and computer-readable recording medium

    WO2015170461A1