Image area generation system and program, image area display device

JP7917115B2Active Publication Date: 2026-09-08EIZO SYST CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024168119
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-06-24
Filing Date
2024-09-27
Publication Date
2026-09-08
Estimated Expiration
2043-06-20

AI Technical Summary

Benefits of technology

【0018】 上述した構成からなる本発明によれば、空間内に観客が入ることにより、各面に表示されている画像領域を鑑賞することができる。この画像領域は、元々全方向動画を6つの面に切り出したものであることから、この空間内の観客は、各面に表示されている画像領域を視認することで、あたかも全方向動画の中心に立っている感覚を楽しむことができる。観客は、各面を視認すると、その視認した面に表示されている画像領域が目に入る。即ち、視認した方向に応じた画像領域が目に入ることから、VRと同様の感覚を得ることができる。しかも、VRを体験する際に必要となる眼鏡型又はゴーグル型の頭部装着型映像表示装置を装着せずに、空間内であたかも観客自身がその場に現実に居るような臨場感を体感することができる。従って、頭部装着型映像表示装置の装着に伴う圧迫感や煩わしさ、装着の手間を無くすことができる。また頭部装着型映像表示装置を介して視覚から得る情報と現実の体が受け取る情報のズレから生じる、いわゆるVR酔い等のような身体に与える影響が及ぶことも無くなる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007917115000001
    Figure 0007917115000001
  • Figure 0007917115000002
    Figure 0007917115000002
  • Figure 0007917115000003
    Figure 0007917115000003
Patent Text Reader

Abstract

To enable a moving image imaged by an omnidirectional imaging device to be transmitted into a virtual space at high speed and displayed with a sense of presence.SOLUTION: Each still image that makes up an acquired moving image is cut into a plurality of image regions in accordance with positional relationships of each surface. The cut image regions are assigned to the respective surfaces. Data including the assigned image regions are transmitted on different channels than one another to each display device that is for displaying an image region on each surface. Here, an adjustment is performed for achieving chronological synchronization among the image regions of the transmitted data.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image region generation system and program, and an image region display apparatus for generating image regions to be displayed on each rectangular surface surrounding a space. [Background Art]

[0002] In recent years, services that allow an operation character (avatar) operated by a user to freely move in a three-dimensional virtual space constructed online have become increasingly widespread. With these services, users can also engage in entertainment activities such as gaming and sightseeing, as well as economic activities such as buying and selling goods and services, making it possible to carry out various activities using the virtual space as a living space. In particular, along with improvements in technologies for VR (Virtual Reality) and AR (Augmented Reality), users are now able to experience virtual spaces that are closer to reality, and it is expected that demand for related services will increase sharply in the future.

[0003] When using services utilizing a virtual space, a user wears a glasses-type or goggle-type head-mounted video display device on the head. Such a head-mounted video display device incorporates a motion sensor, a microphone, and the like, and can freely switch displayed video according to the user's movements and voice. This allows the user to freely move the avatar in the virtual space and freely change the line of sight via the head-mounted video display device, and enjoy various services.

[0004] However, the conventional virtual space services described above required users to wear a head-mounted video display device each time, and there was a need to address the discomfort, inconvenience, and hassle associated with wearing such a device. There were also problems related to the physical effects, such as so-called VR sickness, that arise from the discrepancy between the information received visually through the head-mounted video display device and the information received by the real body. In addition to this, there was a growing demand for a way for multiple people to simultaneously share and view omnidirectional video within a single virtual space, rather than having each user individually wear a head-mounted video display device and view the video independently.

[0005] Conventionally, methods have been proposed to display omnidirectional images in a virtual space without wearing a head-mounted image display device, as shown in Patent Document 1, for example. However, there is no mention of how to actually transmit various video content and video images captured by an omnidirectional imaging device at high speed into the virtual space, or how to display them in a way that gives a sense of presence. Furthermore, although each surface constituting the available virtual space is composed of various aspect ratios, the technology to flexibly create image areas that give a sense of presence according to the shape of the virtual space composed of surfaces with such various aspect ratios has not yet been proposed. [Prior art documents] [Patent Documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2021-177587 [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] Therefore, the present invention was devised in view of the above-mentioned problems, and its objective is to provide an image area generation system, program, and image area display device that, when generating image areas to be displayed on each rectangular surface surrounding a space, allows the user to experience a sense of presence as if they were actually there in the virtual space without having the user wear a head-mounted video display device each time, and allows multiple people to simultaneously share and view omnidirectional images in a single virtual space, and further enables the high-speed transmission of moving images captured by an omnidirectional imaging device into the virtual space for display with a sense of presence. The present invention also aims to provide an image area generation system, program, and image area display device that can flexibly create image areas with a sense of presence according to the shape of a virtual space composed of surfaces of various size ratios. [Means for solving the problem]

[0008] To solve the above-mentioned problems, the present inventors have invented an image region generation system and program, and an image region display device, which extracts each still image constituting an acquired moving image into multiple image regions according to the arrangement relationship of each surface, assigns each extracted image region to each surface, and transmits data containing each assigned image region to each display device for displaying the image region on each surface via different channels, while performing adjustments to synchronize the data transmitted between the image regions in a time series.

[0009] The first invention relates to an image region generation system that generates an image region to be displayed on each rectangular surface surrounding a space, comprising: a moving image and the moving image This includes corresponding audio information, and one or more of the following: distribution contract information, distribution equipment information, customer information regarding distribution audiences, management information for each video or audio information, distribution request information, ranking information, and past playback history information. A means for acquiring video and image data to obtain distribution destination information, and installed in the above space Omnidirectional imaging device capable of simultaneously capturing images in all directions. by at the same time Based on the images of each captured surface, a closed space consisting of six surfaces surrounds the space. Of the four walls and the surfaces corresponding to the ceiling and floor, A means for determining the aspect ratio of the length and width between each adjacent face, Each of the components of the video acquired by the video acquisition means described above Regarding still images, The arrangement of each of the above surfaces,Based on the aspect ratio of the vertical and horizontal dimensions between each of the faces determined by the above determination means, and the distribution destination information obtained by the above moving image acquisition means, a plurality of image regions to be assigned to each face so that the displayed images are continuous between adjacent faces are determined. The image is extracted, and time-series identification information is sequentially assigned to each extracted image region. Image region extraction means and each image region extracted by the image region extraction means Assignment to each of the above surfaces , an assignment means for assigning the above-mentioned acoustic information based on each of the above-mentioned assigned image regions, Each of the above surfaces includes a data transmission means that transmits data containing at least one of the image areas or sound information assigned by the assignment means to each playback device, which includes at least one of the display devices for playing back the image areas or sound devices for playing back the sound information, on different channels in a synchronized manner so that the display timing of each image area and the playback of the sound information coincide; the image area extraction means adjusts the shape of the boundary between the multiple image areas to be extracted to expand or compress based on the aspect ratio of the length and width between the above surfaces, so that each extracted image area becomes rectangular in shape to match the shape of each surface; and the assignment means, in the adjustment for synchronization, if it determines based on the time-series identification information that a particular image area is missing and therefore out of sync, it interpolates pixels in the missing image area based on the preceding and succeeding image areas, or inserts either of the preceding or succeeding image areas to generate the missing image area, and performs adjustments to maintain time-series synchronization between each image area and the sound information. It is characterized by the following.

[0013] Second Invention The image region generation system related to this is First Invention In this configuration, the image region extraction means is characterized by extracting each still image constituting the moving image into multiple image regions based on the acoustic information, such that the displayed images are continuous between adjacent surfaces according to the arrangement relationship of each surface.

[0014] Third Invention The image region generation system related to this is First Invention In this system, the data transmission means is characterized by adjusting the display timing of each image region and the playback of the audio information in the transmitted data so that they coincide on different channels.

[0016] Fourth Invention The image region generation program relates to an image region display space in which an image region is reproduced on each rectangular surface surrounding the space, and the rectangular surface surrounding the space, and the moving image and the moving image This includes corresponding audio information, and one or more of the following: distribution contract information, distribution equipment information, customer information regarding distribution audiences, management information for each video or audio information, distribution request information, ranking information, and past playback history information. A means for acquiring video and image data to obtain distribution destination information, and installed in the above space Omnidirectional imaging device capable of simultaneously capturing images in all directions. by at the same time Based on the images of each captured surface, a closed space consisting of six surfaces surrounds the space. Of the four walls and the surfaces corresponding to the ceiling and floor, A means for determining the aspect ratio of the length and width between each adjacent face, Each of the components of the video acquired by the video acquisition means described above Regarding still images, The arrangement of each of the above surfaces,Based on the vertical and horizontal size ratio between each of the surfaces determined by the determining means and the distribution destination information acquired by the moving image acquiring means, into the plurality of image regions to be allocated to each adjacent surface such that display images are continuous between the adjacent surfaces The image is extracted, and time-series identification information is sequentially assigned to each extracted image region. image region clipping means for clipping each image region clipped by said image region clipping means Assignment to each of the above surfaces , allocating means for allocating said acoustic information based on each of the allocated image regions, and Each of the above surfaces includes a data transmission means that transmits data containing at least one of the image areas or sound information assigned by the assignment means to each playback device, which includes at least one of the display devices for playing back the image areas or sound devices for playing back the sound information, on different channels in a synchronized manner so that the display timing of each image area and the playback of the sound information coincide; the image area extraction means adjusts the shape of the boundary between the multiple image areas to be extracted to expand or compress based on the aspect ratio of the length and width between the above surfaces, so that each extracted image area becomes rectangular in shape to match the shape of each surface; and the assignment means, in the adjustment for synchronization, if it determines based on the time-series identification information that a particular image area is missing and therefore out of sync, it interpolates pixels in the missing image area based on the preceding and succeeding image areas, or inserts either of the preceding or succeeding image areas to generate the missing image area, and performs adjustments to maintain time-series synchronization between each image area and the sound information. characterized by:

[0017] Invention 5 An image region generation program according to , which is an image region generation program for generating image regions to be displayed on respective rectangular surfaces surrounding a space, wherein: On the computer, a moving image acquiring step of acquiring a moving image and said moving image This includes corresponding audio information, and one or more of the following: distribution contract information, distribution equipment information, customer information regarding distribution audiences, management information for each video or audio information, distribution request information, ranking information, and past playback history information. distribution destination information; and a determining step of determining, based on images of respective surfaces captured by Omnidirectional imaging device capable of simultaneously capturing images in all directions. disposed in said space at the same time , a vertical and horizontal size ratio between respective adjacent surfaces of at least two or more of the six surfaces forming a closed space surrounding the space Of the four walls and the surfaces corresponding to the ceiling and floor, ; an image region clipping step of clipping, for a still image, Each of the components of the video acquired by the video acquisition step described above , based on the vertical and horizontal size ratio between each of the surfaces determined in the determining step and the distribution destination information acquired in the moving image acquiring step, The arrangement of each of the above surfaces, into the plurality of image regions to be allocated to each adjacent surface such that display images are continuous between the adjacent surfaces The image is extracted, and time-series identification information is sequentially assigned to each extracted image region. ; and an allocating step of allocating each image region clipped in the image region clipping step to each of said surfaces, and allocating said acoustic information based on each of the allocated image regions, The above-mentioned playback device includes at least one of the display devices for reproducing image regions on each of the above-mentioned surfaces, or at least one of the sound devices for reproducing the above-mentioned sound information, and the above-mentioned data transmission step is performed to transmit data including at least one of the image regions or sound information assigned by the above-mentioned assignment step to each playback device, in a synchronized manner on different channels so that the display timing of each image region and the playback of the sound information match; the above-mentioned image region extraction step adjusts the shape of the boundary between the multiple image regions to be extracted to expand or compress based on the aspect ratio of the length and width between the above-mentioned surfaces, and adjusts each extracted image region to become rectangular in shape to match the shape of each of the above-mentioned surfaces; the above-mentioned assignment step, in the adjustment for synchronization, if it is determined that a particular image region is missing and out of synchronization based on the time-series identification information, it newly generates the missing image region by interpolating pixels based on the image regions before and after it, or by inserting either of the image regions before or after it, and performs adjustments to maintain time-series synchronization between each image region and the sound information. characterized by: Effects of the Invention

[0018] According to the present invention having the above configuration, when an audience member enters the space, they can appreciate the image areas displayed on each surface. Since these image areas are originally obtained by cutting an omnidirectional video into six surfaces, the audience member in the space can enjoy the sensation of standing at the center of the omnidirectional video by visually recognizing the image areas displayed on each surface. When the audience member visually recognizes each surface, the image area displayed on the recognized surface comes into their view. That is, since the image area corresponding to the viewing direction comes into view, they can obtain the same sensation as VR. Moreover, without wearing glasses-type or goggles-type head-mounted video display devices required when experiencing VR, the audience member can experience the presence as if they were actually present at the site in the space. Therefore, the feeling of pressure, inconvenience, and the trouble of wearing caused by wearing a head-mounted video display device can be eliminated. In addition, adverse effects on the body such as the so-called VR sickness, which occurs due to the discrepancy between information obtained visually via the head-mounted video display device and information received by the actual body, is also eliminated.

[0019] Furthermore, according to the present invention, a plurality of audience members can enter the space at the same time and visually recognize the common image area, so that simultaneous sharing of omnidirectional video by multiple people in one virtual space, which cannot be realized by conventional VR, can be achieved.

[0020] Also according to the present invention, each image area can be independently transmitted to the display device through different communication paths, making it possible to provide content to the space at high speed and low cost. Moreover, time-series inconsistencies between the respective image areas, which may occur when each image area is independently transmitted through different communication paths, can be eliminated through synchronization adjustment processing.

[0021] Furthermore, according to the present invention, at least one of the video footage from live video and archived video, audio information corresponding to the video footage, and destination information for distributing the video footage and audio information are acquired. Based on the destination information, the characteristics of each image region extracted into multiple image regions and the spatial characteristics can be determined, and audio information can be assigned to each surface. As a result, data containing at least one of each image region and audio information can be transmitted to each playback device for playing the image region and audio information on each surface via different channels, enabling simultaneous sharing of omnidirectional video footage and audio information by multiple people within a single virtual space.

[0022] Furthermore, according to the present invention, the characteristics of each image region and audience information consisting of one or more of the audience's position, gaze, head orientation, and sounds emitted by the audience are extracted. Therefore, it is possible to interactively assign acoustic information to each surface based on the audience's state in relation to the live video. As a result, data containing at least one of each image region and acoustic information can be transmitted to each playback device for playing image regions and acoustic information on each surface via different channels, enabling simultaneous sharing of omnidirectional video and acoustic information among multiple people in a single virtual space. [Brief explanation of the drawing]

[0023] [Figure 1] Figure 1 is a diagram showing the overall configuration of an image region generation system to which the present invention is applied. [Figure 2] Figure 2 is a perspective view of a space enclosed by six rectangular faces. [Figure 3] Figure 3 shows an example of projecting an image onto a common surface using multiple display devices. [Figure 4] Figure 4 is a detailed block diagram of the control device. [Figure 5] Figure 5 is a flowchart illustrating the operation of each step in the image region generation system. [Figure 6] Figure 6 shows an example of one of the still images that make up an omnidirectional video being plotted on a rectangular plane. [Figure 7] Figure 7 shows examples of each spherical image region that make up an omnidirectional video captured by an omnidirectional imaging device. [Figure 8] Figure 8 shows an example of dividing the still images that make up an omnidirectional video into multiple image regions. [Figure 9] Figure 9 shows an example of adjusting the image area assigned to each face to be rectangular. [Figure 10] Figure 10 shows an image of data from each image region being transmitted sequentially to the display device over time. [Figure 11] Figure 11 shows an example in which each image region extracted from an omnidirectional video is displayed on each surface via each display device. [Figure 12] Figure 12 shows an example of a method for cropping an image area when the display device consists of a projection display device that projects an image area onto a surface. [Modes for carrying out the invention]

[0024] Figure 1 shows an overall configuration diagram of an image region generation system 1 to which the present invention is applied. This image region generation system 1 is centered on a control device 2, and includes a recording module 3 connected thereto. Furthermore, it includes a display device 7 that displays video and an audio device 8 that reproduces sound, which are connected to the control device 2 via a communication network 5, and a moving image storage unit 9 that stores various video content and audio files. In addition, this image region generation system 1 may also include a space 6 in which the display device 7 and audio device 8 are installed. The space 6 may be installed, for example, at multiple locations, each individually or in coordination with others.

[0025] The control device 2 acts as a central control device, controlling the entire image region generation system 1. This control device 2 is embodied, for example, as a personal computer (PC), but is not limited to this; it may also be embodied as a server or dedicated equipment, or as a portable information terminal or tablet terminal.

[0026] The recording module 3 is used to pre-record alternative video footage based on past events, separate from actual events, and includes an omnidirectional imaging device 31 and a microphone 32.

[0027] The omnidirectional imaging device 31 is configured to capture images simultaneously and comprehensively in all directions (360° horizontally and 360° vertically) centered on the imaging device itself. By recording video with this omnidirectional imaging device 31, it is possible to capture video in all directions (hereinafter referred to as omnidirectional video) simultaneously and comprehensively. For example, when imaging an urban space, if vehicles or people move, the moving vehicles or people can be recorded as video images in chronological order across all directions. Furthermore, while the omnidirectional imaging device 31 may be fixed in one place to continuously capture omnidirectional video, the omnidirectional imaging device 31 itself may also be mounted on a moving object such as an unmanned aerial vehicle, vehicle, or helicopter to continuously record omnidirectional video.

[0028] This makes it possible to obtain moving images as if one were riding in such a mobile device and viewing all directions. The omnidirectional video captured by the omnidirectional imaging device 31 is output to the control device 2. The omnidirectional imaging device 31 may be connected not only directly to the control device 2, but may also be connected via a communication network (not shown) consisting of the Internet or a LAN (Local Area Network).

[0029] Microphone 32 collects ambient sound and converts it into an audio signal. Microphone 32 transmits this converted audio signal to the control device 2 via an interface. Microphone 32 is necessary for realizing live video playback, but it is not a particularly essential component and may be omitted.

[0030] The communication network 5 is an internet network or the like, connecting the control device 2, display device 7, and video device 8 via communication lines. Incidentally, if the recording module 3, control device 2, display device 7, and sound device 8 are operated within a certain narrow area, the communication network 5 may be configured as a LAN. This communication network 5 is not limited to a wired communication network, but may also be implemented as a wireless communication network.

[0031] As shown in Figure 2, space 6 is composed of a space enclosed by six rectangular surfaces 61a to 61f. This space 6 consists of surfaces 61a to 61f corresponding to four walls, a ceiling, and a floor, similar to a room. In this case, space 6 may be provided with a door (not shown) that allows people to enter and exit the interior. Furthermore, space 6 is not limited to being a completely closed space enclosed by all six surfaces 61a to 61f; it may be an open space with one or more surfaces 61 omitted, or an open space with only a part of one or more surfaces 6 open. In addition, the interior of space 6 may have various structures other than surfaces 61a to 61f, such as various shapes, irregularities, protrusions, and fixtures.

[0032] The image region generation system 1 reproduces the generated moving image and the corresponding audio information using a playback device. The playback device consists of, for example, a display device 7 and an audio device 8. The display device 7 is a projection display device that projects the image region onto a surface, like a so-called projector. This display device 7 is not limited to being a projection display device, but may also be a display for displaying the image region on a surface, such as a liquid crystal display, an organic EL display, or even an LED display.

[0033] If the display device 7 is equipped with, for example, a speaker that outputs voice or music, it may also function as a sound playback device. The display device 7 may be equipped with various speakers, or it may be linked to, for example, a separate sound device 8. The sound device 8, for example, plays back sound based on the three-dimensional direction, distance, and spread of the sound when recording and playing it back.

[0034] The sound device 8 reproduces sound as immersive, three-dimensional 3D sound by combining multiple elements that constitute the sound. These multiple elements include, for example, "volume difference" which reproduces the sound image localization of the sound source due to the attenuation of volume due to the distance between the sound device 8 and the object (person), or the distance in space, and the difference in intensity between the two ears; "time difference" which reproduces the sound image localization of the sound source due to the time difference in when sound waves reach the object; "change in frequency characteristics" which reproduces the sound image localization of the sound source due to changes in frequency characteristics due to the transmission and shielding of sound waves; "change in phase" which reproduces the sound image localization of the sound source due to changes in phase due to the transmission and shielding of sound waves; and "change in reverberation" which reproduces the sound field of the surrounding environment due to reverberation characteristics.

[0035] The sound device 8 performs 3D sound rendering processing to control the sound field in the three-dimensional space 6, based on multiple elements such as the type of video, such as live video and archived video; the characteristics of each image region as shooting information; and audience information such as the position, gaze, head orientation, and sounds emitted by audience members M in the space. In rendering 3D sound, the sound device 8 performs processing using various known processing and techniques, such as known "feature prediction techniques," "ray method / geometric sound modeling techniques," and "adaptive rectangular decomposition," for example, using known feature prediction methods and ray tracing methods.

[0036] The display device 7 and the sound device 8 function as playback devices that reproduce moving images and sound information, respectively. The control device 2 displays the image area generated by the control device 2 on each of the surfaces 61a to 61f that constitute the space 6 via the display device 7, as shown in Figure 2. In the example in Figure 2, the explanation will take the case where the display devices 7a to 7f are configured as projection display devices such as projectors, and the display device 7g is configured as an LED display. The explanation will also take the case where the sound device 8 is configured as multiple speaker units that reproduce 3D sound or spatial sound, and is installed on the back of each surface of the space 6 (for example, the back of surface 61b). The sound device 8 may also be configured to reproduce sound in conjunction with speakers, etc., if the display devices 7a to 7f are equipped with speakers, etc.

[0037] Display device 7a is mounted near the upper end of surface 61a and projects an image onto surface 61c facing surface 61a. Display device 7b is mounted near the upper end of surface 61b and projects an image onto surface 61d facing surface 61b. Display device 7c is mounted in the middle of surface 61b, and display device 7e is mounted in the middle of surface 61d, and they both project an image onto surface 61e, which is common to both. Display device 7d is mounted near the upper end of surface 61d and projects an image onto surface 61b facing surface 61d. Display device 7f is mounted near the upper end of surface 61c and projects an image onto surface 61a facing surface 61c. Display device 7g, which consists of an LED display, displays an image on surface 61f.

[0038] Furthermore, any combination of display devices 7 used to display an image on any surface 61 constituting space 6 is included, other than the examples described above. Not only can an image be displayed on each surface 61 using one display device 7, but multiple display devices 7 may be used in combination to display the image. In the example in Figure 3, an example is shown in which display devices 7c and 7e project an image onto a common surface 61e. That is, one surface 61e is divided in half, and an image is projected onto one half using display device 7c, and onto the other half using display device 7e. Other surfaces 61 may also be divided in a similar manner, and images may be projected onto them by distributing them among multiple display devices 7. In this case, the sound device 8 may identify the regions of each divided surface and reproduce sound information according to the images displayed in each region.

[0039] The sound device 8 also plays back acoustic information corresponding to the acquired video image. The sound device 8 may, for example, play back acoustic information of the identified features according to the results of the feature identification of each extracted image region. This allows the sound device 8 to play back data containing acoustic information assigned to each playback device and transmitted on different channels as 3D sound.

[0040] The sound device 8, for example, is mounted on the back of surface 61b and reproduces sound as 3D sound for the space 6 enclosed by the six rectangular surfaces 61a to 61f. The sound device 8 may also be configured such that multiple sound devices 8 are installed on the six surfaces 61a to 61f (not shown). This makes it possible to reproduce the three-dimensional direction, distance, and spread of sound in accordance with the moving image displayed in the space 6 enclosed by the six surfaces 61a to 61f.

[0041] The video storage unit 9 is a database for storing at least one of the video images to be displayed via the display device 7, including live video and archived video, and the associated audio information. This video storage unit 9 pre-stores omnidirectional video images, including audio information, that have already been captured by other imaging devices (not shown) other than the omnidirectional imaging device 31 described above. The various video images, including audio information, stored in this video storage unit 9 may be not only the omnidirectional video and audio information described above, but also ordinary two-dimensional video and audio information. The omnidirectional video images stored in this video storage unit 9 are sent to the control device 2 via the communication network 5.

[0042] Next, the detailed block configuration of the control device 2 will be described. As shown in Figure 4, the control device 2 comprises a first video acquisition unit 21, a second video acquisition unit 23, a spatial information acquisition unit 26, an audio data acquisition unit 35, and an operation unit 25. Furthermore, it comprises a control unit 28 to which the first video acquisition unit 21, the second video acquisition unit 23, the spatial information acquisition unit 26, the audio data acquisition unit 35, and the operation unit 25 are connected. In addition, the control unit 28 is connected to I / F (interfaces) 29-1, 29-2, 92-3, ...29-n for transmitting data for each output image region P1, P2, P3, ...Pn. Furthermore, the control unit 28 is connected to I / F 30-1 for transmitting data for the output acoustic information S1. The acoustic information S1 may be configured to be connected to, for example, a plurality of display devices 7a, ..., ~7n and a plurality of acoustic devices 8 (not shown).

[0043] Since the control device 2 is composed of a PC or the like, in addition to these components, it also has a CPU (Central Processing Unit) as a so-called central processing unit for controlling each component, a ROM (Read Only Memory) that stores programs for controlling the hardware resources of the entire control device 2, and a RAM (Random Access Memory) used as a working area for data storage and deployment, as well as an image processing unit for performing various image processing on omnidirectional video and for processing to extract each image region P1 to Pn.

[0044] The first video acquisition unit 21 acquires omnidirectional videos stored in the video storage unit 9 via the communication network 5. The first video acquisition unit 21 may also acquire videos stored in the video storage unit 9 as archived videos, for example. Archived videos may be past videos stored on each video server as a publicly known video provision service on the web or cloud. Videos may include, for example, sound, background at the time of shooting, ambient sounds, music data added by the photographer, or acoustic information (2D / 3D acoustic information, sound source information, acoustic equipment information, sound effect information, setting values, parameters, etc.), and various types of music data and information may be associated with the acoustic information individually or commonly.

[0045] The second video acquisition unit 23 acquires omnidirectional video captured by the omnidirectional imaging device 31. The second video acquisition unit 23 may acquire omnidirectional video from various locations in real time as live video using omnidirectional imaging devices 31 installed in various locations (e.g., fixed cameras, stationary cameras, etc.). The second video acquisition unit 23 may acquire video from inside the space 6 (such as images of people inside or motion sensor information showing the position and movement of people's body parts) using, for example, an omnidirectional imaging device 31 installed in the space 6, or other imaging devices or sensors, or sensors held by spectators (not shown).

[0046] The omnidirectional imaging devices 31 may be installed in multiple spaces 6, for example, and may acquire various types of information as audience information, such as the position, gaze, head orientation, and sounds emitted by people (audience members M) inside, either individually or collectively.

[0047] The second video acquisition unit 23 may also acquire various information and data, such as location information of the place where the omnidirectional imaging device 31 is installed, surrounding environment information, date and time of imaging, and weather. The video acquired by the second video acquisition unit 23 may be stored in the video storage unit 9 via the communication network 5 by the control unit 28, for example.

[0048] The spatial information acquisition unit 26 acquires the space 6 on which the image regions P1 to Pn are actually displayed, as well as various information related to space 6. This spatial information acquisition unit 26 acquires various information related to the shape of space 6, such as the aspect ratio between the length and width of each surface 61 of space 6. The information acquired by this spatial information acquisition unit 26 includes information regarding the arrangement of display devices 7 installed on each surface 61 of space 6, information regarding which display devices 7 will be used to display images on each surface 61, not just when an image is displayed on each surface 61 using one display device 7 as described above, and information regarding the allocation of multiple display devices 7 when multiple display devices 7 are used to display images on each surface 61.

[0049] Furthermore, the spatial information acquisition unit 26 acquires various information about the space 6 through which the acoustic information S1 is actually transmitted. This spatial information acquisition unit 26 acquires various information about the shape, material, and reflectors of the space 6, such as the aspect ratio between the length and width of each surface 61 of the space 6. The information acquired by this spatial information acquisition unit 26 includes information about the arrangement of the acoustic devices 8 installed on each surface 61 of the space 6, information about which acoustic devices 8 will transmit acoustic information to the audience of the space 6, not only when acoustic information is transmitted to each surface 61 by a single acoustic device 8 as described above, the individual acoustic modules (acoustic units) that make up the acoustic device 8, and, in the case when it is transmitted to each surface 61 in combination with multiple display devices 7, various information about their assignment, timing, and directivity.

[0050] If the display device 7 is configured as a playback device combined with a projection display device and an audio device 8, the spatial information acquisition unit 26 may also acquire information such as the arrangement of these devices, or the projection direction and field of view of each projection display device and playback device relative to the surface 61. The spatial information acquisition unit 26 transmits the acquired information regarding the space 6 to the control unit 28.

[0051] Furthermore, if the spatial information acquisition unit 26 is configured as a playback device in combination with one or more sound devices 8 and display devices 7, it may also acquire information such as their arrangement, the directivity of the sound devices 8 relative to the space 6, surface 61, and audience, and the intensity of the 3D sound. The spatial information acquisition unit 26 transmits the acquired information regarding the space 6, surface 61, audience, etc., to the control unit 28. The information regarding the audience, etc., may be pre-set, assumed audience information transmitted to the control unit 28.

[0052] The audio data acquisition unit 35 acquires audio from the microphone 32, etc., and stores it. The method for acquiring audio from the microphone 32 may be, for example, acquisition via wired or wireless connection from a public communication network, or reading audio data recorded on a recording medium and recording it. In addition to audio, the audio data acquisition unit 35 may also acquire various types of music, background music, or sound information and acoustic data from the location where the microphone 32 is installed. The audio data acquisition unit 35 may also acquire audio data from multiple sound sources via multiple microphones 32, etc.

[0053] The audio data acquisition unit 35 may acquire moving images of the inside of space 6 (such as images of people inside or motion sensor information indicating the position and movement of people's body parts) using, for example, a microphone provided on an omnidirectional imaging device 31 installed in space 6, or another microphone 32, a microphone held by an audience member, etc. (not shown). Multiple omnidirectional imaging devices 31 may be installed in multiple spaces 6, and various types of information such as the position, gaze, head orientation, and sounds emitted by people (audience members M) inside may be acquired individually or collectively as audience information.

[0054] The operation unit 25 is implemented via a keyboard or touch panel, and the user inputs execution commands to run the program. When the user inputs an execution command, the operation unit 25 notifies the control unit 28. Upon receiving this notification, the control unit 28 works in coordination with the decision unit 27 and other components to execute the desired processing operation.

[0055] The control unit 28 is a so-called central control unit for controlling each component implemented in the control device 2 by transmitting control signals via an internal bus. The control unit 28 also transmits various control commands via the internal bus in response to operations performed via the operation unit 25. The control unit 28 receives input of various data, including video and audio information, from the first video acquisition unit 21 and the second video acquisition unit 23, respectively.

[0056] As described later, the control unit 28 extracts each still image constituting the received video from the input into multiple image regions P1, P2, ..., Pn. The video data including these extracted image regions P1, P2, ..., Pn is transmitted via I / F 29-1, 29-2, ..., 29-n on different channels from each other. Furthermore, as described later, the control unit 28 transmits each audio data constituting the received audio information via I / F 30-1 on a different channel from the data of different video.

[0057] Each of I / F29-1, 29-2, ..., 29-n, and I / F30-1 serves as an interface for establishing a communication link as a playback device between the control device 2, the display device 7, and the sound device 8. I / F29-1, 29-2, ..., 29-n are not limited to being individually provided for multiple image regions P1, P2, ..., Pn, and I / F30-1 for multiple sound information S1 extracted by the control unit 28, but may also be composed of a common interface unit.

[0058] Next, the operation of the image region generation system 1 to which the present invention, consisting of the above-described configuration, is applied will be explained.

[0059] Figure 5 is a flowchart showing each operation of the image region generation system 1. First, in step S11, the control device 2 acquires a moving image. The acquisition of moving images by the control device 2 is performed via the first moving image acquisition unit 21 and the second moving image acquisition unit 23 described above. That is, when an omnidirectional video stored in the moving image storage unit 9 is sent via the communication network 5, it is acquired via the first moving image acquisition unit 21. Also, when an omnidirectional video is captured via the omnidirectional imaging device 31, it is acquired via the second moving image acquisition unit 23.

[0060] In step S11, the control device 2 acquires not only the omnidirectional video but also audio information and distribution destination information related to the omnidirectional video. The control device 2 may acquire, for example, the audio information related to the video and the distribution destination information for the video and audio information individually, or collectively by including them in the omnidirectional video.

[0061] The control device 2 acquires omnidirectional video (moving images), audio information, and distribution destination information via the first moving image acquisition unit 21 and the second moving image acquisition unit 23 described above. For example, when omnidirectional video stored in the moving image storage unit 9 is sent via the communication network 5, the control device 2 acquires it via the first moving image acquisition unit 21, and when omnidirectional video is captured via the omnidirectional imaging device 31, it acquires it via the second moving image acquisition unit 23.

[0062] If the omnidirectional video and panoramic video acquired by the first video acquisition unit 21 and the second video acquisition unit 23 contain various types of information such as sound information and distribution destination information, this information may also be sent to the control unit 28. The distribution destination information may include, for example, various types of information related to the distribution of the panoramic video. The distribution destination information may include, for example, distribution-enabled contract information (distribution conditions, billing information, point information, etc.), distribution-enabled equipment information (space 6 information, projection equipment information, sound equipment information, lighting information, etc.), and distribution-enabled audience information (member information, gender, age, height, hobby information, group information, etc., and motion information indicating the gaze, posture, and movements of customers who have acquired the panoramic video in space 6 in real time).

[0063] The control unit 28 extracts the image region described below from each still image that constitutes the omnidirectional video sent from the first video acquisition unit 21 and the second video acquisition unit 23 (step S12).

[0064] Figure 6 shows one of the still images that make up the omnidirectional video, illustrated on a rectangular plane. Each still image that makes up the omnidirectional video captured by the omnidirectional imaging device 31 can be divided into spherical image regions Q1-a, Q1-b, Q2-a, Q2-b, Q3-a, Q3-b, Q4-a, Q4-b, Q5, and Q6, which form a sphere overall, as shown in Figure 7. The still images that make up the omnidirectional video shown in Figure 6 are obtained by redrawing these spherical image regions Q1-a, Q1-b, Q2-a, Q2-b, Q3-a, Q3-b, Q4-a, Q4-b, Q5, and Q6 in a planar manner.

[0065] The control unit 28 extracts image regions P1, P2, ..., Pn from each still image that constitutes such an omnidirectional video. In the example in Figure 6, six images, P1 to P6, are extracted. When the control unit 28 extracts image regions P1, P2, ..., Pn, for example, it may also extract various information including distribution destination information and sound information, and each still image that constitutes the video, into multiple image regions according to the arrangement relationship of each plane.

[0066] The control unit 28 further includes, for example, an extraction unit. The extraction unit, for example, identifies the features of the displayed moving image from each extracted image region and extracts audience information consisting of the features of each identified image region and one or more of the following: the real-time position of the audience in space 6, gaze direction, head orientation, and sounds emitted by the audience. The extraction unit acquires motion information indicating the characteristics of various types of information (audience information), such as the position of the audience in space 6, gaze direction, head orientation, and sounds emitted by audience M, using the second moving image acquisition unit 23 and other known sensors, identifies it through processing such as image discrimination and sound discrimination, and extracts motion information indicating the real-time gaze, posture, and movement of each audience member. The extraction unit links each image region (still image or moving image) in space 6 with the real-time motion information of the audience and stores it in the moving image storage unit 9.

[0067] Simultaneously with, or after the completion of, the cropping process of image regions P1 to P6, each cropped image region P1 to P6 is assigned to each surface 61a to 61f (step S13). Image region P1 is assigned to surface 61a as shown in Figure 2, image region P2 is assigned to surface 61b, image region P3 is assigned to surface 61c, image region P4 is assigned to surface 61d, image region P5 is assigned to surface 61e, and image region P6 is assigned to surface 61f. In other words, in this cropping example, each image region P is assigned to each surface 61. Similarly, if multiple image regions P are to be displayed in combination on a single surface 61, each of the multiple image regions P to be displayed is assigned to that single surface 61.

[0068] Furthermore, simultaneously with the assignment of each surface 61a to 61f of each image region P1 to P6, the acoustic information S1 is assigned to the acoustic device 8 shown in Figure 2. If, for example, speakers for each display device 7a to 7n and other sound playback devices (not shown) are set in addition to the acoustic device 8, the control unit 28 divides the acoustic information S1 in accordance with the image region P1 to P6 extraction process and assigns it to each of the multiple acoustic devices.

[0069] Furthermore, the control unit 28 may, for example, assign each image region to each surface based on information such as the characteristics of the extracted image region and sound information, and the characteristics of the space 6, etc. (size, material, number of viewers, characteristics, etc.). In addition, the control unit 28 may individually assign the entire sound information to be played in the space 6 (for a large audience) or a part of it (for specific people, children, adults, paid content, etc.) to the entire space 6 or to each surface.

[0070] The control unit 28 may, for example, cut out each still image constituting the moving image into multiple image regions according to the arrangement relationship of each surface, depending on the length and effect of the acquired sound information, and assign each cut-out image region P1 to P6 to each surface 61a to 61f. Furthermore, the control unit 28 may, for example, determine the directivity of the acquired sound information for each assigned surface 61a to 61f and assign the playback timing, playback pattern, sound effects, etc., of the sound information so that it can be played back in space 6. This makes it possible to reliably play back sound information along with the moving image to the audience M in space 6 with pinpoint accuracy.

[0071] Furthermore, the control unit 28 may, for example, cut out each still image that newly constitutes a moving image into multiple image regions according to the arrangement relationship of each surface, based on the characteristics of each image region extracted by the extraction unit and audience information (motion information) consisting of one or more of the audience's position, gaze, head orientation, and sounds emitted by the audience, and assign each cut-out image region P1 to P6 to each surface 61a to 61f. This makes it possible to interactively assign omnidirectional images and sound information to each surface based on the audience's actions in response to the live video distributed in space 6.

[0072] The still images that make up the omnidirectional video are thus divided into multiple image regions P without leaving any remnants. The shape of the boundaries between these image regions P is determined based on the aspect ratio of the length and width of each surface 61. Similarly, the acoustic information corresponding to the still images is also assigned based on the divided multiple image regions P.

[0073] For example, let's assume that the boundaries of image regions P1 to P6 shown in Figure 8(a) correspond to the aspect ratio of the planes 61a to 61f of a certain space 6. In this case, if another space 6 has a smaller area than space 6, and the aspect ratio of the planes 61a to 61d is different from that of space 6, then, for example, as shown in Figure 8(b), the boundaries of image regions P1 to P4 are expanded vertically, and the boundaries of image regions P5 and P6 are compressed vertically.

[0074] The acoustic information is adjusted, for example, according to the shape and movement of these adjusted image regions, and is then played back by the acoustic device 8. The acoustic information may also be played back through speakers, for example, if each of the display devices 7a to 7n is equipped with speakers. Furthermore, if there are multiple acoustic devices 8, speakers on each of the display devices 7a to 7n, and other acoustic playback devices (not shown) in space 6, the control unit 8 controls and plays back the acoustic information for each of the surfaces 61a to 61f.

[0075] In step S13, the image regions P1 to P6 assigned to each surface 61 may be adjusted so that each image region P1 to P6 is rectangular. In this case, as shown in Figure 9(a), taking image region P2 as an example, by applying image processing that stretches its upper and lower ends in the direction of the arrows in the direction of the dotted line in the figure, it is possible to obtain an image region P2 that has been processed into a rectangular shape as shown in Figure 9(b).

[0076] Next, the process moves to step S14, where the control unit 28 transmits the image regions P1 to Pn generated in this manner via the I / F 29 and acoustic information S1 on different channels. Here, "channel" refers to a communication line. That is, transmitting on different channels means that the data for the image regions P1 to Pn and the acoustic information S1 are transmitted separately on different communication lines. The data for the image regions P1 to Pn and the acoustic information S1, thus separated for each communication line, are then sent independently to the display device 7 and the sound device 8 (or the individual sound modules that make up the sound device 8), respectively.

[0077] In the example shown in Figure 4, image region P1 is transmitted independently to display device 7a, image region P2 is transmitted independently to display device 7b, image region P3 is transmitted independently to display device 7c, and image region Pn is transmitted independently to display device 7n. During this time, the data from each image region P1 to Pn is transmitted to the display device 7 via independent communication paths, without being collected at one location. Furthermore, acoustic information S1 is transmitted independently to the acoustic device 8. Note that each image region P1 to Pn transmitted independently to each of the display devices 7a to 7n may include, for example, acoustic information S1, or individually subdivided acoustic information that constitutes acoustic information S1.

[0078] The original omnidirectional video is composed of a large number of still images per second at a preset frame rate (24fps, 30fps, 60fps, etc.). Continuously transmitting image regions P1 to Pn, which are divided from such a large number of still images, in a time-series manner requires a considerable amount of communication. If such continuous transmission of image regions P1 to Pn were to be performed over a single communication path, it would require a considerable amount of communication time and the communication cost would be excessive. Therefore, in this invention, by transmitting image regions P1 to Pn to a display device 7 and an acoustic device 8 that transmits acoustic information S1, or a playback device composed of the display device 7 and the acoustic device 8, via different communication paths, the transmission rate of image region P in each communication path can be reduced, and as a result, the data of image region P can be transmitted to the display device 7 at high speed and at low cost.

[0079] Furthermore, in addition to using different communication paths, the image regions P1 to Pn and the acoustic information S1 may also be transmitted to the display device 7 and the acoustic device 8 or playback device using different frequency channels. By using different frequency channels to transmit the image regions P1 to Pn to the display device 7 and the acoustic device 8 or playback device, it becomes possible to transmit them with high communication quality without interference.

[0080] Next, the process moves to step S15, where adjustments are made to synchronize the time-series data between each image region of the data to be transmitted to the display device 7 and the sound device 8.

[0081] Figure 10 shows an image of data from each image region P1 to Pn being transmitted sequentially to the display device 7. The data from each image region P1 to Pn, extracted from the still images that make up the omnidirectional video, is transmitted sequentially to the display device 7. Similarly, the image regions P1 to Pn are extracted from the next still image that makes up the omnidirectional video and sent to the display device 7. When this is repeated, the data from each image region P1 to Pn is transmitted from the beginning of the frame along the time axis t, as shown in Figure 10.

[0082] In this manner, time-series identification information may be sequentially assigned to the data streams of each image region P1 to Pn transmitted according to the time axis t. This time-series identification information may be something like a timestamp, and may be assigned in accordance with the time when the image regions P1 to Pn are generated. Alternatively, the time-series identification information may correspond to the frame numbers that are sequentially assigned in chronological order to each still image constituting the omnidirectional video. In other words, the same time-series identification information corresponding to the same frame number may be assigned to image regions P1 to Pn extracted from the same still image frame.

[0083] As a result, image regions P1 to Pn, each assigned time-series identification information corresponding to the same frame number, can synchronize with each other in time, making it possible to display them without any discrepancies between them.

[0084] In Figure 10, for simplicity, the "#" in Pn-# is considered time-series identification information. This time-series identification information is assigned in chronological order from oldest to newest: 1, 2, 3, ..., #, ...

[0085] As shown in Figure 10, for example, in order to synchronize image regions P1, P3, and P4, this time-series identification information is identified between the initially sent image regions P1-1, P3-1, and P4-1 to confirm whether they are time-series consistent with each other.

[0086] For example, if the time-series identification information corresponds to the frame numbers sequentially assigned to the still images that make up the omnidirectional video, then if the time-series identification information corresponding to those frame numbers is common, it can be determined that they are time-series synchronized with each other. Furthermore, assuming that the time-series identification information corresponds to the time when image regions P1 to Pn are generated, and that the generation times between image regions P1 to Pn are always simultaneous without any discrepancies, then if the time-series identification information is common, it can be determined that they are time-series synchronized with each other.

[0087] Initially, when the time-series identification information attached to the end of the image regions P1-1, P3-1, and P4-1 are sent, it can be determined that they are synchronized in time because the time-series identification information is identical to each other. At the next timing, the time-series identification information is the same between image regions P3-2 and P4-2, but it is missing for image region P1. In such a case, it can be determined that the image regions P are not synchronized with each other. Furthermore, at the next timing, if the time-series identification information attached to the end of image region P3-3 does not match that of P1-2 and P4-4, it can also be determined that the image regions P are not synchronized with each other.

[0088] In this way, if the discrimination via time-series identification information determines that the images are not synchronized with each other, adjustments are made to synchronize the time series between each image region P1 to Pn. For example, as described above, if the time-series identification information was common between image regions P3-2 and P4-2, but image region P1 was missing at the same time, and image region P1-2 was linked to image region P3-3 at a later time, synchronization is achieved by linking this image region P1-2 to image regions P3-2 and P4-2, whose time-series identification information is consistent. Alternatively, if image region P1-2, which is at the same time as image regions P3-2 and P4-2, is completely missing, a new image region P1-2 may be generated. In such cases, the missing image region P1-2 may be generated by interpolating pixels based on the preceding and succeeding image regions P1-1 and P1-3 using well-known techniques, or one of the preceding or succeeding image regions P1-1 and P1-3 may be inserted as is.

[0089] The synchronization adjustment using such time-series identification information in step S15 may be performed via a server (not shown) located in the communication network 5, or it may be performed between the display devices 7 that actually receive the data for these image regions P1 to Pn. When the synchronization adjustment is performed between the display devices 7, it may be achieved by the display devices 7 communicating with each other. Alternatively, the synchronization adjustment using time-series identification information may be performed within the control device 2. In any case, since the data for these image regions P1 to Pn are transmitted via different channels, the synchronization adjustment will be performed within the control device 2 before transmission, between the display devices 7 after transmission, or within the communication network 5.

[0090] In step S15, the acoustic information S1 is transmitted to the acoustic device 8 in conjunction with the transmission of the adjusted data for each image region P1 to Pn to the display device 7. The acoustic information S1 is extracted for each effect, for example, at certain time intervals, according to spatial characteristics, sound type, audience information, etc., and sent to the acoustic device 8. By repeatedly performing these steps for the acoustic information S1, the acoustic information S1 is transmitted in a form that aligns with the beginning of each image region P1 to Pn frame on the time axis t, as shown in Figure 10.

[0091] In each image region P1 to Pn and the acoustic information S1, the adjustment for synchronization using such time-series identification information may be performed via a server (not shown) provided in the communication network 5, or it may be performed by the display device 7 and the acoustic device 8 or playback device that actually receive the data for each of these image regions P1 to Pn and the acoustic information S1.

[0092] In this way, the image region P and acoustic information S1 data, which are synchronized with each other in a time series, are sent to each display device 7 that displays its assigned surface 61, and to the acoustic device 8 or playback device.

[0093] Each display device 7 displays an image region P for each surface 61 (step S16). It has already been determined which display devices 7a to 7g will be used to display the image for each surface 61a to 61f. Therefore, the image regions P1 to Pn assigned to each surface 61a to 61f are transmitted to the display devices 7a to 7g that display that surface 61 and displayed. As a result, as shown in Figure 11, each image region P1 to Pn extracted from the omnidirectional video is displayed on each surface 61a to 61f via each display device 7a to 7g.

[0094] Furthermore, the sound device 8 is installed behind each surface 61a to 61f of each display device 7a to 7g so as to transmit sound information S1 to the space 6 and each surface 61. It may be predetermined which display devices 7a to 7g in the space 6 and each surface 61a to 61f will receive sound information via the sound device 8. This allows sound information suitable for the image regions P1 to Pn assigned to the space 6 and each surface 61a to 61f to be transmitted to the display devices 7a to 7g that display the respective surfaces 61. As a result, as shown in Figure 11, sound information synchronized with the space 6 and each surface 61a to 61f can be reproduced via the sound device 8 in accordance with each image region P1 to Pn extracted from the omnidirectional video.

[0095] When the spectator M enters space 6, they can view the image regions P1 to Pn displayed on each of the surfaces 61a to 61g. Since these image regions P1 to Pn are originally cut from an omnidirectional video and divided into six surfaces, the spectator M in space 6 can enjoy the feeling of standing in the center of the omnidirectional video by viewing the image regions P1 to Pn displayed on each of the surfaces 61a to 61g. When the spectator M views each of the surfaces 61a to 61g, the image region P displayed on the surface 61 they viewed comes into their field of vision. In other words, since the image region P corresponding to the direction they viewed comes into their field of vision, they can obtain a sensation similar to that of VR. Moreover, without wearing the glasses-type or goggle-type head-mounted video display device required to experience VR, the spectator M can experience a sense of presence as if they were actually there in space 6. Therefore, the pressure, inconvenience, and hassle of putting on a head-mounted video display device that is associated with wearing one can be eliminated. Furthermore, it eliminates physical effects such as so-called VR sickness, which arise from the discrepancy between the information received visually through the head-mounted video display device and the information received by the body in reality.

[0096] Furthermore, by experiencing the image regions P1 to Pn displayed on each surface 61a to 61g along with the sound information S1, the audience M in space 6 can enjoy the sensation of standing in the center of an omnidirectional video through both visuals and sound (stereophonic sound, 3D sound). By viewing each surface 61a to 61g along with the sound information, the audience M can experience the image region P and sound information S1 displayed on the viewed surface 61. In other words, by viewing the image region P and sound information corresponding to the direction of viewing, the audience M can obtain a sensation in space 6 that is similar to a real-life experience. Moreover, without wearing the glasses-type or goggle-type head-mounted video display device required to experience VR, the 3D sound from the sound device 8 allows the audience M to experience a sense of presence in space 6 as if they were actually there. Therefore, the pressure, inconvenience, and hassle of putting on a head-mounted video display device are eliminated. In addition, physical effects such as so-called VR sickness, which occur due to the discrepancy between the information obtained visually through the head-mounted video display device and the information received by the real body, are eliminated.

[0097] Furthermore, according to the present invention, multiple viewers M can enter the space 6 simultaneously and view a common image area P and sound information S1, enabling simultaneous sharing of omnidirectional video and sound by multiple people within a single virtual space, which was not possible with conventional VR. Moreover, since the common image area P and sound information S1 can be transmitted to other locations via the communication network 5, multiple locations can simultaneously view the common image area P and sound information S1, allowing each to enjoy it in their respective spaces 6.

[0098] Furthermore, according to the present invention, each image region P1 to Pn can be transmitted independently to the display devices 7a to 7f via different communication paths, making it possible to provide content to the space 6 at high speed and low cost. Moreover, any time-series inconsistencies between each image region P1 to Pn that may arise from the independent transmission of each image region P1 to Pn via different communication paths can be resolved through the synchronization adjustment process in step S15.

[0099] Furthermore, the space 6 enclosed by each surface 61 varies in shape and size depending on the actual site conditions, and it becomes necessary to flexibly and freely extract the optimal image region P according to the space 6. According to the present invention, since the image region P to be assigned to each surface 61 can be extracted based on the vertical-horizontal size ratio between each surface 61, it is possible to accommodate the diversity of shapes of such space 6.

[0100] In such a case, if there is a space 6 where a new image region P is to be displayed, an imaging device is installed in space 6. This imaging device may consist of a so-called omnidirectional imaging device that can simultaneously capture images in all directions (360° horizontally and 360° vertically) without omission, with the imaging device body at its center. Alternatively, multiple imaging devices that capture normal planar images may be installed in space 6, and all surfaces 61 may be captured by the multiple imaging devices.

[0101] Using such an imaging device, images of each surface 61 of the space 6 to be displayed are captured, and the aspect ratio of the length and width between each surface 61 is determined using well-known image analysis techniques. Based on the determined aspect ratio of the length and width between each surface 61, the control unit 28 extracts the image region P to be assigned to each surface 61, as described above.

[0102] Furthermore, if the display device 7 consists of a projection display device that projects the image area P onto the surface 61, the image area may be extracted using the following method.

[0103] Figure 12(a) is a side cross-sectional view of a space 6, and Figure 12(b) is a plan view thereof. Display devices 7m and 7w, which consist of projection display devices, and an acoustic device 8 are provided on surfaces 61b and 61d that constitute the side of space 6, respectively. Display devices 7m and 7n project an image area P toward surface 61e that constitutes the ceiling, and the projection direction and field of view θ of the display device 7 are acquired at that time. Display device 7w is provided on surface 61e that constitutes the ceiling, and projects an image area P in four directions toward the four surfaces 61a to 61d that constitute the side, and the projection direction and field of view φ at that time are acquired. Information regarding the arrangement of these display devices 7m, 7n, 7w and the acoustic device 8 is also acquired.

[0104] The projection direction, field of view, and arrangement relationship may be automatically determined not only through input via the operation unit 25, but also through imaging by an imaging device installed in the space 6, as described above.

[0105] Based on the acquired projection direction and field of view or arrangement relationship, the image area to be assigned to each surface may be extracted, and the sound information to be reproduced by the sound device 8 may be assigned to it. In Figures 12(a) and (b), the sound device 8 is installed on the back of surface 61a, but it may also be installed on the back of another surface, or installed in space 6. Furthermore, the sound device 8 may be composed of multiple sound modules (sound components, etc.) combined to form a single sound device 8.

[0106] Furthermore, according to the present invention, the first video acquisition unit 21 and the second video acquisition unit 23 acquire video footage including shooting information of live videos and archived videos, and the audio data acquisition unit 35 acquires audio information corresponding to the video footage. In addition to the shooting information, the distribution destination information may be acquired as appropriate based on, for example, management information for each video or audio information, distribution request information, ranking information, past playback history information, audience information, etc., so that the most suitable distribution destination information is acquired. The distribution destination information may be acquired by pre-specifying it in the shooting information, for example, or it may be stored in the video storage unit 9.

[0107] Furthermore, according to the present invention, the first video acquisition unit 21 and the second video acquisition unit 23 acquire video footage including shooting information of live videos and archived videos, and the audio data acquisition unit 35 acquires audio information corresponding to the video footage. In addition to the shooting information, the distribution destination information may be acquired as appropriate based on, for example, management information for each video or audio information, distribution request information, ranking information, past playback history information, audience information, etc., so that the most suitable distribution destination information is acquired. The distribution destination information may be acquired by pre-specifying it in the shooting information, for example, or it may be stored in the video storage unit 9.

[0108] While embodiments of the present invention have been described, these embodiments are presented as examples only and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their variations are included in the scope and spirit of the invention, as well as in the claims of the invention and its equivalents. [Explanation of symbols]

[0109] 1. Image Region Generation System 2 Control device 3. Recording Module 5. Communication Network 6 Space 7 Display device (playback device) 8 Sound equipment (playback equipment) 9. Video storage unit 21 First video acquisition unit 23 Second video acquisition unit 25 Control section 26 Spatial information acquisition section 27 Judgment Department 28 Control Unit 31 Omnidirectional Imaging Device 32 Mike 35. Audio data acquisition unit 61 sides

Claims

1. In an image region generation system that generates image regions to be displayed on each rectangular surface surrounding a space, A means for acquiring video footage, audio information corresponding to the video footage, and distribution destination information including one or more of the following: distribution contract information, distribution equipment information, customer information regarding distribution audiences, management information for each video or audio information, distribution request information, ranking information, and past playback history information. Based on images of each surface simultaneously captured by an omnidirectional imaging device installed in the above space that is capable of simultaneously capturing images in all directions, a determination means is provided to determine the aspect ratio between the length and width of at least two adjacent surfaces among the four walls and the ceiling and floor surfaces of the enclosed space consisting of six surfaces surrounding the space. Image region extraction means extracts each still image constituting the moving image acquired by the moving image acquisition means into multiple image regions to be assigned to each adjacent surface so that the display images are continuous between them, based on the arrangement relationship of each surface, the aspect ratio between each surface determined by the determination means, and the distribution destination information acquired by the moving image acquisition means, and sequentially assigns time-series identification information to each extracted image region. An assignment means that assigns each image region extracted by the above-mentioned image region extraction means to each of the above-mentioned surfaces, and assigns the above-mentioned acoustic information based on each of the assigned image regions, Each of the above surfaces includes a data transmission means for transmitting data, which includes at least one of the above-mentioned image area or sound device, to each playback device, which includes at least one of the above-mentioned image area or sound device, that has been assigned by the above-mentioned assignment means, in a synchronized manner on different channels such that the display timing of each image area and the playback of the sound information coincide. The above-described image region extraction means adjusts the shape of the boundary between the multiple image regions to be extracted by expanding or compressing it based on the aspect ratio of the length and width between each of the surfaces, and adjusts each extracted image region to be rectangular in shape to match the shape of each of the surfaces. The above-mentioned allocation means, in the adjustment for synchronization, if it is determined that a specific image region is missing and therefore out of synchronization based on the time-series identification information, interpolates the pixels of the missing image region based on the preceding and succeeding image regions, or inserts either of the preceding or succeeding image regions to create a new image region, and performs adjustments to maintain time-series synchronization between each image region and the above-mentioned acoustic information. An image region generation system characterized by the following.

2. The above-mentioned image region extraction means extracts each still image constituting the moving image into multiple image regions based on the acoustic information, such that the displayed images are continuous between adjacent faces according to the arrangement relationship of each face. The image region generation system according to claim 1, characterized by the following:

3. The above data transmission means adjusts the display timing of each image region and the playback of the audio information in the data to be transmitted so that they are synchronized on different channels. The image region generation system according to claim 1, characterized by the following:

4. In an image region display space where an image region is reproduced on each rectangular surface surrounding the space, Each rectangular surface enclosing the space, A means for acquiring video footage, audio information corresponding to the video footage, and distribution destination information including one or more of the following: distribution contract information, distribution equipment information, customer information regarding distribution audiences, management information for each video or audio information, distribution request information, ranking information, and past playback history information. Based on images of each surface simultaneously captured by an omnidirectional imaging device installed in the above space that is capable of simultaneously capturing images in all directions, a determination means is provided to determine the aspect ratio between the length and width of at least two adjacent surfaces among the four walls and the ceiling and floor surfaces of the enclosed space consisting of six surfaces surrounding the space. Image region extraction means extracts each still image constituting the moving image acquired by the moving image acquisition means into multiple image regions to be assigned to each adjacent surface so that the display images are continuous between them, based on the arrangement relationship of each surface, the aspect ratio between each surface determined by the determination means, and the distribution destination information acquired by the moving image acquisition means, and sequentially assigns time-series identification information to each extracted image region. An assignment means that assigns each image region extracted by the above-mentioned image region extraction means to each of the above-mentioned surfaces, and assigns the above-mentioned acoustic information based on each of the assigned image regions, Each of the above surfaces includes a data transmission means for transmitting data, which includes at least one of the above-mentioned image area or sound device, to each playback device, which includes at least one of the above-mentioned image area or sound device, that has been assigned by the above-mentioned assignment means, in a synchronized manner on different channels such that the display timing of each image area and the playback of the sound information coincide. The above-described image region extraction means adjusts the shape of the boundary between the multiple image regions to be extracted by expanding or compressing it based on the aspect ratio of the length and width between each of the surfaces, and adjusts each extracted image region to be rectangular in shape to match the shape of each of the surfaces. The above-mentioned allocation means, in the adjustment for synchronization, if it is determined that a specific image region is missing and therefore out of synchronization based on the time-series identification information, interpolates the pixels of the missing image region based on the preceding and succeeding image regions, or inserts either of the preceding or succeeding image regions to create a new image region, and performs adjustments to maintain time-series synchronization between each image region and the above-mentioned acoustic information. An image region generation device characterized by the following.

5. In an image region generation program that generates image regions to be displayed on each rectangular surface surrounding a space, A video acquisition step involves obtaining video footage, audio information corresponding to the video footage, and distribution destination information including one or more of the following: distribution contract information, distribution equipment information, customer information regarding distribution audiences, management information for each video or audio information, distribution request information, ranking information, and past playback history information, from a computer. Based on images of each surface simultaneously captured by an omnidirectional imaging device installed in the above space that is capable of simultaneously capturing images in all directions, a determination step is made to determine the aspect ratio between the length and width of at least two adjacent surfaces among the four walls and the ceiling and floor surfaces of the enclosed space consisting of six surfaces surrounding the space. Image region extraction step: For each still image constituting the moving image acquired in the moving image acquisition step, based on the arrangement of each surface, the aspect ratio between each surface determined in the determination step, and the distribution destination information acquired in the moving image acquisition step, extracts the image into multiple image regions to be assigned to each surface so that the display images are continuous between adjacent surfaces, and sequentially assigns time-series identification information to each extracted image region. An assignment step in which each image region extracted by the above image region extraction step is assigned to each of the above surfaces, and the above acoustic information is assigned based on each of the assigned image regions, A data transmission step is performed to transmit data, which includes at least one of the display devices for reproducing image regions or the sound devices for reproducing the sound information, to each reproduction device, which includes at least one of the image regions or the sound information assigned in the assignment step, on different channels in a synchronized manner so that the display timing of each image region and the reproduction of the sound information coincide. The above image region extraction step adjusts the shape of the boundary between the multiple image regions to be extracted by expanding or compressing it based on the aspect ratio of the length and width between each of the faces, so that each extracted image region becomes rectangular in shape to match the shape of each of the faces. The above assignment step, in the adjustment for synchronization, if it is determined that a specific image region is missing and therefore out of synchronization based on the above time-series identification information, generates a new image region by interpolating pixels in the missing image region based on the preceding and succeeding image regions, or by inserting either of the preceding or succeeding image regions, and performs adjustments to maintain time-series synchronization between each image region and the above acoustic information. An image region generation program characterized by the following.

Citation Information

Patent Citations

  • Display system, display control apparatus, display apparatus, display method and user interface device

    JP2005099064A

  • Three-dimensional position measurement system, three-dimensional position measurement method, and measurement module

    JP2018009957A

  • Projection type video display device

    JP2018101078A

  • Information processing method, information processing program, information processing system, and information processing device

    JP2019083029A

  • Server device, display device, video display system, and video display method

    JP2019140530A