Three-dimensional image display system, mobile terminal, program thereof, and server and program thereof

The 3D image display system improves compression efficiency and maintains image quality by using a mobile terminal and server to generate and transmit multi-viewpoint images, addressing the low efficiency and degradation issues in conventional systems.

JP2025126977APending Publication Date: 2025-09-01NIPPON HOSO KYOKAI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024023398
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-20
Publication Date
2025-09-01

AI Technical Summary

Technical Problem

Conventional 3D video display systems on mobile devices suffer from low compression efficiency of elemental image groups, leading to image quality degradation due to the high-frequency components in these images.

Method used

A 3D image display system that includes a mobile terminal and a server, where the mobile terminal detects the observer's viewpoint position and receives compressed multi-viewpoint images from the server, which are then expanded into elemental images for display, while the server compresses multi-viewpoint images using methods suitable for 2D video to improve compression efficiency and reduce image quality degradation.

Benefits of technology

The system enhances compression efficiency and maintains image quality by transmitting and displaying 3D images with less degradation compared to conventional methods, allowing for high-quality 3D video on mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025126977000001_ABST
    Figure 2025126977000001_ABST
Patent Text Reader

Abstract

To provide a three-dimensional image display system that enhances a compression efficiency and suppresses an image quality degradation of an element image group.SOLUTION: In a three-dimensional video display system 100, a mobile terminal 2 includes: a camera CIN; a viewpoint position detection part 22 that detects a viewpoint position from a face image and transmits the viewpoint position to the server 1; an image decompression part 23 that decompresses a compressed multi-viewpoint image that is a multi-viewpoint image generated and compressed by the server corresponding to the viewpoint position; an element image group generation part 24 that generates an element image group corresponding to the viewpoint position transmitted from the server corresponding to the compressed multi-viewpoint image; and a three-dimensional display D3D that displays the element image group. The server 1 includes: a multi-viewpoint image generating part 11 that receives a viewpoint position from the mobile terminal 2 and generates a multi-viewpoint image corresponding to the viewpoint position; an image compressing part 12 that generates a compressed multi-viewpoint image; and an information synchronizing part 13 that transmits the compressed multi-viewpoint image and the viewpoint position in association with each other to the mobile terminal 2.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a three-dimensional video display system, a mobile terminal and a program therefor, and a server and a program therefor. [Background technology]

[0002] Three-dimensional images are known as a form of media that can create a high sense of presence. Currently, various methods for displaying three-dimensional images have been proposed, but the ray reproduction 3D method, which reproduces the same light rays as when an object is actually in front of you, is attracting attention. One of these ray reproduction 3D methods is the integral method, which uses a lens array. Furthermore, in recent years, mobile terminals such as smartphones and tablets have become widespread, and it has become common to watch video media on mobile terminals. Therefore, there was a demand for technology that could display high-quality 3D images on mobile devices.

[0003] BACKGROUND ART Conventionally, a system (conventional system) for displaying integral type three-dimensional images on a mobile terminal has been disclosed (see Patent Document 1 and Non-Patent Document 1). This conventional system is configured by connecting a mobile terminal and a rendering server (computer) via a network. The most distinctive feature of this system is its ability to display 3D images with a wide viewing area using viewpoint tracking (for example, Non-Patent Document 2). To support viewpoint tracking, elemental image groups must be generated in real time. However, since it is difficult for currently commercially available mobile terminals to generate multi-viewpoint images from a 3D model in real time and convert them into elemental image groups, the conventional system uses a rendering server. That is, the rendering server of the conventional system receives the viewpoint position of the observer from the mobile device, generates a multi-viewpoint image corresponding to the viewpoint position from the 3D model, converts it into a group of element images, and sends it to the mobile device.The mobile device of the conventional system then receives the group of element images from the rendering server and displays the group of element images on a 3D display, allowing the observer to view a 3D video. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2022-113478 [Non-patent literature]

[0005] [Non-Patent Document 1] Masanori Kano, Naoto Okaichi, Jun Arai, "Viewpoint-tracking 3D image display system using mobile devices", Institute of Image Information and Television Engineers Annual Conference, 23C-4 (2022) [Non-patent document 2] Naoto Okaichi, Hisayuki Sasaki, Masanori Kano, Jun Arai, Masahiro Kawakita, and Takeshi Naemura, “Design of optical viewing zone suitable for eye-tracking integral 3D display,” OSA Continuum, Vol. 4, pp. 1415-1429 (2021) Summary of the Invention [Problem to be solved by the invention]

[0006] In the conventional system, when transmitting elemental images from the rendering server to the mobile terminal, the image is compressed using a compression method for 2D video. However, elemental image groups are composed of many elemental images and contain many high-frequency components. As a result, conventional systems have low compression efficiency for elemental image groups, resulting in degradation of image quality. Therefore, a technology to improve the image quality of elemental image groups has been desired.

[0007] The present invention has been made in consideration of existing demands, and aims to provide a 3D video display system, a mobile terminal and its program, and a server and its program that can improve compression efficiency and suppress deterioration in image quality of elemental images compared to conventional systems. [Means for solving the problem]

[0008] In order to solve the above problem, the 3D image display system of the present invention is a 3D image display system comprising a mobile terminal that allows an observer to view a 3D image corresponding to a viewpoint position, and a server that generates a multi-viewpoint image corresponding to the viewpoint position, wherein the mobile terminal comprises a camera, a viewpoint position detection unit, an image expansion unit, an element image group generation unit, and a 3D display, and the server comprises a multi-viewpoint image generation unit, an image compression unit, and an information synchronization unit.

[0009] In such a configuration, the mobile terminal uses a camera to capture an image of an observer viewing the mobile terminal. The mobile terminal then uses a viewpoint position detection unit to detect the viewpoint position of the observer from the face image of the observer captured by the camera, and transmits the viewpoint position to the server. In addition, the mobile terminal receives from the server, by the image expansion unit, a compressed multi-viewpoint image, which is a multi-viewpoint image generated and compressed by the server in accordance with the viewpoint position, in association with the viewpoint position, and expands the compressed multi-viewpoint image into a multi-viewpoint image. Then, the mobile terminal generates, from the multi-viewpoint image, a group of element images corresponding to the viewpoint positions received from the server in response to the compressed multi-viewpoint image, by the element image group generation unit. The mobile terminal then displays the elemental images on a three-dimensional display. This allows the viewer to view a three-dimensional image.

[0010] In addition, in this configuration, the server receives the viewpoint position from the mobile terminal and generates a multi-viewpoint image of the three-dimensional model corresponding to the viewpoint position by using the multi-viewpoint image generation unit. The server then compresses the multi-viewpoint image using an image compression unit to generate a compressed multi-viewpoint image. The multi-viewpoint image is closer to a two-dimensional natural image than the elemental image group and contains fewer high-frequency components. Therefore, the image compression unit can increase the compression efficiency of the elemental image group and suppress image quality degradation. The server then transmits the compressed multi-viewpoint images and the viewpoint positions to the mobile terminal in association with each other using an information synchronization unit. This allows the mobile terminal to generate a group of elemental images synchronized with the viewpoint position.

[0011] The mobile terminal can be operated by a program that causes a computer to function as the mobile terminal. The server can also operate a computer with a program that causes the computer to function as the server. [Effects of the Invention]

[0012] According to the present invention, it is possible to compress and transmit a multi-viewpoint image according to the viewpoint position. Because this multi-viewpoint image is closer to a natural 2D image, it is possible to increase the compression efficiency and suppress image quality degradation compared to elemental images with many high-frequency components. Therefore, the present invention makes it possible to display elemental images with less degradation than conventional methods on a mobile terminal while maintaining the transmission bit rate, thereby improving the image quality of 3D video. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a schematic diagram illustrating the configuration of a 3D video display system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the configuration of a mobile terminal and a server according to the embodiment of the present invention. [Figure 3] This is an explanatory diagram for explaining the method of generating a multi-viewpoint image, where (a) shows the case where an observer views a three-dimensional model from the center in front, and (b) shows the case where an observer views the three-dimensional model from the left in front. [Figure 4] FIG. 1 is a diagram illustrating an example of a multi-viewpoint image. [Figure 5] FIG. 2 is a diagram illustrating a pixel configuration of a multi-viewpoint image. [Figure 6] 6 is an explanatory diagram for explaining the relationship between the multi-viewpoint image and the elemental image group in FIG. 5. FIG. [Figure 7]This is an explanatory diagram to explain the changes in the group of element images on the display as the viewpoint moves, where (a) shows the state when the viewpoint is in the center in front of the display (before the viewpoint is moved), and (b) shows the state when the viewpoint is moved to the left from (a) (after the viewpoint is moved). [Figure 8] 10 is an explanatory diagram for explaining rearrangement of pixels from a multi-viewpoint image to a group of elemental images after a viewpoint is moved; FIG. [Figure 9] FIG. 10 is a diagram showing an example of an element image group. [Figure 10] FIG. 3 is a sequence diagram showing the operation of the 3D video display system according to the embodiment of the present invention. [Figure 11] FIG. 10 is a block diagram showing the configuration of a 3D video display system according to a modified example of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0014] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. <Configuration of 3D image display system> The configuration of a 3D video display system 100 according to an embodiment of the present invention will be described with reference to FIG.

[0015] The three-dimensional image display system 100 is a system that displays three-dimensional images by tracking the viewpoint position of an observer 5. Note that the three-dimensional image display method will be described here as an example of the integral method. The three-dimensional video display system 100 is configured by connecting a server 1 and a mobile terminal 2 via a network N.

[0016] The server (rendering server) 1 receives viewpoint position information indicating the viewpoint position of the observer 5 from the mobile terminal 2 and the three-dimensional display D of the mobile terminal 2. 3D The computer generates multi-view images of a three-dimensional model to be displayed, photographed by a number of virtual cameras (not shown), based on the specifications of the above. The server 1 transmits the multi-viewpoint images of the three-dimensional model generated in accordance with the viewpoint position of the observer 5 to the mobile terminal 2 via the network N in association with the viewpoint position. At this time, the server 1 compresses the multi-view images to reduce the amount of data to be transmitted.

[0017] The mobile terminal 2 is a device that generates a group of elemental images according to the corresponding viewpoint positions from the multi-viewpoint image generated by the server 1, and displays a 3D video. The mobile terminal 2 is, for example, a smartphone, a tablet terminal, or the like. The mobile terminal 2 has a camera C IN The viewpoint position of the observer 5 is detected from the face image of the observer 5 captured by the camera 1, and is transmitted to the server 1 via the network N as viewpoint position information. The mobile terminal 2 receives the multi-viewpoint image and the corresponding viewpoint position transmitted from the server 1 via the network N, generates a group of element images corresponding to the viewpoint position, and displays the group of element images on the three-dimensional display D. 3D Display in.

[0018] As a result, the 3D video display system 100 can perform the generation of multi-viewpoint images, which involves a heavy load, in the server 1, and display to the observer 5 a 3D video that corresponds to the viewpoint position. The configurations of the server 1 and the mobile terminal 2 that make up the three-dimensional video display system 100 will be described below.

[0019] <server> The configuration of the server 1 will be described with reference to FIG. 2 (and also with reference to FIG. 1 as appropriate). The server 1 includes a communication unit 10, a multi-viewpoint image generation unit 11, an image compression unit 12, and an information synchronization unit 13.

[0020] The communication unit 10 transmits and receives various data to and from the mobile terminal 2 via the network N. The communication unit 10 transmits and receives data via the network N using a router (not shown) or the like. Note that the communication of the communication unit 10 may be wireless communication or wired communication. Here, the communication unit 10 receives viewpoint position information indicating the viewpoint position of the observer 5 from the mobile terminal 2, and outputs the received viewpoint position information to the multi-viewpoint image generation unit 11 and the information synchronization unit 13. Furthermore, the communication unit 10 receives the viewpoint position information and the multi-viewpoint image associated by the information synchronization unit 13 (to be described later), and transmits them to the mobile terminal 2.

[0021] The multi-viewpoint image generating unit 11 receives a viewpoint position from the mobile terminal and generates a multi-viewpoint image of a three-dimensional model corresponding to the viewpoint position. Here, the multi-viewpoint image generating unit 11 receives viewpoint position information indicating the viewpoint position via the communication unit 10, and generates a three-dimensional image D 3D Multi-view images are generated based on the specifications. The viewpoint position information is displayed on the 3D display D 3D are three-dimensional coordinates of the viewpoint position (center of the pupils of both eyes) of the observer 5 in the display coordinate system with the display center as the origin. 3D Display D 3D The specifications of, for example, 3D displays 3D Two-dimensional display D 2D (See Figure 7) pixel count (horizontal and vertical pixel count), pixel pitch, lens array L A (See Figure 7) Lens pitch and focal length.

[0022] The multi-viewpoint image generation unit 11 places a three-dimensional model that is the basis for the three-dimensional image in a computer graphics (CG) space, and generates a multi-viewpoint image by virtually capturing images of the three-dimensional model using a virtual camera array consisting of multiple virtual cameras arranged two-dimensionally around the viewpoint position.

[0023] Here, a method for generating a multi-viewpoint image by the multi-viewpoint image generating unit 11 will be described with reference to FIG. As shown in FIG. 3, the multi-viewpoint image generating unit 11 generates a virtual three-dimensional display D VA three-dimensional model M is placed in a CG space of a three-dimensional coordinate system (x, y, z) with the center of the object as the origin. Here, the horizontal direction is the x-axis, the vertical direction is the y-axis, and the depth direction is the z-axis.

[0024] Observer 5 is using 3D display D 3D When the observer 5 is viewing the virtual 3D display D from the center in front of the virtual 3D display DS, the multi-viewpoint image generator 11 generates a virtual 3D display D at a position (visual distance D) specified by the viewpoint position P of the observer 5 in the CG space, as shown in FIG. V Directly facing the virtual camera array C A Then, the multi-viewpoint image generator 11 uses the virtual cameras c1, c2, ..., c5 to generate a virtual 3D display D V The entire display is photographed as the angle of view.

[0025] When the viewpoint position P of the observer 5 changes, for example, when the observer 5 moves to the 3D display D 3D When the observer 5 is viewing the virtual 3D display D from the front left, the multi-viewpoint image generator 11 generates a virtual 3D display D at a left position (visual distance D) specified by the viewpoint position P of the observer 5 in the CG space, as shown in FIG. V Parallel to the virtual camera array C A 3(a), the multi-viewpoint image generator 11 uses the virtual cameras c1, c2, ..., c5 to generate images on the virtual 3D display D V The entire display is photographed as the angle of view. Here, an example is shown in which five virtual cameras c1, c2, ..., c5 are arranged in the horizontal direction, but this number can be adjusted by adjusting the number of virtual cameras in the horizontal and vertical directions. 3D It is determined according to the specifications.

[0026] The multi-viewpoint image generator 11 generates a multi-viewpoint image by integrating images (viewpoint images) captured by the individual virtual cameras. For example, when there are five virtual cameras in the horizontal direction and five in the vertical direction, the multi-viewpoint image generating unit 11 generates a 5×5 viewpoint image V captured by each virtual camera as shown in FIG. S By connecting these images without gaps, a multi-viewpoint image V MGenerate. For simplicity, the multi-viewpoint images are displayed on the two-dimensional display D of the mobile terminal 2. 2D In other words, the multi-viewpoint image generator 11 generates a virtual 3D display D V The angle of view corresponding to the viewpoint image V S Shoot at a pixel count of 1 / 100. Returning to FIG. 2, the configuration of the server 1 will be further explained. The multi-viewpoint image generator 11 outputs the generated multi-viewpoint image to the image compressor 12.

[0027] Image compression unit 12 compresses the multi-viewpoint images generated by multi-viewpoint image generation unit 11 to generate compressed multi-viewpoint images. Image compression unit 12 compresses the multi-viewpoint images generated for each frame using a general compression method for two-dimensional video. For example, image compression unit 12 compresses the multi-viewpoint images in accordance with standards such as H.264 and H.265. The image compression unit 12 outputs the generated compressed multi-viewpoint image to the information synchronization unit 13.

[0028] The information synchronization unit 13 associates the compressed multi-viewpoint image with the viewpoint position and transmits it to the mobile terminal 2. Here, the information synchronization unit 13 associates the compressed multi-viewpoint image generated by the image compression unit 12 with the viewpoint position information input from the communication unit 10 and synchronizes them. The compressed multi-viewpoint image is generated by the multi-viewpoint image generator 11 in synchronization with the viewpoint position specified by the viewpoint position information. Therefore, the information synchronizer 13 transmits the compressed multi-viewpoint image together with viewpoint position information that specifies the viewpoint position used when generating the multi-viewpoint image to the mobile terminal 2 via the communication unit 10.

[0029] The method for synchronizing the compressed multi-viewpoint images and the viewpoint position information is not particularly limited. For example, the compressed multi-viewpoint images and the viewpoint position information may be associated with each other by assigning the same ID, or the viewpoint position information may be arranged in the header section and the multi-viewpoint images may be arranged in the data section to form a single data structure.

[0030] With the configuration described above, the server 1 can generate the multi-viewpoint image required to display a 3D video corresponding to the viewpoint position of the observer 5 based on the viewpoint position information received from the mobile terminal 2, and transmit it to the mobile terminal 2. At this time, the server 1 compresses the multi-viewpoint images, and therefore the compression efficiency can be improved compared to the conventional case where a group of element images containing many high-frequency components is compressed. The server 1 can be operated by a general computer running a program that causes the computer to function as each of the above-mentioned units.

[0031] <Mobile device> Next, the configuration of the mobile terminal 2 will be described with reference to FIG. 2 (and also with reference to FIG. 1 as appropriate). The mobile terminal 2 has a camera C IN and 3D display D 3D The image processing apparatus also includes a communication unit 20, an image input unit 21, a viewpoint position detection unit 22, an image development unit 23, an element image group generation unit 24, and an image output unit 25 as control units.

[0032] Camera C IN is a general camera (video camera). IN is a 3D display 3D The camera C is mounted on the display side of the camera and captures the face of the observer 5. IN is an in-camera for smartphones, tablet devices, etc. Camera C IN The facial image of the observer 5 captured by the camera is input to the image input unit 21 frame by frame.

[0033] 3D Display D 3D is a display device that displays a group of elemental images and allows the observer 5 to view a three-dimensional image. For example, a three-dimensional display D 3D As an integral type display device, a two-dimensional display such as a liquid crystal display is used. 2D (See Figure 7) and a lens array L arranged on the display surface at a distance equal to the focal length. A(See FIG. 7) For ease of explanation, the lens array will be described here as an example of a square array structure. 3D Display D 3D inputs a group of element images from the image output unit 25 and displays them.

[0034] The communication unit 20 transmits and receives various data to and from the server 1 via the network N. The communication unit 20 transmits and receives data via the network N using a Wi-Fi router (not shown) or the like. Here, the communication unit 20 receives as input viewpoint position information indicating the viewpoint position of the observer 5 detected by the viewpoint position detection unit 22 and transmits it to the server 1. Furthermore, the communication unit 20 receives the associated compressed multi-viewpoint images and viewpoint position information from the server 1, and outputs them to the image development unit .

[0035] The image input unit 21 is a camera C IN A face image is input at a frame cycle. The image input unit 21 outputs the input face image of the observer 5 to the viewpoint position detection unit 22 sequentially for each frame.

[0036] The viewpoint position detection unit 22 detects the camera C IN The viewpoint position of the observer 5 is detected from the face image of the observer 5 captured by the camera, and the viewpoint position is transmitted to the server 1. The viewpoint position detection unit 22 receives a face image from the image input unit 21 and outputs it to the three-dimensional display D 3D Detect the viewpoint position in the display coordinate system. Here, the viewpoint position detection unit 22 detects the camera C from the face image. IN The viewpoint position in the three-dimensional camera coordinate system is detected and displayed on the three-dimensional display D. 3D This is converted into the viewpoint position in the three-dimensional display coordinate system.

[0037] A general method may be used to detect the viewpoint position in the camera coordinate system from the face image. For example, using ARCor described in the following reference 1,IN The viewpoint positions of both eyes of the observer 5 in the camera coordinate system are detected. (Reference 1) “ARCore”,<URL:https: / / developers.google.com / ar?hl=ja>

[0038] Then, the viewpoint position detection unit 22 detects the camera C IN and 3D display D 3D The viewpoint position in the camera coordinate system is converted to the viewpoint position in the display coordinate system based on the known positional relationship between the camera and the display. The viewpoint position of the observer 5 in the display coordinate system is the center of the viewpoint positions of both eyes.

[0039] In addition to the method using ARCor, the gaze point position can also be detected by detecting the two-dimensional positions of both eyes on a face image using Dlib described in Reference 2 below, and converting them into three-dimensional coordinates using the method described in Reference 3 below. (Reference 2) “Dlib”,<URL:http: / / dlib.net / > (Reference 3) Naoto Okaichi, Hisayuki Sasaki, Masanori Kano, Jun Arai, Masahiro Kawakita, and Takeshi Naemura, “Design of optical viewing zone suitable for eye-tracking integral 3D display,” OSA Continuum, Vol. 4, pp. 1415-1429 (2021)

[0040] The viewpoint position detection unit 22 transmits the viewpoint position in the display coordinate system detected from the face image to the server 1 via the communication unit 20.

[0041] The image expansion unit 23 receives from the server 1 a compressed multi-viewpoint image, which is a multi-viewpoint image generated and compressed by the server 1 in accordance with the viewpoint position, in association with the viewpoint position, and expands the compressed multi-viewpoint image into a multi-viewpoint image. Here, image decompression unit 23 receives compressed multi-viewpoint images from communication unit 20 and decompresses the compressed multi-viewpoint images as multi-viewpoint images. The decompression method used by image decompression unit 23 corresponds to the compression (H.264, H.265, etc.) used by image compression unit 12. The image development unit 23 outputs the developed multi-viewpoint image and the viewpoint position information input from the communication unit 20 in association with the compressed multi-viewpoint image to the element image group generation unit 24.

[0042] The element image group generating unit 24 generates, from the multi-viewpoint image developed by the image developing unit 23, an element image group corresponding to the viewpoint position received from the server 1 in correspondence with the compressed multi-viewpoint image. The element image group generation unit 24 displays pixels of the same viewpoint of the multi-viewpoint image according to the viewpoint position (two-dimensional display D 2D (See Figure 7) to generate a group of elemental images.

[0043] Here, the relationship between the multi-viewpoint image and the elemental image group will be described with reference to FIGS. Figure 5 shows the multi-view image V M Here, the pixel configuration of the multi-view image V M is the horizontal M and vertical N viewpoint images V S It is composed of the viewpoint image V S is composed of horizontal U pixels and vertical V pixels. S The same pixel positions are indicated by the same shapes (circle, triangle, etc.).

[0044] Figure 6 shows the element image group E G The pixel configuration of elemental image group E is shown. G is a configuration in which element images E are arranged in accordance with the number of lenses (element lenses) in the lens array. Here, for example, the upper left element image E is the multi-viewpoint image V in FIG. M Each viewpoint image V S Similarly, the other element images E are arranged in the multi-viewpoint image V M Each viewpoint image V SThe pixel structure is such that pixels at the same pixel positions are arranged.

[0045] Next, with reference to FIG. 7, changes in the position and size of elemental images on the display accompanying movement of the viewpoint will be described. Figure 7(a) shows the 3D display D 3D The viewpoint position Pa is located at a distance of a viewing distance D (mm) in the depth direction (z direction) from the origin [(x,y,z)=(0,0,0)] of the display coordinate system. In other words, the viewpoint position Pa is Pa=[0 0 D] T (T is transposition). Also, the element image E is A It corresponds one-to-one to the lens L.

[0046] Here, the two-dimensional display D 2D The pixel pitch is P P (mm), Lens array L A The lens pitch is P L (mm), Lens array L A The focal length is F (mm). In this case, the two-dimensional display D 2D The size of the element image E above is the lens pitch P L The image is projected after being enlarged by a magnification factor α. The magnification factor α is expressed by the following equation (1).

[0047]

number

[0048] That is, two-dimensional display D 2D The size of the element image E above is αP L (mm). When this is expressed in image pixels, the size of element image E is αP L / P P (pixels).

[0049] FIG. 7(b) shows the state where the viewpoint has moved from the viewpoint Pa in FIG. 7(a) to the viewpoint Pb. Here, the three-dimensional display D 3DFrom the origin of the display coordinate system, x in the horizontal direction (x direction) E (mm), and the viewpoint position Pb is located at a distance of a viewing distance D' (mm) in the depth direction (z direction). That is, the viewpoint position Pb is defined as Pb=[x E 0 D′] T (T is transpose). In this case, the two-dimensional display D 2D The size of the element image E above is the lens pitch P L The image is projected by enlarging it by a magnification factor α'. The magnification factor α' is expressed by the following equation (2).

[0050]

number

[0051] That is, two-dimensional display D 2D The size of the element image E above is α′P L (mm). When this is expressed in image pixels, the size of element image E is α′P L / P P (pixels).

[0052] Furthermore, when the viewpoint position moves horizontally or vertically, the position of the separator (the center of the element image in FIG. 7(b)) moves in element image E, as shown in FIG. 7(b). Note that FIG. 7(b) only shows movement in the horizontal direction. As shown in Figure 7(b), the viewpoint position is x in the x direction. E (mm) If the change occurs, the 2D display D 2D In the above, Δx shown in the following equation (3) E It is necessary to move the separator of element image E by (mm) in the direction opposite to the movement direction of the viewpoint position.

[0053]

number

[0054] When this is expressed in terms of image pixels, Δu shown in the following equation (4) is EThe division of element image E is moved by (pixels).

[0055]

number

[0056] That is, when the viewpoint position is on the z-axis of the display coordinate system as shown in FIG. 7(a), the element image group generation unit 24 generates the multi-viewpoint image V shown in FIG. M The pixels are grouped into the element image group E shown in FIG. G The size of each element image E is αP L / P P By converting the image into the size of the element image, a group of element images for display is generated.

[0057] 7B, when the viewpoint position moves in the horizontal direction from the z-axis of the display coordinate system, the element image group generation unit 24 generates the multi-viewpoint image V shown in FIG. M The pixels are grouped into the element image group E shown in FIG. G , and rearrange the pixels into Δu as shown in Figure 8. E (pixels) and change the size of each element image E to α′P L / P P By converting it into the size of G Generate.

[0058] Although only the movement of the viewpoint position in the horizontal direction has been described here, the same applies to the movement in the vertical direction. That is, the element image group generation unit 24 changes the size of the element images by moving the viewpoint position of the observer 5 in the z direction (depth direction), and changes the division of the element images by moving in the x direction and y direction (horizontal direction and vertical direction), thereby generating an element image group for display. Returning to FIG. 2, the configuration of the mobile terminal 2 will be further described.

[0059] The element image group generating unit 24 generates, for example, the multi-viewpoint image V shown in FIG. M From the above, element image group E consisting of element image E shown in FIG.G Generate. The element image group generating unit 24 generates the element image group E G is output to the image output unit 25.

[0060] The image output unit 25 outputs the element image group generated by the element image group generation unit 24 to a three-dimensional display D 3D The output is Here, the image output unit 25 is a three-dimensional display D 3D Two-dimensional display D 2D The element images are written to the display memory (see FIG. 7). This allows for a 3D display. 3D The lens array L A (See Figure 7) to display integral 3D images.

[0061] With the configuration described above, the mobile terminal 2 can transmit the viewpoint position of the observer 5 to the server 1, receive a multi-viewpoint image corresponding to the viewpoint position from the server 1, and display a three-dimensional video. The mobile terminal 2 can be operated by a program that causes a general computer to function as each of the above-mentioned units.

[0062] <Operation of the 3D image display system> The operation of the 3D video display system 100 according to the embodiment of the present invention will be described with reference to FIG. 10 (and as appropriate, with reference to FIG. 2 for the configuration).

[0063] First, the mobile terminal 2 performs the following steps S1 to S3. In step S1, the image input unit 21 receives an image from the camera C IN The face image of the observer 5 taken by the camera is input. In step S2, the viewpoint position detection unit 22 detects the position of the three-dimensional display D from the face image input in step S1. 3D Detect the viewpoint position in the display coordinate system. Here, the viewpoint position detection unit 22 detects the camera C from the face image. INThe viewpoint positions of both eyes of the observer 5 in the camera coordinate system of the camera C are detected. IN and 3D display D 3D Based on the known positional relationship, the viewpoint position in the camera coordinate system is converted to the viewpoint position in the display coordinate system. In step S3, the communication unit 20 transmits to the server 1 via the network N viewpoint position information indicating the viewpoint position of the viewer 5 detected in step S2.

[0064] Then, the server 1 performs the following steps S4 to S8. In step S4, the communication unit 10 receives the viewpoint position information transmitted from the mobile terminal 2 via the network N. In step S5, the multi-viewpoint image generator 11 calculates the viewpoint position information input in step S4 and the three-dimensional display D 3D Multi-viewpoint images are generated based on the specifications. That is, the multi-viewpoint image generation unit 11 generates a multi-viewpoint image by placing a three-dimensional model that is the basis for the three-dimensional image within a CG space and virtually photographing the three-dimensional model with a virtual camera array consisting of multiple virtual cameras arranged two-dimensionally centered on the viewpoint position.

[0065] In step S6, the image compression unit 12 compresses the multi-view images generated in step S5 to generate compressed multi-view images. For example, the image compression unit 12 compresses the multi-view images using a general compression method for two-dimensional video, such as H.264 or H.265. In step S7, the information synchronization unit 13 synchronizes (corresponds) the compressed multi-viewpoint image generated in step S6 with the viewpoint position information input in step S4. In step S8, the communication unit 10 transmits the compressed multi-viewpoint image and the viewpoint position information associated with each other in step S7 to the mobile terminal 2 via the network N.

[0066] Then, the mobile terminal 2 performs the following steps S9 to S12. In step S9, the communication unit 20 receives the compressed multi-viewpoint image and the viewpoint position information transmitted from the server 1 via the network N. In step S10, the image decompression unit 23 decompresses the compressed multi-viewpoint image received in step S9 into the original multi-viewpoint image. The decompression in step S10 corresponds to the compression method in step S6.

[0067] In step S11, the element image group generation unit 24 generates an element image group for the three-dimensional display D from the multi-viewpoint images developed in step S10 and the viewpoint positions indicated by the corresponding viewpoint position information received in step S9. 3D Generate a group of element images to be displayed. The elemental image group generating unit 24 generates an elemental image group by projecting pixels of the same viewpoint of the multi-viewpoint image onto the display surface according to the viewpoint position. In step S12, the image output unit 25 outputs the elemental images generated in step S11 to the three-dimensional display D 3D Display in.

[0068] By repeating the operations of steps S1 to S12 described above in a frame cycle, the observer 5 can view a three-dimensional image corresponding to the viewpoint position. Through the above operations, the 3D video display system 100 can generate the multi-viewpoint images, which is the most demanding process for generating elemental images, on the server 1.

[0069] Furthermore, in the 3D video display system 100, multi-viewpoint images are compressed and transmitted from the server 1 to the mobile terminal 2, thereby improving compression efficiency and suppressing degradation in image quality compared to conventional methods of compressing and transmitting groups of elemental images. Although the embodiment of the present invention has been described above, the present invention is not limited to this embodiment.

[0070] <<Modification of 3D image display system>> In the 3D image display system 100, the 3D image to be displayed is controlled only based on the viewpoint position of the observer 5, but the 3D image to be displayed may also be controlled based on the operation of the observer 5 or the attitude of the mobile terminal 2. A modified example in which the displayed three-dimensional video is controlled by the operation of the observer 5 or the attitude of the mobile terminal 2 will be described with reference to FIG.

[0071] The 3D image display system 100B is a system that displays 3D images by tracking the viewpoint position of the observer 5. Furthermore, the 3D image display system 100B controls the 3D images to be displayed according to the operation of the observer 5 and the attitude of the mobile terminal 2B. The three-dimensional video display system 100B is configured by connecting a server 1B and a mobile terminal 2B via a network N.

[0072] The server 1B includes a communication unit 10, a multi-viewpoint image generation unit 11B, an image compression unit 12, and an information synchronization unit 13B. The mobile terminal 2B has a camera C IN and 3D display D 3D The device also includes a communication unit 20, an image input unit 21, a viewpoint position detection unit 22, an image development unit 23, an element image group generation unit 24B, an image output unit 25, an operation input unit 26, and a terminal attitude detection unit 27 as control units. The same components as those in the three-dimensional image display system 100 are denoted by the same reference numerals and the description thereof will be omitted.

[0073] <Observer manipulation> As a configuration for controlling the displayed three-dimensional image by the operation of the observer 5, an operation input unit 26 is newly provided, and the functions of the multi-viewpoint image generator 11 are added to form a multi-viewpoint image generator 11B.

[0074] The operation input unit 26 is used to input operations by the observer 5. For example, the operation input unit 26 inputs swipes, pinches in, pinches out, and the like as operations by the observer 5 using a touch panel or the like (not shown). For example, a swipe operation is used to rotate the three-dimensional model in a specified direction, and a pinch-in / pinch-out operation is used to input instructions to reduce / enlarge the three-dimensional model. The operation input unit 26 transmits operation information indicating the operation content to the server 1B via the communication unit 20.

[0075] In addition to the functions of multi-viewpoint image generation unit 11, multi-viewpoint image generation unit 11B performs processes such as rotation, reduction / enlargement, etc. on the three-dimensional model based on operation information input from communication unit . For example, when rotation is instructed, multi-viewpoint image generating unit 11B rotates three-dimensional model M by a specified amount in a specified direction and places it in the CG space shown in FIG. Furthermore, when the three-dimensional model M is reduced / enlarged, for example, the multi-viewpoint image generating unit 11B reduces / enlarges the three-dimensional model M at a specified magnification and places it in the CG space shown in FIG. Then, the multi-viewpoint image generating unit 11B generates a multi-viewpoint image according to the viewpoint position for the three-dimensional model placed after the operation. As a result, the 3D image display system 100B can display 3D images in response to operations by the viewer 5, in addition to the effects of the 3D image display system 100.

[0076] <Device orientation> As a configuration for controlling the 3D image displayed depending on the attitude of the mobile terminal 2B, a new terminal attitude detection unit 27 is provided, and the functions of the multi-viewpoint image generation unit 11, information synchronization unit 13 and element image group generation unit 24 are added to form a multi-viewpoint image generation unit 11B, an information synchronization unit 13B and an element image group generation unit 24B.

[0077] The terminal attitude detection unit 27 detects the attitude of the portable terminal 2 B. For example, the terminal attitude detection unit 27 detects whether the portable terminal 2 B is facing horizontally or vertically using a gyro sensor (not shown) or the like. The terminal orientation detection unit 27 transmits orientation information indicating the terminal orientation, either landscape or portrait, via the communication unit 20 to the server 1B.

[0078] The multi-viewpoint image generating unit 11B generates a multi-viewpoint image based on the posture information input from the communication unit 10 in addition to the function of the multi-viewpoint image generating unit 11 or the function of the operation input described above. When the orientation information indicates landscape orientation, the multi-viewpoint image generating unit 11B performs the process described above. Also, when the orientation information indicates portrait orientation, the multi-viewpoint image generating unit 11B generates a virtual three-dimensional display D V The orientation of the virtual camera array C is arranged vertically. A By setting the above and photographing the 3D model M, a multi-viewpoint image is generated.

[0079] Information synchronization unit 13B associates the compressed multi-viewpoint image with the viewpoint position and the terminal attitude, and transmits them to mobile terminal 2. Here, information synchronization unit 13B associates and synchronizes the compressed multi-viewpoint image generated by image compression unit 12 with the viewpoint position information and attitude information input from communication unit 10. The information synchronization unit 13B transmits the associated compressed multi-viewpoint image, the viewpoint position information, and the attitude information to the mobile terminal 2B via the communication unit .

[0080] The element image group generating unit 24B generates an element image group for the three-dimensional display D from the multi-viewpoint images developed by the image developing unit 23, the viewpoint positions indicated by the corresponding viewpoint position information, and the posture information. 3D The orientation information is input from the communication unit 20 in synchronization with the multi-viewpoint images. The element image group generating unit 24B generates an element image group from the multi-viewpoint image in accordance with the orientation of the posture information. As a result, the 3D video display system 100B can display 3D video in accordance with the orientation of the mobile terminal 2, in addition to the effects of the 3D video display system 100.

[0081] Although the embodiment of the present invention has been described above, the present invention is not limited to this embodiment. Although the example described here uses an integral system with parallax in both the horizontal and vertical directions, a lenticular system with parallax only in the horizontal direction may also be used. In this case, a cylindrical lens array is used as the lens array, and the multi-viewpoint image is an image with different viewpoints in the horizontal direction. However, if a lenticular system that has parallax only in the horizontal direction is used, the control based on the terminal attitude of the above-described three-dimensional image display system 100B will not be supported.

[0082] Although the lens array has been described here as having a square array structure, other lens array structures such as a honeycomb structure may also be used. In this case, the element image group generating units 24 and 24B may generate element image groups corresponding to the lens array structure from the multi-viewpoint images. [Explanation of symbols]

[0083] 100 3D image display system 1 server 10. Communications Department 11 Multi-viewpoint image generation unit 12 Image Compression Unit 13 Information Synchronization Department 2. Mobile devices 20 Communications Department 21 Image input unit 22 Viewpoint position detection unit 23 Image development section 24 Element image group generation unit 25 Image output unit 26 Operation input section 27 Terminal orientation detection unit C IN camera D 3D 3D display N Network V M Multi-view images E G Elemental images

Claims

1. A three-dimensional video display system including a mobile terminal that allows an observer to view a three-dimensional video corresponding to a viewpoint position, and a server that generates a multi-viewpoint image corresponding to the viewpoint position, The mobile terminal a camera for photographing the observer; a viewpoint position detection unit that detects a viewpoint position of the observer from a face image of the observer captured by the camera and transmits the viewpoint position to the server; an image expansion unit that receives from the server a compressed multi-viewpoint image that is a multi-viewpoint image generated and compressed by the server in correspondence with the viewpoint position, in association with the viewpoint position, and expands the compressed multi-viewpoint image into a multi-viewpoint image; an element image group generation unit that generates, from the multi-viewpoint image, an element image group corresponding to the viewpoint positions received from the server in response to the compressed multi-viewpoint image; a three-dimensional display that displays the elemental images; The server a multi-viewpoint image generating unit that receives the viewpoint position from the mobile terminal and generates a multi-viewpoint image of a three-dimensional model corresponding to the viewpoint position; an image compression unit that compresses the multi-viewpoint image to generate a compressed multi-viewpoint image; an information synchronization unit that associates the compressed multi-viewpoint image with the viewpoint position and transmits the image to the mobile terminal; A three-dimensional image display system comprising:

2. the mobile terminal further includes an operation input unit that inputs an instruction for an operation of rotating, reducing, or enlarging the three-dimensional model of the three-dimensional image and transmits the instruction to the server; The three-dimensional video display system according to claim 1, characterized in that the multi-viewpoint image generation unit generates the multi-viewpoint image by rotating, shrinking or enlarging the three-dimensional model in accordance with the operation instruction.

3. the mobile terminal further includes a terminal orientation detection unit that detects the orientation of the mobile terminal and transmits the detected orientation to the server; the multi-viewpoint image generation unit generates a portrait or landscape multi-viewpoint image in accordance with the orientation; the information synchronization unit transmits the compressed multi-viewpoint image, the viewpoint position, and the attitude to the mobile terminal in association with each other; The three-dimensional video display system according to claim 1, characterized in that the element image group generation unit generates an element image group from the multi-viewpoint image corresponding to the viewpoint position and the posture received from the server in response to the compressed multi-viewpoint image.

4. A portable terminal in a three-dimensional video display system including: a portable terminal that allows an observer to view a three-dimensional video corresponding to a viewpoint position; and a server that generates a multi-viewpoint image corresponding to the viewpoint position, a camera for photographing the observer; a viewpoint position detection unit that detects a viewpoint position of the observer from a face image of the observer captured by the camera and transmits the viewpoint position to the server; an image expansion unit that receives from the server a compressed multi-viewpoint image that is a multi-viewpoint image generated and compressed by the server in correspondence with the viewpoint position, in association with the viewpoint position, and expands the compressed multi-viewpoint image into a multi-viewpoint image; an element image group generation unit that generates, from the multi-viewpoint image, an element image group corresponding to the viewpoint positions received from the server in response to the compressed multi-viewpoint image; a three-dimensional display that displays the elemental images; A mobile terminal comprising:

5. A program for causing a computer to function as the mobile terminal according to claim 4.

6. A server in a three-dimensional video display system including a mobile terminal that allows an observer to view a three-dimensional video corresponding to a viewpoint position, and a server that generates a multi-viewpoint image corresponding to the viewpoint position, a multi-viewpoint image generating unit that receives the viewpoint position from the mobile terminal and generates a multi-viewpoint image of a three-dimensional model corresponding to the viewpoint position; an image compression unit that compresses the multi-viewpoint image to generate a compressed multi-viewpoint image; an information synchronization unit that associates the compressed multi-viewpoint image with the viewpoint position and transmits the image to the mobile terminal; A server comprising:

7. A program for causing a computer to function as the server according to claim 6.

Citation Information

Patent Citations

  • Integral stereoscopic display system and method thereof

    JP2022113478A