Stereoscopic Image Generation Device and Its Program

The stereoscopic image generation device addresses the challenge of implementing virtual camera work in three-dimensional video by converting user operations into camera parameters, enhancing immersion and reducing eye fatigue through high-quality three-dimensional image display.

JP7701843B2Active Publication Date: 2025-07-02NIPPON HOSO KYOKAI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021155217
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-24
Publication Date
2025-07-02
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

Existing interactive content systems lack a straightforward method to implement camera work for virtual cameras in three-dimensional video production, particularly in applications like video games, which are often viewed for extended periods, leading to potential eye fatigue.

Method used

A stereoscopic image generation device that includes an operation signal conversion unit, a multi-viewpoint image generation unit, and a stereoscopic image conversion unit, which converts user operations into camera parameters for a virtual camera array to generate and display stereoscopic images.

Benefits of technology

Enables easy reflection of user operations in virtual camera work, reducing eye fatigue by providing high-quality, immersive three-dimensional images with improved depth range and reduced spatial resolution issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701843000005
    Figure 0007701843000005
  • Figure 0007701843000006
    Figure 0007701843000006
  • Figure 0007701843000007
    Figure 0007701843000007
Patent Text Reader

Abstract

To provide a stereoscopic image generation device which can easily realize a camera work of a virtual camera.SOLUTION: A stereoscopic image generation device 3 comprises: an operation signal conversion unit 31 which converts an input operation signal into a position, posture and size of a virtual display on the basis of a prescribed conversion rule; a multi-viewpoint image generation unit 32 which calculates a camera parameter of a virtual camera array with a rigid body transformation matrix in accordance with the position, posture and size of the virtual display, and images the virtual display with the virtual camera array to generate a multi-viewpoint image; and a stereoscopic image conversion unit 33 which converts the multi-viewpoint image generated by the multi-viewpoint image generation unit 32 into a stereoscopic image.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a stereoscopic image generation device for generating a stereoscopic image as interactive content and a program thereof.

Background Art

[0002] Interactive content using video is utilized in various fields such as entertainment and education. Currently, a typical example of interactive content that has become widespread is video games. Video games are very popular because they can be easily enjoyed alone or with multiple people. In recent years, with the spread of the Internet, it has become possible to enjoy video games together with people in remote locations while staying at their respective homes. Besides video games, interactive content using video is also used in various scenes of daily life.

[0003] Note that interactive content refers to content that reflects user operations, that is, content that has two-way interaction between the user and the producer. For example, examples of interactive content include digital signage and two-way TVs in addition to video games.

[0004] Recently, three-dimensional video has attracted attention as a video medium. Since three-dimensional video has binocular parallax and motion parallax, it can provide a stronger sense of immersion and presence than general two-dimensional video. Especially in interactive content such as video games, immersion and presence are very important elements. Against the above background, a device that can experience interactive content using three-dimensional video is desired.

[0005] Various three-dimensional display methods have been proposed. Here, for interactive content such as video games, since it is often watched continuously for a long time, it is preferable to apply a three-dimensional method that causes less eye fatigue. Therefore, a volumetric image reproduction method capable of displaying a natural three-dimensional image has been proposed (Non-Patent Documents 1 and 2). This volumetric image reproduction method is said to be less likely to cause eye fatigue due to the mismatch between the convergence and adjustment of the eyeballs because it can reproduce light rays almost the same as those emitted by real objects (Non-Patent Document 3).

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0007] In the above-described interactive content, it is necessary to realize the camera work of a virtual camera that shoots a three-dimensional video in a virtual space for video production. However, in interactive content, no prior art for easily realizing the camera work of a virtual camera has been proposed.

[0008] Therefore, an object of the present invention is to provide a stereoscopic image generation device and a program thereof that can easily realize the camera work of a virtual camera.

Means for Solving the Problems

[0009] In order to solve the above problems, a stereoscopic image generation device according to the present invention is a stereoscopic image generation device that generates a stereoscopic image of a spatial image reproduction method as interactive content in which a user's operation is reflected, and includes an operation signal conversion unit, a multi-viewpoint image generation unit, and a stereoscopic image conversion unit.

[0010] According to such a configuration, the operation signal conversion unit receives an operation signal input by the user and converts the input operation signal into the position, posture, and size of a spatial image reproduction type virtual display that virtually displays a 3D CG scene based on a predetermined conversion rule.

[0011] The multi-viewpoint image generation unit calculates the camera parameters of a virtual camera array in which virtual cameras are arranged by a rigid body transformation matrix according to the position, posture, and size of the virtual display, and generates a multi-viewpoint image by shooting the virtual display on which the 3D CG scene is displayed with the virtual camera array. The stereoscopic image conversion unit converts the multi-viewpoint image generated by the multi-viewpoint image generation unit into a stereoscopic image.

[0012] The camera work of the virtual camera is considered to be equivalent to changing the position, orientation, and size of the virtual display. Therefore, the stereoscopic image generation device reflects the user's operation in the position, orientation, and size of the virtual display, and photographs this virtual display with a virtual camera array. In this way, the stereoscopic image generation device can easily reflect the user's operation in the camera work of the virtual camera.

[0013] Note that the present invention can also be realized by a program for causing a computer to function as the stereoscopic image generation device described above.

Effects of the Invention

[0014] According to the present invention, it is possible to provide a stereoscopic image generation device and a program thereof that can easily realize the camera work of a virtual camera.

Brief Description of the Drawings

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Embodiments for Carrying Out the Invention

[0016] Hereinafter, each embodiment of the present invention will be described with reference to the drawings. However, each embodiment described below is for embodying the technical idea of the present invention, and the present invention is not limited to the following unless specifically described. Also, the same means may be denoted by the same reference numerals, and the description may be omitted.

[0017] [System Configuration of Content Experience System] Before detailing each embodiment, the overall configuration of an interactive content experience system will be described. Hereinafter, the "interactive content experience system" will be simply abbreviated as the "content experience system". As described below, the content experience system has various system configurations according to the number of users and the presence or absence of a viewpoint tracking function.

[0018] <First Configuration: Single User, No Viewpoint Tracking Function> Referring to FIG. 1, the system configuration when the user U is alone and there is no viewpoint tracking function will be described. The content experience system 1 is a system for the user U to experience interactive content such as games. The content experience system 1 in FIG. 1 is the most basic system configuration among the system configurations described below. As shown in FIG. 1, the content experience system 1 includes a controller 2, a stereoscopic image generation device 3, and a stereoscopic display 4. Here, the controller 2, the stereoscopic image generation device 3, and the stereoscopic display 4 are connected to each other by a wired cable or wireless communication.

[0019] The controller 2 is for the user U to perform operations on interactive content. When the interactive content is a video game, the user U operates the player character using the controller 2. For example, the controller 2 is a game pad equipped with a stick and buttons. Also, the controller 2 outputs an operation signal by the user to the stereoscopic image generation device 3.

[0020] The stereoscopic image generation device 3 generates a stereoscopic image in a spatial image reproduction method as interactive content (for example, a video game) in which the operations of the user U are reflected. In the present embodiment, the stereoscopic image generation device 3 generates an integral type stereoscopic image (element image group) (Non-Patent Documents 1, 2, Reference Document 1). Then, the stereoscopic image generation device 3 outputs the generated stereoscopic image to the stereoscopic display 4.

[0021] Reference Document 1: Mikoshiba, "Outline of the Integral Method", Japan Broadcasting Corporation, NHK Technical Research and Development / No. 144, March 2014

[0022] The stereoscopic display 4 is a spatial image reproduction type display that displays the stereoscopic image input from the stereoscopic image generation device 3. In the present embodiment, the stereoscopic display 4 includes a lens array and a display element (not shown) similar to a general integral type display. For example, the lens array has element lenses corresponding to the elemental images arranged in a two-dimensional direction. For example, as the display element, general flat panel displays such as a liquid crystal display, an organic EL (Electro Luminescence) display, and a plasma display can be mentioned.

[0023] <Second Configuration: Multiple Users, No Viewpoint Tracking Function> Referring to FIGS. 2 and 3, the system configuration when there are multiple users U and no viewpoint tracking function will be described. In the spatial image reproduction type display, not only the main lobe which is the viewing area formed in the front thereof, but also the side lobes formed on both sides of the main lobe enable stereoscopic viewing. As shown in FIG. 2, in the content experience system 1, for two users U (U1, U2), a common 3D CG scene M can be displayed using the side lobes, and thus it is suitable for, for example, a fighting game. When there are multiple users U, as shown in FIG. 3, the content experience system 1 includes a controller 2 (21,..., 2 N ) for each user U, while only one stereoscopic display 4 common to all users U is provided. Here, N is an integer of 1 or more indicating the number of users U. In other respects, since the content experience system 1 of the second configuration is the same as the first configuration, further description is omitted.

[0024] <Third Configuration: Single User, With Viewpoint Tracking Function> Referring to FIG. 4, the system configuration when there is one user U and the viewpoint tracking function will be described. As shown in FIG. 4, the content experience system 1B includes a controller 2, a stereoscopic image generation device 3B, a stereoscopic display 4, and a camera 5. Since the parts other than the stereoscopic image generation device 3B and the camera 5 are the same as those in the first configuration, the description thereof is omitted.

[0025] The stereoscopic image generation device 3B generates a stereoscopic image in a spatial image reproduction method, similar to the stereoscopic image generation device 3 in FIG. 1. At this time, the stereoscopic image generation device 3B detects the viewpoint position (both-eye position) of the user U by the viewpoint tracking function, and displays the stereoscopic image only in the viewing area of the user U. With this viewpoint tracking function, in the stereoscopic image generation device 3B, the stereoscopic display 4 can use a lens array with a long focal length, and the depth range determined by the light density can be expanded. Furthermore, the stereoscopic image generation device 3B can also expand the viewing area by dynamically tracking the viewing area to the user U.

[0026] The camera 5 is a camera for detecting the viewpoint position. For example, the camera 5 is a general web camera that captures an image of the face area of the user U. Then, the camera 5 outputs the captured face area image to the stereoscopic image generation device 3B.

[0027] <Fourth Configuration: Multiple Users, with Viewpoint Tracking Function> Referring to FIG. 5, the system configuration when there are multiple users U and the viewpoint tracking function will be described. As shown in FIG. 5, the content experience system 1B includes a controller 2 (21,..., 2 N ) and a stereoscopic display 4 (41,..., 4 N ) and a camera 5 (51,..., 5 N ) for each user U. The stereoscopic image generation device 3B in FIG. 5 generates a stereoscopic image for each user U. In other respects, since the content experience system 1B of the fourth configuration is the same as that of the third configuration, the description is omitted.

[0028] As described above, the content experience systems 1 and 1B can accommodate the number of users U and the presence or absence of the viewpoint tracking function with only simple configuration changes, so the versatility can be improved. That is, the content experience systems 1 and 1B can flexibly change the system configuration according to the number of users U and the desired video quality.

[0029] Furthermore, the spatial image reproduction type stereoscopic display 4 can reproduce light rays that are almost the same as the light rays emitted by an object, and it is said that eye fatigue due to the inconsistency between the convergence and adjustment of the eyes is less likely to occur (Non-Patent Document 3). In this way, the content experience systems 1 and 1B can reduce eye fatigue even in interactive content with a lot of long-term viewing such as games by adopting the spatial image reproduction type stereoscopic display 4 that can display natural stereoscopic images.

[0030] Here, in a conventional TV game using 2D video, the player character is only displayed on a flat screen, and visually, it is impossible to feel the thickness of the player character's body, jump out of the screen, or move deeper. That is, it is difficult to strongly feel the sense of immersion and presence in a conventional TV game using 2D video. On the other hand, since the content experience systems 1 and 1B can obtain a three-dimensional effect compared to a 2D video TV game, in addition to the improvement of the sense of immersion and the presence of the character, new effects using binocular parallax and motion parallax are also possible.

[0031] (First Embodiment) [Configuration of Stereoscopic Image Generation Device] With reference to FIG. 6, the configuration of the stereoscopic image generation device 3 will be described. As shown in FIG. 6, the stereoscopic image generation device 3 includes a storage unit 30, an operation signal conversion unit 31, a multi-viewpoint image generation unit 32, and a stereoscopic image conversion unit 33. Note that the stereoscopic image generation device 3 in FIG. 6 corresponds to the system configuration (the first configuration in FIG. 1) when the user U is alone and there is no viewpoint following function.

[0032] The storage unit 30 is a storage device such as a memory, a HDD (Hard Disk Drive), or an SSD (Solid State Drive) that stores the 3D CG scene M of the interactive content. Here, it is assumed that the user of the content experience system 1 stores a desired 3D CG scene M in the storage unit 30 in advance. For example, in the case of a combat fighting game, the 3D CG scene M is composed of two or more player character models and background models. In the following description, the processing related to the camera work of the virtual camera will be mainly described, and the description of the general production method of interactive content will be omitted.

[0033] The operation signal conversion unit 31 receives an operation signal from the user U and converts the input operation signal into the position, orientation, and size of a space image reproduction type virtual display that virtually displays the 3D CG scene M based on a predetermined conversion rule. In the present embodiment, the operation signal conversion unit 31 receives an operation signal from the controller 2 operated by each user U and converts this operation signal into the position, orientation, and size of the virtual display. Then, the operation signal conversion unit 31 outputs the converted position, orientation, and size of the virtual display to the multi-viewpoint image generation unit 32.

[0034] <Virtual Display and Virtual Camera in 3D CG Scene> Hereinafter, the virtual display 90 and the virtual camera array 91 in the 3D CG scene M will be described.

[0035] As shown in FIG. 7, a spatial image reproduction type virtual display 90 is arranged in a 3D CG scene M to be displayed as a stereoscopic image. Then, by photographing this virtual display 90 with a virtual camera array 91, a stereoscopic image (element image E) can be generated. For example, the virtual display 90 is a virtual display including a display and a lens array, similar to the integral method. Further, the virtual camera array 91 has a plurality of virtual cameras 92 arranged in a two-dimensional direction. Here, the virtual camera array 91 is arranged at equal intervals at a position separated by a preset viewing distance D in the depth direction from the display surface (virtual display surface) of the virtual display 90. Note that the viewing distance D is the distance from the user U to the virtual display 90.

[0036] In FIG. 7, the virtual camera array 91 is composed of five virtual cameras 921 to 925 arranged in the horizontal direction. It is assumed that all the virtual cameras 921 to 925 are arranged facing the virtual display 90, and the internal parameters are preset so that the viewing angle can coincide with the area of the virtual display 90 for photographing.

[0037] Here, the midpoint (viewpoint position P) of both eyes of the user U coincides with the center of the virtual camera array 91. Also, in the depth direction, the center of the virtual camera array 91 coincides with the center of the virtual display 90. In FIG. 7, the horizontal direction is the x-axis, the vertical direction is the y-axis, and the depth direction is the z-axis.

[0038] When shooting a general 2D video, camera work consists of dolly (position change), pan·tilt·roll (posture change), and zoom in·zoom out (focal length change). When performing these operations in a 3D virtual space, the camera work is equivalent to the operation of the virtual display 90. As shown in Fig. 8(a), the dolly of the virtual camera 92 corresponds to the position change of the virtual display 90. In Fig. 8, a person and a dog are illustrated as the 3D CG scene M. Also, the field of view angle α of the virtual camera 92 before the camera work is illustrated by a dashed line, and the field of view angle β of the virtual camera 92 after the camera work is illustrated by a solid line. That is, the field of view angles α and β of the virtual camera 92 represent the changes in the position, posture, and size of the virtual display 90 before and after the camera work. As shown in Fig. 8(b), the pan, tilt, or roll of the virtual camera 92 corresponds to the posture change of the virtual display 90. As shown in Fig. 8(c), the zoom in or zoom out of the virtual camera 92 corresponds to the reduction or enlargement of the size of the virtual display 90.

[0039] As described above, the camera work of the virtual camera 92 is considered to be equivalent to the operation of the virtual display 90. Therefore, the operation signal conversion unit 31 converts the operation signal from the controller 2 into the operation (position, posture, and size) of the virtual display 90 in order to reflect the operation by the user U in the interactive content.

[0040] The conversion rule for converting the operation signal of the user U into the operation of the virtual display 90 can be arbitrarily set. For example, the conversion rule may be a rule that treats the operation of the user U as the operation of the virtual display 90 as it is in the case of the first-person perspective or the third-person perspective. Also, the conversion rule may be a rule that switches to the perspective corresponding to the position when the user U moves the player character to a predetermined position in the game stage in the case of a fixed perspective within the game stage.

[0041] Returning to Fig. 6, the description of the stereoscopic image generation device 3 will be continued. The multi-viewpoint image generation unit 32 calculates the camera parameters of the virtual camera array 91 with the virtual cameras 92 arranged according to the position, orientation, and size of the virtual display 90 by means of a rigid body transformation matrix. Then, the multi-viewpoint image generation unit 32 generates a multi-viewpoint image by photographing the virtual display 90 on which the 3D CG scene M is being displayed with the virtual camera array 91.

[0042] <Generation of Multi-Viewpoint Image Reflecting Camera Work> With reference to Fig. 9, the generation of a multi-viewpoint image reflecting camera work will be described. The camera parameters of the virtual camera 92 consist of its position, orientation, and internal parameters. Also, the camera parameters of the virtual camera 92 change according to the position, orientation, and size of the virtual display 90. Hereinafter, when generating the multi-viewpoint image V, the position and orientation of the virtual display 90 and the size of the virtual display 90 will be considered separately.

[0043] Fig. 9 shows each coordinate system within the virtual space. The world coordinate system Σ W is a 3D coordinate system serving as a reference within the virtual space, with a desired position as the origin. Also, the virtual display coordinate system Σ D is a 3D coordinate system with the virtual display 90 as a reference, for example, with the center of the virtual display surface as the origin. Also, the camera coordinate system Σ C is a coordinate system with one virtual camera 92 constituting the virtual camera array 91 as a reference, for example, with the optical principal point of the virtual camera 92 as the origin.

[0044] Generally, coordinate systems are convertible by means of a rigid body transformation matrix. Here, the transformation from the coordinate system Σ A to the coordinate system Σ B is represented by the rigid body transformation matrix B M A In this case, the coordinates A X of the 3D point in the coordinate system Σ A are, according to the following equation (1), the 3D coordinates B in the coordinate system Σ BIt can be converted into X. In the three-dimensional coordinates X, the upper-left subscript represents the coordinate system, and the tilde above represents homogeneous coordinates.

[0045]

Mathematics

[0046] Also, for the three-dimensional coordinates X = [X, Y, Z] T the homogeneous coordinates are represented by the following equation (2). Here, the subscript T represents the transpose of the vector.

[0047]

Mathematics

[0048] In FIG. 9, the rigid body transformation matrix D M W represents the coordinate transformation from the world coordinate system Σ W to the virtual display coordinate system Σ D . Also, the rigid body transformation matrix C M D represents the coordinate transformation from the virtual display coordinate system Σ D to the camera coordinate system Σ C .

[0049] First, consider the position and orientation of the virtual display 90. The position D P of the camera in the virtual display coordinate system Σ D can be converted into the position W P of the camera in the world coordinate system Σ W as shown in the following equation (3).

[0050]

Mathematics

[0051] Here, if no camera work is performed, the rigid body transformation matrix D M Wremains a fixed value without change. However, as described above, since camera work is performed, the rigid body transformation matrix D M W changes according to the position and orientation of the virtual display 90. For example, since the rigid body transformation matrix D M W is represented by a rotation matrix and a translation vector, the translation vector of the rigid body transformation matrix D M W is changed according to the position of the virtual display 90, and the rotation matrix of the rigid body transformation matrix D M W is changed according to the orientation of the virtual display 90.

[0052] Next, consider the size of the virtual display 90. Changing the size of the virtual display 90 is equivalent to changing the lens pitch and focal length of the virtual display 90. Also, the viewing distance changes proportionally accordingly. Therefore, the position of the virtual camera 92 in the virtual display coordinate system Σ D changes.

[0053] As shown in FIGS. 10 and 11, when the size of the virtual display 90 changes, the position of the virtual camera 92 changes. As shown in FIG. 10, assume that the virtual display 90 has a predetermined size, and the virtual cameras 921,..., 92 K ..., 92 N are arranged at regular intervals (where 1 < K < N). Also, assume that the viewing distance is D and the camera interval of each virtual camera 92 is W. At this time, in the depth direction, the center of the virtual camera array 91 (virtual camera 92 K ) coincides with the center of the virtual display 90. As shown in FIG. 11, when the size of the virtual display 90 is multiplied by a, the viewing distance becomes aD and the camera interval of each virtual camera 92 becomes aW. At this time, the orientation and internal parameters of the virtual camera 92 do not change.

[0054] As described above, the multi-viewpoint image generation unit 32 calculates camera parameters for each virtual camera 92 according to the position, orientation, and size of the virtual display 90 by using the rigid transformation matrix of Equation (3). Then, the multi-viewpoint image generation unit 32 generates a multi-viewpoint image by photographing the virtual display 90 with each virtual camera 92, and outputs the generated multi-viewpoint image to the stereoscopic image conversion unit 33.

[0055] Returning to FIG. 6, the description of the stereoscopic image generation apparatus 3 will be continued. The stereoscopic image conversion unit 33 converts the multi-viewpoint image generated by the multi-viewpoint image generation unit 32 into a stereoscopic image. In the present embodiment, as shown in FIG. 12, the stereoscopic image conversion unit 33 converts the multi-viewpoint image V into an elemental image E by rearranging the pixel positions of the multi-viewpoint image V input from the multi-viewpoint image generation unit 32. Then, the stereoscopic image conversion unit 33 outputs the converted elemental image E to the stereoscopic display 4. When this elemental image E is displayed on the stereoscopic display 4, the user U can view a stereoscopic image.

[0056] [Operation of the Stereoscopic Image Generation Apparatus] Referring to FIG. 13, the operation of the stereoscopic image generation apparatus 3 will be described. Here, it is assumed that the 3D CG scene M is stored in the storage unit 30 in advance.

[0057] In step S1, the operation signal conversion unit 31 receives an operation signal from the user U and converts the input operation signal into the position, orientation, and size of the virtual display 90 based on a predetermined conversion rule. For example, the conversion rule may be a rule that treats the operation of the user U as the operation of the virtual display 90 as it is in the case of the first-person viewpoint or the third-person viewpoint. Also, the conversion rule may be a rule that switches to a viewpoint corresponding to the position when the user U moves the player character to a predetermined position in the game stage in the case of a fixed viewpoint within the game stage.

[0058] In step S2, the multi-viewpoint image generation unit 32 calculates the camera parameters of the virtual camera array 91 in which the virtual cameras 92 are arranged by a rigid body transformation matrix according to the position, orientation, and size of the virtual display 90. Here, the multi-viewpoint image generation unit 32 calculates the camera parameters for each virtual camera 92 according to the position, orientation, and size of the virtual display 90 by using the rigid body transformation matrix of the above-described formula (3). Then, the multi-viewpoint image generation unit 32 generates a multi-viewpoint image V by photographing the virtual display 90 with each virtual camera 92.

[0059] In step S3, the stereoscopic image conversion unit 33 converts the multi-viewpoint image V generated by the multi-viewpoint image generation unit 32 into a stereoscopic image. In the present embodiment, as shown in FIG. 12, the stereoscopic image conversion unit 33 rearranges the pixel positions of the multi-viewpoint image V input from the multi-viewpoint image generation unit 32 to convert the multi-viewpoint image V into an elemental image E.

[0060] [Operation and Effect] The camera work of the virtual camera 92 is considered to be equivalent to the change in the position, orientation, and size of the virtual display. Therefore, the stereoscopic image generation apparatus 3 according to the first embodiment reflects the operation of the user U in the position, orientation, and size of the virtual display 90, and photographs this virtual display 90 with the virtual camera array 91. Thereby, the stereoscopic image generation apparatus 3 can easily reflect the operation of the user U in the camera work of the virtual camera for the production of interactive contents such as video games.

[0061] (Second Embodiment) [Configuration of Stereoscopic Image Generation Apparatus] Referring to FIG. 14, the differences from the first embodiment in the configuration of the stereoscopic image generation apparatus 3 according to the second embodiment will be described. The integral type three-dimensional display 4 has the characteristic that the spatial resolution decreases as the distance from the display surface increases. In this case, when trying to reproduce an object away from the display surface, it is difficult to visually recognize it correctly, resulting in a very difficult-to-view three-dimensional image. That is, the depth range of the three-dimensional CG scene M that can be reproduced as a three-dimensional image by the three-dimensional display 4 is narrow. Therefore, the three-dimensional image generation device 3 is different from the first embodiment in that it compresses the depth of the three-dimensional CG scene M according to the three-dimensional display 4.

[0062] As shown in FIG. 14, the three-dimensional image generation device 3 includes a storage unit 30, an operation signal conversion unit 31, a multi-viewpoint image generation unit 32, a three-dimensional image conversion unit 33, and a depth compression unit 34. Since each means other than the multi-viewpoint image generation unit 32 and the depth compression unit 34 is the same as in the first embodiment, the description thereof is omitted.

[0063] The depth compression unit 34 compresses the depth of the three-dimensional CG scene M using a preset depth compression function. In this embodiment, the depth compression unit 34 compresses the three-dimensional CG scene M stored in the storage unit 30. Here, the depth compression unit 34 may use either a linear depth compression function or a non-linear depth compression function (References 2, 3). At this time, it is set in the depth compression unit 34 which depth compression function the user of the content experience system 1 uses. The linear depth compression function is a function that uniformly compresses the entire three-dimensional CG scene M in the depth direction and is represented by, for example, a linear function. The non-linear depth compression function is a function that compresses the three-dimensional CG scene M so that the compression rate becomes smaller in the range where the spatial resolution is high in the depth direction and is represented by, for example, a sigmoid function. Then, the depth compression unit 34 outputs the three-dimensional CG scene M after depth compression to the multi-viewpoint image generation unit 32.

[0064] Reference 2: Y. Sawahata and T. Morita, “Estimating Depth Range Required for 3-D Displays to Show Depth-Compressed Scenes Without Inducing Sense of Unnaturalness”, IEEE Transactions on Broadcasting, Vol. 64, No. 2, pp. 488 - 497 (2018) Reference 3: Sawahata, “Integral Stereo Display Technology Based on Spatial Cognition Characteristics”, Japan Broadcasting Corporation, NHK Technical Research and Development / No. 164, August 2017

[0065] The multi-viewpoint image generation unit 32 captures a virtual display 90 on which a three-dimensional CG scene M after depth compression is being displayed with a virtual camera array 91. In the present embodiment, since the multi-viewpoint image generation unit 32 is the same as that in the first embodiment except that the three-dimensional CG scene M is input from the depth compression unit 34, further description thereof is omitted.

[0066] [Operation and Effect] As described above, similar to the first embodiment, the stereoscopic image generation device 3 according to the second embodiment can easily reflect the operation of the user U in the camera work of the virtual camera. Furthermore, since the stereoscopic image generation device 3 compresses the three-dimensional CG scene M in depth, even for interactive content having a wide depth range, a decrease in spatial resolution can be suppressed and high-quality stereoscopic images can be displayed.

[0067] [Modification Example 1] In the first and second embodiments, the content experience system 1 is described as including only one controller 2 and one stereoscopic display 4 and corresponding to the system configuration of FIG. 1, but the present invention is not limited thereto. For example, the content experience system 1 may be provided with a plurality of controllers 2 and may have a system configuration corresponding to a plurality of users U (the second configuration in FIG. 3). In this case, the operation signal conversion unit 31 averages the operation signals from each controller 2 and converts the operations by the plurality of users U into camera work. Further, the operation signal conversion unit 31 may convert the operation signal into camera work so that the operation signal of the user U who has performed the operation first is reflected.

[0068] (Third Embodiment) [Configuration of Stereo Image Generation Device] Referring to FIG. 15, the differences between the configuration of the stereo image generation device 3B according to the third embodiment and the first embodiment will be described. The stereo image generation device 3B according to the third embodiment is different from the first embodiment in that it has a viewpoint tracking function. Note that the stereo image generation device 3B in FIG. 15 corresponds to a system configuration (the third configuration in FIG. 4) when the user U has a viewpoint tracking function alone.

[0069] As shown in FIG. 15, the stereo image generation device 3B includes a storage unit 30, an operation signal conversion unit 31, a multi-viewpoint image generation unit 32B, a stereo image conversion unit 33B, and a viewpoint position detection unit 35. Since each means other than the multi-viewpoint image generation unit 32 and the viewpoint position detection unit 35 is the same as that in the first embodiment, the description thereof is omitted.

[0070] The viewpoint position detection unit 35 detects the viewpoint position for each user U. In the present embodiment, the intermediate position between the two eyes of the user U is detected as the viewpoint position from the face region image of each user U input from the camera 5 using a known method (References 4, 5). In the present embodiment, since there is one user U, the viewpoint position detection unit 35 detects the viewpoint position for one person. Then, the viewpoint position detection unit 35 outputs the detected viewpoint position of the user U to the multi-viewpoint image generation unit 32B.

[0071] Reference 4: Japanese Unexamined Patent Application Publication No. 2019-213127 Reference 5: Okoshi, "Development of a viewpoint-following integral 3D video display system", Journal of the Institute of Image Information and Television Engineers, Vol. 75, No. 1, pp. 125-130 (2021)

[0072] The multi-viewpoint image generation unit 32B changes the rigid transformation matrix according to the viewpoint position input from the viewpoint position detection unit 35, and generates a multi-viewpoint image V for each user U using the rigid transformation matrix for each user U. Here, the multi-viewpoint image generation unit 32B obtains the rigid transformation matrix from the viewpoint positions of each user U, respectively, and generates a multi-viewpoint image V corresponding to the viewpoint position of each user U. In this embodiment, since there is one user U, the multi-viewpoint image generation unit 32B detects a multi-viewpoint image for one person.

[0073] <Generation of multi-viewpoint images according to the viewpoint position> Referring to FIG. 16, the generation of the multi-viewpoint image V according to the viewpoint position will be described. When there is a viewpoint following function, the positions and internal parameters of each virtual camera 92 of the virtual camera array 91 change according to the viewpoint position of the user U. As shown in FIG. 16, when the viewpoint position P of the user U moves, the positions and internal parameters of each virtual camera 92 of the virtual camera array 91 are set according to the moved viewpoint position P.

[0074] Specifically, when there is a viewpoint following function, the rigid transformation matrix D from the virtual display coordinate system Σ C to the camera coordinate system Σ C M D changes. At this time, the position of the virtual camera 92 changes, but the orientation of the virtual camera 92 does not change. That is, in the rigid transformation matrix C M D , as shown in the following equation (4), the translation vector C T changes according to the viewpoint position, while the rotation matrix C R D becomes the identity matrix.

[0075]

Equation

[0076] From the above, the multi-viewpoint image generation unit 32B uses Equation (4) to change the rigid transformation matrix C M D according to the viewpoint position of the user U, and sets the virtual camera array 91 for each user U. After that, the multi-viewpoint image generation unit 32B generates a multi-viewpoint image in the same manner as in the first embodiment, and outputs the generated multi-viewpoint image to the stereoscopic image conversion unit 33B.

[0077] Returning to FIG. 15, the description of the configuration of the stereoscopic image generation apparatus 3B will be continued. The stereoscopic image conversion unit 33B converts a multi-viewpoint image into a stereoscopic image for each user U. Here, since there is one user U, the stereoscopic image conversion unit 33B converts a stereoscopic image for one person. Note that since the method of converting the stereoscopic video by the stereoscopic image conversion unit 33B is the same as that in the first embodiment, the description thereof is omitted. Then, the stereoscopic image conversion unit 33B outputs the stereoscopic image (element image E) generated for each user U to the stereoscopic display 4 corresponding to each user U.

[0078] [Operation and Effect] As described above, the stereoscopic image generation apparatus 3B according to the third embodiment can easily reflect the operation of the user U in the camera work of the virtual camera, similar to the first embodiment. Furthermore, the stereoscopic image generation apparatus 3B can display a stereoscopic image only in the region where both eyes of the user U are located by the viewpoint tracking function. As a result, since the stereoscopic display 4 can use a lens array with a long focal length, the depth range can be expanded by the light ray density. Furthermore, the stereoscopic image generation apparatus 3B can expand the viewing field by the viewpoint tracking function and display a high-quality stereoscopic image.

[0079] (Modification 2) Note that the viewpoint position detection unit 35 may determine whether there are a plurality of users U based on the number of detected viewpoint positions. For example, the viewpoint position detection unit 35 may divide the number of eyes included in the face region image by 2 (rounding down the decimal part) and use the obtained value as the number of users U.

[0080] And when it is determined that there are a plurality of users U, the multi-viewpoint image generation unit 32B may generate a multi-viewpoint image V for each user U. Further, when it is determined that there are a plurality of users U, the stereoscopic image conversion unit 33B may convert the multi-viewpoint image V into a stereoscopic image (element image E) for each user U.

[0081] In this way, it is possible to switch the presence or absence of the viewpoint following function according to the number of users U, and the content experience system 1B can have a more flexible system configuration and further improve versatility.

[0082] (Fourth Embodiment) [Configuration of Stereoscopic Image Generation Device] Referring to FIG. 17, the differences between the configuration of the stereoscopic image generation device 3B according to the fourth embodiment and the third embodiment will be described. The stereoscopic image generation device 3B according to the fourth embodiment is different from the third embodiment in that it compresses the depth of the 3D CG scene M.

[0083] As shown in FIG. 17, the stereoscopic image generation device 3B includes a storage unit 30, an operation signal conversion unit 31, a multi-viewpoint image generation unit 32B, a stereoscopic image conversion unit 33B, a depth compression unit 34, and a viewpoint position detection unit 35.

[0084] The multi-viewpoint image generation unit 32B captures the virtual display 90 in which the 3D CG scene M after depth compression is being reproduced with the virtual camera array 91. In this embodiment, since the multi-viewpoint image generation unit 32B is the same as that in the third embodiment except that the 3D CG scene M is input from the depth compression unit 34, further description thereof is omitted.

[0085] Note that since the storage unit 30, the operation signal conversion unit 31, the stereoscopic image conversion unit 33B, and the viewpoint position detection unit 35 are the same as those in the third embodiment, description thereof is omitted. Also, since the depth compression unit 34 is the same as that in the second embodiment, description thereof is omitted.

[0086] [Operation and Effect] As described above, the three-dimensional image generation device 3B according to the third embodiment can easily reflect the operation of the user U in the camera work of the virtual camera, similar to the third embodiment. Furthermore, similar to the third embodiment, the three-dimensional image generation device 3B can expand the depth range, expand the viewing field by the viewpoint tracking function, and display a high-quality three-dimensional image. Furthermore, similar to the second embodiment, the three-dimensional image generation device 3B can suppress a decrease in spatial resolution and display a high-quality three-dimensional image.

[0087] (Modification 3) Note that in the third and fourth embodiments, the content experience system 1B is described as including only one controller 2 and one three-dimensional display 4 and corresponding to the system configuration of FIG. 4, but the present invention is not limited to this. For example, the content experience system 1B may include a plurality of controllers 2 and three-dimensional displays 4 and have a system configuration corresponding to a plurality of users U (the fourth configuration in FIG. 5).

[0088] (Other modifications) Although each embodiment has been described in detail above, the present invention is not limited to the above-described embodiments, and also includes design changes and the like within a range not departing from the gist of the present invention. In each of the above-described embodiments, the spatial image reproduction method has been described as the integral method, but the present invention is not limited to this. For example, the spatial image reproduction method may be the lenticular method. In this case, the three-dimensional display may include a lenticular lens array in which ridged lenticular lenses are arranged in a one-dimensional direction and a display element. In this case, the multi-viewpoint image generation unit may use a virtual camera array in which virtual cameras are arranged in a one-dimensional direction.

[0089] In each of the above-described embodiments, the stereoscopic image generation device has been described as being independent hardware, but the present invention is not limited to this. For example, the present invention can also be realized by a program for causing hardware resources such as a CPU, a memory, and a hard disk included in a computer to function as the above-described stereoscopic image generation device. This program may be distributed via a communication line, or may be written on a recording medium such as a CD-ROM or a flash memory and distributed.

Explanation of Signs

[0090] 1,1B Content experience system 2 Controller 3,3B Stereoscopic image generation device 4 Stereoscopic display 5 Camera 30 Storage unit 31 Operation signal conversion unit 32,32B Multi-viewpoint image generation unit 33,33B Stereoscopic image conversion unit 34 Depth compression unit 35 Viewpoint position detection unit 90 Virtual display 91 Virtual camera array 92 Virtual camera E Element image M 3D CG scene V Multi-viewpoint image

Claims

1. A stereoscopic image generation device that generates a stereoscopic image in a spatial image reproduction method as interactive content that reflects a user's operation, comprising: an operation signal conversion unit that receives an operation signal from the user and converts the input operation signal into the position, orientation, and size of a spatial image reproduction type virtual display that virtually displays a 3D CG scene based on a predetermined conversion rule; a multi-viewpoint image generation unit that calculates camera parameters of a virtual camera array with virtual cameras arranged according to the position, orientation, and size of the virtual display, and generates a multi-viewpoint image by photographing the virtual display displaying the 3D CG scene with the virtual camera array; a stereoscopic image conversion unit that converts the multi-viewpoint image generated by the multi-viewpoint image generation unit into the stereoscopic image; A stereoscopic image generation device characterized by comprising the above.

2. further comprising a depth compression unit that compresses the depth of the 3D CG scene using a preset depth compression function, The multi-viewpoint image generation unit photographs the virtual display displaying the 3D CG scene after depth compression with the virtual camera array. The stereoscopic image generation device according to claim 1.

3. further comprising a viewpoint position detection unit that detects the viewpoint position for each user, The multi-viewpoint image generation unit changes the rigid body transformation matrix according to the viewpoint position, and generates the multi-viewpoint image for each user using the rigid body transformation matrix for each user. The stereoscopic image conversion unit converts the multi-viewpoint image into the stereoscopic image for each user. The stereoscopic image generation device according to claim 1 or claim 2.

4. The viewpoint position detection unit determines whether there are multiple users based on the number of detected viewpoint positions, When it is determined that there are multiple users, the multi-viewpoint image generation unit generates the multi-viewpoint image for each user. When it is determined that there are multiple users, the stereoscopic image conversion unit converts the multi-viewpoint image into the stereoscopic image for each user. The stereoscopic image generation device according to claim 3.

5. A program for causing a computer to function as the stereoscopic image generation device according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Program, pose derivation method, and pose derivation device

    JP2018084954A

  • Camera calibration device and camera calibration program

    JP2018189580A

  • Stereoscopic image generation device and program of the same

    JP2020170979A

  • Information processing apparatus, and information processing method

    JP2021005399A

  • Pan tilt angle calculation device and program therefor

    JP2021124395A