Adjustable stereoscopic image display for viewers by cameras sharing a common portion of the camera's field of view

The image correction process addresses the issue of unsatisfactory 3D viewing experiences in stereoscopic imaging by generating stereoscopic images and videos from a common field of view shared by two capture cameras, regardless of their orientation, resulting in an adaptable and comfortable 3D viewing experience.

JP2025518850APending Publication Date: 2025-06-19CUBICSPACE TECHNOLOGIES INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024571333
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-03
Filing Date
2023-06-01
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Conventional stereoscopic imaging techniques often result in an unsatisfactory 3D viewing experience when images acquired for a specific viewing geometry, such as a movie theater, are viewed at home, due to misalignment and non-optimal orientation of capture cameras.

Method used

An image correction process that generates stereoscopic images and videos from a common field of view shared by two capture cameras, without the need for specific camera orientations, using a method that involves generating left and right polygon meshes, defining virtual camera orientations, and applying texture mapping and rendering techniques.

Benefits of technology

This process allows for the creation of stereoscopic videos that adapt in real time to a viewer's dynamic input, overcoming differences in camera viewpoints and providing a comfortable 3D viewing experience across various camera orientations and misalignments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025518850000001_ABST
    Figure 2025518850000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image correction process that is used to generate a selected type of stereoscopic video of a common field of view shared by two cameras without the need for pre - settings and arrangements normally required for stereoscopic display. This process can use the left and right video streams of images generated from a wide range of types of cameras that share some portion of the field of view of each camera, without being limited to a particular relative orientation, i.e., convergence, parallel, or divergence. Using this process, a stereoscopic video to be displayed to a viewer can be adapted in real - time with the viewer's dynamic input being constantly updated. Using the proposed method, various devices such as web cameras, cameras, scanners, etc. that may come out of a production line with imperfections in the capture cameras can be optimized, corrected, and / or made usable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the priority of U.S. Provisional Patent Application No. 63 / 348,743, filed on June 3, 2022, the content of which is incorporated herein by reference.

[0002] This application relates to stereoscopic image display.

Background Art

[0003] In conventional stereoscopic imaging, to provide a comfortable 3D viewing experience, it is necessary to acquire left - eye and right - eye images at a reference viewing geometry, for example, the position and orientation of the cameras corresponding to a person sitting in the center of a movie theater. When a stereoscopic video acquired for a movie theater is to be viewed at home, the 3D experience is often unsatisfactory. The depth of objects may look unnatural and can cause stress or fatigue.

[0004] Alternative methods of performing image processing to re - format and / or optimize stereoscopic images have been previously presented. The applicant has previously proposed a technique for adjusting stereoscopic images acquired from parallel cameras, as described in U.S. Patent No. 10,917,623, which has been demonstrated to reduce eye stress and improve the 3D viewing experience. However, this technique imposes a restrictive constraint that the capture cameras are parallel.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Means for Solving the Problems

[0006] In this patent, the term "capture camera" refers to a physical camera used to capture images or videos of scenes selected in the real environment (i.e., not a virtual environment).

[0007] The present disclosure relates to an image correction process that is used to generate selected types of stereoscopic images and / or videos of a common field of view shared by two capture cameras without the need for settings or preparations normally required for stereoscopic display. This process can use left and right videos and / or image streams of images generated from a wide range of types of cameras that share some portion of the field of view of each camera, which is not limited to a particular relative orientation (converging, parallel, or diverging), and this process can be used to adapt in real time the stereoscopic video to be displayed to a viewer with the viewer's dynamic input always updated.

[0008] The applicant has discovered that it is possible to overcome differences in the viewpoints of images sharing a common portion of the field of view of the cameras due to misalignment and / or non-optimal orientation of the capture cameras.

[0009] A method of processing images using a graphics processing unit (GPU) to stereoscopically display to a viewer images captured by a left capture camera and a right capture camera having a common portion of the capture field of view, the method comprising: generating left and right polygon meshes using the specifications of the left and right capture cameras having a central axis; defining the positions of virtual cameras with respect to the meshes so as to correspond to the fields of view of the capture cameras; pasting textures onto the left and right polygon meshes respectively using at least one image from the left and right capture cameras to create left and right virtual 3D objects respectively; rendering left-eye and right-eye images using left and right virtual cameras respectively oriented to capture at least a portion of the left and right virtual 3D objects, wherein the orientation parameters of the left and right virtual cameras are defined by: (a) defining a reference line extending between the position of the left capture camera and the position of the right capture camera in the capture space; (b) defining the orientation in the capture space such that the left and right virtual cameras are in a mutually parallel orientation on a selected horizontal plane of view; (c) determining left and right orientation offset vectors that define the change between the orientation of the left and right virtual cameras and the orientation of the optical axes corresponding to the central axes of the left and right capture cameras; (d) using the left and right orientation vectors to align the orientations of the left and right polygon meshes with respect to the left and right virtual cameras respectively about the central axis in the virtual space; and (e) defining frustum parameters of the left and right virtual cameras using display parameters including the dimensions of the display, the distance to the display, the interocular distance, and the viewing direction.

[0010] In some embodiments, the display parameters include a magnification, and the frustum parameters are defined using the magnification.

[0011] In some embodiments, step (d) includes defining lens shift parameters of the left and right virtual cameras.

[0012] In some embodiments, the directions of the optical axes of the left and right capture cameras are in the same plane.

[0013] In some embodiments, the optical axes of the left and right capture cameras are parallel, the left and right capture cameras have different fields of view, and the reference line is not perpendicular to at least one of the optical axes of the left and right capture cameras.

[0014] In some embodiments, the raw image may be a still image.

[0015] In some embodiments, the raw image may be a video image.

[0016] In some embodiments, the left-eye image and the right-eye image are displayed to the viewer on a single screen of a stereoscopic display.

[0017] In some embodiments, the left-eye image and the right-eye image are generated as an anaglyph-formatted image, as a column-interleaved-formatted image for display on a naked-eye stereoscopic display, as a sequence of page-flip images for viewing with shutter glasses, or as a sequence of line-interleaves for a polarized display.

[0018] In some embodiments, the stereoscopic display includes a left image display and a right image display, the left-eye image is displayed to the viewer's left eye by the left image display, and the right-eye image is displayed to the viewer's right eye by the left image display.

[0019] In some embodiments, the viewer includes a plurality of viewers, and the interocular distance is selected to be the minimum interocular distance among the plurality of viewers.

[0020] In some embodiments, the relative orientation between the left capture camera and the right capture camera is divergent.

[0021] In some embodiments, the relative orientation between the left capture camera and the right capture camera is parallel.

[0022] In some embodiments, the relative orientation between the left capture camera and the right capture camera is convergent.

[0023] In some embodiments, at least one of the left-eye image and the right-eye image has an aspect ratio different from the aspect ratio defined by the dimensions of the stereoscopic display, and at least one of the left-eye image and the right-eye image is provided with a boundary line for display to the viewer.

[0024] In some embodiments, the capture field of view includes a panoramic image, and the method further includes defining a selected optical axis inside a portion of the panoramic image included in the common portion of the capture field of view, and this portion of the panoramic image and the selected optical axis will be processed by this method.

[0025] Another broad aspect is a device for processing a stereoscopic image for display to a viewer, the device comprising a graphics processing unit, a central processor, and a memory readable by the central processor, the memory storing instructions for performing the previously defined method.

[0026] Another broad aspect is a computer program product comprising a non-transitory memory storing instructions for a processor connected to a GPU for performing the previously defined method.

[0027] The present invention will be better understood from the following detailed description of embodiments of the invention with reference to the accompanying drawings.

Brief Description of the Drawings

[0028]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 1E

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 2E

Figure 2F

Figure 2G

Figure 2H

Figure 2I

Figure 2J

Figure 3A

Figure 3B

Figure 3C

Figure 4A

Figure 4B

Figure 4C

Figure 4D

Figure 4E

Figure 4F

Figure 4G

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Figure 10A

Figure 10B

Figure 11

Figure 12

Best Mode for Carrying Out the Invention

[0029] Image processing for stereoscopic display will be better understood by the following detailed description of the embodiments with reference to the accompanying drawings.

[0030] Using the method described herein, a pair of rendered stereoscopic images that can be generated from the images captured by the capture cameras, which share a common portion of the field of view of cameras that may not be intended for stereoscopic viewing, and also displaying such finally rendered stereoscopic images using a selected stereoscopic technique (such as autostereoscopy, anaglyph, line interleaving, column interleaving, etc.) and a display device, a 3D experience can be provided.

[0031] This method can include the step of generating a polygon mesh in a virtual environment corresponding to the type of lens (standard, wide-angle, fisheye, etc.) distortion specification, magnification, position and / or orientation of the capture camera for each capture camera.

[0032] In the virtual environment, using the captured images mentioned above, a texture can be pasted onto the corresponding polygon mesh to create a virtual object (3D object) that can exist in 3D. This virtual environment can enable positioning the virtual object relative to the corresponding virtual camera at a selected position and orientation according to the position specification of the capture camera, or vice versa.

[0033] Using a virtual camera, selected portions of each virtual object of the camera can be captured (projective transformation). The final stereoscopic image can be displayed to the viewer after operations that may include clipping and rendering of the selected portions of the virtual objects.

[0034] Such virtual cameras with frustums can be sized by the aspect ratio of the stereoscopic display, the interocular distance of the viewer, vergence-accommodation, the viewer's distance from the stereoscopic display, and the magnification input and zoom selected by the viewer.

[0035] The lens shift of the virtual camera can be defined by the interocular distance of the viewer, the viewer's distance from the stereoscopic display, the dimensions of the stereoscopic display, convergence / accommodation compensation, the viewer's off-axis field of view (the viewer's region of interest), or a combination thereof.

[0036] In some embodiments, the rendered images can be further processed to generate an optimized final pair of stereoscopic images by a method as disclosed in U.S. Patent No. 10,917,623B2.

[0037] Figure 1A shows a schematic top view of a scene characterized in that a sphere is located in the center of the ground, the sun at infinity (D∞) is in the center of the sky, and a cube is placed on the right, slightly further away from the sphere. In the center, the left and right eyes of an observer located at the origin of the axis and the scene are also shown as if the observer were looking at an object at infinity, i.e., the sun in this case, without the direction of convergence motion (parallel lines).

[0038] Figure 1B presents a view of the same scene as if captured by a camera placed at the origin, which is useful for comparison with what can be considered a passive visualization of the scene for subsequent stereograms.

[0039] Figures 1C and 1D respectively illustrate the common areas of the scenes observed by the left and right eyes of the observer.

[0040] Figure 1E presents a diagram in which the previous images of the scenes seen by the left eye (thinly shown here from Figure 1C) and the right eye (shown by a dashed line here from Figure 1D) of the observer are superimposed. This superimposition can be performed so that an object (the sun) on which the viewer may be focusing appears as a single object without differences, thereby showing various differences in other objects within the scene.

[0041] As pointed out by those skilled in the art, generally, practical stereoscopic images are realized by general stereoscopic techniques described in the prior art using limited types and configurations of capture cameras. For example, in order to generate a stereoscopic image similar to the anaglyph diagram of Figure 1E using capture cameras, usually, two identical capture cameras having a lens arrangement specific to stereoscopic capture / photography, and further capture cameras that need to be accurately positioned and oriented using extremely specific settings (such as the eyes of the viewer in Figure 1A) are required to obtain a practical stereoscopic image of the scene. In fact, if these criteria are not strictly considered, the viewer may misperceive the depth and / or cause harmful physiological reactions.

[0042] However, the applicant has developed an image processing technique that enables the generation of practical stereoscopic images having a geometric correspondence (the convergence point of the viewer's eyes) between reality and stereoscopic display, which allows the use of images captured using a wide range of types of capture cameras (standard, wide-angle, fisheye, etc.) without the need for limited pre-selected positioning as long as the capture cameras share the common part of each field of view.

[0043] In fact, as pointed out by those skilled in the art, the type of the capture camera and the relative positioning of the cameras may typically be specific to those considered to be optimal stereoscopic imaging conditions or practical correspondence relationships. Using some embodiments of the developed process, it is possible to generate a stereoscopic video that conforms to the viewer's perspective and dynamic parameters in real time, or to speed up some of the computational processes of existing reformatting techniques for stereoscopic images.

[0044] Here, it should be noted that a "practical" stereoscopic image can be defined as a stereoscopic image that allows the viewer to correctly recognize depth, size, scale, etc. without causing significant harmful physiological reactions such as dizziness, headache, discomfort, etc.

[0045] The diagram of FIG. 2A shows a schematic top view of a wide-angle left capture camera 20L and a standard right capture camera 20R on the same plane that record a common scene from different positions and possible orientations. In the example of FIG. 2A, the right capture camera can be located on the x-axis and have an optical axis 22R with an orientation parallel to the z-axis.

[0046] The orientation and distance of the relative reference axis 26 between the center points of the two capture cameras can be extracted from the static parameters of the capture cameras using the relative three-dimensional positions of the cameras. The relative orientations 30L or 30R of the optical axes 22L or 22R of the capture cameras with respect to the relative reference axis 26 can be measured or calculated using the capture camera parameters. Then, some of the elements included in the virtual environment can be positioned using some of the characteristics of such a real environment, including the relative orientation, relative reference axis, etc.

[0047] FIG. 2B shows an isometric view of the fields of view of the two capture cameras. As can be seen from this figure, the proposed method can enable the use of a pair of various capture cameras that do not need to have the same type of lens or field of view. In this example, the left capture camera 20L may have a field of view that is narrower than the field of view of the right capture camera 20R.

[0048] As shown in FIG. 2C, the optical axes 22L and 22R do not have to be parallel. The cameras 20 may even have a roll angle between the cameras, which means that the captured horizontal lines in the images of the cameras may be at various angles. The parameter data from the capture cameras 20L and 20R can be recorded from measurements made by the camera operator or by analyzing the images acquired from the capture cameras. The parameter data (capture camera parameters) of the capture cameras can include specifications of the viewing angle and shape (e.g., rectangular, circular, elliptical), as well as the relative orientation and relative position between the capture cameras. Note that the optical axes 22L and 22R do not have to be in the same plane.

[0049] Note that the reference axis 26 is defined herein as the line passing through the capture camera positions 20L and 20R. In the virtual space, the reference axis 26 can be used as a reference axis for positioning various virtual components (virtual cameras, meshes, 3D objects, etc.) relative to the reference axis. In a preferred embodiment, the virtual cameras can be virtually positioned such that the reference axis 26 passes through both virtual cameras.

[0050] As described below, a GPU can be used to process the capture camera images and provide an image suitable for display to the viewer. In the case of parallel cameras having a common baseline, International Application Publication No. WO 2019 / 041035, a PCT patent application publication published on March 7, 2019 by the applicant, describes the use of a graphics processing unit (GPU) to perform image processing to provide an image suitable for display to the viewer.

[0051] One way to process an image with a GPU is to create a virtual surface or object, paste a texture onto the object using the captured image, and then use a virtual camera that actually performs the desired scaling, shifting, and optionally distortion correction according to the specifications of the virtual surface or object and the virtual camera parameters to render the image. Using the GPU according to this method, real-time video processing can be easily performed.

[0052] This technique can be extended in this application to cameras where the capture properties are not limited by the capture orientation. The problem of this limitation can be solved, as also shown in FIG. 2C, by defining a reference axis 26 in the capture space and defining the position of the virtual camera to be at the location of the central axis of the capture camera on the reference axis 26 and having virtual camera axes parallel to each other. The direction of the virtual camera, i.e., the left-right angle and the up-down angle, is the direction of the line of sight. The distance between the virtual cameras 21 can correspond to the interocular distance of the viewer. This is shown in FIG. 2D, and a more detailed explanation of the angular orientation offset will be described below with reference to FIG. 2J.

[0053] FIGS. 2E and 2F schematically show how the virtual cameras 21L and 21R can be positioned relative to their respective 3D surfaces or objects 25. The 3D objects 25L and 25R are shown as flat 3D surfaces, but it will be understood that, as described below with reference to FIGS. 3B and 4A, the camera optics can result in an image that is not uniform and the corresponding 3D object 25 may not be flat.

[0054] As shown in FIGS. 2E and 2F, in the virtual space, the 3D object 25 with the texture pasted using the original captured camera image can be positioned at a predetermined location. The GPU virtual camera can be used as a tool for extracting a display image that collects views of a part of the 3D object 25. Many GPUs further support "lens shift" parameters for defining the properties of the virtual camera 21. The lens shift can potentially be a parameter that results in the frustum of the camera being skewed (i.e., asymmetric). The lens shift can, for example, shift the lens horizontally or vertically from the center. The value of the lens shift can be a multiple of the sensor size. For example, if it is shifted by 0.5 in the X-axis direction, the sensor is offset by half of the horizontal size of the sensor.

[0055] A part of the process of this embodiment is shown in FIG. 2G. This process can include step 201 of being able to read the captured camera parameters, step 202 of being able to determine the 3D object 25 from the captured camera parameters and establish it within the GPU, step 203 of being able to load the original image into the GPU and the GPU pasting the texture onto the object 25 using each original image, and step 204 of obtaining the reference axis 26 that can be pre-determined in the capture space and used in the virtual space, based on whether it may be given as part of the captured camera parameters or based on the captured camera parameters (e.g., the location of the captured camera in 3D space).

[0056] The viewer parameters can include the distance to the display (i.e., the distance between the viewer and the screen), the dimensions of the display, the interpupillary distance of the viewer, and / or the direction of the line of sight. These may be supplied or may be measured.

[0057] In some embodiments, step 204 can be used to place / position the virtual camera 20 on the reference axis 26 (e.g., at each end of the reference axis) using viewer parameters and / or capture camera parameters, and / or to determine / position the orientation of the parallel virtual camera axes 23 that may extend in the same direction from the reference axis 26. This can enable aligning the orientation of the optical axis of the virtual camera 23 within the virtual space.

[0058] In some embodiments, step 205 can be used to calculate the relative orientation between the virtual camera axis 23 and each 3D object (i.e., a 3D mesh textured using the captured image / texture).

[0059] In some embodiments, step 206 can be used to determine or calculate each frustum of the virtual camera's field of view and / or lens shift, which can be done using viewer parameters and / or capture camera parameters.

[0060] In some embodiments, step 207 can be used, where the parameters of the virtual camera 21 can be passed to the GPU, which can then process the rendering of the left and right eye images for the viewer, and the images can then be displayed to the viewer through a display screen using various methods and devices.

[0061] As shown in the exemplary embodiment of FIG. 2H, the left virtual camera 21L, the right virtual camera 21R, and the polygon mesh 25 (geometry) can be positioned using some of the features of the real environment. It will be appreciated that the previously mentioned orientation calculations may be similar to a simple change of coordinate system.

[0062] For example, in the embodiment of FIG. 2H, each of the virtual cameras 21L / 21R can be fixed to the corresponding ends of the relative reference axis 26 to position and align these virtual cameras. In some embodiments, the distance between both virtual cameras 21L / 21R can be positioned such that it has the same relative distance as the actual distance between the capture cameras 20L / 20R.

[0063] Most importantly, the virtual cameras 21L / 21R can be given optical axes 23L / 23R having an orientation perpendicular to the relative reference axis 26, and the optical axes can be parallel to each other. The polygon meshes of each camera (25L for the left camera and 25R for the right camera) can be positioned in the direction of the virtual object axis 22', which is equivalent to the optical axis 22 of the capture camera, and can be oriented by the same relative orientation 30 (30L for the left virtual camera 21L and 30R for the right virtual camera 21R) together with the relative reference axis 26.

[0064] All of these requirements can be integrated or transformed by the virtual orientation vector 28 (28L for the left virtual camera 21L and 28R for the right virtual camera 21R) between the optical axis 23 of the virtual camera 21L / 21R and the virtual object direction 22' (equivalent to the optical axis 22L / 22R of the capture camera). It will be understood that the virtual orientation vector may be in a direction equal to the difference between the orientation of the optical axis 23 of the virtual camera with respect to the relative reference axis and the relative orientation 30 of the orientation.

[0065] Therefore, in some embodiments, the virtual camera can be oriented by the virtual orientation vector 28 from the orientation (direction) of the virtual object, or similarly, the orientation (direction) of the virtual object can be oriented by the virtual orientation 28 from the virtual camera.

[0066] In some embodiments, the distance between the virtual camera 21 and the mesh 25 corresponding to the camera may be adjusted in the direction of the virtual object axis using a part of the parameters of the corresponding capture camera, taking into account the distance between the viewer and the stereoscopic display (e.g., the display screen) or the equivalent of the stereoscopic display.

[0067] One skilled in the art will understand that, in the case of the schematic example of FIG. 2H, the raw images for attaching textures to the respective polygon meshes, herein called texture 24, captured by the capture cameras can only be seen from the respective virtual cameras. In other words, in order to prevent the texture and mesh of one camera from blocking the field of view of the other camera, as shown in FIG. 2I, the left virtual camera 21L can only recognize the 3D object 39L of the left capture camera 20L (i.e., the left 3D mesh 25L textured with the left image / texture 24L), as if the left 3D mesh and the right 3D mesh were in a different virtual environment, while the right virtual camera 21R can only recognize the 3D object 39R of the right capture camera 20R (i.e., the right 3D mesh 25R textured with the right image / texture 24R).

[0068] In some embodiments, to establish the geometric correspondence between the real environment and the virtual environment, the specifications of the capture cameras 20L / 20R can be used to position the virtual cameras 21L / 21R relative to the corresponding polygon meshes 25L / 25R, and vice versa (i.e., the virtual cameras have positions equivalent to the positions of the corresponding capture cameras within the virtual environment).

[0069] In some embodiments, to establish such a geometric correspondence, the virtual cameras 21L / 21R can be positioned on an axis equivalent to the optical axis of the capture camera (e.g., an axis perpendicular to the center of the polygon mesh), which can be extracted or calculated using the specifications of the capture camera.

[0070] In some embodiments, to establish such geometric correspondences, the virtual cameras 21L / 21R can be positioned at a distance from the polygon meshes 25L / 25R such that a virtual camera having a field of view equivalent to that of the corresponding capture cameras 20L / 20R (i.e., having the same angular dimensions and extractable or calculable using the capture camera specifications) can accurately view the entire polygon mesh (i.e., neither more nor less) when facing the polygon mesh from this distance (i.e., such that the outer perimeter of this field of view exactly coincides with the sides or outer perimeter of the polygon mesh).

[0071] FIG. 2J presents a three-dimensional schematic of various elements used to calculate the virtual orientation vectors 28 for each of the left and right capture cameras, which can be used to establish geometric correspondences between the real environment and the virtual environment, as previously presented. The illustrated embodiment presents representations of the relative reference axis 26 (extending between the two capture cameras 20L / 20R), the base axis 29 perpendicular to the relative reference axis 26, and the optical axis 22 of the capture camera 20. As previously explained, some characteristics of these elements (e.g., position and orientation) can be extracted, calculated, or derived from the static parameters of the capture camera (part of the capture camera parameters) using the relative three-dimensional position of the capture camera and the relative orientation of the optical axis of the capture camera.

[0072] The base axis 29 can be defined in some embodiments to have a selected pitch related to the pitch component 200p of the field of view horizontal 200 and also to have a yaw perpendicular to the relative reference axis 26. The relative reference axis 26 can have the direction of the straight line between the center points of the two capture cameras, calculated using the static parameters of the capture camera (relative three-dimensional position). In some embodiments, the virtual orientation 28 (the orientation between the base axis 29 and the optical axis 23 of the virtual camera) can be calculated using the static parameters (specifications) of the capture camera.

[0073] In some embodiments, the components (roll 28r, pitch 28p, or yaw 28y) of the virtual orientation 28 shown in FIG. 2J can be calculated. In some embodiments, the pitch component 28p (pitch orientation) of the virtual orientation 28 can correspond to the difference in pitch (any pitch of the capture camera 20) between the optical axes with respect to the base axis 29.

[0074] It will be appreciated that in some embodiments, if the base axis 29 can have a pitch 29p in a reference coordinate system of static parameters used to determine, for example, the pitch 30p of the virtual camera, the pitch of the virtual orientation 28p can alternatively be calculated as 28p = |29p - 30p|.

[0075] In some embodiments, the yaw component (yaw orientation) 28y of the virtual orientation can correspond to the difference in yaw between the optical axis 22 and the base axis 29, and can be calculated as 28y = |90° - 30y| using the yaw 30y between the optical axis 22 and the relative reference axis 26.

[0076] In some embodiments, the roll component (roll orientation) 28r of the virtual orientation, which can correspond to the roll between the virtual object and the virtual camera, can be equivalent to the roll difference 28r between the capture camera and the visual field horizontal 200.

[0077] The base axis 29 can be in the same plane as the optical axis 22 such that the pitch component 28p of the virtual axis is zero in some embodiments. In some embodiments, the pitch 200p of the visual field horizontal can be directly defined using the direction of the viewer's line of sight.

[0078] In some embodiments, the pitch of the base axis 29 can be directly defined using the direction of the viewer's line of sight. In some embodiments, the pitch component (looking up or down) of the direction of the viewer's line of sight can be used to define the pitch component 200p of the visual field horizontal or the pitch component of the base axis.

[0079] One of ordinary skill in the art will understand that the virtual orientation 28 can be calculated and characterized using alternative calculations, equations, coordinate systems, etc.

[0080] FIG. 3A shows a schematic top view of the scene shown in FIG. 1A, and in this embodiment, an exemplary diverging wide-angle lens capture camera 20 can be used to capture the scene.

[0081] FIG. 3B shows a side view of a schematic 3D view of a geometry diagram of the field of view of this exemplary diverging wide-angle lens capture camera 20. This geometry diagram may be similar to a geometry 25, also known as a polygon mesh in the art, that is generated to correct an image captured by the wide-angle lens capture camera used in the embodiment shown in FIG. 3A.

[0082] FIG. 3C shows a 2D view of a polygon mesh 25 having a deformed grid pattern (mesh) 32 of the lens of the capture camera as viewed from the front. In this embodiment, a 3D object 39 (virtual image) can be obtained by applying a texture to the polygon mesh (geometry) 25 using the texture 24 (captured raw image) of the wide-angle lens capture camera. One of ordinary skill in the art will understand that the process of this embodiment is not limited to the type of capture camera (standard, wide-angle, fisheye, etc.), and that virtual 3D objects can be generated for any of these.

[0083] In a virtual environment, the field of view or viewing frustum of a virtual camera is typically called a frustum, and the frustum can be defined by two parameters, namely the aspect ratio and the angular dimension (a field of view having an aspect ratio that matches or conforms to the aspect ratio of the display screen). The aspect ratio of the frustum can be adjusted according to the aspect ratio of the display area (the height and width of the corresponding display screen), but the angular dimension (field of view) of the frustum can be adjusted according to the interocular distance of the viewer, vergence / accommodation compensation, the distance between the viewer and the stereoscopic display, the magnification (zoom) input selected by the viewer, or a combination of these.

[0084] The interocular distance can be selected to be a given value (e.g., the average interocular distance of an adult or a child) or fixed to a given value in some embodiments. In some embodiments, multiple viewers can be considered, and the interocular distance can be selected to be the minimum interocular distance among the multiple viewers.

[0085] FIGS. 4A and 4B are schematic top views showing an example of a view conversion step of the applicant's method in a virtual environment, where a cross-section of a virtual object 39 (a polygon mesh 25 pasted with a texture using an image 24 of a capture camera 20) can be selected (i.e., captured) using a virtual camera 21 (within a frustum 44 of the virtual camera). One of the view conversions can be applied, according to some embodiments, to adjust and modify the field of view (frustum 44) of the virtual camera in consideration of relevant dynamic parameters such as the zoom selected by the viewer, the aspect ratio and / or resolution of the display screen, the interocular distance of the user, and the distance between the user and the stereoscopic display.

[0086] FIG. 4A presents an exemplary embodiment showing such a change in the dimensions of the frustum 44, where the angular dimension of the width 45 (horizontal field of view) of the frustum can be adjusted to be a wider width 45', and thus have a larger frustum 44', which can be obtained by a horizontal dimension change (i.e., magnification) from the image sensor 42 to a dimension-changed image sensor 42" or something equivalent to the dimension change.

[0087] According to some embodiments, one of the view transformations can be applied to adjust and modify the symmetry of the frustum 44 geometry to make the frustum oblique (asymmetric). Such a view transformation may be referred to in the art as a lens shift of the virtual camera 21, which may be similar to a lateral translation. The lens shift 40 may include a translation 40' to an alternative position (see 42') of the image sensor 42 of the virtual camera 21, as shown in the exemplary embodiment of FIG. 4B. As a result of the movement, the viewing frustum 44 is shifted (translated), resulting in an asymmetric viewing frustum 44". As shown in FIG. 4B, it will be understood that the lens shift may not cause a change in the optical axis 22 of the virtual camera 21, and the optical axis may be the same for both the frustum 44 and 44".

[0088] The well-known lens shift operation in the art is typically used in virtual cameras and can be used, for example, in Unity, a 3D development platform, to "shift the lens in any axial direction to skew the camera frustum".

[0089] The selected lens shift can be defined by the interocular distance of the viewer, the distance between the viewer and the stereoscopic display, the dimensions of the stereoscopic display, vergence / accommodation compensation, the off-axis viewing angle of the viewer (the region of interest of the viewer), or a combination thereof. The lens shift can be applied in some embodiments to adjust various aspects of the stereoscopic image, including but not limited to placing infinity at the center on the 3D object, corresponding to infinity in the captured image, and / or establishing a correspondence between the real and virtual environments.

[0090] Once the necessary view transformation is complete, the virtual camera 21 can capture a virtual image within the display area (the projection of the viewing frustum 44 onto the virtual object) by rendering the pixels within this region of interest, also known as clipping. This can enable the rendering of practical stereoscopic images for the viewer's left and right eyes, such as the images shown in FIGS. 4C and 4D respectively. In some embodiments, these re-captured images can be further processed by simple operations and image adjustment techniques before being displayed to the viewer using a selected device. This may be similar to the method described in U.S. Patent No. 10,917,623B2 in some embodiments.

[0091] FIG. 4E presents a figure in which the previous left (same as FIG. 4C, shown thinly here) and right (same as FIG. 4D, shown as a dashed line here) final stereoscopic images (rendered re-captured images 62) are overlaid such that the sun appears as a single object with no difference between the two images to more appropriately visualize the difference between these two images. These left and right re-captured images 62 (62 Left for the left virtual camera 21L and 62 Right ) can enable an optimized three-dimensional (3D) viewing perspective of the scene when displayed at an image interval equal to the viewer's interocular distance with respect to the distance between the viewer and the display for the viewer's corresponding eye.

[0092] FIG. 4F shows an anaglyph figure of the overlay of the left and right re-captured images 62 having a horizontal offset corresponding to the viewer's interocular distance (Io) so as to be displayable on the display screen.

[0093] Figure 4G presents an exemplary embodiment of a pair of final stereoscopic images 62 having a common area that is smaller than the display screen 48 or has a different aspect ratio. The area of the image displayed outside the common area, corresponding to the hatched area of the display screen in Figure 4G, may be excluded from the rendering process or may be displayed with intermediate color pixels (such as white, black, gray-scale, selected alternative solid colors, gradients of selected colors, etc.). On the other hand, the dimensions of the re-captured images 62 can be kept unchanged.

[0094] The method described in the presented embodiment is innovative by creating a virtual environment for re-capturing the captured images using parallel virtual cameras, and using this method, a geometric correspondence can be established between the capture mode and the generated practical stereoscopic images of a part of the common field of view of the capture cameras.

[0095] In some embodiments, the possible distortion associated with the lenses of the capture cameras may be corrected and nullified using the geometry generated and used to create virtual 3D objects in the virtual environment. This enables the optimization of the camera's frustum, the optimization of the geometric correspondence between the capture environment and the images for the stereoscopic display, and, if necessary, the efficient and rapid re-capture of images using parallel virtual cameras that can be dynamically updated to correct lens and viewpoint distortion during the process. Those skilled in the art will understand that the process of this embodiment is not limited to the type of capture cameras and that various stereoscopic display outputs can be created by adjusting the rendering mode (the rendering process).

[0096] This integrated re-capture method using a fixed-position parallel virtual camera is one of the important aspects of the applicant's image processing method, which enables the correction of possible deviations between the relative positions of cameras using a wide range of capture camera types and arrangements without the need to specifically pre-configure the cameras for stereoscopic applications. In some embodiments, a virtual image 39 (geometry textured with raw captured images within a virtual environment) can be generated and saved in advance, then re-uploaded and used later to quickly generate stereoscopic images and update them in real time using new or updated viewer dynamic parameters, thereby reducing the processing load.

[0097] FIG. 5A presents an exemplary flowchart of a method for processing variables and inputs according to the viewer's dynamic parameters V and the capture camera parameters II (e.g., the static parameters of the capture camera) to generate and display a stereoscopic image to the viewer. This dynamic parameter V can be treated as a variable input considered in various steps. The step of positioning and generating a polygon mesh in a virtual environment, step III, can be performed using the capture camera parameters II, such as lens type, aspect ratio, projection type (orthographic cylindrical projection, spherical, planar, etc.), media name, horizontal field of view, vertical field of view, toe-in, and the relative positions of the left and right cameras. The capture cameras can have their own unique parameters, which means that they do not have to be the same cameras as long as they share a common portion of the field of view. In some embodiments, all parameters (display parameters and viewer parameters) can be acquired simultaneously. In some embodiments, the display parameters IV (aspect ratio, stereoscopic rendering mode, etc.) can be used to adjust a part of the virtual camera parameters (e.g., frustum) in step VI, while a part of the dynamic parameters of the display can be adjusted using a part of the viewer's dynamic parameters V (interpupillary distance, distance from the eyes to the display screen, etc.). A part of the viewer's dynamic parameters V (e.g., zoom-in, shift) can also be used to perform an increase / decrease in the zoom factor (frustum of the virtual camera), a shift or change in the orientation of the frustum of the virtual camera following steps VII, VIII, and IX. In some embodiments, additional processing such as methods and parameters similar to those disclosed in U.S. Patent No. 10,917,623B2, which is performed within step VIII, may be added.

[0098] In some embodiments, the left and right images may be processed simultaneously (each using, for example, a partition and texture of its own graphics processing unit), and the values of the dynamic parameters can be updated frame by frame. In other embodiments, the update frequency of the dynamic parameters can be adjusted (lowered or increased).

[0099] It will be appreciated that this method can be used to generate various modes of a stereoscopic display mode, such as autostereoscopic, anaglyph, line interleaving, column interleaving, etc., as required. The rendering method may have no restrictions on any color encoding and may vary depending on the device. For example, in the case of an autostereoscopic device, the final shader will depend on the display device and may be provided by the device manufacturer or information regarding the specifications of the stereoscopic display may be required to appropriately map the rendered image to achieve optimal results.

[0100] In some embodiments, for example, during clipping step X, the selection region of the left virtual camera can be compared with the selection region of the right virtual camera to exclude (leave blank / darken) the pixels corresponding to the area of the selection region outside the initially captured common area from the rendering process. In some embodiments, if there is a possibility that a stereoscopic image that may be smaller than the intended display screen is generated, a border can be provided for the image so that the image surely fills the stereoscopic display before display.

[0101] FIG. 5B presents a block diagram of an input and parameters that can include the parameters mentioned in FIG. 5A, and the input and parameters are used to set elements of a virtual environment (such as a mesh, texture, 3D object, frustum, position and orientation of a virtual camera, etc.), covering and summarizing the previously described relationship between the input and the steps of the method in some embodiments.

[0102] In some of the embodiments, the proposed method can be used by using some of the following steps. Step 501 of capturing an image using a capture camera having a common portion of the FOV of the camera, step 502 of defining and generating one or more polygon meshes, which may require the use of parameters 552 of one or more capture cameras, step 503 of generating a virtual 3D object (e.g., pasting a texture on a polygon mesh using an image captured by a capture camera), which may require the use of the image 553 captured by the capture camera, step 504 of positioning the virtual 3D object relative to the virtual camera, which may require the use of the position and orientation 554 of the optical axis of the capture camera with respect to a line extending between two capture cameras, step 505 of defining the orientation of the virtual camera (e.g., calculating the angle between the orientation of the virtual object and the optical axis), which may require the use of various parameters 555 (e.g., the relative position and orientation between two capture cameras), step 506 of sizing the aspect ratio of a frustum, which may require the use of the aspect ratio 556 of the display, step 507 of sizing the frustum (e.g., the width and height of the display area), which may require the use of various parameters 557 (e.g., interpupillary distance, convergence / divergence movement, distance between the viewer and the screen, and / or zoom selected by the viewer), step 508 of defining and applying a lens shift of the virtual camera, which may require the use of various parameters 558 (e.g., any of interpupillary distance, convergence / divergence movement, distance between the viewer and the screen, dimensions of the screen, and / or off-axis field of view angle of the viewer), and step 509 of formatting (e.g., clipping) and generating (e.g., rendering) a stereoscopic image for stereoscopic display to the viewer.

[0103] FIG. 6 also shows an example of a re-captured image 62 corresponding to a selection area of a virtual object 39 within a frustum projection 60, also referred to as the display area of a virtual camera. The frustum projection 60 can be defined, in some embodiments, as a shape or area corresponding to the projection of a frustum (the field of view or viewing cone of a virtual camera) onto a given virtual object 39. The frustum projection 60 can be made rectangular in the exemplary embodiment shown in FIG. 6, having an aspect ratio that may be the same as the aspect ratio of a display screen having a width (Ws) and a height (Hs), where the width (Wp) and height (Hp) of the frustum projection are values corresponding to the corresponding dimensions (Ws or Hs) of the display screen multiplied by a scale factor (SF) between the real world and the virtual world. Using this re-captured image 62, clipping of the virtual media (selectively enabling or disabling the rendering operation of pixels within a defined region of interest) can be performed prior to rendering, whereby only the pixels of the re-captured image 62 corresponding to the selection areas (within the frustum 44) of the left and right virtual objects 39 can be rendered to optimized GPU resources. In some embodiments, the resolution of the re-captured image within the display area (frustum projection) to be rendered can be selected to be the selected screen resolution so that the GPU can render the correct number of pixels (subdivision of the image to be rendered).

[0104] In some embodiments, the viewer can select the screen resolution to be the maximum resolution or the minimum resolution, or both. The resolution can vary and can include values such as standard definition (480p) having 640×480 pixels, high definition (720p) having 1280×720 pixels, full high definition (1080p) having 1920×1080 pixels, etc., but is not limited thereto.

[0105] Fig. 7 shows a 3D schematic diagram illustrating some of the changes induced on the viewing frustum that may be associated with a viewing transformation similar or similar to that previously presented in Fig. 4A. The viewing transformation may include, in the exemplary embodiment of Fig. 7, increasing the angular height 46 of the frustum, also called the field of view normal, to another modified angular height 46' while maintaining the aspect ratio of the frustum, thereby expanding the projection 60 of the original viewing frustum, which corresponds to the original viewing frustum 44, to a modified viewing frustum projection 60', which corresponds in this embodiment to the modified viewing frustum 44'. Such a viewing transformation may thus correct the output image recaptured using the virtual camera to an image now displayed as an adjusted final image 62'.

[0106] Those skilled in the art will appreciate that the parameters affecting the view frustum are not limited to the specifications described in this example, and various parameters can be added to the process as needed.

[0107] In these embodiments, it may be possible to use a selected technique to arrange the corrected left and right image pairs to generate a more appropriate stereoscopic display for the viewer. L,R must always be parallel to each other, it may be similar to the viewer-accommodated stereoscopic image display method described in U.S. Pat. No. 10,917,623 B2, but may utilize one or a combination of simple operations such as cropping, scaling, translation, and rotation. The images may be similar to the references in this patent in some embodiments, for example, but for stereoscopic content, the impression of depth may be created using parallax, which may often contradict other visual cues. One of the main problems may arise from the difference between two important pieces of information: the vergence of the eyes (the yaw orientation between the left and right eyes) and the distance at which the eyes focus (e.g., the distance between the display screen and the viewer's eyes), called accommodation. Those skilled in the art will appreciate that the vergence can be calculated using the viewer's interocular distance (Io) and the viewer's distance to the display screen (Ds).

[0108] The brain updates these two pieces of information regularly, takes them into account, and compares them with each other, enabling clear vision and a more appropriate interpretation of visual information. In the real world, the brain is accustomed to a specific learned correspondence between convergence movement and eye accommodation (for example, when one changes, the other piece of information changes in a specific way), but in some embodiments, these two pieces of information may conflict (be inconsistent) with this correspondence. For example, the viewer's eyes may remain focused on the display screen while the convergence movement may change. Those skilled in the art know that when there is a conflict between convergence movement and eye accommodation, many adverse effects may occur, such as a certain loss of depth perception or discomfort, pain, and double vision. To address and cope with this problem, as described in the above patent references, it is possible to determine the maximum or farthest distance (recognized as "inside" the screen) and the minimum or shortest distance (recognized as "outside" the screen) considering the angular constraints. In fact, in this U.S. Patent No. 10,917,623B2, these maximum (Df) and minimum (Dc) distances are determined and can be expressed as a function of the optical base (Bo) to enable the viewer's depth perception to correspond to such differences between convergence movement information.

[0109] In some embodiments, the stereoscopic image is further scaled and / or positioned using a relative base offset so that the farthest object appears closer to the screen and / or the closest object appears closer to the single or multiple screens. Such further scaling and / or positioning can be used to maintain the appearance that the objects appearing at a given depth appear to be at approximately the same depth while reducing the possible eye stress caused by the difference between the eye accommodation for focusing on the single or multiple screens and the eye accommodation for focusing on the near and / or far objects. In some embodiments, the process of clipping can be used to control the maximum and / or minimum distances and enable convergence-divergence movement - accommodation.

[0110] In some embodiments, a stereoscopic zoom that suggests changing the optical base (Bo) defined as the distance between the centers of two images on the screen can be calculated using the following formula with the interocular distance (Io) of the viewer, the vergence / accommodation movement coefficient (V) which may be an average value, and the screen / viewer distance (Ds).

[0111]

Equation

[0112] Here, the convergence-divergence motion coefficient is a selected angular offset (e.g., about 1 degree or 2 degrees) of the eye convergence motion that can be used to control the maximum angular change required for focusing between objects at the maximum distance and / or objects at the minimum distance, thereby enabling compensation for eye convergence motion / accommodation that can reduce the conflict with some visual cues, and thus providing a better stereoscopic experience for the viewer. The convergence-divergence motion coefficient can be selected or fixed to a given value in some embodiments. The convergence-divergence motion coefficient may be modified or actively adjusted as a function of the screen / viewer distance (Ds) in some embodiments. In embodiments where the convergence-divergence motion coefficient may be negligible and approximable to zero (e.g., when Ds is less than about 3 meters), it will be understood that the optical base can also have a negligible value (v = 0 → Bo = 0). As previously mentioned, in some embodiments, such an optical base can be used to determine the maximum (Df) distance or minimum (Dc) distance of the stereoscopic image using the following exemplary relationships. Df = Ds / (1 - Bo / Io) or Dn = Ds / (Bo / Io + 1). In some embodiments, the value of the optical base can be used to control and adjust the width and / or depth (distance) recognized by the viewer by modifying the frustum of a cone, more specifically the lens shift. The dimensions of the frustum of a cone (FoV) may be further adjusted in some embodiments to take into account the optical base while maintaining the same aspect ratio. In some embodiments, the dimensions of the frustum of a cone can be changed to adjust the dimensions of the projection of the frustum of a cone onto a virtual object, herein referred to as the width (Wp) and height (Hp) of the projection of the frustum of a cone, which can be calculated using the following formula using the width (Ws) or height (Hs) of the display screen and the scale factor (SF) between the real world and the virtual world.

[0113] [Number]

[0114] [Number]

[0115] Wp or Hp can also be calculated using the aspect ratio (Hs / Ws) of the screen, for example, Hp = Wp·Hs / Ws, so that it can be understood that the aspect ratio can also be reliably considered. The dimensional change of this frustum can be completed in some embodiments by defining and adjusting the angular width (Wa) of the frustum, also called the horizontal line of sight of the field of view, and the angular height (Ha) of the frustum, also called the vertical line of sight of the field of view, and in some embodiments, it can be calculated by the following formula.

[0116]

Number

[0117]

Number

[0118] In some embodiments, the lens shift is adjusted to enable the vertical shift (shift_PF) of the projection of the frustum using the optical base and the width (Ws) of the display screen, such as shift_PF = Bo·Wp / Ws = Bo·SF, etc., so that the proportional relationship between the real world (capture environment) and the virtual environment can be maintained more appropriately.

[0119] When all the final corrections are made, the pair of the generated left optimized image and right optimized image can then be projected onto the selected display device using the selected techniques and technologies for the benefit of the viewer. The stereoscopic image can be displayed in some embodiments using separate outputs (one for each eye), for example, using the screens of a projector or a virtual reality headset.

[0120] FIG. 8 presents a block diagram of the following hardware electronic components used to process and execute the method presented for generating a stereoscopic display for a viewer. It will be understood that all the components presented can be interconnected with a circuit such as a motherboard, enabling communication between many components of the system. Input data 80 can provide information regarding raw images of a capture camera, respective parameters of the camera, input parameters of the viewer, and / or other parameters from various capturers or sensors (controller / mouse / touch, accelerometer, gyroscope, etc.). The data storage device 82 can store multiple pieces of information, such as instructions for executing the presented method, raw images 24 of the capture camera 20 and parameters related to the raw images, pre-generated polygon meshes 25 and parameters of the virtual environment of the polygon meshes, stereoscopic videos obtained from previously regenerated and recorded examples of the method, panoramic images of scenes from two different positions that can be cropped to extract a selected region, etc. It will be understood that the computer processing device 84 can include a central processing unit (CPU) and a random-access memory (RAM) that can enable computer processing of the necessary commands. According to the CPU instructions, the graphics processing unit (GPU) 86 can be used to computer process and render the stereoscopic images to be displayed to the viewer. The display 88 used to present the stereoscopic images to the viewer can be of various types, such as a variety of screens, projectors, or head-mounted displays. It will be understood that the program product storing all the instructions for executing the presented method can be stored in a data storage unit, which may be an external data storage unit for processing later on another device including the components mentioned above.

[0121] In some embodiments, various types of captured images taken by various types of capture cameras are used to process and generate viewer-adjustable stereoscopic images where the viewer's perspective can be dynamic, similar to within a virtual reality (VR) environment. FIG. 9A presents a schematic of a dynamic head-mounted 3D VR stereoscopic display with two screens (left screen 48L and right screen 48R), or a single screen split into two, such as a smartphone equipped with compatible VR glasses, where each screen presents the corresponding images of a pair of processed stereoscopic images (left image 94L and right image 94R) for each of the viewer's eyes (left eye 90L has a corresponding left-eye field of view 96L and right eye 90R has a corresponding right-eye field of view 96R). FIG. 9B shows the transition state between frames, where a certain change in the orientation of the VR device (here represented as a clockwise rotation) corresponds to a change in the viewer's region of interest, i.e., a change in the viewer's dynamic parameters (step V of FIG. 5A). In FIG. 9B, the images 94L / 94R of the previous frame can be displayed using the previous dynamic parameters (the same as in FIG. 9A) to help visualize the change in the orientation of the VR device and the parameter changes to be implemented for the next image. FIG. 9C presents the next adjusted stereoscopic images 94’L / 94’R displayed on the display screen 48’ L;R corrected to match the perspective of the new dynamic parameter input, with all frames processed up to the viewer's new perspective 90’ L;R Note that the viewer's dynamic parameters can be extracted from any type of controller, such as the position of a computer mouse, accelerometers and gyroscopes, or a video game controller, but are not limited thereto.

[0122] It will be appreciated that the proposed method can capture raw images to be processed within a smartphone using most smartphones equipped with two rear cameras that share a common portion of the camera's field of view, regardless of the type of camera or the camera's magnification, and generate stereoscopic images that can be visualized using the smartphone's display.

[0123] FIG. 10A shows a view of a camera 20 (e.g., a smartphone camera or an additional camera lens that can be clipped to the smartphone) coupled to the back of the smartphone 100, where the line A-A of the cross-section can be seen. FIG. 10B shows a cross-sectional view A-A on the right side of FIG. 10A, and it is possible to see the positions of the two types of cameras 20L and 20R coupled inside the housing 101. In this embodiment, similar to what was previously described for FIGS. 2A, 2H, and 2I, considering and correcting using the applicant's method, there may be a slight shift between the cameras by only the gap 27 that can be considered and corrected.

[0124] One skilled in the art will understand that such a gap may be a minor imperfection on the order of a few micrometers that affects the viewer's stereoscopic experience, and it is not necessary to be as large as the gap 27 in this example. The proposed method can compensate for and correct such imperfections (e.g., gaps and orientation offsets) between the cameras as needed.

[0125] In some embodiments, the orientation (optical axis) of the capture cameras 22 may be parallel, nearly parallel (i.e., intended to be parallel in design but actually slightly offset from parallel, for example, due to manufacturing tolerances and / or imperfections), or non-parallel. The proposed method can compensate for and correct such offsets between the camera orientations as needed.

[0126] This method can also compensate for and correct differences in magnification between the cameras as needed.

[0127] In some embodiments, when the smartphone is held horizontally, the stereoscopic image can be displayed in the left and right halves of the screen and can be observed through or using a set of lenses for visualizing stereoscopic images that may be mounted on the smartphone, such as Google VR Cardboard.

[0128] FIG. 10B can also be understood to show a camera assembly / apparatus of any alternative pair of cameras sharing a common portion of the capture field, integrated into an apparatus for capturing images / videos, such as a 3D web camera, a 3D scanner, a 3D camera, or any other such apparatus known in the art.

[0129] The proposed method may require the completion of a feature evaluation of any of the capture camera parameters (e.g., position, orientation, field of view, lens distortion / deformation) of the capture cameras, or the completion of the feature evaluation may result in better results. Using this feature evaluation, various small changes and / or imperfections up to large changes and / or imperfections (e.g., position and / or angle offsets) between the capture camera parameters of similar apparatuses (e.g., between the same camera models and / or between camera assemblies / apparatuses) that are theoretically expected to be identical to each other can be characterized, measured, and / or identified, so that the values of the capture camera parameters to be used in the proposed method can be adjusted or fine-tuned.

[0130] It will be understood that a "defective" 3D product / apparatus (e.g., a 3D web camera having an undesirable angular and / or positional offset of the capture camera from the desired device, which may occur during production) may not be usable to provide an optimal or, in some cases, a practical stereoscopic experience to the viewer.

[0131] Using the proposed method, it is possible to enable the use of "defective" products, such as apparatuses coming off the production line with camera parameters that are slightly different from those designed and expected. Thus, it will be understood that using the proposed method, various apparatuses (web cameras, cameras, scanners, etc.) that may come off the production line with imperfections in the capture camera can be optimized, corrected, and / or made usable.

[0132] This feature evaluation of the capture camera parameters can be completed once (e.g., once for a new device) before first using the proposed method according to the user's requirements, up to and including before completing each new capture of an image (e.g., before taking a new photo), or not performed at all.

[0133] Figure 11 shows a block diagram of the main steps of an embodiment of this method, where the original image data 110 can be processed using the viewer's parameters 115 to render a stereoscopic image for a stereoscopic display 119. The capture parameters 111 can be used in some embodiments in steps 112 to generate a polygon mesh, which can be performed according to the parameters of the capture camera that can be attached to the original image data, step 113 to paste a texture onto the polygon mesh, which can be performed using an image source (texture) to obtain a virtual image, and step 114 to position a virtual camera, which can be performed for a virtual object. The parameters of the virtual camera can be adjusted in step 116 with respect to the display parameters and the viewer's parameters 115. After step 117 to format (e.g., clip) the selected area of the virtual object, the stereoscopic image can be rendered 118 in the selected stereoscopic format before step 119 to display it stereoscopically.

[0134] Figure 12 shows a block diagram of some steps that may be required to ensure that each of the virtual cameras can be appropriately set to render the stereoscopic image in a practical, optimal, and faithful manner and have a geometric correspondence with the real world. The steps can include calculating the previously proposed position, orientation, lens shift, dimensions of the frustum (view frustum), etc. from various parameters and values, which may include various parameters of the capture camera, various parameters of the display (resolution, aspect ratio, type of stereoscopy, and / or number of screens), and various parameters of the viewer (static parameters and / or dynamic parameters).

Claims

1. A method for processing an image for stereoscopic display to a viewer using a graphics processing unit (GPU), wherein the image is captured by a left capture camera and a right capture camera having a common portion of the capture field, and the method comprises: generating a left polygon mesh and a right polygon mesh using the specifications of the left capture camera and the right capture camera having a central axis; defining the position of a virtual camera with respect to the mesh so as to correspond to the field of view of the capture camera; pasting a texture onto each of the left polygon mesh and the right polygon mesh using at least one image from the left capture camera and the right capture camera to create a left virtual 3D object and a right virtual 3D object respectively; rendering a left-eye image and a right-eye image using a left virtual camera and a right virtual camera respectively oriented to capture at least a portion of the left virtual 3D object and the right virtual 3D object; and the orientation parameters of the left virtual camera and the right virtual camera are: a) defining a reference line extending between the position of the left capture camera and the position of the right capture camera in the capture space; b) defining the orientation of the left virtual camera and the right virtual camera in the capture space such that the left virtual camera and the right virtual camera are in a field-of-view horizontal plane selected to be parallel to each other; c) determining a left orientation offset vector and a right orientation offset vector that define a change between the orientation of the left virtual camera and the right virtual camera and the orientation of the optical axis corresponding to the central axis of the left capture camera and the right capture camera; d) using the left orientation vector and the right orientation vector to align the orientations of the left polygon mesh and the right polygon mesh with respect to the left virtual camera and the right virtual camera respectively about the central axis in the virtual space; e) defining the frustum parameters of the left and right virtual cameras using display parameters including the dimensions of the display, the distance to the display, the interocular distance, and the viewing direction A method defined thereby.

2. The method according to claim 1, wherein the display parameter includes a magnification, and the frustum parameter is defined using the magnification.

3. The method according to claim 1 or 2, wherein aligning includes defining lens shift parameters of the left and right virtual cameras.

4. The method according to any one of claims 1 to 3, wherein the orientation of the optical axis of the left capture camera is in the same plane as the orientation of the optical axis of the right capture camera.

5. The method according to any one of claims 1 to 4, wherein the optical axis of the left capture camera is parallel to the optical axis of the right capture camera, the left capture camera and the right capture camera have different fields of view, and the reference line is not perpendicular to at least one of the optical axis of the left capture camera and the optical axis of the right capture camera.

6. The method according to any one of claims 1 to 5, wherein the left-eye image and the right-eye image are still images.

7. The method according to any one of claims 1 to 6, wherein the left-eye image and the right-eye image are video images.

8. The method according to any one of claims 1 to 7, wherein the left-eye image and the right-eye image are displayed to the viewer on a single screen of the stereoscopic display.

9. The method according to any one of claims 1 to 8, wherein the left-eye image and the right-eye image are generated as anaglyph images.

10. The method according to any one of claims 1 to 8, wherein the left-eye image and the right-eye image are generated as column-interleaved images for display on a naked-eye stereoscopic display.

11. The method according to any one of claims 1 to 8, wherein the left-eye image and the right-eye image are generated as a sequence of page-flip images for viewing with shutter glasses.

12. The method according to any one of claims 1 to 8, wherein the left-eye image and the right-eye image are generated as a sequence of line-interleaved images for a polarized display.

13. The method according to any one of claims 1 to 7, wherein the stereoscopic display includes a left image display and a right image display, the left-eye image is displayed to the left eye of the viewer by the left image display, and the right-eye image is displayed to the right eye of the viewer by the left image display.

14. The method according to any one of claims 1 to 12, wherein the viewer includes a plurality of viewers, and the interocular distance is selected to be the minimum interocular distance among the plurality of viewers.

15. The method according to any one of claims 1 to 14, wherein the relative orientation between the left capture camera and the right capture camera is divergent.

16. The method according to any one of claims 1 to 14, wherein the relative orientation between the left capture camera and the right capture camera is parallel.

17. The method according to any one of claims 1 to 14, wherein the relative orientation between the left capture camera and the right capture camera is convergent.

18. At least one of the left-eye image and the right-eye image has an aspect ratio different from the aspect ratio defined by the dimensions of the display, and a boundary line for display to the viewer is provided for at least one of the left-eye image and the right-eye image. The method according to any one of claims 1 to 17.

19. The capture field of view includes a panoramic image, and the method further includes defining an optical axis selected inside a part of the panoramic image included in the common part of the capture field of view, and the part of the panoramic image and the selected optical axis are to be processed by the method. The method according to any one of claims 1 to 18.

20. A device for processing a stereoscopic image for display to a viewer, the device comprising a graphics processing unit, a central processor, and a memory readable by the central processor, the memory storing instructions for executing the method according to any one of claims 1 to 19.

21. A computer program product, wherein a processor connected to a GPU comprises a non-transitory memory storing instructions for executing the method according to any one of claims 1 to 20.

Citation Information

Patent Citations

  • 3D imaging

    JP2007507945A

  • Image processing apparatus, image processing method, and program

    JP2011223482A

  • Method and Apparatus for Capturing, Streaming and / or Playing Content

    JP2017535985A

  • Stereoscopic imaging

    US20050117215A1

  • Card reader device

    US20200117862A1