Stereoscopic image pair from monoscopic image saliency
The autostereoscopic display system processes monoscopic images to determine saliency and disparity, generating a stereoscopic image pair that positions regions of interest at the display plane, improving 3D viewing without glasses.
Patent Information
- Application Number
- PCT/US2024/042082
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-22
- Filing Date
- 2024-08-13
- Publication Date
- 2025-07-31
AI Technical Summary
Existing autostereoscopic displays struggle to effectively generate stereoscopic images from monoscopic inputs without requiring special eyewear, as they lack efficient methods to determine regions of interest and accurately position depth planes for optimal 3D perception.
An autostereoscopic display system processes monoscopic images to determine saliency information and disparity characteristics, generating a disparity map and convergence plane to create a stereoscopic image pair, ensuring that regions of interest are positioned at or near the display plane for optimal 3D viewing without glasses.
The system enhances 3D image perception by accurately positioning regions of interest at the display plane, providing clear and immersive 3D viewing experiences without the need for additional eyewear.
Smart Images

Figure US2024042082_31072025_PF_FP_ABST
Abstract
Description
STEREOSCOPIC IMAGE PAIR FROM MONOSCOPIC IMAGE SALIENCYCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 623,742, filed January 22, 2024, which is incorporated by reference herein in its entirety.FIELD OF THE DISCLOSURE
[0002] This document relates generally to display systems, and more specifically relates to multiview displays, three-dimensional displays, or autostereoscopic displays.BACKGROUND OF THE DISCLOSURE
[0003] A multiview display can provide different views of a multiview image to a viewer. A stereoscopic display can provide two different views of a three-dimensional image to the two eyes of a viewer. An autostereoscopic display can provide the two different views to the two eyes of the viewer without requiring the viewer to wear special glasses or eyewear. There is ongoing effort to improve autostereoscopic displays.SUMMARY
[0004] In an example, an autostereoscopic display system can comprise processing circuitry that can perform operations. The operations can comprise: receiving data specifying a monoscopic image; determining, from the monoscopic image, saliency information that indicates a region of interest of the monoscopic image; determining a convergence plane for a stereoscopic representation of the monoscopic image, the convergence plane being a function of disparity characteristics of the region of interest of the monoscopic image; generating a disparity map based on the monoscopic image and the convergence plane; and generating a stereoscopic image pair using the monoscopic image and the disparity map.
[0005] In an example, a method can include generating a stereoscopic image pair from a monoscopic image. The stereoscopic image pair can optionally be displayed using an autostereoscopic system. The method can comprise: receiving, with processing circuitry, data specifying the monoscopic image; determining, with the processingcircuitry, from the monoscopic image, saliency information that indicates a region of interest of the monoscopic image; determining, with the processing circuitry, a convergence plane for a stereoscopic representation of the monoscopic image, the convergence plane being a function of disparity characteristics of the region of interest of the monoscopic image; generating, with the processing circuitry, a disparity map based on the monoscopic image and the convergence plane; and generating, with the processing circuitry, the stereoscopic image pair using the monoscopic image and the disparity map.
[0006] This Summary is intended to provide an overview of subject matter of the present document. It is not intended to provide an exclusive or exhaustive explanation of the invention. The detailed description is included to provide further information about the present subject matter.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] FIG. 1 shows an exploded, perspective-view schematic drawing of an example of a multi view display system.
[0008] FIG. 2 shows a front-view drawing of an example of a display panel that includes an array of light-emitting diodes.
[0009] FIG. 3 shows a front-view drawing of an example of a display panel that includes a backlight and a light valve array.
[0010] FIG. 4 shows a front-view drawing of an example of a parallaxgenerating optic that includes a lenticular lens.
[0011] FIG. 5 shows a cross-sectional view of the lenticular lens of FIG. 4.
[0012] FIG. 6 shows a front-view drawing of an example of a parallaxgenerating optic that includes a parallax barrier.
[0013] FIG. 7 shows a cross-sectional view of the parallax barrier of FIG. 6 having transmissive slits.
[0014] FIG. 8 shows an example of generating a stereoscopic image pair from a monoscopic image.
[0015] FIG. 9 shows a flowchart of an example of a method for generating the stereoscopic image pair from the monoscopic image.
[0016] FIG. 10 shows an example of a monoscopic image that includes a plus sign (“+”) on a background and a corresponding saliency map.DETAILED DESCRIPTION
[0017] In general, the quantity disparity can represent a horizontal distance between left and right images of a stereoscopic image pair for the same point in a scene. In other words, the disparity can correspond to a distance between a pixel in the left (right) image and its horizontal match in the right (left) image. It is assumed that a viewer is oriented upright, such that the left and right eyes of the viewer are at the same height and that an axis connecting the left and right eyes is horizontal.
[0018] A disparity map can quantify the disparity of a stereoscopic image pair. For example, the disparity map can be an array, such as with a same size as the images of the stereoscopic image pair. The value of the disparity map at a particular pixel can correspond to the disparity value at the particular pixel of the image of the stereoscopic image pair.
[0019] A convergence plane is the plane in the scene at which the disparity equals zero. For objects in the scene on one side of the convergence plane, the disparity is positive. For objects in the scene on the other side of the convergence plane, the disparity is negative.
[0020] When an autostereoscopic display displays the stereoscopic image pair, objects in the scene at the convergence plane appear to be located at a plane of the autostereoscopic display. Objects in the scene that are closer to the viewer than the convergence plane appear to be located above the autostereoscopic display. Objects in the scene that are farther away from the viewer than the convergence plane appear to be located below the autostereoscopic display.
[0021] In general, a disparity map can include the same amount of information as a depth map, such that one can be created from the other. In addition, for three data structures that include a left image, a right image, and a disparity map, any two of the three data structures can be used to create the third of the three data structures.
[0022] Further, there exists image processing software, executable by processing circuitry, that can use artificial intelligence to generate an artificial or simulated disparity map from a (single) monoscopic image. For example, the processing circuitry can use a fully convolutional neural network model. The neural network model can include an encoder based on a residual neural network (ResNet). The neural network model can include a decoder configuration having skip connection that are similar to those in a U-Net. The architecture of a U-Net can include an encoder path, also known as the contracting path, and a decoder path, also known as the expanding path. The encoder can extract features from an input image. The decoder can project those features onto pixel space to produce a dense classification. The encoder and decoder may be symmetrical and connected by paths, which can give the model a “U” shape. The convolutional neural network model can include millions (or more) of trainable parameters, such as can be trained on multiple (e.g., thousands, millions) of images that have known disparity characteristics. Processing circuitry can use the model to process a (single) monoscopic image and provide a corresponding disparity map.
[0023] In general, the quantity saliency can refer to data about visually important or prominent areas within an image. A saliency map can quantify the saliency of a (single) monoscopic image or one or both images of a stereoscopic image pair. For example, the saliency map can be an array, such as with a same size as the monoscopic image or the left and right images of the stereoscopic image pair. Image processing software, executable by the processing circuitry, can analyze the monoscopic image or one or both images of the stereoscopic image pair, such as by examining one or more of the color variance, the luminance variance, or the texture variance, to determine saliency information about the image and generate the corresponding saliency map.
[0024] For example, to generate a saliency map of an image, the image processing software can compute a local variance of colors around each pixel in the image. The software can identify areas with high color variance, which typically correspond to edges or borders of image objects. The software can deem areas that are of high visual perceptual importance to a human viewer to be salient objects or salient areas. The software can perform saliency analysis on a coarse scale. For example, the software can generate sizeable clusters of information-rich objects or objects that stand out from their surroundings. In performing the coarse analysis, the software can highlight the most salient areas with relatively high granularity, rather than focusing on finer details in the image.
[0025] The autostereoscopic display system can use saliency characteristics in part to generate a stereoscopic image pair from a monoscopic image. For example, an autostereoscopic display system can use saliency information from a monoscopic image to identify a region of interest in the monoscopic image. The autostereoscopic display system can determine a location of a convergence plane based on disparity informationin the region of interest. When the autostereoscopic display system displays the stereoscopic image pair on the autostereoscopic display, a viewer can perceive the region of interest as being located roughly at or near the plane of autostereoscopic display, rather than too close to the viewer (e.g., in front of the autostereoscopic display) or too far from the viewer (e.g., behind the autostereoscopic display).
[0026] As a specific example, FIG. 10 shows an example of a monoscopic image 1002 that includes a plus sign 1004 (“+”) on a background and a corresponding saliency map 1008. The autostereoscopic display system can determine that the plus sign 1004 forms the region of interest 1006, such that a perimeter of the plus sign 1004 defines a perimeter of the region of interest 1006. The autostereoscopic display system can generate a saliency map 1008 that can quantify the saliency information of the monoscopic image 1002. The saliency map 1008 can include saliency values (S) that are relatively large (such as being positive) in the region of interest 1006 and saliency values (S) that are relatively small (such as equaling zero) outside the region of interest 1006. The autostereoscopic display system can generate the stereoscopic image pair based on disparity information in the region of the plus sign 1004 (such as without using information from the background). When the autostereoscopic display system displays the stereoscopic image pair on the autostereoscopic display, the plus sign 1004 can be perceived as being at or near the plane of the autostereoscopic display, rather than far above or far behind the plane of the autostereoscopic display.
[0027] The preceding paragraphs merely summarize some aspects of the image generation and display technique described in detail below, and should not be construed as limiting in any way.
[0028] FIG. 1 shows an exploded, perspective-view schematic drawing of an example of a multiview display system 100 that includes a multiview display 110. The configuration of FIG. 1 is but one example of a multiview display system 100; other configurations can be used.
[0029] The sign conventions shown in FIG. 1 and used below assume that the multiview display 110 extends in an (x, y) plane, and that a z-axis extends away from the multiview display 110 and generally toward a viewer 42, along a direction that is orthogonal to a plane of the multiview display 110. Other sign conventions can be used.
[0030] The multiview display system 100 can include a multiview display 110. The multiview display 110 can provide different views of a multiview image to theviewer 42. For example, as the viewer 42 moves in space, the multiview display 110 can direct different views of the multiview image to the left and right eyes of the viewer 42, so that the viewer 42 can observe the different views of the multiview image from different locations or orientations. In some configurations, the multiview display 110 can provide the multiple views at respective fixed location regions in space, so that the multiview display 110 can operate without using eye tracking. In other configurations, such as the autostereoscopic configurations described in detail below, the multiview display system 100 can use eye tracking to dynamically and continuously (or at relatively frequent discrete times) determine a location of the viewer 42, and in response, can dynamically and continuously control how the multiview display 110 displays the multiview image so that the multiple views follow the viewer 42 or follow the tracked eye location(s) of the viewer 42 as the viewer 42 moves in space relative to a position of the multi view display 110.
[0031] For configurations in which the multiview display 110 provides just two different views of the multiview image, the multiview display 110 can be an autostereoscopic display or three-dimensional (3D) display. The autostereoscopic display can provide a left image to a left eye of the viewer 42 and a right image to a right eye of the viewer 42. The left image and the right image can correspond to different views of an object or a scene, and can allow the viewer 42 to perceive the object or scene in 3D with just the viewer’s naked eyes, without the use of additional glasses or headgear.
[0032] The multiview display system 100 can include a viewer tracker 120 that can dynamically determine the location of the viewer 42. The multiview display system 100 can use the determined location of the viewer 42 to direct the left image to the left eye of the viewer 42 and the right image to the right eye of the viewer 42. Because the viewer’s location can vary as the viewer 42 moves in space, using eye tracking can allow the multiview display system 100 to follow the viewer 42, so that the autostereoscopic display can automatically direct the left image to the left eye at the viewer 42’ s (dynamically varying) location and automatically direct the right image to the right eye at the viewer 42’ s (dynamically varying) location. The viewer tracker 120 can provide a tracked position of the viewer 42, such as a tracked position of a head of the viewer 42, of one or both eyes of the viewer 42, or of another anatomical feature of the viewer 42. The viewer tracker 120 can be coupled to the processing circuitry 130 (described below)or controller, such as by providing viewer 42 location data (shown in FIG. 1 as coordinates xv, yv, and zv) that represents a measured position or location of the viewer 42. The viewer tracker 120 can provide the viewer 42 location data at regular or irregular intervals to the processing circuitry 130. In a specific example of a viewer tracker 120, a camera can capture an image of the viewer 42. The viewer tracker 120 can further include an image processor (or general-purpose computer programmed as an image processor) that can determine a position of the viewer 42 within the captured image to provide the tracked position. In some examples, the processing circuitry 130 can include the image processor of the viewer tracker 120. In other examples, the processing circuitry 130 can be separate from the image processor of the viewer 42 tracker. Other suitable viewer trackers can be used, including viewer trackers based on lidar (e.g., using time-of-flight of reflected light over a scene to determine distances to one or more objects in the scene, such as a viewer’s head or a viewer’s eyes) or other technologies. The processing circuitry 130 can use an output of the viewer tracker 120, among other data, to perform one or more downstream calculations involved with providing the left view or left image to the left eye of the viewer 42 and the right view or right image to the right eye of the viewer 42.
[0033] The autostereoscopic display can be a lenticular autostereoscopic display. In a lenticular autostereoscopic display, a display panel 112 can display the multiview image, and a parallax-generating optic 118 can direct light from the display panel 112 to the viewer 42 such that a left image can be visible from the left eye of the viewer 42 and a right image can be visible from the right eye of the viewer 42. During use of the lenticular autostereoscopic display, the processing circuitry 130 can track the location of the viewer 42, and can use the tracked location to dynamically determine how to distribute content of the multiview image over a surface area of the display panel 112 (e.g., using pixels distributed over the display panel 112) such that a left image remains visible from the left eye of the viewer 42 and a right image remains visible from the right eye of the viewer 42, even as the viewer 42 changes location. In this manner, the location tracking and the image content distribution can be performed in software, such that following the location of the viewer 42 may not involve physically moving any components of the lenticular autostereoscopic display with respect to one another. Examples of suitable display panels and examples of suitable parallax-generating optics are described below.
[0034] In an example, a display panel 112 can display the multiview image. The display panel 112 can have an array of subpixels 114 that can display an image according to stereo mapping coordinates associated with the viewer 42. The subpixels 114 can be located at subpixel locations in a grid having grid axes (for example, the x-axis and y- axis). Each subpixel 114 can generate light having a specified color. For example, the subpixels 114 can include red subpixels, green subpixels, and blue subpixels, which generate red light, green light, and blue light, respectively. Other color / wavelength schemes can be used. The subpixels 114 can be grouped into pixels, with each pixel including at least two subpixels 114 that produce light of different colors. The display panel 112 can receive, from the processing circuitry 130 (described below), a display panel driving electrical signal 138 that can specify how the content of the multiview image is distributed over the pixels and / or subpixels 114 of the display panel 112. Two possible configurations for the display panel 112 are described below and shown in FIGS. 2 and 3. Other configurations can be used.
[0035] FIG. 2 shows a front-view drawing of an example of a display panel 112A that includes an array 202 of light-emitting diodes 204, such as an array 202 of organic light-emitting diodes. Each light-emitting diode 204 can correspond to a subpixel. The array 202 of light-emitting diodes 204 can include red light-emitting diodes 204R, green light-emitting diodes 204G, and blue light-emitting diodes 204B, which correspond to the red subpixels, green subpixels, and blue subpixels, respectively. Each light-emitting diode 204 can controllably generate light in response to an electrical signal provided by the processing circuitry 130, such as display panel-driving electrical signal 138, or by suitable light-emitting diode-driving circuitry in communication with the processing circuitry 130. The processing circuitry 130 can cause a specified lightemitting diode 204 to be directly powered with a power that varies as a function of an intensity in a corresponding location in the image. The power delivered to a lightemitting diode 204 can optionally be pulse-width modulated at a modulation frequency that is greater than can be perceived by a human eye. Using pulse-width modulation can simplify a design of a light-emitting diode array controller, because it can generate an arbitrary average power level from a relatively small number of instantaneous power levels by varying a duty cycle of the power. In some examples, the array 202 of lightemitting diodes 204 can be arranged in a rectangular or square repeating pattern over a surface area 206 of the array 202. For example, the array 202 can have grid axes 208that are orthogonal to each other. In some examples, the grid axes 208 can be parallel to edges 210 of the array 202 of light-emitting diodes 204.
[0036] FIG. 3 shows a front-view drawing of an example of a display panel 112B that includes a backlight 302 and a light valve array 304. Although FIG. 3 shows the backlight 302 and the light valve array 304 as being separated, in practice, the backlight 302 and the light valve array 304 may be in contact or may be located as close together as is practical. The backlight 302 can provide illumination having a uniform or substantially uniform intensity over a surface area of the backlight 302. The backlight 302 can provide illumination having a relatively broad spectrum, such as including most or all of the visible portion of the electromagnetic spectrum. The backlight 302 can provide the illumination into a continuum of propagation angles toward the light valve array 304. The backlight 302 can provide unmodulated illumination to the light valve array 304. The light valve array 304 can include light valves 306 that are individually controllable or controllable in one or more groups by the processing circuitry 130 (described below). Each light valve 306 can controllably attenuate the illumination from the backlight 302, such as in response to an electrical signal provided by the processing circuitry 130, such as display panel-driving electrical signal 138, or by suitable light valve driving circuitry in communication with the processing circuitry 130. Each light valve 306 can have a corresponding color filter that allows only a portion of the electromagnetic spectrum to pass through the light valve. For example, the light valves 306 can include red light valves 306R that have a red filter that allows only red light to pass through the red light valves 306R, green light valves 306G that have a green filter that allows only green light to pass through the green light valves 306G, and blue light valves 306B that have a blue filter that allows only blue light to pass through the blue light valves 306B. Other color schemes and numbers of colors can be used. Suitable light valves can include liquid crystal light valves, electrophoretic light valves, light valves based on electrowetting, and others. In some examples, the light valves 306 of the light valve array 304 can be arranged in a rectangular or square repeating pattern over a surface area 308 of the light valve array 304. For example, the light valve array 304 can have grid axes 208 that are orthogonal to each other. In some examples, the grid axes 208 can be parallel to edges 312 of the light valve array 304.
[0037] Referring again to FIG. 1, the autostereoscopic display can include a parallax-generating optic 118 that can direct light from the display panel 112 to theviewer 42, such that a left view or a left image can be visible from a left eye of the viewer 42 and a right view or a right image can be visible from a right eye of the viewer 42. Two possible configurations for the parallax-generating optic 118 are described below and shown in FIGS. 4 and 5 and in FIGS. 6 and 7. Other configurations can be used. Each of the configurations of FIGS. 4 and 5 and in FIGS. 6 and 7 can be used in combination with any of the configurations of the display panel 112 shown in FIGS. 2 and 3 (e.g., the array of light-emitting diodes 204 in FIG. 2 or the backlight 302 and light valve array 304 in FIG. 3).
[0038] FIG. 4 shows a front-view drawing of an example of a parallaxgenerating optic 118A that includes a lenticular lens 402. FIG. 5 shows a cross-sectional view of the lenticular lens 402 of FIG. 4. The lenticular lens 402 can include a plurality of cylindrical lenses 504 that are equally spaced apart. The lenticular lens 402 can have a focal plane coincident with the display panel 112. The lenticular lens 402 can be positioned to receive light from the display panel 112 and at least partially focus the received light to direct the light to specified regions proximate the viewer’s eyes.
[0039] FIG. 6 shows a front-view drawing of an example of a parallaxgenerating optic 118B that includes a parallax barrier 602. The parallax barrier can include a plurality of transmissive slits 704 that are equally spaced apart. FIG. 7 shows a cross-sectional view of the parallax barrier 602 (FIG. 6) having transmissive slits 704. The parallax barrier 602 can include an array of opaque strips 706 and thin transmissive slits 704 arranged to occlude portions of a displayed image in left and right viewing regions. The transmissive slits 704 can be spatially arranged to ensure that the left / right image portions are only visible in the corresponding left / right viewing regions for which they are intended. The parallax barrier 602 can be provided by a static physical layer in which the slits are precisely positioned, or electronically generated on an adaptive intermediate liquid crystal display layer.
[0040] The parallax-generating optic 118 can be invariant along an optical axis (OA) that is angled with respect to the grid axes (e.g., the x-axis and the -axis), such as at a slant angle (a) of 45 degrees or about 45 degrees with respect to the grid axes 208. For example, the parallax-generating optic 118 can have transmissive features, such as the cylindrical lenses or the transmissive slits, that are invariant along the optical axis (OA) and are periodic along an axis that is orthogonal to the optical axis (OA).
[0041] Referring again to FIG. 1, the autostereoscopic display can include a material 116 disposed between the display panel 112 and the parallax-generating optic 118. In some examples, the material 116 may extend fully between the display panel 112 and the parallax-generating optic 118, such that a light ray originating at the display panel 112 passes only through the material 116 (and does not pass through any air or unfilled volume) before arriving at the parallax-generating optic 118. In other examples, the material 116 may occupy only a portion of the volume between the display panel 112 and the parallax-generating optic 118, such that a light ray originating at the display panel 112 passes through at least some of the material 116 and passes through a volume of air before arriving at the parallax-generating optic 118. The material 116 may have a refractive index denoted by quantity n. The value of the refractive index n may be between about 1.3 and about 2, although other suitable values may be used. Suitable materials 116 can include glass, plastic, a transparent optical adhesive, and others. In some examples, the material 116 can be dispensed in a liquid form, then cured in place, such as by exposure to ultraviolet light or heat. In other examples, the material 116 can be manufactured as a solid unit and placed in its location in the autostereoscopic display. For example, the material 116 can function as a cover glass for the display panel 112. In some examples, the material 116 can function as a relatively precise spacing element. For example, the material 116 can be manufactured to have a specified thickness to within a specified thickness tolerance and can set the spacing between the display panel 112 and parallax-generating optic 118 to have a value equal to the specified thickness when the autostereoscopic display is assembled.
[0042] As illustrated in FIG. 1, the multiview display system 100 can include processing circuitry 130. The processing circuitry 130 can include a non-transitory computer-readable storage medium 132, such as a hard disk, a solid-state hard drive, memory, optical media, magnetic media, semiconductor media, punch cards, or others. The non-transitory computer-readable storage medium 132 can be included locally with the processing circuitry 130 or can be located remotely from the processing circuitry 130 and be accessible through a wired or wireless connection. The non-transitory computer- readable storage medium 132 can store instructions 134 for performing a particular task or series of tasks or executing some or all steps of a method.
[0043] The instructions 134, when executed by the processing circuitry 130, can cause the processing circuitry 130 to perform operations 136. For example, theinstructions 134 may be for generating a stereoscopic image pair from a monoscopic image. Suitable operations 136 for generating a stereoscopic image pair from a monoscopic image can include, among other operations: receiving, with processing circuitry 130, data specifying the monoscopic image; determining, with the processing circuitry 130, from the monoscopic image, saliency information that indicates a region of interest of the monoscopic image; determining, with the processing circuitry 130, a convergence plane for a stereoscopic representation of the monoscopic image, the convergence plane being a function of disparity characteristics of the region of interest of the monoscopic image; generating, with the processing circuitry 130, a disparity map based on the monoscopic image and the convergence plane; and generating, with the processing circuitry 130, the stereoscopic image pair using the monoscopic image and the disparity map. By executing these operations 136, the processing circuitry 130 can cause the multiview display system 100 to perform generating a stereoscopic image pair from a monoscopic image. The operations 136 can further include using the multiview display system 100 to display the stereoscopic image pair. These operations 136 are described in detail below.
[0044] FIG. 8 shows an example of generating a stereoscopic image pair from a monoscopic image. The quantities 800 can be data structures, such as arrays, matrices, or variables. The quantities 800 can be stored in memory. The processing circuitry 130 can read one or more of the quantities 800 from memory and write one or more of the quantities 800 to memory when the processing circuitry 130 executes the operations 136. The quantities 800 are but mere examples of quantities that can be used to generate a stereoscopic image pair from a monoscopic image. Other suitable quantities can be used.
[0045] The processing circuitry 130 can receive data that specifies a monoscopic image 802. The data can be received from memory, a local storage device, or a remote storage device. The data can be included as one or more frames of a video signal, which can specify the monoscopic image 802 as one of a series of sequential monoscopic images.
[0046] The processing circuitry 130 can determine, from the monoscopic image 802, saliency information that indicates a region of interest 806 of the monoscopic image 802.
[0047] To quantify the saliency information, the processing circuitry 130 can generate a saliency map 804 based on the monoscopic image 802. The saliency map 804can be an array or matrix, such as with the same dimensions (e.g. size in pixels) as the monoscopic image 802. The saliency map 804 can include values that correspond to importance to a human visual system and determine the region of interest 806 of the monoscopic image 802. The saliency map 804 can have saliency values (S) in an array that corresponds to the monoscopic image 802. A relatively large value of S, for a pixel of the saliency map 804, can indicate a relatively high saliency for the corresponding pixel in the monoscopic image 802. A relatively small value of S (such as zero), for a pixel of the saliency map 804, can indicate a relatively low saliency for the corresponding pixel in the monoscopic image 802. In some examples, the values of S can take on a finite or infinite number of values. In some examples, the saliency map 804 can be binary, such as with a (single) positive or non-zero value of S in the region of interest and a zero value of S outside the region of interest. For example, the values of the saliency map 804 can include a non-zero value in the region of interest 806 and a zero value outside the region of interest 806. In other examples, the values can be any suitable value, such as with a magnitude that corresponds to a relative importance to the human visual system. The region of interest 806 can include less than all of the monoscopic image 802. The region of interest 806 can include one or more regions of deemed importance. In other words, the region of interest 806 can include one contiguous region in the monoscopic image 802 or multiple non-contiguous regions in the monoscopic image 802.
[0048] The processing circuitry 130 can determine a convergence plane for a stereoscopic representation of the monoscopic image 802. The convergence plane can be a function of disparity characteristics of the region of interest 806 of the monoscopic image 802. To determine the convergence plane, the processing circuitry 130 can generate a provisional disparity map 808 based on the monoscopic image 802. As explained above, the processing circuitry 130 can using artificial intelligence to generate the provisional disparity map 808, which can be an artificial or simulated disparity map. The provisional disparity map 808 can include values that correspond to depth information within the monoscopic image 802 and include the disparity characteristics of the monoscopic image 802. The provisional disparity map 808 can extend over a full extent of the monoscopic image 802, such as includes regions both within the region of interest 806 and outside the region of interest 806.
[0049] To confine the disparity characteristics to just the region of interest 806, the processing circuitry 130 can generate a weighted disparity map 810. The weighted disparity map 810 can use values of the provisional disparity map 808 that are weighted by respective values of the saliency map 804. For a saliency map 804 that is binary or a saliency map 804 that has a zero value outside the region of interest 806, such a weighting can ensure that the weighted disparity map 810 uses values of the provisional disparity map 808 in the region of interest 806 and does not use values of the provisional disparity map 808 outside the region of interest 806.
[0050] The convergence plane may be located using disparity information from the region of interest 806. In general, a disparity map defines the convergence plane as a plane for which values of the disparity map equal zero. To ensure that the convergence plane is located at a plane of the autostereoscopic display, the processing circuitry 130 can subtract a mean value of disparity in the region of interest 806, so that the average disparity in the region of interest 806 becomes zero. Subtracting the mean value in this manner can shift the convergence plane from an (unspecified) initial location to a final location at or near the plane of the autostereoscopic display. The provisional disparity map 808 can represent the disparities in the region of interest 806 before the mean-value correction; the final disparity map 812 can represent the disparities in the region of interest 806 after the mean-value correction.
[0051] The processing circuitry 130 can generate the final disparity map 812 based on the monoscopic image 802 and the convergence plane. More specifically, the processing circuitry 130 can generate the final disparity map 812 based on the provisional disparity map 808 and the saliency map 804. As a specific example, the processing circuitry 130 can perform the following three operations to generate the final disparity map 812. First, the processing circuitry 130 can multiply the values of the provisional disparity map 808 by corresponding values of the saliency map 804 to form the weighted disparity map 810. Second, the processing circuitry 130 can average values of the weighted disparity map 810 to form a mean value. Third, the processing circuitry 130 can subtract the mean value from the values of the weighted disparity map 810 to form the final disparity map 812. The three operations detailed above are but mere examples of how to generate the final disparity map 812 based on the provisional disparity map 808 and the saliency map 804. Other operations can be used.
[0052] The processing circuitry 130 can generate a stereoscopic image pair 814 using the monoscopic image 802 and the final disparity map 812. The stereoscopic image pair 814 can include a left image 816 and a right image 818. When an autostereoscopic display displays the stereoscopic image pair 814 to a viewer, the autostereoscopic display directs the left image 816 to a left eye of the viewer and directs the right image 818 to a right eye of the viewer.
[0053] For the method 900 discussed below, the monoscopic image 802 is the input and the stereoscopic image pair 814 is the output. The stereoscopic image pair 814 can be displayed on an autostereoscopic display, such as the multiview display system 100 of FIG. 1
[0054] FIG. 9 shows a flowchart of an example of a method 900 for generating a stereoscopic image pair from a monoscopic image. The method 900 can be executed on, apply to, or can be used by or used with, an autostereoscopic display system, such as the multiview display system 100 of FIG. 1. The method 900 can be stored as instructions 134 on a non-transitory computer-readable storage medium 132. The instructions 134, when executed by the processing circuitry 130, can cause the processing circuitry 130 to perform the operations 136. The operations 136 can include the operations detailed below for the method 900. The method 900 is but one method for generating a stereoscopic image pair from a monoscopic image; other suitable methods may be used.
[0055] At operation 902, the processing circuitry 130 can receive data specifying a monoscopic image 802.
[0056] At operation 904, the processing circuitry 130 can determine, from the monoscopic image 802, saliency information that indicates a region of interest 806 of the monoscopic image 802. The region of interest 806 can include less than all of the monoscopic image 802. The region of interest 806 can include one or more regions of deemed importance, such as to the human visual system.
[0057] At operation 906, the processing circuitry 130 can determine a convergence plane for a stereoscopic representation of the monoscopic image 802 based on disparity characteristics of the region of interest 806 of the monoscopic image 802.
[0058] At operation 908, the processing circuitry 130 can generate the final disparity map 812 based on the monoscopic image 802 and the convergence plane.
[0059] At operation 910, the processing circuitry 130 can generate the stereoscopic image pair 814 using the monoscopic image 802 and the final disparity map 812
[0060] The autostereoscopic display, such as the multiview display 110, can display the stereoscopic image pair 814 such that when the stereoscopic image pair 814 is displayed on the autostereoscopic display: the convergence plane coincides with a plane of the autostereoscopic display (and the region of interest), a first portion of the stereoscopic image pair 814 appears to be located above the plane of the autostereoscopic display, and a second portion of the stereoscopic image pair 814 appears to be located below the plane of the autostereoscopic display. A viewer tracker, such as the viewer tracker 120, can determine a location of a viewer, such as the viewer 42
[0061] The autostereoscopic display can include a display panel, such as the display panel 112. The display panel can have an array of subpixels, such as the array of subpixels 114, that can display the stereoscopic image pair 814. The autostereoscopic display can include a parallax-generating optic, such as the parallax-generating optic 118, that can direct light from the display panel to the viewer. The parallax-generating optic can include one of a lenticular lens, such as the lenticular lens 402, or a parallax barrier having transmissive slits, such as the parallax barrier 602 having transmissive slits 704. The parallax-generating optic can be invariant along an optical axis, such as the optical axis (CM), that has a slant angle relative to the display panel. The slant angle can be within a specified angular tolerance of forty-five degrees, such as within one degree, two degrees, five degrees, ten degrees, or another suitable angular tolerance. The parallax-generating optic can be periodic along an axis orthogonal to the optical axis. The processing circuitry 130 can arrange the left image 816 and the right image 818 of the stereoscopic image pair 814 on the array of subpixels such that the parallaxgenerating optic directs the left image 816 and the right image 818 to respective eyes of the viewer at the location of the viewer.
[0062] To further illustrate the system and method disclosed herein, a nonlimiting list of examples is provided below. Each of the following non limiting examples can stand on its own or can be combined in any permutation or combination with any one or more of the other examples.
[0063] In Example 1, an autostereoscopic display system can comprise processing circuitry configured to perform operations. The operations can comprise: receiving data specifying a monoscopic image; determining, from the monoscopic image, saliency information that indicates a region of interest of the monoscopic image; determining a convergence plane for a stereoscopic representation of the monoscopic image, the convergence plane being a function of a disparity characteristic of the region of interest of the monoscopic image; generating a disparity map based on the monoscopic image and the convergence plane; and generating a stereoscopic image pair using the monoscopic image and the disparity map.
[0064] In Example 2, the autostereoscopic display system of Example 1 can optionally be configured such that the operations further comprise: generating a provisional disparity map based on the monoscopic image; generating a saliency map based on the monoscopic image; and generating the disparity map based on the provisional disparity map and the saliency map.
[0065] In Example 3, the autostereoscopic display system of any one of Examples 1-2 can optionally be configured such that the operations further comprise: the provisional disparity map includes values that correspond to depth information within the monoscopic image and include the disparity characteristic of the monoscopic image; the saliency map includes values that correspond to importance to a human visual system and determine the region of interest of the monoscopic image; and the disparity map defines the convergence plane as a plane for which values of the disparity map equal zero.
[0066] In Example 4, the autostereoscopic display system of any one of Examples 1-3 can optionally be configured such that the operations further comprise: generating the disparity map using values of the provisional disparity map that are weighted by respective values of the saliency map.
[0067] In Example 5, the autostereoscopic display system of any one of Examples 1-4 can optionally be configured such that the operations further comprise: generating the disparity map using values of the provisional disparity map in the region of interest and without using values of the provisional disparity map outside the region of interest.
[0068] In Example 6, the autostereoscopic display system of any one of Examples 1-5 can optionally be configured such that: values of the saliency map arebinary; values of the saliency map include a non-zero value in the region of interest and a zero value outside the region of interest; and the operations further comprise: multiplying the values of the provisional disparity map by corresponding values of the saliency map to form a weighted disparity map; averaging values of the weighted disparity map to form a mean value; and subtracting the mean value from the values of the weighted disparity map to form the disparity map.
[0069] In Example 7, the autostereoscopic display system of any one of Examples 1-6 can optionally be configured such that the region of interest includes less than all of the monoscopic image.
[0070] In Example 8, the autostereoscopic display system of any one of Examples 1-7 can optionally be configured such that the region of interest includes one or more regions of deemed importance.
[0071] In Example 9, the autostereoscopic display system of any one of Examples 1-8 can optionally further comprise: an autostereoscopic display configured to display the stereoscopic image pair such that when the stereoscopic image pair is displayed on the autostereoscopic display: the convergence plane coincides with a plane of the autostereoscopic display; a first portion of the stereoscopic image pair appears to be located above the plane of the autostereoscopic display; and a second portion of the stereoscopic image pair appears to be located below the plane of the autostereoscopic display; and a viewer tracker configured to determine a location of a viewer.
[0072] In Example 10, the autostereoscopic display system of any one of Examples 1-9 can optionally be configured such that the autostereoscopic display comprises: a display panel having an array of subpixels configured to display the stereoscopic image pair; and a parallax-generating optic configured to direct light from the display panel to the viewer, the parallax-generating optic including one of a lenticular lens or a parallax barrier having transmissive slits, the parallax-generating optic being invariant along an optical axis having a slant angle relative to the display panel, the slant angle being within a specified angular tolerance of forty -five degrees, the parallaxgenerating optic being periodic along an axis orthogonal to the optical axis; and wherein the operations further comprise: arranging left and right images of the stereoscopic image pair on the array of subpixels such that the parallax-generating optic directs the left and right images to respective eyes of the viewer at the location of the viewer.
[0073] In Example 11, a method for generating a stereoscopic image pair from a monoscopic image can comprise: receiving, with processing circuitry, data specifying the monoscopic image; determining, with the processing circuitry, from the monoscopic image, saliency information that indicates a region of interest of the monoscopic image; determining, with the processing circuitry, a convergence plane for a stereoscopic representation of the monoscopic image, the convergence plane being a function of a disparity characteristic of the region of interest of the monoscopic image; generating, with the processing circuitry, a disparity map based on the monoscopic image and the convergence plane; and generating, with the processing circuitry, the stereoscopic image pair using the monoscopic image and the disparity map.
[0074] In Example 12, the method of Example 11 can optionally further comprise: generating, with the processing circuitry, a provisional disparity map based on the monoscopic image; generating, with the processing circuitry, a saliency map based on the monoscopic image; and generating, with the processing circuitry, the disparity map based on the provisional disparity map and the saliency map.
[0075] In Example 13, the method of any one of Examples 11-12 can optionally be configured such that: the provisional disparity map includes values that correspond to depth information within the monoscopic image and include the disparity characteristic of the monoscopic image; the saliency map includes values that correspond to importance to a human visual system and determine the region of interest of the monoscopic image; and the disparity map defines the convergence plane as a plane for which values of the disparity map equal zero.
[0076] In Example 14, the method of any one of Examples 11-13 can optionally further comprise: generating the disparity map using values of the provisional disparity map that are weighted by respective values of the saliency map.
[0077] In Example 15, the method of any one of Examples 11-14 can optionally further comprise: generating, with the processing circuitry, the disparity map using values of the provisional disparity map in the region of interest and without using values of the provisional disparity map outside the region of interest.
[0078] In Example 16, a non-transitory computer-readable storage medium can store instructions for generating a stereoscopic image pair from a monoscopic image. The instructions, when executed by processing circuitry, can cause the processing circuitry to perform operations. The operations can comprise: receiving, with theprocessing circuitry, data specifying the monoscopic image; determining, with the processing circuitry, from the monoscopic image, saliency information that indicates a region of interest of the monoscopic image; determining, with the processing circuitry, a convergence plane for a stereoscopic representation of the monoscopic image, the convergence plane being a function of a disparity characteristic of the region of interest of the monoscopic image; generating, with the processing circuitry, a disparity map based on the monoscopic image and the convergence plane; and generating, with the processing circuitry, the stereoscopic image pair using the monoscopic image and the disparity map.
[0079] In Example 17, the non-transitory computer-readable storage medium of Example 16 can optionally be configured such that the operations further comprise: generating, with the processing circuitry, a provisional disparity map based on the monoscopic image; generating, with the processing circuitry, a saliency map based on the monoscopic image; and generating, with the processing circuitry, the disparity map based on the provisional disparity map and the saliency map.
[0080] In Example 18, the non-transitory computer-readable storage medium of any one of Examples 16-17 can optionally be configured such that: the provisional disparity map includes values that correspond to depth information within the monoscopic image and include the disparity characteristic of the monoscopic image; the saliency map includes values that correspond to importance to a human visual system and determine the region of interest of the monoscopic image; and the disparity map defines the convergence plane as a plane for which values of the disparity map equal zero.
[0081] In Example 19, the non-transitory computer-readable storage medium of any one of Examples 16-18 can optionally be configured such that the operations further comprise: generating the disparity map using values of the provisional disparity map that are weighted by respective values of the saliency map.
[0082] In Example 20, the non-transitory computer-readable storage medium of any one of Examples 16-19 can optionally be configured such that the operations further comprise: generating, with the processing circuitry, the disparity map using values of the provisional disparity map in the region of interest and without using values of the provisional disparity map outside the region of interest.
[0083] The above detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings show, by way of illustration, specific embodiments in which the invention can be practiced. These embodiments are also referred to herein as "examples." Such examples can include elements in addition to those shown or described. However, examples are contemplated in which only those elements shown or described are provided. Moreover, other examples can any combination or permutation of those elements shown or described (or one or more aspects thereof), either with respect to a particular example (or one or more aspects thereof), or with respect to other examples (or one or more aspects thereof) shown or described herein.
[0084] In this document, the terms "a" or "an" are used, as is common in patent documents, to include one or more than one, independent of any other instances or usages of "at least one" or "one or more." In this document, the term "or" is used to refer to a nonexclusive or, such that "A or B" can include "A but not B," "B but not A," and "A and B," unless otherwise indicated. In the appended claims, the terms "including" and "in which" are used as the plain-English equivalents of the respective terms "comprising" and "wherein". Also, in the following claims, the terms "including" and "comprising" are open-ended, that is, a system, device, article, or process that includes elements in addition to those listed after such a term in a claim are still deemed to fall within the scope of that claim. Moreover, in the following claims, the terms "first," "second," and "third," etc. are used merely as labels, and are not intended to impose numerical requirements on their objects.
[0085] The above description is intended to be illustrative, and not restrictive. For example, the above-described examples (or one or more aspects thereof) can be used in combination with each other. Other embodiments can be used, such as by one of ordinary skill in the art upon reviewing the above description. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Also, in the above Detailed Description, various features can be grouped together to streamline the disclosure. This should not be interpreted as intending that an unclaimed disclosed feature is essential to any claim. Rather, inventive subject matter can lie in less than all features of a particular disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separate embodiment, and it is contemplated that such embodiments canbe combined with each other in various combinations or permutations. The scope of the invention should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. An autostereoscopic display system, comprising: processing circuitry configured to perform operations, the operations comprising: receiving data specifying a monoscopic image; determining, from the monoscopic image, saliency information that indicates a region of interest of the monoscopic image; determining a convergence plane for a stereoscopic representation of the monoscopic image, the convergence plane being a function of a disparity characteristic of the region of interest of the monoscopic image; generating a disparity map based on the monoscopic image and the convergence plane; and generating a stereoscopic image pair using the monoscopic image and the disparity map.
2. The autostereoscopic display system of claim 1, wherein the operations further comprise: generating a provisional disparity map based on the monoscopic image; generating a saliency map based on the monoscopic image; and generating the disparity map based on the provisional disparity map and the saliency map.
3. The autostereoscopic display system of claim 2, wherein: the provisional disparity map includes values that correspond to depth information within the monoscopic image and include the disparity characteristic of the monoscopic image; the saliency map includes values that correspond to importance to a human visual system and determine the region of interest of the monoscopic image; and the disparity map defines the convergence plane as a plane for which values of the disparity map equal zero.
4. The autostereoscopic display system of claim 3, wherein the operations further comprise: generating the disparity map using values of the provisional disparity map that are weighted by respective values of the saliency map.
5. The autostereoscopic display system of claim 3, wherein the operations further comprise: generating the disparity map using values of the provisional disparity map in the region of interest and without using values of the provisional disparity map outside the region of interest.
6. The autostereoscopic display system of claim 3, wherein: values of the saliency map are binary; values of the saliency map include a non-zero value in the region of interest and a zero value outside the region of interest; and the operations further comprise: multiplying the values of the provisional disparity map by corresponding values of the saliency map to form a weighted disparity map; averaging values of the weighted disparity map to form a mean value; and subtracting the mean value from the values of the weighted disparity map to form the disparity map.
7. The autostereoscopic display system of claim 3, wherein the region of interest includes less than all of the monoscopic image.
8. The autostereoscopic display system of claim 3, wherein the region of interest includes one or more regions of deemed importance.
9. The autostereoscopic display system of claim 1, further comprising: an autostereoscopic display configured to display the stereoscopic image pair such that when the stereoscopic image pair is displayed on the autostereoscopic display: the convergence plane coincides with a plane of the autostereoscopic display; a first portion of the stereoscopic image pair appears to be located above the plane of the autostereoscopic display; and a second portion of the stereoscopic image pair appears to be located below the plane of the autostereoscopic display; and a viewer tracker configured to determine a location of a viewer.
10. The autostereoscopic display system of claim 9, wherein the autostereoscopic display comprises: a display panel having an array of subpixels configured to display the stereoscopic image pair; and a parallax-generating optic configured to direct light from the display panel to the viewer, the parallax-generating optic including one of a lenticular lens or a parallax barrier having transmissive slits, the parallax-generating optic being invariant along an optical axis having a slant angle relative to the display panel, the slant angle being within a specified angular tolerance of forty -five degrees, the parallax-generating optic being periodic along an axis orthogonal to the optical axis; and wherein the operations further comprise: arranging left and right images of the stereoscopic image pair on the array of subpixels such that the parallax-generating optic directs the left and right images to respective eyes of the viewer at the location of the viewer.
11. A method for generating a stereoscopic image pair from a monoscopic image, the method comprising: receiving, with processing circuitry, data specifying the monoscopic image; determining, with the processing circuitry, from the monoscopic image, saliency information that indicates a region of interest of the monoscopic image; determining, with the processing circuitry, a convergence plane for a stereoscopic representation of the monoscopic image, the convergence plane being a function of a disparity characteristic of the region of interest of the monoscopic image; generating, with the processing circuitry, a disparity map based on the monoscopic image and the convergence plane; and generating, with the processing circuitry, the stereoscopic image pair using the monoscopic image and the disparity map.
12. The method of claim 11, further comprising: generating, with the processing circuitry, a provisional disparity map based on the monoscopic image; generating, with the processing circuitry, a saliency map based on the monoscopic image; and generating, with the processing circuitry, the disparity map based on the provisional disparity map and the saliency map.
13. The method of claim 12, wherein: the provisional disparity map includes values that correspond to depth information within the monoscopic image and include the disparity characteristic of the monoscopic image; the saliency map includes values that correspond to importance to a human visual system and determine the region of interest of the monoscopic image; and the disparity map defines the convergence plane as a plane for which values of the disparity map equal zero.
14. The method of claim 13, further comprising: generating the disparity map using values of the provisional disparity map that are weighted by respective values of the saliency map.
15. The method of claim 13, further comprising: generating, with the processing circuitry, the disparity map using values of the provisional disparity map in the region of interest and without using values of the provisional disparity map outside the region of interest.
16. A non-transitory computer-readable storage medium storing instructions for generating a stereoscopic image pair from a monoscopic image, the instructions, when executed by processing circuitry, cause the processing circuitry to perform operations, the operations comprising: receiving, with the processing circuitry, data specifying the monoscopic image; determining, with the processing circuitry, from the monoscopic image, saliency information that indicates a region of interest of the monoscopic image; determining, with the processing circuitry, a convergence plane for a stereoscopic representation of the monoscopic image, the convergence plane being a function of a disparity characteristic of the region of interest of the monoscopic image; generating, with the processing circuitry, a disparity map based on the monoscopic image and the convergence plane; and generating, with the processing circuitry, the stereoscopic image pair using the monoscopic image and the disparity map.
17. The non-transitory computer-readable storage medium of claim 16, wherein the operations further comprise: generating, with the processing circuitry, a provisional disparity map based on the monoscopic image; generating, with the processing circuitry, a saliency map based on the monoscopic image; and generating, with the processing circuitry, the disparity map based on the provisional disparity map and the saliency map.
18. The non-transitory computer-readable storage medium of claim 17, wherein: the provisional disparity map includes values that correspond to depth information within the monoscopic image and include the disparity characteristic of the monoscopic image; the saliency map includes values that correspond to importance to a human visual system and determine the region of interest of the monoscopic image; and the disparity map defines the convergence plane as a plane for which values of the disparity map equal zero.
19. The non-transitory computer-readable storage medium of claim 18, wherein the operations further comprise: generating the disparity map using values of the provisional disparity map that are weighted by respective values of the saliency map.
20. The non-transitory computer-readable storage medium of claim 18, wherein the operations further comprise: generating, with the processing circuitry, the disparity map using values of the provisional disparity map in the region of interest and without using values of the provisional disparity map outside the region of interest.
Citation Information
Patent Citations
Method and system for converting 2d image data to stereoscopic image data
US20110050853A1
Establishment method of 3D Saliency Model Based on Prior Knowledge and Depth Weight
US20180182118A1
Image Processing Method and Device
US20210118111A1
Method, system and computer program product for encoding disparities between views of a stereoscopic image
US20220377362A1
Discontinuous warping for 2D-to-3D conversions
US8666146B1