View synthesis system and method using a depth map

The method of ray casting and fragment shaders optimizes the conversion of single-view to multi-view images, addressing computational intensity and quality issues, enabling efficient real-time generation of high-quality multi-view images for multi-view displays.

JP7702567B2Active Publication Date: 2025-07-03LEIA INC

Patent Information

Application Number
JP2024506192
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-03
Filing Date
2022-07-28
Publication Date
2025-07-03
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently convert single-view images into multi-view images in real-time for multi-view displays, leading to computational intensity and quality issues such as holes and artifacts at object edges.

Method used

A method utilizing ray casting and fragment shaders to synthesize multi-view images from a color image and depth map, optimizing ray progression with predetermined intervals and threshold comparisons to reduce computational load and improve image quality.

Benefits of technology

Enables real-time generation of high-quality multi-view images with reduced computational requirements, minimizing holes and artifacts, suitable for mobile devices and multi-view displays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702567000001
    Figure 0007702567000001
  • Figure 0007702567000002
    Figure 0007702567000002
  • Figure 0007702567000003
    Figure 0007702567000003
Patent Text Reader

Abstract

In multi-view image generation and display, a computing device may synthesize a view image of a multi-view image of a scene from a color image and a depth map. Each view image may include a color value at a respective pixel location. The computing device may render the view images of the multi-view image on a multi-view display. Synthesizing the view images may include, for a pixel location in the view image, the following operations: The computing device may cast a ray from the pixel location toward the scene in a direction corresponding to the view direction of the view image. The computing device may determine a ray intersection location where the ray intersects a virtual surface specified by the depth map. The computing device may set a color value of the view image at the pixel location to correspond to a color of the color image at the ray intersection location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 229,054, filed Aug. 3, 2021, which is hereby incorporated by reference in its entirety.

[0002] Statement Regarding Federally Sponsored Research or Development Not applicable

Background Art

[0003] A scene in three - dimensional (3D) space can be viewed from multiple viewpoints depending on the viewing angle. Additionally, when viewed by a user, multiple views representing different viewpoints of the scene can be simultaneously perceived, effectively creating a sense of depth that can be perceived by the user. A multi - view display is capable of rendering and displaying multi - view images such that multiple views can be simultaneously perceived. Some content can be natively captured as multi - view images or multi - view videos, while multi - view images or multi - view videos can be generated from a variety of other sources.

[0004] Various features of the examples and embodiments according to the principles described herein can be more readily understood with reference to the following detailed description in conjunction with the accompanying drawings, where the same reference numbers indicate the same structural elements.

Brief Description of the Drawings

[0005]

Figure 1

[0006]

Figure 2

[0007]

Figure 3

[0008]

Figure 4

[0009]

Figure 5

[0010]

Figure 6

DETAILED DESCRIPTION OF THE INVENTION

[0011] Certain examples and embodiments have other features in addition to, and instead of, the features illustrated in the figures mentioned above. These and other features are detailed below with reference to the figures mentioned above.

[0012] Examples and embodiments according to the principles described herein provide techniques for generating multi-view images from a single view image and a depth map in a real-time rendering pipeline. This enables visual content (e.g., images or videos) to be converted on-the-fly into a multi-view format and presented to the user. As will be described below, embodiments are involved in real-time view synthesis using color images and depth inputs.

[0013] According to an embodiment, the shader program implements a subroutine to calculate pixel values for each synthesized view. The subroutine may use ray casting to synthesize a view from a monochromatic (or grayscale) image and a depth map. Each pixel in the synthesized view may be determined by casting individual rays at an angle corresponding to the view position onto the depth map. In a particular view, all the rays of the view have the same direction such that the rays are parallel. Each ray may correspond to a pixel that is to be rendered as part of the synthesized view. For each ray, the subroutine steps through points along the ray to determine where the ray intersects the virtual surface by reading a depth value from the depth map. At each step along the ray, the horizontal position is incremented by a predetermined original interval, and the depth position is incremented by a predetermined depth interval. In some embodiments, the vertical position may remain constant. The location where the ray intersects the virtual surface specifies the location of the coordinates used to read a color from a color map corresponding to the pixel of the view, and the pixel position corresponds to the origin of the ray. In some embodiments, when the ray for a particular pixel falls below the virtual surface by a threshold amount, the pixel may be set as a hole rather than representing a color.

[0014] The shader program may be a program executed by a graphics processing unit (GPU). The GPU may include one or more vector processors that execute instructions configured to perform various subroutines in parallel. In this regard, a single subroutine may be configured to calculate the individual pixel values of the pixels of a view image as it is being synthesized. Some instances of the subroutine may be executed in parallel to calculate the pixel values of all the pixels of a view.

[0015] Embodiments are directed to synthesizing a view image, which may include accessing a color image and a corresponding depth map to synthesize the view image, where the depth map defines a virtual surface. Synthesizing the view image may further include stepping through a plurality of points along a ray at a predetermined horizontal interval, a predetermined depth interval, and a constant vertical value until one of the plurality of points along the ray is identified as being located on the virtual surface, where the ray includes a ray origin and a ray direction, the ray origin defines the pixel position of a pixel of the view image to be rendered, and the ray direction corresponds to the view position of the view image. Synthesizing the view image may also include rendering a pixel of the view image by sampling a color value of the color image at the hit point.

[0016] More specifically, the color image may be an RGB image, in which case the color image specifies pixel values for different pixels having respective coordinates. The pixel values may be a value indicating an amount of red, a value indicating an amount of green, a value indicating an amount of blue, or any combination thereof. The depth map may have a similar format to the color image, but instead of specifying color, the depth map specifies depth values for different pixels having respective coordinates. Thus, the depth map may define a virtual surface that varies along the depth axis as a function of horizontal and vertical position. In other words, from a top-down orientation, the depth may vary along the depth axis (e.g., into or out of the screen), while the position may vary horizontally (e.g., left or right) or vertically (e.g., up or down).

[0017] One or more view images can be synthesized from a color image and a depth map. The view images correspond to specific viewpoints that are different from each viewpoint of the other view images. A set of view images forms a multi-view image, and the multi-view image represents an object or a scene that includes different viewpoints. Each view image of the multi-view image corresponds to a respective view position. For example, if the multi-view image includes four views, the view positions can range from left to right such that they include the leftmost view, the center-left view, the center-right view, and the rightmost view. The distance between each view can be referred to as a gain or a baseline.

[0018] The subroutine can be executed to calculate pixel values (e.g., RGB colors) when each pixel in the view image is being synthesized and rendered. The subroutine can include stepping through a plurality of points along a ray at a predetermined horizontal interval. In this regard, the subroutine moves along a ray cast from a ray origin in a particular ray direction, and the ray is cast towards a virtual surface. The ray can be defined as a line in a space defined by coordinates indicating a particular direction. When synthesizing a particular view image, parallel rays are cast from various origins towards the virtual surface of the depth map to detect depth information for the view image when the view image is being synthesized. Instead of reading out all possible depth values along the ray path from the origin to the bottom of the scene, embodiments are directed to stepping along the ray at points defined by a predetermined horizontal interval, a predetermined depth interval, and a constant vertical value. Further, the subroutine automatically ends or is otherwise interrupted when the subroutine identifies a point on the ray where the ray is located on the virtual surface (e.g., the hit point where the ray intersects the virtual surface). In other words, the subroutine steps along the ray at a predetermined interval until the ray intersects or otherwise hits the virtual surface which is the hit point. In this regard, the depth at the point which is the hit point is equal to the corresponding depth of the virtual surface. The hit point can be slightly deeper than the virtual surface or within the depth tolerance near the virtual surface. When using a predetermined depth interval, the quantized depth can approximate the location of the intersection of the ray and the virtual surface.

[0019] When this hit point is identified, the location is recorded from the perspective of coordinates. To render the pixel, the color is sampled from the color image using the recorded location of the hit point. Thus, the subroutine renders the pixel of the view image by sampling the color value of the color image at the hit point coordinates.

[0020] In some embodiments, the hit point is identified as being located on the virtual surface by comparing the depth value readout of the hit point with a threshold depth. For example, as the subroutine moves to the next point on the ray away from the origin, the subroutine reads the depth at a particular point to obtain a depth value readout. The hit point is determined by comparing the depth value readout with the threshold depth. In some embodiments, the threshold depth is decreased by a predetermined depth interval at each step. In this regard, the threshold depth can be equal to or separately derived from the depth coordinate of the corresponding point.

[0021] In some embodiments, synthesizing the view image further includes detecting that additional rays among the plurality of rays fall below the virtual surface by a predetermined depth level. Some rays intersect the virtual surface while other rays can hit the vertical edge of the virtual surface that can be considered a "hole" depending on the steepness threshold. In some embodiments, when the hit point is on a steep surface, the subroutine a picture corresponding to that position element is a hole can be specified. When the color image is a frame of the video stream, these holes can be filled using pixel information from temporally adjacent video frames. In the case of a still image, the surrounding color can be used for hole filling.

[0022] In some embodiments, the fragment shader is configured to step through a plurality of points along the ray. The fragment shader can be configured to combine the view image with at least one other view image. Combining involves spatially multiplexing the pixels of different views to match the combined multi-view format of a multi-view display. The fragment shader can be executed by a processor that renders a particular color to the output pixel coordinates. The color can be determined by sampling the color image at a location determined by processing the depth map.

[0023] In some embodiments, the predetermined horizontal spacing is one pixel, and the predetermined depth spacing is a function of the baseline of the plurality of synthesized views. In other embodiments, the predetermined horizontal spacing is a function of the baseline of the plurality of synthesized views. By using the baseline to control the granularity of the steps analyzed along the ray, the subroutine can be optimized to reduce the number of depth reads while maintaining the minimum level of image quality when synthesizing the view images.

[0024] In some embodiments, the subroutine can read depth values from the depth texture of the depth map. The subroutine can start from the position of the output pixel coordinates of the ray origin (or a predetermined offset from this position, which can be specified by the convergence input parameter). The subroutine can perform per-point reads of depth values in one direction (e.g., from left to right) or the opposite direction (e.g., from right to left) depending on the view position.

[0025] At each point along the ray, the extracted depth value is compared with a threshold depth that starts from 1 and progresses towards zero in steps. Such a search for the hit point ends when the depth value read is equal to or exceeds the threshold depth. The coordinates at which the search ends are used to sample the color from the color image and return it as the result of the subroutine. The fragment shader uses this sampled color to render the pixels of the synthesized view.

[0026] The maximum number of horizontal steps (e.g., the predetermined horizontal spacing) can be modified or controlled by the baseline input parameter. The more steps there are, the more generated views branch from the input color image. The threshold depth reduction step may be equal to 1 divided by the number of horizontal steps.

[0027] Additionally, to improve anti-aliasing at the edges, a predetermined horizontal interval can be divided into sub-intervals such as the horizontal interval, and as a result, the coordinates for sampling colors can be adjusted to be more precise. Adjacent depth value readings can be used with linear interpolation between them, and these adjacent depth value readings can be used during comparison. A predetermined depth interval can be divided into sub-intervals. The resulting coordinates of the hit points (e.g., the depth value readings are equal to or exceed the depth threshold) can be passed to a shader texture read function that supports built-in color interpolation.

[0028] Additionally, if it is necessary to create holes at opposite edges (as a result, the rendering appears as if the separate objects are moving independently), the difference between the depth value reading and the depth threshold can be compared, and if it is greater than a predetermined depth level, holes can be created for subsequent impainting.

[0029] In an example, a computing device can receive a color image of a scene and a depth map of the scene. The computing device can synthesize view images of a multi-view image of the scene from the color image and the depth map. The view images can represent the scene from different view directions. Each view image can include a pixel position and respective color values at the pixel position. For example, a view image can have a two-dimensional (2D) grid of pixels, and each pixel has a location within the image. The computing device can render the view images of the multi-view image on a multi-view display of the computing device. Synthesizing the view images of the multi-view image can include casting a ray in a direction corresponding to the view direction of the view image from the pixel position towards the scene for pixel positions within the view image, determining a ray intersection position where the ray intersects a virtual surface specified by the depth map, and setting the color value of the view image at the pixel position to correspond to the color of the color image at the ray intersection position.

[0030] As used herein, a "two-dimensional display" or "2D display" is defined as a display configured to provide views of an image that are substantially the same regardless of the direction in which the image is viewed (i.e., within a predefined viewing angle or within the range of the 2D display). Conventional liquid crystal displays (LCDs) found in many smartphones and computer monitors are examples of 2D displays. In contrast, as used herein, a "multi-view display" is defined as an electronic display or display system configured to provide different views of a multi-view image in different viewing directions or from different viewing directions. In particular, the different views may represent different perspectives of a scene or object in the multi-view image. The use of the one-sided backlighting and one-sided multi-view displays described herein includes, but is not limited to, cellular phones (e.g., smartphones), watches, tablet computers, mobile computers (e.g., laptop computers), personal computers and computer monitors, automotive display controllers, cameras, displays, and various other portable as well as substantially non-portable display applications and devices.

[0031] FIG. 1 illustrates a perspective view of a multi-view display 10 in an example according to an embodiment consistent with the principles described herein. As illustrated in FIG. 1, the multi-view display 10 includes a screen 12 for displaying a multi-view image to be viewed. The screen 12 can be, for example, a display screen of a phone (e.g., a cellular phone, smartphone, etc.), a tablet computer, a laptop computer, a computer monitor of a desktop computer, a camera display, or an electronic display of substantially any other device.

[0032] The multi-view display 10 provides different views 14 of a multi-view image in different view directions 16 with respect to the screen 12. The view directions 16 are illustrated as arrows extending from the screen 12 in various different major angular directions, and the different views 14 are illustrated as shaded polygonal boxes at the ends of the arrows (i.e., depicting the view directions 16), and only 4 views 14 and 4 view directions 16 are illustrated, all of which are examples and not limitations. Although the different views 14 are illustrated in FIG. 1 as being above the screen, it should be noted that the views 14 actually appear on or near the screen 12 when the multi-view image is displayed on the multi-view display 10. Depicting the views 14 above the screen 12 is for simplicity of illustration only and is meant to represent looking at the multi-view display 10 from each of the view directions 16 corresponding to the particular views 14. A 2D display may be substantially similar to the multi-view display 10 except that, in contrast to the different views 14 of the multi-view image provided by the multi-view display 10, a 2D display is generally configured to provide a single view of the displayed image (e.g., one view similar to the view 14).

[0033] A light beam having a view direction, or equivalently, a direction corresponding to the view direction of a multi-view display, generally has a major angular direction given by angular components {θ,φ} according to the definitions herein. The angular component θ is referred to herein as the "elevation angle component" or "elevation angle" of the light beam. The angular component φ is referred to as the "azimuth angle component" or "azimuth angle" of the light beam. By definition, the elevation angle θ is an angle in a vertical plane (e.g., perpendicular to the plane of the multi-view display screen), while the azimuth angle φ is an angle in a horizontal plane (e.g., parallel to the multi-view display screen plane).

[0034] Figure 2 illustrates a graphic representation of the angular components {θ, φ} of a light beam 20 having a particular principal angular direction corresponding to the view direction (e.g., view direction 16 in FIG. 1) of a multi-view display in an example, according to an embodiment consistent with the principles described herein. Additionally, the light beam 20 is, by definition within this specification, emitted from or originates at a particular point. That is, by definition, the light beam 20 has a central ray associated with a particular point of origin within the multi-view display. FIG. 2 also illustrates the light beam (or view direction) point of origin O.

[0035] Furthermore, as used herein, the article “a” is intended to have its original meaning in patent technology, i.e., “one or more.” For example, “a computing device” means one or more computing devices, and as such, “the computing device” means “the computing devices” herein. Also, any reference in this specification to “up,” “down,” “above,” “below,” “upward,” “downward,” “front,” “back,” “first,” “second,” “left,” or “right” is not intended to be limiting within this specification. As used herein, the term “about,” when applied to a value, generally means within the tolerance range of the equipment used to generate that value, or, unless otherwise explicitly specified, can mean plus or minus 10%, or plus or minus 5%, or plus or minus 1%. Further, the term “substantially,” as used herein, means most, or almost all, or all, or an amount within the range of about 51% to about 100%. Additionally, the examples within this specification are intended to be merely illustrative and are presented for purposes of discussion and not limitation.

[0036] FIG. 3 shows a block diagram of an example of a system 100 that can perform multi-view image generation and display, according to an embodiment consistent with the principles described herein. The system 100 may include a computing device 102, such as a smartphone, tablet, laptop computer, and the like. FIG. 6 and the accompanying text below describe an example of the computing device 102 in detail. The computing device 102 may be configured to execute a computer-implemented method of multi-view image generation and display, as described herein.

[0037] According to various embodiments, a computer-implemented method of multi-view image generation when executed by a computing device 102 includes receiving a color image 104 of a scene 106 and a depth map 108 of the scene 106. In FIG. 3, the scene 106 is depicted as a cat, although other suitable scenes may be used. The color image 104 may include intensity and color data representing the appearance of the scene 106. The color image data may optionally be arranged to correspond to a rectangular array of locations, such as pixels. In some examples, the intensity and color data may include a first value representing the intensity of red light, a second value representing the intensity of green, and a third value representing the intensity of blue. In some examples, for a scene that is monochromatic, such as in some cases, the intensity and color data may include a value representing intensity (e.g., including intensity information but no color information). The depth map 108 may account for the relative distance between locations within the scene 106 and the computing device 102. In the case of the example of the cat in the scene 106, the depth map 108 corresponding to the scene 106 may specify that the tip of the cat's tail may be farther from the computing device 102 than the cat's right hind leg. In some examples, the color image 104 and the depth map 108 may be received from a server or storage device via a wired or wireless connection. In some examples, the color image 104 and the depth map 108 may be generated by a camera and a depth map generation device included with the computing device 102. In some examples, the depth map generation device may utilize time-of-flight reflections in different directions to map the distance from the computing device 102 for various propagation directions away from the computing device 102.

[0038] In various embodiments, a computer-implemented method of multi-view image generation when executed by a computing device 102 further includes synthesizing view images 110A, 110B, 110C, 110D (collectively referred to as view images 110) of a multi-view image of a scene 106 from a color image 104 and a depth map 108. The view images 110 may represent the scene 106 from different view directions. In the case of an example of a cat in the scene 106, the view images 110 may represent how the cat would look when viewed from the view direction corresponding to the view image 110. Each view image 110 may include a pixel position and respective color values at the pixel position. In the example of FIG. 3, there are four view images 110. In other examples, more or fewer than four view images 110 may be used. FIG. 4 and the accompanying text below provide further details regarding how the view images 110 are synthesized.

[0039] In various embodiments, a computer-implemented method of multi-view image generation when executed by a computing device 102 further includes rendering view images 110 of the multi-view image onto a multi-view display 112 of the computing device 102. The view images 110A, 110B, 110C, and 110D are each viewed from respective view directions 114A, 114B, 114C, and 114D. For example, the computing device 102 may be configured as a smartphone, and the display of the smartphone may be configured as the multi-view display 112. In some examples, the multi-view image may be included as a frame within a video such that the computing device 102 can synthesize the view images 110 and render the view images 110 at a suitable video frame rate, such as 60 frames per second, 30 frames per second, or the like.

[0040] In this specification, a set of computer-implemented instructions, such as those referred to as shaders, can synthesize view images of a multi-view image. A detailed description of the shader follows, but an overview of the shader follows here. The shader can cast a ray in a direction corresponding to the view direction of the view image from the pixel position towards the scene. The shader is configured to determine a ray intersection position where the ray intersects a virtual surface specified by the depth map. The shader is further configured to set the color value of the view image at the pixel position to correspond to the color of the color image at the ray intersection position.

[0041] In some examples, determining the ray intersection position can include the following operations shown as the first to fifth operations for simplicity. In the first operation, the shader can determine sequential tentative positions along the ray between the pixel position and the specified plane such that the virtual surface is between the pixel position and the specified plane. In the second operation, the shader can identify a particular one of the tentative positions along the ray. In the third operation, the shader can repeatedly determine that the identified particular tentative position is between the pixel position and the specified plane, and the shader can advance the identified particular tentative position to the next tentative position along the ray. In the fourth operation, the shader can virtual plane determine that is between the pixel position and the identified particular tentative position. In the fifth operation, the shader can set the ray intersection position to comprehensively correspond to a position between the identified particular tentative position and the adjacent previously identified tentative position.

[0042] In some of the above examples, determining the ray intersection position may include the following operations shown as the sixth to tenth operations for simplicity. The sixth to tenth operations may effectively repeat the first to fifth operations, but may have different (e.g., finer) resolutions. In the sixth operation, the shader may determine sequential second tentative positions along the ray between the identified tentative position and an adjacent previously identified tentative position. In the seventh operation, the shader may identify one of the second tentative positions among the second tentative positions along the ray. In the eighth operation, the shader may repeatedly determine that the identified second tentative position is between the pixel position and the specified plane, and may advance the identified second tentative position to the next second tentative position along the ray. In the ninth operation, the shader may virtual plane determine that it is between the pixel position and the identified second tentative position. In the tenth operation, the shader may set the ray intersection position to comprehensively correspond to the position between the identified second tentative position and an adjacent previously identified second tentative position.

[0043] In some examples, the tentative positions may be equally spaced along the ray. In some examples, the view image may define a horizontal direction parallel to the upper and lower edges of the view image, a vertical direction within the plane of the view image and orthogonal to the horizontal direction, and a depth orthogonal to both the horizontal and vertical directions. In some examples, the tentative positions may be spaced such that the horizontal component of the interval between adjacent tentative positions corresponds to a specified value. In some examples, the specified value may correspond to the horizontal interval between adjacent pixels within the view image.

[0044] In some examples, different view directions are at the upper and lower edges of the view image parallel toIt may be in the horizontal plane. In some examples, one or both of the first to fifth operations and the first to tenth operations may not result in an executable outcome. Since ray casting alone may not be able to obtain a suitable color value for the pixel position, the shader may perform additional operations to obtain a suitable color value. These additional operations are shown as the eleventh to fifteenth operations for simplicity and are detailed below.

[0045] In the eleventh operation, the shader may cast a ray in a direction corresponding to the view direction of the view image from the pixel position towards the depth map representing the scene. In the twelfth operation, the shader may determine that the ray does not intersect the virtual surface specified by the depth map. In the thirteenth operation, the shader may set the color value of the view image at the pixel position to correspond to the acquired color information. In the fourteenth operation, the shader may determine that the ray does not intersect the virtual surface specified by the depth map by determining that the ray has propagated away from the pixel position by a distance exceeding a threshold distance.

[0046] In some examples, the view image may correspond to a time-series image of a video signal. In some examples, the color information may be acquired from pixel positions of at least one temporally adjacent video frame of the video signal. The following description relates to the details of computer-implemented operations that can generate and render a view image, such as including or using a shader. The use of the computer-implemented operations is to create an image with multiple views, such as four or more views. The multiple views can be arranged as tiles in various patterns, such as a 2×2 pattern, a 1×4 pattern, and a 3×3 pattern, although not limited thereto. The computer-implemented operations can synthesize multiple views from a color image and an accompanying depth map. The depth map can be formed as a grayscale image whose luminance or intensity represents proximity to the viewer or camera.

[0047] Computer-implemented operations can synthesize multiple views such that when a human views a pair of views, the human can perceive a stereo effect. Since humans typically have horizontally separated eyes, computer-implemented operations can synthesize multiple views that can be viewed from different horizontally separated locations. Computer-implemented operations can reduce or eliminate artifacts at the edges of shapes or images. Such artifacts can interfere with the stereo effect.

[0048] Since the synthesized image is included in a video image, computer-implemented operations can synthesize multiple views relatively quickly. Further, computer-implemented operations can synthesize multiple views to be compatible with a rendering pipeline such as an OpenGL rendering pipeline. Further, computer-implemented operations can synthesize multiple views without performing heavy computations that can load a mobile device, cause a thermal throttle on a mobile device, or both.

[0049] The computer-implemented operations discussed herein can be more efficient and robust than another technique called "forward mapping", or point cloud rendering using horizontal skew for parallel projection cameras and camera matrices. In forward mapping, each view is created by moving the color of a two-dimensional grid of colored points left and right, or both, according to corresponding depth values. The computer-implemented operations discussed herein can avoid holes, separated points, and empty regions at one or both of the left and right edges of views generated by forward mapping. Further, the computer-implemented operations discussed herein do not require allocating more points than are present on a display device.

[0050] The computer-implemented operations discussed herein can be more efficient and robust than another technique called "backward mapping" in which the color value of a point is replaced with the color value of some neighboring points according to the depth of the point. Backward mapping can create an illusion of depth but does not accurately depict the edges of shapes. For example, the foreground portion of the edge of a shape can be unexpectedly covered by the background.

[0051] The computer-implemented operations discussed herein can be more efficient and robust than equivalent techniques using a three-dimensional mesh. Such a mesh is computationally demanding and thus may not be suitable for real-time generation of multiple views.

[0052] The computer-implemented operations discussed herein can use a modified form of ray casting. Ray casting can cast rays from some center point in 3D space through each point on a virtual screen and can determine the hit point on the surface to set the color of the corresponding point on the screen. For view synthesis, the computer-implemented operations can determine the color of the hit point without any additional ray tracing. Further, the texture can determine the surface without placing voxel data in memory. Ray casting can provide accuracy and the ability to render perspective views. Without the modifications discussed herein, ray casting can be computationally intensive, especially for high-quality images. However, it is a power-consuming operation, especially for high-quality images. Without the modifications discussed herein, ray casting can divide the ray path into hundreds of steps, with a check for each step. Using one or more of the modifications discussed herein, ray casting can use a structure for marking so-called safe zones (e.g., zones without geometry) to accelerate ray progression.

[0053] In some examples, the computer-implemented operations discussed herein may perform raycasting using a single full-screen quad rendering by a relatively simple fragment (pixel) shader. In some examples, the computer-implemented operations discussed herein may achieve real-time performance and may be fed into the rendering chain of a video-multi-viewer workflow.

[0054] The computer-implemented operations discussed herein may use a fragment shader. The fragment shader may be a program that determines the color for a single pixel and stores the determined color information in an output buffer. The fragment shader may be executed multiple times in parallel for each pixel, using the corresponding entry in the output buffer to determine the color of all output pixels.

[0055] An example of pseudocode that may implement the computer-implemented operations discussed herein is as follows. x, y = output_coord view_id = pick one from [-1.5, -0.5, +0.5, +1.5] direction = if (view_id < 0) -1 else +1 total_steps = abs (view_id * gain_px) x = x - convergence * view_id * gain_px z = 1 for _ in total_steps: if (read_depth (x, y) <= z) return read_color (x, y) x += direction z - = 1.0 / total_steps return read_color (x, y)

[0056] Figure 4 shows a graphic representation 400 of a simplified example of a computer-implemented operation discussed herein, according to an embodiment consistent with the principles described herein. It should be understood that the simplified example of Figure 4 shows a one-dimensional example where light rays propagate at a propagation angle (corresponding to the viewing direction) within a single plane. In reality, the depth map can extend two-dimensionally, and the propagation direction can have three dimensions (e.g., two dimensions and depth), etc.

[0057] In the simplified example of Figure 4, the depth away from the viewer is represented as the height from the bottom horizontal line 402. The bottom horizontal line 402 represents the foreground boundary of the depth map. The output buffer 404 is shown below the bottom horizontal line 402. The top horizontal line 406 represents the rear, background plane, or the boundary of the depth map.

[0058] In the simplified example of Figure 4, the depth map is shown as extending across a series of pixels. The depth map is shown as a series of height values with one value per horizontal pixel. In the simplified example of Figure 4, the depth map is quantized to have one of 11 possible values ranging from 0 to 1 inclusively. In reality, the actual depth map values can have a relatively large number of possible values (such as 256) between the foreground and background boundaries. In the specific example of Figure 4, from the leftmost edge to the rightmost edge of the view image, the depth increases from 0.1 to 0.4, then decreases to 0.0, then increases to 0.8, and then decreases to 0.6 at the rightmost edge of the view image.

[0059] In the simplified example of Figure 4, the coordinate x runs along the horizontal direction. The coordinate x can represent an output buffer pixel. The shader can be executed once per pixel, optionally in parallel with other shader executions corresponding to other pixels. The shader can write the output of the shader (e.g., a color value) to the coordinate x. In some examples, the shader output (color value) is written to the coordinate x at which the shader is executed. In other examples, the coordinate x can be a variable that is optionally shifted (positively or negatively) by an initial shift determined by a convergence value.

[0060] In the simplified example of FIG. 4, the variable z is a variable for each iteration placed on the grid relative to the depth value. In the simplified example of FIG. 4, corresponding to 6 pixels, there are a total of 6 steps 408. The value 6 can be calculated from the view shift multiplied by the gain. The initial value of the variable z is 1.0. The initial value of the coordinate x is 7 (for example, cell number 7, cells are numbered sequentially starting from zero).

[0061] At iteration number 1, at coordinate x = 7, the shader reads the depth (of the depth map) as 0.0. The shader compares the value of z (initial value 1.0) with the depth (value 0.0) at coordinate x = 7. Since the z value is greater than or equal to the depth, the shader decreases the z value by an amount equal to the reciprocal of the total number of steps. For a total of 6 steps, the decreased value of z becomes 0.833. Since the coordinate x is incremented by 1 pixel, x becomes 8.

[0062] At iteration number 2, at coordinate x = 8, the shader reads the depth (of the depth map) as 0.0. The shader compares the value of z (0.833) with the depth (value 0.0) at coordinate x = 8. Since the z value is greater than or equal to the depth, the shader decreases the z value by an amount equal to the reciprocal of the total number of steps. For a total of 6 steps, the decreased value of z becomes 0.667. Since the coordinate x is incremented by 1 pixel, x becomes 9.

[0063] At iteration number 3, at coordinate x = 9, the shader reads the depth (of the depth map) as 0.3. The shader compares the value of z (0.667) with the depth (value 0.3) at coordinate x = 9. Since the z value is greater than or equal to the depth, the shader decreases the z value by an amount equal to the reciprocal of the total number of steps. For a total of 6 steps, the decreased value of z becomes 0.5. Since the coordinate x is incremented by 1 pixel, x becomes 10.

[0064] At iteration number 4, at coordinate x = 10, the shader reads the depth (of the depth map) as 0.4. The shader compares the z value (0.5) with the depth (value 0.4) at coordinate x = 10. Since the z value is greater than or equal to the depth, the shader decreases the z value by an amount equal to the reciprocal of the total number of steps. For a total of 6 steps, the decreased value of z becomes 0.333. Since the coordinate x is incremented by only 1 pixel, x becomes 11.

[0065] At iteration number 5, at coordinate x = 11, the shader reads the depth (of the depth map) as 0.5. The shader compares the z value (0.333) with the depth (value 0.5) at coordinate x = 11. Since the z value is less than or equal to the depth, the shader reads the color from the x coordinate of x = 11. The shader assigns the color value from x = 11 to cell number 7 of the output buffer. In other words, the preceding operations determined the color value (such as from x = 11) for the pixel (located at cell number 7) of the view image. There is no iteration number 6 in the simplified example of Figure 4.

[0066] The shader may optionally perform additional comparisons by interpolating between adjacent read depth values. These additional comparisons may occur at locations between adjacent pixels within the x coordinate. The shader may use the following quantities as input: color texture, depth texture, x and y coordinates of the output pixel, gain (a single scalar value), and convergence (another single scalar value).

[0067] The computer-implemented operations discussed in this specification can generate multiple views within a single output buffer. For example, in a configuration where the computer-implemented operation generates four view images, the output buffer can be divided into four regions corresponding to a 2×2 tile, for example, by horizontal and vertical lines. In a particular example, view 0 is assigned to the upper left tile, view 1 is assigned to the upper right tile, view 2 is assigned to the lower left tile, and view 3 is assigned to the lower right tile. The computer-implemented operations discussed in this specification can present views 0 through 3 for different view angles arranged in a horizontal row. In some examples, the computer-implemented operations discussed in this specification can set the maximum offset distance of each feature of the original view image relative to the width of the view to a specified value that can be comfortable for the user. For example, the specified value can comprehensively be from 10 pixels to 20 pixels. In some examples, peripheral views (such as views 0 and 3) can receive the maximum offset, and the x and y coordinates can wrap to cover the quadrants of the selected view.

[0068] When the view identifier, or view id (0, 1, 2, or 3) is known, the computer-implemented operations discussed in this specification can select a view shift value from an array of default shift values, such as an array consisting of [-1.5, -0.5, +0.5, +1.5]. A user or a routine of another computer implementation can provide a convergence value. The convergence value can help to anchor the depth plane in place by reducing the pixel shift to zero. For example, when view images are to be shown sequentially, such as in an animation, a convergence value of 0 can make it appear that the background is fixed in place while the foreground can move from left to right. Similarly, a convergence value of 1 can make it appear that the foreground is fixed in place while the background can move from left to right. A convergence value of 0.5 can fix the middle plane in place such that the background and foreground move in opposite directions. Other values that can fix the plane at a depth that is not between the background and the foreground can also be used.

[0069] A user or another computer-implemented routine may provide a gain value. The gain value may increase or decrease the relative motion between the background and the foreground, as described above. Numerically, an example of gain implementation may be, in units of pixels, view_shift_px = shift_value_array[view_id] * gain_px.

[0070] The gain value can be positive or negative. The absolute value of the gain can determine how many depths the shader can use to perform its iterations. For example, the shader can use the following total number of steps: total_steps = abs(view_shift_px). To apply convergence, the shader can modify the x coordinate according to x = x - convergence * view_shift_px. The shader can initialize the variable z to a value of 1.0. The shader can initialize the variable Nx to the value of the x coordinate that can already be wrapped within the view quadrant. At each step, the shader can read the depth value from the depth texture using Nx,y. At each step, the shader can increase or decrease the variable Nx by one value (depending on the sign of view_shift_px). At each step, the shader can decrease the z variable by 1.0 divided by total_steps. When the z variable becomes less than the read depth value, the shader interrupts the iteration and returns the color value read from the color texture at Nx,y.

[0071] In this way, the shader can generate a forward-mapped view within the output buffer. The forward-mapped view may lack the problem of the background covering the foreground, may lack holes or separated pixels, or may lack empty sides. Since the shader can allow the texture to mirror non-empty parameter sides in the resulting view, the shader can sample outside the boundaries of the texture and can solve problems on the sides of the view.

[0072] Since the depth texture can contain non-zero (e.g., non-flat) values, the shader can interrupt iterations earlier than the value of total_steps. Further, since the value of total_steps can be specified by a comfortable limit (inclusively 10 pixels to 20 pixels, etc.), the shader can perform relatively few iterations per pixel. Since the shader can perform relatively few iterations per pixel, the shader can use a relatively small number of calculations to obtain color values for the pixels in the view image, thereby reducing the computational load required for real-time performance. Generally, increasing the number of views can increase the size of the output buffer and also increase the computational load required for real-time performance.

[0073] The computer-implemented operations discussed herein can reduce or remove artifacts that occur at the edges of objects, such as in regions with relatively rapid changes in depth values. For example, the shader can perform additional intermediate iterations, and the number of steps is increased by a specified oversampling factor such as 4. For example, the iterations can extract the depth every fourth step using previously read depth values for linear interpolation. Such intermediate iterations can produce a higher-quality view image without requiring additional reads from the depth texture.

[0074] The computer-implemented operations discussed herein can optionally select a view id for each output pixel, which need not directly correspond to the quadrants discussed above. Selecting the view id in this manner can help remove downstream processing of the buffer, such as rearranging the pixels into a pattern compatible with a multi-view display device. For example, a lenticular lens array covering the display can receive a set of thin slices of different views under each lens. By changing the comparison read_depth(x,y)<=z to a more advanced difference check, such as a check using a threshold, the shader can leave holes for further impainting if impainting is available and preferred over the stretched edges.

[0075] Generally, the computer-implemented operations discussed herein can achieve improved performance using relatively low gain values and relatively high contrast values. Generally, the computer-implemented operations discussed herein can achieve improved performance for depth maps that are normalized or at least not darkened. Generally, the computer-implemented operations discussed herein can achieve improved performance when shifting is provided in only one direction, such as horizontally or vertically.

[0076] The computer-implemented operations discussed in this specification can operate in real time. The computer-implemented operations discussed in this specification can operate as part of an OpenGL-based rendering pipeline (shader). The computer-implemented operations discussed in this specification can generate many relatively high-quality forward-mapped views that are horizontally shifted relative to each other. The computer-implemented operations discussed in this specification can use the input of a color image and a corresponding depth map. The computer-implemented operations discussed in this specification can utilize gain parameters and convergence parameters. The computer-implemented operations discussed in this specification can relate to the form of ray casting. The computer-implemented operations discussed in this specification can relate to synthesizing different views in a forward mapping format using a fragment (pixel) shader (e.g., with the intention of using the views for stereoscopic applications). The computer-implemented operations discussed in this specification can successfully balance image quality and computational speed. The computer-implemented operations discussed in this specification can create any number of orthographically horizontally-skewed views from a color image and a depth map in real time, without preprocessing, and without allocating a 3D mesh, using a single shader pass. The computer-implemented operations discussed in this specification can be executed on a mobile device. The computer-implemented operations discussed in this specification can operate with a relatively light computational load, using relatively few iterations and relatively few texture reads. The computer-implemented operations discussed in this specification can generate a view image without holes or isolated pixels. The computer-implemented operations discussed in this specification can have a simpler internal workflow than typical rendering by ray casting. The computer-implemented operations discussed in this specification can have better performance than simple forward mapping based on individual point movement.

[0077] FIG. 5 shows a flowchart of an example of a method 500 for performing multi-view image generation and display, according to an embodiment consistent with the principles described herein. Method 500 may be executed on system 100, or any other suitable system capable of performing multi-view image generation and display.

[0078] In operation 502, the system may receive a color image of the scene and a depth map of the scene. In operation 504, the system may synthesize view images of a multi-view image of the scene from the color image and the depth map. The view images may represent the scene from different view directions. Each view image may include a pixel position and a respective color value at the pixel position. In operation 506, the system may render the view images of the multi-view image on a multi-view display.

[0079] Synthesizing view images of a multi-view image may include the following operations for pixel positions within the view images. The system may cast a ray in a direction corresponding to the view direction of the view image from the pixel position towards the scene. The system may determine a ray intersection position where the ray intersects a virtual surface specified by the depth map. The system may set the color value of the view image at the pixel position to correspond to the color of the color image at the ray intersection position.

[0080] In some examples, determining sequential tentative positions along a ray between a pixel position and a specified plane such that the virtual surface is between the pixel position and the specified plane may include the following operations. The system may identify one of the tentative positions along the ray. The system may repeatedly determine that the identified tentative position is between the pixel position and the specified plane and advance the identified tentative position to the next tentative position along the ray. The system may virtual plane determine that is between the pixel position and the identified tentative position. The system may set the ray intersection position to generally correspond to a position between the identified tentative position and an adjacent previously identified tentative position.

[0081] In some examples, determining the ray intersection position may further include the following operations. The system may determine successive second tentative positions along the ray between the identified tentative position and a previously identified tentative position adjacent thereto. The system may identify one of the second tentative positions among the second tentative positions along the ray. The system may repeatedly determine that the identified second tentative position is between the pixel position and the specified plane, and may advance the identified second tentative position to the next second tentative position along the ray. The system may virtual plane determine that it is between the pixel position and the identified second tentative position. The system may set the ray intersection position to generally correspond to a position between the identified second tentative position and a previously identified second tentative position adjacent thereto.

[0082] In some examples, the tentative positions may be equally spaced along the ray. In some examples, the view image may define a horizontal direction parallel to the upper and lower edges of the view image, a vertical direction within the plane of the view image and orthogonal to the horizontal direction, and a depth orthogonal to both the horizontal and vertical directions. In some examples, the tentative positions may be spaced such that the horizontal component of the spacing between adjacent tentative positions corresponds to the horizontal spacing between adjacent pixels within the view image.

[0083] FIG. 6 is a schematic block diagram depicting an example of a computing device capable of performing multi-view image generation and display according to an embodiment consistent with the principles described herein. Computing device 1000 may include a system of components that perform various computing operations for a user of computing device 1000. Computing device 1000 may be a laptop, tablet, smartphone, touch screen system, intelligent display system, other client device, server, or other computing device. Computing device 1000 may include various components, such as, for example, processor 1003, memory 1006, input / output (I / O) component 1009, display 1012, and possibly other components. These components may be coupled to bus 1015, which serves as a local interface to enable the components of computing device 1000 to communicate with each other. Although the components of computing device 1000 are shown as being included within computing device 1000, it should be understood that at least some of the components may be coupled to computing device 1000 through external connections. For example, the components may be externally connected into or otherwise connected to computing device 1000 via an external port, socket, plug, connector, or wireless link.

[0084] Processor 1003 may include a processor circuit such as a central processing unit (CPU), a graphics processing unit (GPU), or any other integrated circuit that performs computational processing operations or any combination thereof. Processor 1003 may include one or more processing cores. Processor 1003 comprises a circuit for executing instructions. The instructions include, for example, computer code, programs, logic, or other machine-readable instructions that are received and executed by Processor 1003 to perform the computing functions embodied in the instructions. Processor 1003 may execute instructions to operate on data or to generate data. For example, Processor 1003 may receive input data (e.g., an image), process this input data according to a set of instructions, and generate output data (e.g., a processed image). As another example, Processor 1003 may receive instructions and generate new instructions for subsequent execution. Processor 1003 may be provided with hardware for implementing a graphics pipeline (e.g., the graphics pipeline schematically shown in FIG. 3) for rendering video, images, or frames generated by an application. For example, Processor 1003 may include one or more GPU cores, vector processors, scalar processes, decoders, or hardware accelerators.

[0085] Memory 1006 may include one or more memory components. Memory 1006 is defined herein as including either or both of volatile and non-volatile memory. Volatile memory components do not retain information upon loss of power. Volatile memory may include, for example, random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), magnetic random access memory (MRAM), or other volatile memory structures. System memory (e.g., main memory, cache, etc.) may be implemented using volatile memory. System memory refers to high-speed memory that can temporarily store data or instructions for fast read and write access to assist the processor 1003. Images (e.g., still images, video frames) may be stored or loaded into memory 1006 for subsequent access.

[0086] Non-volatile memory components retain information upon loss of power. Non-volatile memory includes read-only memory (ROM), hard disk drives, solid state drives, USB flash drives, memory cards accessed by a memory card reader, floppy disks accessed by an associated floppy disk drive, optical disks accessed by an optical disk drive, magnetic tapes accessed by a suitable tape drive. ROM may include, for example, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other similar memory devices. Storage memory may be implemented using non-volatile memory to provide long-term retention of data and instructions. According to various embodiments, the multi-view video cache may be implemented using volatile, non-volatile, or a combination of volatile and non-volatile memory.

[0087] Memory 1006 can refer to a combination of volatile and non-volatile memories used to store instructions and data. For example, data and instructions can be stored in non-volatile memory and loaded into volatile memory for processing by processor 1003. The execution of instructions can include, for example, a compiled program that is translated into machine code in a form that can be loaded from non-volatile memory into volatile memory and then executed by processor 1003, source code that is converted into a suitable form such as object code that can be loaded into volatile memory for execution by processor 1003, or source code that is interpreted by another executable program to generate instructions in volatile memory and executed by processor 1003. Instructions can be stored or loaded in any part or component of memory 1006, including, for example, RAM, ROM, system memory, storage, or any combination thereof.

[0088] Memory 1006 is shown as being separate from the other components of computing device 1000, but it should be understood that memory 1006 can be at least partially embedded in or otherwise integrated with one or more components. For example, processor 1003 can include on-board memory registers or caches for performing processing operations.

[0089] I / O component 1009 includes, for example, a touch screen, speaker, microphone, button, switch, dial, camera, sensor, accelerometer, or other component that receives user input or generates output directed to the user. I / O component 1009 can receive user input and convert it into data for storage in memory 1006 or for processing by processor 1003. I / O component 1009 can receive data output by memory 1006 and processor 1003 and convert them into a form perceivable by the user (e.g., sound, tactile response, visual information, etc.).

[0090] One type of I / O component 1009 is the display 1012. The display 1012 can include a multi-view display (e.g., multi-view display 112), a multi-view display combined with a 2D display, or any other display that presents graphic content. A capacitive touch screen layer that serves as an I / O component 1009 can be layered within the display to enable the user to provide input simultaneously while perceiving the visual output. The processor 1003 can generate data that is formatted as an image or frame for presentation on the display 1012. The processor 1003 can execute instructions for rendering an image or frame on the display 1012 for the user. The camera I / O component 1009 can be used for a video capture process that captures video that can be converted into multi-view video.

[0091] The bus 1015 facilitates the communication of instructions and data among the processor 1003, the memory 1006, the I / O components 1009, the display 1012, and any other components of the computing device 1000. The bus 1015 can include an address translation device, an address decoder, a fabric, conductive traces, conductive wires, ports, plugs, sockets, and other connectors for enabling the communication of data and instructions.

[0092] Instructions within memory 1006 can be embodied in various forms in a manner that implements at least a portion of the software stack. For example, the instructions can be embodied as the operating system 1031, an application 1034, a device driver (e.g., display driver 1037), firmware (e.g., display firmware 1040), or other software components. The operating system 1031 is a software platform that supports the basic functions of the computing device 1000, such as scheduling tasks, controlling the I / O components 1009, providing access to hardware resources, managing power, and supporting applications 1034.

[0093] Application 1034 runs on operating system 1031 and can obtain access to the hardware resources of computing device 1000 via operating system 1031. In this regard, the execution of application 1034 is at least partially controlled by operating system 1031. Application 1034 can be a user-level software program that provides high-level functions, services, and other functionality to the user. In some embodiments, application 1034 can be a downloadable or separately accessible dedicated "app" on computing device 1000 by the user. The user can launch application 1034 via the user interface provided by operating system 1031. Application 1034 can be developed by a developer and can be defined in various source code formats. Application 1034 can be developed using several programming or scripting languages such as, for example, C, C++, C#, Objective C, Java®, Swift, JavaScript®, Perl, PHP, Visual Basic®, Python®, Ruby, Go, or other programming languages. Application 1034 can be compiled into object code by a compiler or interpreted by an interpreter for execution by processor 1003. The various embodiments discussed herein can be implemented as at least a part of application 1034.

[0094] For example, a device driver such as display driver 1037 includes instructions that enable operating system 1031 to communicate with various I / O components 1009. Each I / O component 1009 may have its own device driver. The device drivers can be installed such that they are stored in storage and loaded into system memory. For example, during installation, display driver 1037 translates high-level display instructions received from operating system 1031 into low-level instructions to be executed by display 1012 to display an image.

[0095] Firmware such as display firmware 1040 may include machine code or assembly code that enables I / O component 1009 or display 1012 to perform low-level operations. The firmware can convert electrical signals of a particular component into higher-level instructions or data. For example, display firmware 1040 may control how display 1012 activates individual pixels at a low level by adjusting voltage or current signals. The firmware can be stored in non-volatile memory and executed directly from the non-volatile memory. For example, display firmware 1040 may be embodied in a ROM chip coupled to display 1012, and as a result, the ROM chip is separate from other storage and the system memory of computing device 1000. Display 1012 may include processing circuitry for executing display firmware 1040.

[0096] The operating system 1031, the application 1034, the driver (e.g., the display driver 1037), the firmware (e.g., the display firmware 1040), and perhaps other instruction sets may each include instructions executable by the processor 1003 or other processing circuitry of the computing device 1000 to perform the functionality and operations discussed above. The instructions described herein may be embodied as software or code executed by the processor 1003 as discussed above, but alternatively, the instructions may also be embodied in dedicated hardware or a combination of software and dedicated hardware. For example, the functionality and operations performed by the instructions discussed above may be implemented as a circuit or state machine employing any one or combination of several techniques. These techniques may include, but are not limited to, discrete logic circuits having logic gates for performing various logical functions when one or more data signals are applied, application specific integrated circuits (ASICs) having appropriate logic gates, field programmable gate arrays (FPGAs), or other components, etc.

[0097] In some embodiments, the instructions for performing the functionality and operations discussed above may be embodied in a non-transitory computer-readable storage medium. The computer-readable storage medium may or may not be part of a computing system such as the computing device 1000. The instructions may include, for example, statements, code, or declarations that can be fetched from a computer-readable medium and executed by a processing circuit (e.g., the processor 1003). In the context of this disclosure, a "computer-readable medium" may be any medium that can contain, store, or maintain the instructions described herein for use by or in connection with an instruction execution system such as the computing device 1000.

[0098] A computer-readable medium may comprise any one of many physical media such as, for example, magnetic, optical, or semiconductor media. More specific examples of suitable computer-readable media may include, but are not limited to, magnetic tape, magnetic floppy disk, magnetic hard drive, memory card, solid state drive, USB flash drive, or optical disk. Also, the computer-readable medium may be a random access memory (RAM) including, for example, static random access memory (SRAM) and dynamic random access memory (DRAM), or magnetic random access memory (MRAM). In addition, the computer-readable medium may be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other type of memory device.

[0099] Computing device 1000 may perform any of the operations or implement the functionality described above. For example, the flowcharts and process flows discussed above may be implemented by computing device 1000 that executes instructions and processes data. Computing device 1000 is shown as a single device, but embodiments are not so limited. In some embodiments, computing device 1000 may offload the processing of instructions in a distributed fashion such that multiple computing devices 1000 may operate together to execute instructions that may be stored or loaded in a distributed arrangement of computing components. For example, at least some instructions or data may be stored, loaded, or executed in a cloud-based system that operates in conjunction with computing device 1000.

[0100] In particular, a non-transitory computer-readable storage medium may store executable instructions that, when executed by a processor of a computer system, perform operations for generating and displaying a multi-view image. According to various embodiments, the operations may include receiving a color image of a scene and a depth map of the scene. The operations may further include synthesizing view images of the multi-view image of the scene from the color image and the depth map. In various embodiments, the view images may represent the scene from different view directions, and each view image may include a pixel position and a respective color value at the pixel position. Further, the operations may include rendering the view images of the multi-view image on a multi-view display.

[0101] According to various embodiments, synthesizing view images of the multi-view image includes casting a ray in a direction corresponding to the view direction of the view image from the pixel position toward the scene for pixel positions within the view image. Synthesizing the view images further includes determining a ray intersection position where the ray intersects a virtual surface specified by the depth map, and setting the color value of the view image at the pixel position to correspond to the color of the color image at the ray intersection position.

[0102] In some embodiments, determining the ray intersection position may include determining sequential provisional positions along the ray between the pixel position and a specified plane such that the virtual surface is between the pixel position and the specified plane. Determining the ray intersection position may further include identifying one of the provisional positions along the ray. In particular, identifying the provisional position may include repeatedly determining that the identified provisional position is between the pixel position and the specified plane, and advancing the identified provisional position to the next provisional position along the ray.

[0103] According to these embodiments, determining the ray intersection position is virtual planemay further include determining that it is between the pixel position and the identified provisional position. Further, determining the ray intersection position may further include setting the ray intersection position to correspond to a position between the identified provisional position and a previously identified provisional position adjacent to the identified provisional position.

[0104] According to some embodiments, determining the ray intersection position further includes determining sequential second provisional positions along the ray between the identified provisional position and a previously identified provisional position adjacent to the identified provisional position, and identifying one of the second provisional positions among the second provisional positions along the ray. In particular, identifying the second provisional position may include repeatedly determining that the identified second provisional position is between the pixel position and the specified plane, and advancing the identified second provisional position to the next second provisional position along the ray.

[0105] In some of these embodiments, determining the ray intersection position virtual plane may further include determining that it is between the pixel position and the identified second provisional position. Further, determining the ray intersection position may include setting the ray intersection position to correspond to a position between the identified second provisional position and a previously identified second provisional position adjacent to the identified second provisional position.

[0106] In some embodiments, the provisional positions are equally spaced along the ray. In some embodiments, the view image may define a horizontal direction that is parallel to the upper and lower edges of the view image. Here, the vertical direction may be within the plane of the view image and orthogonal to the horizontal direction, and the depth may be orthogonal to the horizontal and vertical directions. In some embodiments, the provisional positions are spaced such that the horizontal component of the spacing between adjacent provisional positions corresponds to the horizontal spacing between adjacent pixels within the view image.

[0107] Thus, for example, examples and embodiments of generating and using a multi-view video cache for multi-view video rendering are described. The cache may include at least a pair of cache data entries corresponding to a target timestamp, and the first and second cache data entries of the pair may each include a first and a second image frame group, respectively. The first image frame group may correspond to a first multi-view frame preceding the target timestamp, and the second image frame group may correspond to a second multi-view frame after the target timestamp. The view of a particular multi-view frame corresponding to the target timestamp may be generated from the cache using information from the first and second image frame groups. It should be understood that the examples described above are merely illustrative of some of the many specific examples that represent the principles described herein. Clearly, those skilled in the art may readily devise numerous other arrangements without departing from the scope as defined by the following claims. It should be noted that the following aspects are disclosed in this specification. [Aspect 1] A computer-implemented method for generating and displaying multi-view images, comprising: receiving, using a computing device, a color image of a scene and a depth map of the scene; synthesizing, using the computing device, a plurality of view images of the multi-view image of the scene from the color image and the depth map, wherein the plurality of view images represent the scene from a plurality of different view directions, and each view image includes a plurality of pixel positions and respective color values at the plurality of pixel positions; rendering the plurality of view images of the multi-view image on a multi-view display of the computing device; synthesizing the view images of the multi-view image comprises, for pixel positions within the view image: casting a ray in a direction corresponding to the view direction of the view image from the pixel position towards the scene; determining a ray intersection position where the ray intersects a virtual surface specified by the depth map; and setting the color value of the view image at the pixel position to correspond to the color of the color image at the ray intersection position. [Aspect 2] Determining the ray intersection position comprises: determining sequential tentative positions along the ray between the pixel position and a specified plane such that the virtual surface is between the pixel position and the specified plane; identifying a tentative position among the tentative positions along the ray, determining that the identified tentative position is between the pixel position and the specified plane, and advancing the identified tentative position to the next tentative position along the ray; determining that the specified plane is between the pixel position and the identified tentative position; and setting the ray intersection position to generally correspond to a position between the identified tentative position and an adjacent previously identified tentative position. [Aspect 3] Determining the light ray intersection position includes determining successive second tentative positions along the light ray between the identified tentative position and the adjacent previously identified tentative position, identifying one of the second tentative positions among the second tentative positions along the light ray, determining that the identified second tentative position is between the pixel position and the specified plane, and identifying, including advancing the identified second tentative position to the next second tentative position along the light ray, determining that the specified plane is between the pixel position and the identified second tentative position, and further including setting the light ray intersection position to correspond generally to a position between the identified second tentative position and an adjacent previously identified second tentative position, the computer-implemented method according to aspect 2. [Aspect 4] The computer-implemented method according to aspect 2, wherein the tentative positions are equally spaced along the light ray. [Aspect 5] The view image defines a horizontal direction parallel to the upper and lower edges of the plurality of view images, a vertical direction within the plane of the view image and orthogonal to the horizontal direction, and a depth orthogonal to the horizontal and vertical directions, The computer-implemented method according to aspect 4, wherein the tentative positions are spaced such that a horizontal component of the spacing between adjacent tentative positions corresponds to a specified value. [Aspect 6] The computer-implemented method according to aspect 5, wherein the specified value corresponds to a horizontal spacing between adjacent pixels in the view image. [Aspect 7] Synthesizing the view images of the multi-view image for pixel positions in the view image includes casting a light ray in a direction corresponding to the view direction of the view image from the pixel position towards the scene, determining that the light ray does not intersect a virtual surface specified by the depth map, acquiring color information from at least one temporally adjacent video frame of the plurality of view images of the multi-view image, and setting the color value of the view image at the pixel position to correspond to the acquired color information, the computer-implemented method according to aspect 1. [Aspect 8] Determining that the light ray does not intersect the virtual surface specified by the depth map includes determining that the light ray has propagated away from the pixel position by a distance exceeding a threshold distance, the computer-implemented method according to aspect 7. [Aspect 9] The plurality of view images correspond to time-series images of a video signal, and the color information is obtained from the pixel positions of the at least one temporally adjacent video frame of the video signal, the computer-implemented method according to aspect 7. [Aspect 10] The different view directions are in a horizontal plane including the upper and lower edges of the plurality of view images, the computer-implemented method according to aspect 1. [Aspect 11] A system configured to perform multi-view image generation and display, A multi-view display, A central processing unit, A memory storing a plurality of instructions that, when executed, cause the central processing unit to perform operations, the operations including: Receiving a color image of a scene and a depth map of the scene, Synthesizing a plurality of view images of a multi-view image of the scene from the color image and the depth map, the plurality of view images representing the scene from a plurality of different view directions, and each view image including a pixel position and respective color values at the pixel position, Rendering the plurality of view images of the multi-view image on the multi-view display, Synthesizing the view images of the multi-view image includes, for pixel positions in the view images, Casting a light ray in a direction corresponding to the view direction of the view image from the pixel position towards the scene, Determining a light ray intersection position where the light ray intersects a virtual surface specified by the depth map, and Setting the color value of the view image at the pixel position to correspond to the color of the color image at the light ray intersection position, a system. [Aspect 12] Determining the light ray intersection position includes Determining sequential provisional positions along the light ray between the pixel position and a specified plane such that the virtual surface is between the pixel position and the specified plane, Identifying one of the provisional positions along the light ray, Determining that the identified tentative position is between the pixel position and the specified plane, and Identifying, including advancing the identified tentative position to the next tentative position along the ray, Determining that the specified plane is between the pixel position and the identified tentative position, and The system according to aspect 11, further comprising setting the ray intersection position to correspond to a position between the identified tentative position and an adjacent previously identified tentative position. [Aspect 13] Determining the ray intersection position comprises Determining successive second tentative positions along the ray between the identified tentative position and the adjacent previously identified tentative position, Identifying one of the second tentative positions among the second tentative positions along the ray, Determining that the identified second tentative position is between the pixel position and the specified plane, and Identifying, including advancing the identified second tentative position to the next second tentative position along the ray, Determining that the specified plane is between the pixel position and the identified second tentative position, and The system according to aspect 12, further comprising setting the ray intersection position to correspond to a position between the identified second tentative position and an adjacent previously identified second tentative position. [Aspect 14] The system according to aspect 12, wherein the tentative positions are equally spaced along the ray. [Aspect 15] The view image defines a horizontal direction parallel to the upper and lower edges of the view image, a vertical direction within the plane of the view image and orthogonal to the horizontal direction, and a depth orthogonal to the horizontal and vertical directions. The system according to aspect 14, wherein the tentative positions are spaced such that a horizontal component of the spacing between adjacent tentative positions corresponds to a horizontal spacing between adjacent pixels within the view image. [Aspect 16] The system according to aspect 11, wherein the different view directions are in a horizontal plane including the upper and lower edges of the plurality of view images. [Aspect 17] A non-transitory computer-readable storage medium storing executable instructions that, when executed by a processor of a computer system, perform operations for multi-view image generation and display, the operations comprising Receiving a color image of the scene and a depth map of the scene; Synthesizing a plurality of view images of a multi-view image of the scene from the color image and the depth map, wherein the plurality of view images represent the scene from different view directions, and each view image includes a plurality of pixel positions and respective color values at the plurality of pixel positions; Rendering the plurality of view images of the multi-view image on a multi-view display; and Synthesizing the view images of the multi-view image includes, for pixel positions within the view images, Casting a ray in a direction corresponding to the view direction of the view image from the pixel position towards the scene; Determining a ray intersection position where the ray intersects a virtual surface specified by the depth map; and Setting a color value of the view image at the pixel position to correspond to a color of the color image at the ray intersection position. A non-transitory computer-readable storage medium. [Aspect 18] Determining the ray intersection position includes Determining sequential provisional positions along the ray between the pixel position and a specified plane such that the virtual surface is between the pixel position and the specified plane; Identifying a provisional position among the provisional positions along the ray, repeatedly Determining that the identified provisional position is between the pixel position and the specified plane; and Advancing the identified provisional position to the next provisional position along the ray, including identifying; Determining that the specified plane is between the pixel position and the identified provisional position; and Setting the ray intersection position to correspond to a position between the identified provisional position and an adjacent previously identified provisional position. The non-transitory computer-readable storage medium according to Aspect 17. [Aspect 19] Determining the ray intersection position includes Determining sequential second provisional positions along the ray between the identified provisional position and the adjacent previously identified provisional position; Identifying one second provisional position among the second provisional positions along the ray, repeatedly Determining that the identified second tentative position is between the pixel position and the specified plane, and Identifying, including advancing the identified second tentative position along the ray to the next second tentative position along the ray, Determining that the specified plane is between the pixel position and the identified second tentative position, The non-transitory computer-readable storage medium according to aspect 18, further comprising setting the ray intersection position to correspond generally to a position between the identified second tentative position and an adjacent previously identified second tentative position. [Aspect 20] The tentative positions are equally spaced along the ray, The view image defines a horizontal direction parallel to the upper and lower edges of the view image, within the plane of the view image, a vertical direction orthogonal to the horizontal direction, and a depth orthogonal to the horizontal and vertical directions, The non-transitory computer-readable storage medium according to aspect 18, wherein the tentative positions are spaced such that a horizontal component of the spacing between adjacent tentative positions corresponds to a horizontal spacing between adjacent pixels within the view image.

Explanation of Symbols

[0108] 10 Multi-view Display 12 Screens 14 Views 16 View Directions 20 Light Beams 100 System 102 Computing Device 104 Color Image 106 Scene 108 Depth Map 110, 110A, 110B, 110C, 110D View Images 114A, 114B, 114C, 114D View Directions 400 Graphic Representation 402 Bottom Horizontal Line 404 Output Buffer 406 Top Horizontal Line 1000 Computing Device 1003 Processor 1006 Memory 1009 Input / Output (I / O) Component 1012 Display 1015 Bus 1031 Operating System 1034 Application 1037 Display Driver 1040 display firmware

Claims

1. A computer-implemented method for generating and displaying a multi-view image, comprising: Receiving, using a computing device, a color image of a scene and a depth map of the scene; Synthesizing, using the computing device, a plurality of view images of the multi-view image of the scene from the color image and the depth map, wherein the plurality of view images represent the scene from a plurality of different view directions, and each view image includes a plurality of pixel positions and respective color values at the plurality of pixel positions; Rendering, on a multi-view display of the computing device, the plurality of view images of the multi-view image; Synthesizing a view image among the plurality of view images of the multi-view image for a pixel position among the plurality of pixel positions within the view image, comprising: Casting a ray in a direction corresponding to the view direction of the view image from the pixel position towards the scene; Determining a ray intersection position where the ray intersects a virtual surface specified by the depth map; and Setting a color value among the respective color values of the view image at the pixel position to correspond to the color of the color image at the ray intersection position; Determining the ray intersection position comprises: Determining sequential tentative positions along the ray between the pixel position and a specified plane such that the virtual surface is between the pixel position and the specified plane; Identifying a tentative position among the sequential tentative positions along the ray, comprising: Determining that the identified tentative position is between the pixel position and the specified plane; and Advancing the identified tentative position to the next tentative position along the ray; Determining that the virtual surface is between the pixel position and the identified tentative position; and Comprehensively setting the ray intersection position to correspond to a position between the identified tentative position and an adjacent previously identified tentative position; A computer-implemented method.

2. Determining the ray intersection position comprises: Determining sequential second tentative positions along the ray between the identified tentative position and the adjacent previously identified tentative position; identifying a second provisional position among the successive second provisional positions along the ray, determining that the identified second provisional position is between the pixel position and the specified plane, and identifying, including advancing the identified second provisional position along the ray to the next second provisional position, determining that the virtual surface is between the pixel position and the identified second provisional position, further comprising setting the ray intersection position to generally correspond to a position between the identified second provisional position and an adjacent previously identified second provisional position, the computer-implemented method of claim 1. **Claim 3** The computer-implemented method of claim 1, wherein the provisional positions are equally spaced along the ray. **Claim 4** The view images define a horizontal direction parallel to the top and bottom edges of the plurality of view images, a vertical direction within the plane of the view images and orthogonal to the horizontal direction, and a depth orthogonal to the horizontal and vertical directions, The computer-implemented method of claim 3, wherein the provisional positions are spaced such that a horizontal component of the spacing between adjacent provisional positions corresponds to a specified value. **Claim 5** The computer-implemented method of claim 4, wherein the specified value corresponds to a horizontal spacing between adjacent pixels within the view image. **Claim 6** A computer-implemented method for multi-view image generation and display, comprising: receiving, using a computing device, a color image of a scene and a depth map of the scene; synthesizing, using the computing device, a plurality of view images of a multi-view image of the scene from the color image and the depth map, the plurality of view images representing the scene from a plurality of different view directions, each view image including a plurality of pixel positions and respective color values at the plurality of pixel positions; rendering the plurality of view images of the multi-view image on a multi-view display of the computing device; and synthesizing a view image among the plurality of view images of the multi-view image for pixel positions among the plurality of pixel positions within the view image, Casting a ray from the pixel position towards the scene in a direction corresponding to the view direction of the view image; Determining a ray intersection position where the ray intersects a virtual surface specified by the depth map; and Setting a color value among the respective color values of the view image at the pixel position so as to correspond to the color of the color image at the ray intersection position, Synthesizing a view image among the plurality of view images of the multi-view image for the pixel positions within the view image, Casting the ray from the pixel position towards the scene in the direction corresponding to the view direction of the view image; Determining that the ray does not intersect the virtual surface specified by the depth map; Obtaining color information from at least one temporally adjacent video frame of the plurality of view images of the multi-view image; Setting the color value among the respective color values of the view image at the pixel position so as to correspond to the obtained color information, a computer-implemented method.

7. Determining that the ray does not intersect the virtual surface specified by the depth map includes determining that the ray has propagated away from the pixel position by a distance exceeding a threshold distance, the computer-implemented method according to claim 6.

8. The plurality of view images correspond to time-series images of a video signal, and the color information is obtained from the pixel positions of the at least one temporally adjacent video frame of the video signal, the computer-implemented method according to claim 6.

9. The different view directions are in a horizontal plane parallel to the upper and lower edges of the plurality of view images, the computer-implemented method according to claim 1.

10. A system configured to perform multi-view image generation and display, A multi-view display; A central processing unit; A memory storing a plurality of instructions that, when executed, cause the central processing unit to perform operations, the operations including Receiving a color image of a scene and a depth map of the scene; Synthesizing a plurality of view images of the multi-view image of the scene from the color image and the depth map, wherein the plurality of view images represent the scene from a plurality of different view directions, and each view image includes a plurality of pixel positions and respective color values at the plurality of pixel positions; Rendering the plurality of view images of the multi-view image on the multi-view display; Synthesizing the view images among the plurality of view images of the multi-view image is for the pixel positions among the plurality of pixel positions in the view image; Casting a ray in a direction corresponding to the view direction of the view image from the pixel position towards the scene; Determining a ray intersection position where the ray intersects a virtual surface specified by the depth map; and Setting the color value among the respective color values of the view image at the pixel position so as to correspond to the color of the color image at the ray intersection position; Determining the ray intersection position; Determining sequential provisional positions along the ray between the pixel position and a specified plane such that the virtual surface is between the pixel position and the specified plane; Identifying one of the sequential provisional positions along the ray, Determining that the identified provisional position is between the pixel position and the specified plane, and Identifying, including advancing the identified provisional position to the next provisional position along the ray; Determining that the virtual surface is between the pixel position and the identified provisional position; A system, including setting the ray intersection position to generally correspond to a position between the identified provisional position and an adjacent previously identified provisional position.

11. Determining the ray intersection position; Determining sequential second provisional positions along the ray between the identified provisional position and the adjacent previously identified provisional position; Identifying one of the sequential second provisional positions along the ray, Determining that the identified second provisional position is between the pixel position and the specified plane; and Identifying, including advancing the identified second tentative position along the light ray to the next second tentative position along the light ray; Determining that the virtual surface is between the pixel position and the identified second tentative position; The system of claim 10, further comprising setting the light ray intersection position to generally correspond to a position between the identified second tentative position and an adjacent previously identified second tentative position.

12. The system of claim 10, wherein the tentative positions are equally spaced along the light ray.

13. The view image defines a horizontal direction parallel to the upper and lower edges of the view image, a vertical direction within the plane of the view image and orthogonal to the horizontal direction, and a depth orthogonal to the horizontal and vertical directions. The system of claim 12, wherein the tentative positions are spaced such that a horizontal component of the spacing between adjacent tentative positions corresponds to a horizontal spacing between adjacent pixels within the view image.

14. The system of claim 10, wherein the different view directions are in a horizontal plane parallel to the upper and lower edges of the plurality of view images.

15. A non-transitory computer-readable storage medium storing executable instructions that, when executed by a processor of a computer system, perform operations for multi-view image generation and display, the operations comprising: Receiving a color image of a scene and a depth map of the scene; Synthesizing a plurality of view images of a multi-view image of the scene from the color image and the depth map, the plurality of view images representing the scene from different view directions, each view image including a plurality of pixel positions and respective color values at the plurality of pixel positions; Rendering the plurality of view images of the multi-view image on a multi-view display; and Synthesizing a view image of the plurality of view images of the multi-view image for pixel positions of the plurality of pixel positions within the view image, Casting a light ray from the pixel position towards the scene in a direction corresponding to the view direction of the view image; Determining a light ray intersection position where the light ray intersects a virtual surface specified by the depth map; and setting color values among the respective color values of the view image at the pixel position so as to correspond to the color of the color image at the light intersection position, determining the light intersection position, determining sequential provisional positions along the light ray between the pixel position and a specified plane such that the virtual surface is between the pixel position and the specified plane, identifying a provisional position among the sequential provisional positions along the light ray, repeatedly determining that the identified provisional position is between the pixel position and the specified plane, and identifying including advancing the identified provisional position to the next provisional position along the light ray, determining that the virtual surface is between the pixel position and the identified provisional position, setting the light intersection position to correspond to a position between the identified provisional position and an adjacent previously identified provisional position, A non-transitory computer-readable storage medium.

16. determining the light intersection position, determining sequential second provisional positions along the light ray between the identified provisional position and the adjacent previously identified provisional position, identifying one second provisional position among the sequential second provisional positions along the light ray, repeatedly determining that the identified second provisional position is between the pixel position and the specified plane, and identifying including advancing the identified second provisional position to the next second provisional position along the light ray, determining that the virtual surface is between the pixel position and the identified second provisional position, The non-transitory computer-readable storage medium according to claim 15, further comprising setting the light intersection position comprehensively to correspond to a position between the identified second provisional position and an adjacent previously identified second provisional position.

17. the provisional positions are equally spaced along the light ray, the view image defines a horizontal direction parallel to the upper and lower edges of the view image, within the plane of the view image, a vertical direction orthogonal to the horizontal direction, and a depth orthogonal to the horizontal and vertical directions, The non-transitory computer-readable storage medium according to claim 15, wherein the provisional positions are spaced apart such that a horizontal component of an interval between adjacent provisional positions corresponds to a horizontal interval between adjacent pixels in the view image.

Citation Information

Patent Citations

  • Element image group generation device and program therefor

    JP2017073710A

  • Image display apparatus, video display system including apparatus, image display method and program for displaying image

    JP2019133214A

Cited By

  • View synthesis system and method using depth map

    US12567201B2