System and method for detecting multiview file formats
The system automatically detects multiview image formats using cross-correlation and a trained classifier, addressing inaccuracies in manual and automated methods, ensuring precise processing and rendering.
Patent Information
- Application Number
- JP2023551667
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-02-25
- Filing Date
- 2021-05-08
- Publication Date
- 2026-02-25
- Estimated Expiration
- 2041-05-08
AI Technical Summary
Existing multiview display systems rely on manual specification of image formats, which can disrupt workflow and produce inaccurate results, especially for images with repeating patterns or textures, and automated methods are not reliable.
Automatically detect multiview image formats by extracting features such as cross-correlation values and aspect ratio, generating a cross-correlation map, and using a trained classifier to classify the format.
Provides highly accurate and autonomous multiview format classification, enabling proper processing and rendering of multiview images.
Smart Images

Figure 0007820398000001 
Figure 0007820398000002 
Figure 0007820398000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 153,917, filed February 25, 2021, which is incorporated herein by reference in its entirety.
[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT none [Background technology]
[0003] Multiview display is an emerging display technology that offers a more immersive viewing experience compared to traditional 2D content. Multiview content can include a still image or a series of video frames, with each multiview image (e.g., a picture or frame) containing a different view of a scene. End devices that receive and process multiview content are responsible for converting the received multiview content into a native or local format so that it can be displayed on the end device. Multiview content can be received from a variety of sources. [Brief explanation of the drawings]
[0004] Various features of examples and embodiments according to the principles described herein can be more readily understood by reference to the following detailed description in conjunction with the accompanying drawings, in which like reference numerals indicate like structural elements.
[0005] [Figure 1] 1 illustrates a multi-view image in an example, according to an embodiment consistent with principles described herein.
[0006] [Figure 2] 1 illustrates an example of a multi-view display, according to one embodiment consistent with principles described herein.
[0007] [Figure 3]1 illustrates an example of accessing a multi-view image that includes multiple tile-view images, according to an embodiment consistent with principles described herein.
[0008] [Figure 4] 1 illustrates an example of generating a cross-correlation map from multi-view images according to one embodiment consistent with principles described herein.
[0009] [Figure 5] 10 shows an example of various cross-correlation maps generated from different multi-view formats, according to one embodiment consistent with principles described herein.
[0010] [Figure 6] 1 illustrates an example of sampling a cross-correlation map, according to an embodiment consistent with principles described herein.
[0011] [Figure 7] FIG. 1 illustrates an example of detecting the multiview format of a multiview image, according to an embodiment consistent with principles described herein.
[0012] [Figure 8] 1 illustrates an example of rendering a multiview image on a multiview display based on a multiview format, according to an embodiment consistent with principles described herein.
[0013] [Figure 9] 1 is a flowchart providing an example of the functionality of a multi-view display system configured to detect multi-view formats, according to one embodiment consistent with principles described herein.
[0014] [Figure 10] FIG. 1 is a schematic block diagram illustrating an exemplary diagram of a multi-view display system according to one embodiment consistent with principles described herein. DETAILED DESCRIPTION OF THE INVENTION
[0015] Particular examples and embodiments have other features in addition to, and in place of, those shown in the above-referenced figures. These and other features are described in detail below with reference to the above-referenced figures.
[0016] Examples and embodiments according to principles described herein provide techniques for detecting the multiview format of a multiview image to enable a multiview display system to properly process and render the multiview image. Multiview image files can arrange the view images in a variety of formats. End devices responsible for processing, rendering, and displaying the multiview image may require information about how the view images are arranged. Relying on users to manually specify the multiview format can disrupt the workflow of rendering a multiview image. Furthermore, automating the process in general can produce inaccurate results, especially for images with repeating patterns or textures. Embodiments relate to automatically detecting the multiview format by extracting features of the multiview image and classifying the features. The results of such embodiments can provide highly accurate and autonomous multiview format classification. For example, in some embodiments, one feature can be a set of cross-correlation values. The cross-correlation values are determined by performing an autocorrelation operation on the multiview image to generate a cross-correlation map. The cross-correlation map is a matrix that can be visualized as an image representing the locations of repetitions within the multiview image. The cross-correlation map can be sampled at predetermined locations to generate features based on the cross-correlation values. Other features may be extracted from the multi-view images, including, for example, aspect ratio, and a classifier is trained to classify the multi-view images using this feature set in order to detect the multi-view format at runtime.
[0017] FIG. 1 illustrates an example multi-view image according to one embodiment consistent with principles described herein. The multi-view image 103 may be a still image or a video frame from a multi-view video stream. The multi-view image 103 has multiple view images 106 (e.g., views). Each of the view images 106 corresponds to a different principal angular direction 109 (e.g., left view, right view, etc.). The view images 106 are rendered on a multi-view display 112. Each view image 106 represents a different viewing angle of the scene in the multi-view image 103. Thus, the different view images 106 have a certain degree of parallax relative to each other. A viewer can perceive a different view image 106 with their left eye while perceiving a certain view image 106 with their right eye. This allows the viewer to perceive the different view images 106 simultaneously, thereby experiencing a three-dimensional (3D) effect.
[0018] In some embodiments, as a viewer physically changes their viewing angle relative to the multi-view display 112, their eyes may capture different views of the multi-view image 103. As a result, the viewer can interact with the multi-view display 112 to see different view images 106 of the multi-view image 103. For example, as the viewer moves to the left, the viewer can see more of the left side of the scene in the multi-view image 103. The multi-view image 103 may have multiple view images 106 along a horizontal plane, or multiple view images 106 along a vertical plane, or both. Thus, as the user changes their viewing angle to see different view images 106, the viewer can obtain additional visual details of the scene represented in the multi-view image 103.
[0019] As described above, each view image 106 is presented by the multiview display 112 at a different corresponding principal angular direction 109. When presenting the multiview images 103 for display, the view images 106 may actually appear on or near the multiview display 112. In this regard, the rendered multiview content may be referred to as light field content. A characteristic of observing light field content is the ability to observe different views simultaneously. Light field content includes visual images that may appear behind the screen as well as in front of it to convey a sense of depth to the viewer.
[0020] A 2D display may be substantially similar to a multi-view display 112, except that it is generally configured to provide a single view (e.g., only one of the views) as opposed to different views of a multi-view image 103. As used herein, a "two-dimensional display" or "2D display" (i.e., within a predetermined viewing angle or range of the 2D display) is defined as a display configured to provide substantially the same view of an image regardless of the direction from which the image is viewed. Conventional liquid crystal displays (LCDs) found on many smartphones and computer monitors are examples of 2D displays. In contrast, as used herein, a "multi-view display" is defined as an electronic display or display system configured to provide different views of a multi-view image at or from different viewing directions simultaneously from a user's viewpoint. In particular, the different view images 106 may represent different perspectives of the multi-view image 103.
[0021] The multi-view display 112 can be implemented using a variety of techniques that support the presentation of different image views so that they are perceived simultaneously. One example of a multi-view display is one that uses multi-beam elements that scatter light to control the predominant angular direction of the different view images 106. According to some embodiments, the multi-view display 112 may be a light field display that presents multiple light beams of different colors and different directions that correspond to the different views. In some examples, the light field display is a so-called "glasses-free" three-dimensional (3D) display that can use multi-beam elements (e.g., diffraction gratings) to provide an autostereoscopic representation of the multi-view images without the need to wear special eyewear to perceive depth.
[0022] FIG. 2 illustrates an example of a multi-view display according to one embodiment consistent with principles described herein. The multi-view display 112 can generate light field content when operating in a multi-view mode. In some embodiments, the multi-view display 112 renders a multi-view image and a 2D image depending on its operating mode. For example, the multi-view display 112 can include multiple backlights for operating in different modes. The multi-view display 112 can be configured to provide wide-angle emitted light during 2D mode using a wide-angle backlight 115. Additionally, the multi-view display 112 can be configured to provide directional emitted light during multi-view mode using a multi-view backlight 118 having an array of multi-beam elements, where the directional emitted light includes multiple directional light beams provided by each multi-beam element of the multi-beam element array. In some embodiments, multi-view display 112 may be configured to time-multiplex the 2D mode and the multi-view mode using mode controller 121, sequentially activating wide-angle backlight 115 during a first consecutive time interval corresponding to the 2D mode and sequentially activating multi-view backlight 118 during a second consecutive time interval corresponding to the multi-view mode. The directions of the directional light beams may correspond to different view directions of multi-view image 103. Mode controller 121 may generate mode selection signal 124 to activate wide-angle backlight 115 or multi-view backlight 118.
[0023] In 2D mode, a wide-angle backlight 115 can be used to generate images such that the multi-view display 112 behaves like a 2D display. By definition, "wide-angle" emitted light is defined as light that has a cone angle that is greater than the cone angle of the multi-view image or view of the multi-view display. In particular, in some embodiments, the wide-angle emitted light can have a cone angle that is greater than about 20 degrees (e.g., >±20°). In other embodiments, the wide-angle emitted light cone angle can be greater than about 30 degrees (e.g., >±30°), or greater than about 40 degrees (e.g., >±40°), or about It may be greater than 50 degrees (e.g., >±50°). For example, the cone angle of the wide-angle emitted light is about 60 degrees. It could be bigger (e.g., >±60°) 。
[0024] The multi-view mode may use a multi-view backlight 118 instead of the wide-angle backlight 115. The multi-view backlight 118 may have a multi-beam element array on its top or bottom surface that scatters light as multiple directional light beams with different principal angular directions. For example, if the multi-view display 112 operates in the multi-view mode to display a multi-view image having four views, the multi-view backlight 118 may scatter light into four directional light beams, each corresponding to a different view. The mode controller 121 may sequentially switch between the 2D mode and the multi-view mode such that the multi-view image is displayed using the multi-view backlight for a first consecutive time interval and the 2D image is displayed using the wide-angle backlight for a second consecutive time interval. The directional light beams may be at predetermined angles, with each directional light beam corresponding to a different view of the multi-view image.
[0025] In some embodiments, each backlight of the multi-view display 112 is configured to direct light within a light guide as guided light. As used herein, a "light guide" is defined as a structure that uses total internal reflection (TIR) to guide light within the structure. In particular, a light guide may include a core that is substantially transparent at the operating wavelength of the light guide. In various examples, the term "light guide" generally refers to a dielectric light guide that uses total internal reflection to guide light at the interface between the dielectric material of the light guide and the material or medium surrounding the light guide. By definition, the condition for total internal reflection is that the refractive index of the light guide is greater than the refractive index of the surrounding medium adjacent to the surface of the light guide material. In some embodiments, a light guide may include a coating in addition to or instead of the aforementioned refractive index difference to further promote total internal reflection. The coating may be, for example, a reflective coating. The light guide may be any of several light guides, including, but not limited to, a plate or slab guide and one or both of a strip guide. The light guide may be shaped like a plate or slab. The light guide may be edge-lit by a light source (e.g., a light-emitting device).
[0026] In some embodiments, the multi-view backlight 118 of the multi-view display 112 is configured to scatter portions of the guided light as directional emitted light using multibeam elements of a multibeam element array, where each multibeam element of the multibeam element array includes one or more of a diffraction grating, a micro-refractive element, and a micro-reflective element. In some embodiments, the diffraction grating of the multibeam element can include multiple individual sub-gratings. In some embodiments, the micro-reflective element is configured to reflectively combine or scatter the guided light portions as multiple directional light beams. The micro-reflective element may have a reflective coating to control how the guided light is scattered. In some embodiments, the multi-beam element includes a micro-refractive element configured to combine or scatter the guided light portions as multiple directional light beams by or using refraction (i.e., refractively scattering the guided light portions).
[0027] The multi-view display 112 may also include a light valve array disposed above a backlight (e.g., above the wide-angle backlight 115 and above the multi-view backlight 118). The light valves of the light valve array may be, for example, liquid crystal light valves, electrophoretic light valves, light valves based on or using electrowetting, or any combination thereof. When operating in 2D mode, the wide-angle backlight 115 emits light toward the light valve array. This light may be diffuse light emitted at a wide angle. Each light valve, when illuminated by light emitted by the wide-angle backlight 115, is controlled to achieve a particular pixel valve to display a 2D image. In this regard, each light valve corresponds to a single pixel. In this regard, a single pixel may include different colored pixels (e.g., red, green, blue) that make up a single pixel cell (e.g., an LCD cell).
[0028] When operating in multi-view mode, the multi-view backlight 118 emits directional light beams to illuminate the light valve array. The light valves may be grouped together to form a multi-view pixel. For example, in a four-view multi-view format, the multi-view pixel may include four different pixels, each corresponding to a different view. Each pixel within the multi-view pixel may further include a different color pixel.
[0029] Each light valve in the multi-view pixel arrangement may be illuminated by one of the light beams having a primary angular direction. A multi-view pixel is thus a group of pixels that provide different views of a pixel in a multi-view image. In some embodiments, each multi-beam element of the multi-view backlight 118 is dedicated to a multi-view pixel of the light valve array.
[0030] The multi-view display 112 comprises a screen for displaying the multi-view image 103. The screen may be, for example, the display screen of a telephone (e.g., a mobile phone, a smartphone, etc.), a tablet computer, a laptop computer, a computer monitor of a desktop computer, a camera display, or an electronic display of virtually any other device.
[0031] As used herein, the article "a" is intended to have its ordinary meaning in the patent art, namely, "one or more." For example, "a processor" means "one or more processors," and thus, as used herein, "the memory" means "one or more memory components."
[0032] FIG. 3 illustrates an example of accessing a multi-view image including multiple tile-view images according to one embodiment consistent with principles described herein. FIG. 3 illustrates an example of multi-view format detection by accessing a multi-view image 204 including multiple tile-view images. The multi-view image 204 may be similar to the multi-view image 103 of FIGS. 1 and 2. The multi-view image 204 may be formatted as or contained in a file. The multi-view image 204 may be received from a source such as a remote server, a streaming service, a repository, or a local application. A multi-view display system may be configured to process and render the multi-view image 204 for display. For example, the multi-view display system may execute an application 207 that causes the multi-view image 204 to be displayed. The application 207 may be a media player application, a video game, a social media application, or any other application that displays multi-view content. A view image is considered tiled because it is positioned adjacent to another view image in a tiled configuration. Thus, the multi-view image is formatted such that each view image is arranged as a grid of tiles along the horizontal direction, vertical direction, or both.
[0033] A multi-view image 204 can include one or more view images organized in a tiled arrangement. The multi-view format refers to the number of view images as well as the way the view images are spatially arranged as tiles. In this regard, the multi-view image 204 can be visualized as a 2D image with various view images formed as an array of tiles, with each tile corresponding to a different view image. Thus, the multi-view image 204 is arranged according to a particular multi-view format 210. FIG. 3 illustrates different examples of multi-view formats. The first multi-view format shown is a single tile format 210a (e.g., a single-view image). The single tile format 210a includes a single view (e.g., one tile) and is therefore not a true multi-view image. However, embodiments relate to distinguishing the single tile format from other multi-view formats 210 having multiple tiles. The second multi-view format shown is a 2×1 tile format 210b. The 2×1 tile format may also be referred to as a side-by-side format, in which two different view images are arranged side by side along the horizontal direction. The 2x1 tile format 210b may be used such that one view image is intended for the left eye and the other view image is intended for the right eye. As described above, the multi-view image 204 includes different views of the same scene, with the separate view images being mostly similar but with some degree of parallax. The third multi-view format shown is the 1x2 tile format 210c. The 1x2 tile format 210c also includes two different view images, but arranged adjacently along the vertical direction. The fourth multi-view format shown is the 2x2 tile format 210d. The 2x2 tile format 210d includes four different view images arranged two tiles horizontally and two tiles vertically. The 2x2 tile format 210d is sometimes referred to as a quad format.
[0034] The multiview image 204 may conform to one of various types of multiview formats 210 having any number of tiles in any of a variety of arrangements (e.g., a 1x4 tile format, a 4x1 tile format, an 8-view format, etc.). The application 207 may be unaware of the type of multiview format 210 of the multiview image 204. Embodiments relate to a multiview format detector 213 that may be used to analyze the multiview image 204 and detect the multiview format 210 of the multiview image 204. The multiview format detector 213 may generate a multiview format label 216 that indicates the multiview format 210 of the multiview image. The application 207 may use the multiview format label 216 to render or process the multiview image 204. This allows the application 207 to process a multiview image 204 having an initially unknown or unknown multiview format 210.
[0035] The multiview format detector 213 begins by accessing a multiview image 204 that includes multiple tile-view images. The multiview image 204 may be accessed from memory or may be provided by a source that sends the multiview image 204 to an application 207 running on the multiview display system.
[0036] FIG. 4 illustrates an example of generating a cross-correlation map from a multiview image 204 according to one embodiment consistent with principles described herein. For example, upon accessing the multiview image 204, the multiview format detector 213 may perform operations including generating a cross-correlation map from the multiview image by autocorrelating the multiview image with a shifted copy of the multiview image. FIG. 4 illustrates an autocorrelation operation in which the multiview image 204 is correlated with a shifted copy 219 (e.g., autocorrelation) of the multiview image. The shifted copy 219 may be convolved or slid horizontally and vertically over the multiview image 204 as part of the autocorrelation process. The autocorrelation operation may include a dot product calculation between a portion of the shifted copy 219 and the multiview image 204. The autocorrelation operation may be performed in the pixel domain or the frequency domain. The autocorrelation operation generates a cross-correlation map 222. The cross-correlation map 222 may be a two-dimensional matrix of elements. The cross-correlation map 222 can be represented or visualized as an image, where each pixel value contains the cross-correlation value for a given position of the shifted copy 219 relative to the multi-view image 204 .
[0037] FIG. 5 shows examples of various cross-correlation maps generated from different multi-view formats according to one embodiment consistent with principles described herein. A multi-view image 231 having a single tile format can yield a cross-correlation map 234. As shown in FIG. 5, the cross-correlation map 234 has low cross-correlation values, indicated as areas with sparse cross-hatching, and high cross-correlation values, indicated as areas with denser cross-hatching. The majority of the cross-correlation map has sparse cross-hatching, indicating low levels of cross-correlation in most locations. However, the cross-correlation map 234 has a single hot spot in the center. A hot spot is a region within the cross-correlation map that contains relatively high cross-correlation values. Specifically, the center of the cross-correlation map 234 is densely cross-hatched, meaning that there is significant cross-correlation when the shifted copy 219 is placed on top of (e.g., centered on) the multi-view image 204. Additionally, there are no strong indications of cross-correlation outside the center of the cross-correlation map 234.
[0038] A multi-view image 237 with a 2x1 tile format can result in a cross-correlation map 240 with a hot spot in the center and two hot spots 243 along the left and right edges, indicating three strong cross-correlation locations as the shifted copy 219 slides along the horizontal direction.
[0039] A multiview image 246 with a 2x2 tile format can result in a cross-correlation map 249 with a hotspot in the center, as well as several hotspots along the left and right edges and hotspots along the corners 252. This indicates that there are strong cross-correlation values in these regions when the shifted copy 219 slides along both the horizontal and vertical directions.
[0040] FIG. 6 illustrates an example of sampling a cross-correlation map according to one embodiment consistent with principles described herein. For example, upon generating cross-correlation map 222, multi-view format detector 213 may perform operations including sampling the cross-correlation map at multiple predetermined locations on the cross-correlation map to identify a set of cross-correlation values. FIG. 6 illustrates predetermined locations 260a-d having specific coordinates on cross-correlation map 222. In some embodiments, at least one of the predetermined locations includes a corner region of the cross-correlation map (e.g., predetermined location 260a). In some embodiments, at least one of the predetermined locations includes a midpoint along the length or width of cross-correlation map 222 (e.g., predetermined locations 260b, 260c). Sampling the corners may enable detection of vertical tiles, and sampling the midpoints may enable detection of horizontal tiles. The number of samples may be selected based on the number of expected view images. For example, as the number of potential views increases, more sample points may be used to distinguish between multi-view images.
[0041] In some embodiments, each of the cross-correlation values is determined based on a single cross-correlation value. For example, sampling the cross-correlation map 222 may include identifying the value of a single element (e.g., pixel) of the cross-correlation map 222. The predetermined location may be a particular x-y coordinate of the cross-correlation map 222. In some embodiments, each of the cross-correlation values is determined by averaging the values within the corresponding predetermined location. In this example, each predetermined location may include a range of x-y coordinates of the cross-correlation map 222. The cross-correlation values within each range may be averaged or otherwise aggregated. Thus, rather than sampling one element that may be in a hotspot, several pixels within the hotspot may be sampled to generate an averaged cross-correlation value.
[0042] Once sampled, a set of cross-correlation values 263 is determined. The set of cross-correlation values 263 may be a relatively small array of elements, each element being a cross-correlation value. The set of cross-correlation values 263 may be used as a feature in a feature set. To detect a multi-view format with a larger number of views, more samples can be used. In the example of FIG. 6, four samples are acquired, each corresponding to a respective predetermined location 260a-d. Furthermore, in this example, the first cross-correlation value is 4 and is sampled from a corner region (predetermined location 260a). The second cross-correlation value is 8 and is sampled from the top horizontal midpoint (predetermined location 260b). The third cross-correlation value is 121 and is sampled from the left vertical midpoint (predetermined location 260c). The fourth cross-correlation value is 118 and is sampled from the center (predetermined location 260d). This feature of the cross-correlation values 263 indicates the possibility of autocorrelation along the horizontal direction that matches the signature pattern of a 2x1 format. Specifically, the first two elements of the cross-correlation value 263 are relatively low, indicating low autocorrelation as the shifted copy slides vertically, and the last two elements of the cross-correlation value 263 are relatively high, indicating high autocorrelation as the shifted copy slides horizontally.
[0043] 7 illustrates an example of detecting the multiview format of a multiview image, according to one embodiment consistent with principles described herein. For example, upon determining the set of cross-correlation values 263, the multiview format detector 213 may perform operations including detecting the multiview format of the multiview image by classifying a feature set of the multiview image that includes the set of cross-correlation values. The multiview format detector 213 may include a feature extractor 266 that extracts features from the multiview image 204 and generates a feature set 269 of the multiview image 204.
[0044] As mentioned above, with respect to Figure 6, one feature may be a set of cross-correlation values 263. In some embodiments, As shown in Figure 7,The feature extractor 266 extracts the aspect ratio of the multi-view image 204. 272 and the feature set 269 can perform other operations, including determining the aspect ratio 272 Aspect Ratio 272 may be the ratio of the height to the width of the multiview image 204 in pixels. The feature extractor 266 may identify any number of other features that characterize the multiview image 204.
[0045] The multiview format detector 213 may include a classifier 275. The classifier 275 may be a support vector machine. In some embodiments, the classifier 275 is trained using training data, where the training data includes a training set of multiview images and corresponding training labels. A feature set with known multiview formats may be input to the classifier 275 as training data for training the classifier 275. When training the classifier 275, the training labels may include labels indicating a single tile format, a 1x2 tile format, a 2x1 tile format, and a 2x2 tile format. The training data may include any multiview format 210. The classifier 275 may be trained during construction time and then executed at runtime to classify unknown multiview formats of the multiview image 204.
[0046] For example, the training data can include feature vectors with labels that provide the ground truth. The training data can be formatted as a file or table that correlates a particular set of feature values to the ground truth. The training data can be input to the classifier 275, for example, using an application programming interface. The classifier 275 can use the training data to train a classifier by building a classification or regression model in the form of a tree structure, by separating the training data by class and calculating the conditional probability of each class, by generating training weights from the training data to be used in a convolution operation, by converting the training data into data points in a feature vector space, or by using other training methods to configure the classifier.
[0047] The multiview format detector 213 can input a feature set of the multiview image to a classifier, where the feature set includes a set of cross-correlation values. The classifier 275 is provided with a feature set 269 extracted from the multiview image 204, where the feature set 269 includes cross-correlation values sampled at predetermined locations. The classifier 275 generates a multiview format label 216. The application 207 can retrieve a label from the classifier, where the label indicates the multiview format of the multiview image. The label may be the multiview format label 216.
[0048] 8 illustrates an example of rendering a multiview image on a multiview display based on a multiview format, according to one embodiment consistent with principles described herein. For example, an application 207 may retrieve a multiview format label 216 to enable rendering of a multiview image 204 using the multiview display 112. FIG. 8 provides an example of how an application 207 may use the multiview format label 216 to render a multiview image 204. The embodiment illustrated in FIG. 8 relates to an application 207 rendering a multiview image 204 using a rendering pipeline 276.
[0049] The rendering pipeline 276 may include one or more components that process image content for presentation on a display (e.g., the multiview display 112). The rendering pipeline 276 may include components executable by a central processing unit (CPU), a graphics processing unit (GPU), or any combination thereof. The application 207 may use function calls or an application programming interface (API) to control the rendering pipeline 276. 276 In some embodiments, the rendering pipeline 276 may include a view compositor 277, an interlacer 278, or other components that render graphics.
[0050] The view synthesizer 277 performs view synthesis to generate a target number of view images from the original view images included in the multiview image 204. For example, if the multiview image 204 includes fewer views than the target number of views supported by the multiview display system, the view synthesizer can generate additional view images based on the original view images of the multiview image 204. View synthesis includes the act of interpolating or extrapolating one or more original view images to generate new view images. View synthesis can include one or more of forward warping, depth testing, and inpainting techniques to sample nearby regions, such as to fill in unoccluded regions. Forward warping is an image distortion process that applies a transformation to a source image. Pixels from the source image can be processed in scanline order, and the result is projected into the target image. Depth testing is a process in which a fragment of an image that is or is to be processed by a shader has its depth value tested against the depth of the sample being written. The fragment is discarded when the test fails. Additionally, a depth buffer is updated with the output depth of the fragment when the test passes. Inpainting refers to filling in missing or unknown regions of an image. Some techniques include predicting pixel values based on nearby pixels or reflecting nearby pixels into unknown or missing regions. Missing or unknown regions of an image may result from scene disocclusion, which refers to a scene object that is partially obscured by another scene object. In this regard, reprojection may involve image processing techniques for constructing a new perspective of a scene from an original perspective. View images can be synthesized using a trained neural network. Additionally, the view synthesizer 277 can generate additional view images using a depth map of the multi-view image. The depth map may be a file or image indicating the depth of each pixel in a particular view image.
[0051] The rendering pipeline 276 may also include an interlacer 278 (e.g., an interlacing program, an interlacing shader, etc.). The interlacer 278 may be configured to interlace pixels of the tile view images of the multi-view image (along with any composite view images). The interlacer 278 may be a graphics shader and may be part of the application 207 or may otherwise be invoked by the application 207.
[0052] The multiview format label 216 specifies the number of view images and their orientation. The view compositor 277 can generate additional view images based on the multiview format label 216. For example, assume that the target number of views is four views and the multiview format label 216 indicates that the multiview image is in a 2×1 format. In this example, the view compositor 277 can identify the specific positions of the two original views within the multiview image 204 and generate two additional view images to achieve the target number of view images. Thus, in this example, the rendering pipeline 276 converts the 2×1 format multiview image into a four-view multiview image by synthesizing two additional view images from the two original view images present in the 2×1 format multiview image. The rendering pipeline 276 determines the coordinates or positions of the two original view images within the tiled multiview image by using the multiview format label 216 to determine that the multiview image is a 2×1 format multiview image. Knowing the multiview format allows the original view images to be positioned, separated, and new view images to be synthesized.
[0053] In some embodiments, if the multiview format label 216 indicates that no view images need to be synthesized, the view compositor 277 can be bypassed. For example, if the target number of views is four and the multiview format label 216 indicates that the multiview image has four views in a particular orientation (e.g., a quad orientation), then view synthesis is not required. In some embodiments, if the multiview format label 216 indicates that there are more view images in the multiview image 204 than are supported by the multiview display system, view images may be removed to achieve the target number of view images supported by the multiview display system.
[0054] Thus, the multiview format label 216 can be used to identify the location of the view images and to determine whether additional views need to be synthesized or whether original view images need to be removed to achieve a target number of views. Once the view images have been identified based on the multiview format label (and potentially new view images synthesized), the interlacer 278 generates an interlaced multiview image 279.
[0055] An interlacer 278 (e.g., an interlacing program) that performs the rendering can receive the multi-view image 204 and any composite views and generate an interlaced multi-view image 279 as output. The interlaced multi-view image 279 can be loaded into a graphics memory buffer or otherwise painted to the multi-view display 112. The interlacing program can format the interlaced multi-view image 279 according to a format specific and native to the multi-view display 112. In some embodiments, this includes interlacing pixels of each view image to co-locate pixels of the various views based on their corresponding positions. For example, the top-left pixel of each view image may be co-located when generating the interlaced multi-view image 279. In other words, the pixels of each view image are spatially multiplexed or otherwise interlaced to generate the interlaced multi-view image 279. The multi-view format label 216 allows the rendering pipeline 276 to identify the positions of the view images. For example, knowing the multi-view format allows the rendering pipeline 276 to determine the location of each pixel in each tile (e.g., view image). In other words, knowing the multi-view format allows the rendering pipeline 276 to identify the start and end pixels of each view image in the multi-view image 204. In this example, the multi-view display system presents four views (e.g., view 1, view 2, view 3, and view 4). The multi-view format label 216 indicates the locations of these view images and whether new view images need to be synthesized.
[0056] As shown in FIG. 8, an interlaced multiview image 279 has spatially multiplexed or otherwise interlaced views. FIG. 8 shows pixels corresponding to one of four views, where the pixels are interlaced (e.g., interleaved or spatially multiplexed). Pixels belonging to view 1 are represented by the number 1, pixels belonging to view 2 are represented by the number 2, pixels belonging to view 3 are represented by the number 3, and pixels belonging to view 4 are represented by the number 4. The views of the interlaced multiview image 279 are interlaced horizontally, pixel by pixel, along each row. The interlaced multiview image 279 has rows of pixels represented by uppercase letters A through E and columns of pixels represented by lowercase letters a through h. FIG. 8 shows the location of one of the multiview pixels 282 in row E, columns e through h. The multiview pixel 282 is an array of pixels taken from each of the four views. In other words, multi-view pixel 282 is the result of spatially multiplexing and interlacing the individual pixels of each of the four views. Although Figure 8 shows pixels of different views being interlaced horizontally, pixels of different views can be interlaced vertically and both horizontally and vertically.
[0057] The interlaced multiview image 279 may result in a multiview pixel 282 having one pixel from each of the four views, assuming the target number of views is four. As mentioned above, the multiview format label allows the rendering pipeline 276 to find the corresponding pixel in each view image. In some embodiments, the multiview pixels may be staggered in a particular direction, as shown in FIG. 8, where the multiview pixels are horizontally aligned while being staggered vertically. In other embodiments, the multiview pixels may be staggered horizontally and vertically aligned. The particular way in which the multiview pixels are interlaced and staggered may depend on the design of the multiview display 112. The interlaced multiview image 279 is created by interlacing pixels and arranging the pixels into multiview pixels to map them to the physical pixels of the multiview display 112 (e.g., a light valve array). 284 ) In other words, pixel coordinates of interlaced multi-view image 279 correspond to physical locations on multi-view display 112. Multi-view pixels 282 have a mapping 287 to a particular set of light valves in light valve array 284. Light valve array 284 is controlled to modulate light according to interlaced multi-view image 279.
[0058] FIG. 9 is a flowchart providing example functionality of a multi-view display system configured to detect multi-view formats, according to one embodiment consistent with principles described herein. The flowchart of FIG. 9 provides an example of different types of functionality implemented by a computing device (e.g., the multi-view display system of FIG. 10) executing a set of instructions. For example, the multi-view display system may include a multi-view display configured to display a multi-view image (e.g., multi-view display 112). The multi-view display system may further include a processor and a memory storing a plurality of instructions that, when executed by the processor, cause the processor to perform various operations illustrated in the flowchart. Alternatively, the flowchart of FIG. 9 may be considered to illustrate example steps of a method implemented on a computing device according to one or more embodiments. FIG. 9 may also represent a computer program, such as a computer program on a transient or non-transitory computer-readable storage medium, storing executable instructions that, when executed by a processor of the computer system, perform operations of multi-view format detection. An example computer system may be the multi-view display system described below with respect to FIG. 8. These operations are illustrated in a flowchart.
[0059] The instructions may be executed to cause a processor to access (304) a multi-view image including multiple tile-view images. For example, the multi-view image may be accessed by a multi-view format detector. The multi-view format detector may be a software module or routine configured to detect the multi-view format of the multi-view image. The multi-view image may be accessed from a memory location or may be provided to the multi-view format detector by an application configured to render the multi-view image. The application may invoke or otherwise include the multi-view format detector to classify the multi-view image based on its multi-view format.
[0060] The instructions may be executed to cause a processor to generate a cross-correlation map from the multiview image by autocorrelating the multiview image with a shifted copy of the multiview image (307). The shifted copy may be replicated from the multiview image to be identical to the multiview image. The shifted copy is shifted across the multiview image to calculate cross-correlation values. The cross-correlation values are stored as a cross-correlation map. The cross-correlation map may be a two-dimensional array of elements, each element having a cross-correlation value and a location (e.g., coordinates) within the cross-correlation map. The location of the element corresponds to the location of the shifted copy relative to the multiview image. Thus, the size of the cross-correlation map may be larger than the multiview image as the shifted copy slides from an upper corner of the multiview image to a lower corner of the multiview image.
[0061] To identify the set of cross-correlation values, instructions 310 can be executed that cause the processor to sample the cross-correlation map at multiple predetermined locations in the cross-correlation map. The predetermined locations can be specified as coordinates or coordinate ranges within the cross-correlation map. The predetermined locations can be corner regions, midpoints along the length of the cross-correlation map, midpoints along the height of the cross-correlation map, the center of the cross-correlation map, or other predetermined locations. The predetermined locations correspond to locations where hotspots are likely to appear based on the arrangement of the tiled view images.
[0062] The instructions may be executed to cause the processor to generate 313 a feature set. In some embodiments, when generating the feature set, the instructions may cause the processor to determine an aspect ratio of the multi-view image, the feature set further including the aspect ratio. The feature set may include individual features extracted from the multi-view image.
[0063] The processor may execute instructions 316 to input a feature set of the multi-view image to a classifier, the feature set including a set of cross-correlation values. For example, the feature set may include an array of sampled cross-correlation values. The feature set is input (e.g., passed) to a trained classifier configured to generate a label. The label may include a list of potential multi-view formats with corresponding probabilities that the input feature set is a particular multi-view format. The label may identify the multi-view format with the highest probability based on the input feature set. Prior to execution of the classifier, embodiments relate to training the classifier to classify the feature set using training data, the training data including a training set of multi-view images and corresponding training labels.
[0064] Instructions 319 can be executed to cause the processor to retrieve a label from the classifier, the label indicating a multiview format of the multiview image. The label may be sent to an application responsible for rendering the multiview image. The label can indicate whether the multiview format is a 1x2 tile format, a 2x1 tile format, or a 2x2 tile format. In some embodiments, the label can indicate when the multiview image is not a true multiview image, such that the multiview image is in a single tile format.
[0065] The instructions may be executed to cause a processor to render a multi-view image on a multi-view display based on the multi-view format (322). In this regard, the multi-view image is configured to be rendered on the multi-view display based on the multi-view format. The rendering may be performed by a rendering pipeline. Rendering the multi-view image may include, upon detecting the multi-view format, interlacing pixels of tile-view images of the multi-view image. For example, by knowing the multi-view format, the rendering pipeline can arrange the individual view images of the multi-view image and determine whether new view images need to be synthesized to achieve a target number of views. Upon identifying the view images and potentially synthesizing new view images, an interlacer can interlace the view images, whether they were originally part of the multi-view image or newly synthesized. Thus, an application can use the label to instruct the rendering pipeline to render the multi-view image.
[0066] The flowchart of FIG. 9 described above may illustrate an implementation of a system or method for multi-view format detection and an executable instruction set. When implemented in software, each box may represent a module, segment, or portion of code containing instructions for implementing a specified logical function. The instructions may be embodied in the form of source code containing human-readable instructions written in a programming language, object code compiled from the source code, or machine code containing numerical instructions recognizable by a suitable execution system such as a processor or computing device. The machine code may be converted from the source code. When implemented in hardware, each block may represent a circuit or several interconnected circuits for implementing a specified logical function.
[0067] 9 shows a particular order of execution, it is understood that the order of execution may differ from that shown. For example, the order of execution of two or more boxes may be scrambled relative to the order shown. Also, two or more boxes shown may be executed simultaneously or with partial concurrence. Furthermore, in some embodiments, one or more of the boxes may be skipped or omitted.
[0068] FIG. 10 is a schematic block diagram illustrating an exemplary diagram of a multi-view display system according to one embodiment consistent with principles described herein. The multi-view display system 1000 may include a system of components that perform various computational operations to detect the multi-view format of a multi-view image. The multi-view display system 1000 may be a laptop, tablet, smartphone, touchscreen system, intelligent display system, or other client device. The multi-view display system 1000 may include various components, such as, for example, a processor 1003, memory 1006, input / output (I / O) components 1009, a display 1012, and potentially other components. These components may be coupled to a bus 1015 that functions as a local interface allowing the components of the multi-view display system 1000 to communicate with each other. While the components of the multi-view display system 1000 are shown as contained within the multi-view display system 1000, it should be understood that at least some of the components may be coupled to the multi-view display system 1000 via external connections. For example, components may be externally plugged into or otherwise connected to the multi-view display system 1000 via external ports, sockets, plugs, wireless links, or connectors.
[0069] The processor 1003 may be a central processing unit (CPU), a graphics processing unit (GPU), any other integrated circuit that performs computational operations, or any combination thereof. The processor 1003 may include one or more processing cores. The processor 1003 comprises circuitry that executes instructions. Instructions include, for example, computer code, programs, logic, or other machine-readable instructions that are received and executed by the processor 1003 to perform computational functions embodied in the instructions. The processor 1003 may execute instructions that operate on data. For example, the processor 1003 may receive input data (e.g., images or frames), process the input data according to an instruction set, and generate output data (e.g., processed images or frames). As another example, the processor 1003 may receive instructions and generate new instructions for subsequent execution. The processor 1003 may comprise hardware that implements a graphics pipeline for processing and rendering light field content. For example, the processor 1003 may include one or more GPU cores, vector processors, scalar processes, or hardware accelerators.
[0070] The memory 1006 may include one or more memory components. The memory 1006 is defined herein to include either or both volatile and nonvolatile memory. A volatile memory component is one that does not retain information upon loss of power. Volatile memory may include, for example, random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), magnetic random access memory (MRAM), or other volatile memory structures. System memory (e.g., main memory, cache, etc.) may be implemented using volatile memory. System memory refers to high-speed memory that can temporarily store data or instructions for quick read and write access to support the processor 1003.
[0071] Nonvolatile memory components are those that retain information upon loss of power. Nonvolatile memory includes read-only memory (ROM), hard disk drives, solid-state drives, USB flash drives, memory cards accessed via memory card readers, floppy disks accessed via associated floppy disk drives, optical disks accessed via optical disk drives, and magnetic tapes accessed via appropriate tape drives. ROM can include, for example, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other similar memory devices. Storage memory may be implemented using nonvolatile memory to provide long-term retention of data and instructions.
[0072] Memory 1006 may refer to a combination of volatile and non-volatile memory used to store instructions and data. For example, data and instructions may be stored in non-volatile memory and loaded into volatile memory for processing by processor 1003. Execution of instructions may include, for example, a compiled program that is loaded from non-volatile memory into volatile memory and then converted into machine code in a format that can be executed by processor 1003; source code that is converted into a suitable format such as object code that can be loaded into volatile memory for execution by processor 1003; or source code that is interpreted by another executable program to generate instructions in volatile memory and executed by processor 1003. Instructions may be stored in or loaded into any portion or component of memory 1006, including, for example, RAM, ROM, system memory, storage, or any combination thereof.
[0073] Although memory 1006 is shown as separate from other components of multi-view display system 1000, it should be understood that memory 1006 may be embedded or otherwise integrated, at least in part, into one or more components. For example, processor 1003 may include on-board memory registers or cache for performing processing operations. Device firmware or drivers may include instructions stored in dedicated memory devices.
[0074] The I / O components 1009 include, for example, a touchscreen, speakers, microphones, buttons, switches, dials, cameras, sensors, accelerometers, or other components that receive user input or generate output directed to a user. The I / O components 1009 can receive user input and convert it into data for storage in memory 1006 or for processing by processor 1003. The I / O components 1009 can receive data output by memory 1006 or processor 1003 and convert it into a format that can be perceived by the user (e.g., sound, tactile response, visual information, etc.).
[0075] A particular type of I / O component 1009 is a display 1012. The display 1012 can include a multi-view display (e.g., multi-view display 112), a multi-view display combined with a 2D display, or any other display that presents images. A capacitive touchscreen layer functioning as the I / O component 1009 can be laminated within the display to allow a user to provide input while simultaneously perceiving visual output. The processor 1003 can generate data formatted as an image for presentation on the display 1012. The processor 1003 can execute instructions to render the image on the display as perceived by the user.
[0076] The bus 1015 facilitates the communication of instructions and data between the processor 1003, the memory 1006, the I / O component 1009, the display 1012, and any other components of the multiview display system 1000. The bus 1015 may include address translators, address decoders, fabric, conductive traces, conductors, ports, plugs, sockets, and other connectors to enable the communication of data and instructions.
[0077] The instructions in memory 1006 may be embodied in various forms in a manner that implements at least a portion of a software stack. For example, the instructions may be embodied as part of an operating system 1031, applications 1034, device drivers (e.g., display driver 1037), firmware (e.g., display firmware 1040), other software components, or any combination thereof. The operating system 1031 is a software platform that supports basic functionality of the multiview display system 1000, such as scheduling tasks, controlling the I / O components 1009, providing access to hardware resources, managing power, and supporting the applications 1034.
[0078] The applications 1034 run on the operating system 1031 and can access the hardware resources of the multi-view display system 1000 through the operating system 1031. In this regard, the execution of the applications 1034 is controlled, at least in part, by the operating system 1031. The applications 1034 may be user-level software programs that provide high-level functionality, services, and other features to the user. In some embodiments, the applications 1034 may be dedicated “apps” that are downloadable or accessible to users on the multi-view display system 1000. A user can launch the applications 1034 through a user interface provided by the operating system 1031. The applications 1034 are developed by developers and may be defined in a variety of source code formats. The applications 1034 may be developed using numerous programming or scripting languages, such as, for example, C, C++, C#, Objective C, Java, Swift, JavaScript, Perl, PHP, Visual Basic, Python, Ruby, Go, or other programming languages. The application 1034 may be compiled into object code by a compiler or interpreted by an interpreter for execution by the processor 1003. The application 1034 may be the application 207 described above with respect to FIG.
[0079] Device drivers, such as display driver 1037, contain instructions that allow operating system 1031 to communicate with various I / O components 1009. Each I / O component 1009 may have its own device driver. Device drivers may be installed so that they are stored in storage and loaded into system memory. For example, upon installation, display driver 1037 translates high-level display instructions received from operating system 1031 into low-level instructions that are executed by display 1012 to display images.
[0080] Firmware, such as display firmware 1040, may include machine code or assembly code that enables I / O component 1009 or display 1012 to perform low-level operations. Firmware may translate the electrical signals of a particular component into higher-level instructions or data. For example, display firmware 1040 may control how display 1012 activates individual pixels at a low level by adjusting voltage or current signals. Firmware may be stored in and executed directly from non-volatile memory. For example, display firmware 1040 may be embedded in a ROM chip coupled to display 1012 such that the ROM chip is separate from other storage and system memory of multi-view display system 1000. Display 1012 may include processing circuitry for executing display firmware 1040.
[0081] The operating system 1031, applications 1034, drivers (e.g., display driver 1037), firmware (e.g., display firmware 1040), and potentially other instruction sets may each include instructions executable by the processor 1003 or other processing circuitry of the multiview display system 1000 to perform the functions and operations described above. The instructions described herein may be embodied in software or code executed by the processor 1003 as described above, although alternatively, the instructions may also be embodied in dedicated hardware or a combination of software and dedicated hardware. For example, the functions and operations performed by the instructions described above may be implemented as circuits or state machines using any one or combination of several technologies. These technologies may include, but are not limited to, discrete logic circuits having logic gates for implementing various logical functions upon the application of one or more data signals, application specific integrated circuits (ASICs) having appropriate logic gates, field programmable gate arrays (FPGAs), or other components, etc.
[0082] In some embodiments, instructions for performing the functions and operations described above may be embodied in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium may or may not be part of the multi-view display system 1000. The instructions may include, for example, statements, code, or declarations that may be fetched from the computer-readable medium and executed by a processing circuit (e.g., processor 1003). As defined herein, a "non-transitory computer-readable storage medium" is defined as any medium that can contain, store, or maintain the instructions described herein for use by or in connection with an instruction execution system, such as the multi-view display system 1000, and further excludes transitory media, including, for example, carrier waves.
[0083] The non-transitory computer-readable medium may include any one of many physical media, such as magnetic, optical, or semiconductor media. More specific examples of suitable non-transitory computer-readable media include, but are not limited to, magnetic tape, magnetic floppy diskettes, magnetic hard drives, memory cards, solid-state drives, USB flash drives, or optical disks. The non-transitory computer-readable medium may also be, for example, random access memory (RAM), including static random access memory (SRAM) and dynamic random access memory (DRAM), or magnetic random access memory (MRAM). Furthermore, the non-transitory computer-readable medium may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other types of memory devices.
[0084] The multi-view display system 1000 can perform any of the operations or implement the functions described above. For example, the process flows described above may be performed by the multi-view display system 1000 executing instructions and processing data. Although the multi-view display system 1000 is shown as a single device, the embodiments are not so limited. In some embodiments, the multi-view display system 1000 can offload processing of instructions in a distributed manner, such that multiple multi-view display systems 1000 or other computing devices operate together to execute instructions that may be stored or loaded in a distributed manner. For example, at least some instructions or data may be stored, loaded, or executed on a cloud-based system operating in conjunction with the multi-view display system 1000.
[0085] Thus, examples and embodiments for detecting the multi-view format of a multi-view image have been described. The operations include performing autocorrelation on the multi-view image and generating features based on sampling specific locations of the cross-correlation map. A classifier is trained based on at least the features to detect the multi-view format. It should be understood that the above examples are merely illustrative of some of the many specific examples that illustrate the principles described herein. Clearly, those skilled in the art can readily devise numerous other configurations without departing from the scope defined by the following claims. It should be noted that the present specification discloses the following aspects. [Aspect 1] 1. A computer-implemented method for multiview format detection, comprising: accessing a multi-view image comprising a plurality of tile-view images; generating a cross-correlation map from the multiview image by autocorrelating the multiview image with shifted copies of the multiview image; sampling the cross-correlation map at a plurality of predetermined locations in the cross-correlation map to identify a set of cross-correlation values; detecting a multiview format of the multiview image by classifying a feature set of the multiview image comprising the set of cross-correlation values; The method, wherein the multi-view image is configured to be rendered on a multi-view display based on the multi-view format. [Aspect 2] determining an aspect ratio of the multi-view image, the set of features further comprising the aspect ratio. 2. The computer-implemented method of multi-view format detection according to embodiment 1. [Aspect 3] 2. The computer-implemented method of multi-view format detection of embodiment 1, wherein at least one of the predetermined locations includes a corner region of the cross-correlation map. [Aspect 4] 2. The computer-implemented method of multi-view format detection of embodiment 1, wherein at least one of the predetermined locations comprises a midpoint along one of a length and a width of the cross-correlation map. [Aspect 5] 2. The computer-implemented method of multi-view format detection of aspect 1, wherein each cross-correlation value of the set of cross-correlation values is determined based on a single cross-correlation value. [Aspect 6] 2. The computer-implemented method of multi-view format detection of aspect 1, wherein each cross-correlation value in the set of cross-correlation values is determined by averaging values within a corresponding predetermined position. [Aspect 7] and training a classifier to classify the feature set according to training data, the training data including a training set of multi-view images and corresponding training labels. 2. The computer-implemented method of multi-view format detection according to embodiment 1. [Aspect 8] 8. The computer-implemented method of multi-view format detection of embodiment 7, wherein the training labels include labels indicating a single tile format, a 1x2 tile format, a 2x1 tile format, and a 2x2 tile format. [Aspect 9] the multi-view display is configured to provide wide-angle emitted light in a two-dimensional mode (2D mode) using a wide-angle backlight; the multi-view display is configured to provide directional emitted light in a multi-view mode using a multi-view backlight having an array of multi-beam elements, the directional emitted light comprising a plurality of directional light beams provided by each multi-beam element of the multi-beam element array; the multi-view display is configured to time-multiplex the 2D mode and the multi-view mode using a mode controller, sequentially activating the wide-angle backlight during a first consecutive time interval corresponding to the 2D mode and sequentially activating the multi-view backlight during a second consecutive time interval corresponding to the multi-view mode; 2. The computer-implemented method of multi-view format detection of embodiment 1, wherein directions of the directional light beams correspond to different view directions of the multi-view image. [Aspect 10] The multi-view display is configured to guide the light in the light guide as guided light; A computer-implemented method of multi-view format detection as described in aspect 1, wherein the multi-view display is configured to scatter a portion of the guided light as directional radiation using multi-beam elements of an array of multi-beam elements, and each multi-beam element of the array of multi-beam elements includes one or more of a diffraction grating, a micro-refractive element, and a micro-reflective element. [Aspect 11] 1. A multi-view display system, comprising: a multi-view display configured to display a multi-view image; a processor; and a memory for storing a plurality of instructions, the plurality of instructions, when executed, causing the processor to: accessing a multi-view image comprising a plurality of tile-view images; generating a cross-correlation map from the multiview image by autocorrelating the multiview image with shifted copies of the multiview image; sampling the cross-correlation map at a plurality of predetermined locations in the cross-correlation map to identify a set of cross-correlation values; inputting a feature set of the multi-view images into a classifier, the feature set including the set of cross-correlation values; deriving a label from the classifier that indicates the multi-view format of the multi-view image; and rendering the multiview image on a multiview display based on the multiview format. [Aspect 12] The instructions, when executed, cause the processor to: 12. The multi-view display system of claim 11, further comprising determining an aspect ratio of the multi-view image, wherein the set of features further comprises the aspect ratio. [Aspect 13] 12. The multi-view display system of claim 11, wherein at least one of the predetermined locations comprises a corner region of the cross-correlation map. [Aspect 14] A multi-view display system as described in aspect 11, wherein at least one of the predetermined locations includes a midpoint along one of a length and a width of the cross-correlation map. [Aspect 15] 12. The multi-view display system of claim 11, wherein each cross-correlation value in the set of cross-correlation values is determined based on a single cross-correlation value. [Aspect 16] 12. The multiview display system of claim 11, wherein the classifier is trained using training data, the training data including a training set of multiview images and corresponding training labels. [Aspect 17] 17. The multi-view display system of embodiment 16, wherein the training labels include labels indicating a single tile format, a 1x2 tile format, and a 2x1 tile format. [Aspect 18] 1. A computer program comprising executable instructions that, when executed by a processor of a computer system, perform an operation of multiview format detection, the operation comprising: accessing a multi-view image comprising a plurality of tile-view images; generating a cross-correlation map from the multiview image by autocorrelating the multiview image with shifted copies of the multiview image; sampling the cross-correlation map at a plurality of predetermined locations in the cross-correlation map to identify a set of cross-correlation values; detecting a multiview format of the multiview image by classifying a feature set of the multiview image comprising the set of cross-correlation values; and upon detecting the multi-view format, interlacing pixels of the tile-view images of the multi-view image. [Aspect 19] The operation is 20. The computer program of embodiment 18, further comprising determining an aspect ratio of the multi-view image, wherein the set of features further comprises the aspect ratio. [Aspect 20] 20. The computer program of embodiment 18, wherein the multiview format is a 1x2 tile format, a 2x1 tile format, or a 2x2 tile format. [Explanation of symbols]
[0086] 103, 204, 231, 237, 246, 279 Multi-view images 106 ViewsImage 109 Principal angular direction 112 Multi-view display 115 Wide-angle backlight 118 Multi-view backlight 121 Mode Controller 124 Mode selection signal 207, 1034 Applications 210 multiview format 210a Single Tile Format 210b 2x1 tile format 210c 1x2 tile format 210d 2x2 tile format 213 Multi-view format detector 216 Multiview Format Labels 219 Shifted Copy 222, 234, 240, 249 Cross-correlation maps 243, 252 Hotspots 260a~d designated position 263 Cross-correlation value 266 Feature Extractor 267, 276 Rendering Pipeline 269 Feature Set 275 Classifier 277 viewsCompositor 278 Interlacer 282 multi-view pixels 284, 285 Light valve array 287 Mapping 1000 Multi-view Display System 1003 processor 1006 memory 1009 I / O components 1012 display 1015 Bus 1031 Operating Systems 1037 Display Driver 1040 Display Firmware
Claims
1. 1. A computer-implemented method for multiview format detection, comprising: accessing a multi-view image comprising a plurality of tile-view images; generating a cross-correlation map from the multiview image by autocorrelating the multiview image with shifted copies of the multiview image; sampling the cross-correlation map at a plurality of predetermined locations in the cross-correlation map to identify a set of cross-correlation values; detecting a multiview format of the multiview image by classifying a feature set of the multiview image comprising the set of cross-correlation values; the multi-view image is configured to be rendered on a multi-view display based on the multi-view format; A method wherein at least one of the predetermined locations comprises a corner region of the cross-correlation map.
2. determining an aspect ratio of the multi-view image, the set of features further comprising the aspect ratio. The computer-implemented method of multiview format detection of claim 1 .
3. The computer-implemented method of multi-view format detection of claim 1 , wherein at least one of the predetermined locations comprises a midpoint along one of a length and a width of the cross-correlation map.
4. The computer-implemented method of multi-view format detection of claim 1 , wherein each cross-correlation value of the set of cross-correlation values is determined based on a single cross-correlation value.
5. The computer-implemented method of multi-view format detection of claim 1 , wherein each cross-correlation value in the set of cross-correlation values is determined by averaging values within a corresponding predetermined location.
6. and training a classifier to classify the feature set according to training data, the training data including a training set of multi-view images and corresponding training labels. The computer-implemented method of multiview format detection of claim 1 .
7. 7. The computer-implemented method of multi-view format detection of claim 6, wherein the training labels include labels indicating a single tile format, a 1x2 tile format, a 2x1 tile format, and a 2x2 tile format.
8. the multi-view display is configured to provide wide-angle emitted light in a two-dimensional mode (2D mode) using a wide-angle backlight; the multi-view display is configured to provide directional emitted light in a multi-view mode using a multi-view backlight having an array of multi-beam elements, the directional emitted light comprising a plurality of directional light beams provided by each multi-beam element of the multi-beam element array; the multi-view display is configured to time-multiplex the 2D mode and the multi-view mode using a mode controller, sequentially activating the wide-angle backlight during a first consecutive time interval corresponding to the 2D mode and sequentially activating the multi-view backlight during a second consecutive time interval corresponding to the multi-view mode; The computer-implemented method of multi-view format detection of claim 1 , wherein directions of the directional light beams correspond to different view directions of the multi-view image.
9. The multi-view display is configured to guide the light in the light guide as guided light; 2. The computer-implemented method of multi-view format detection of claim 1, wherein the multi-view display is configured to scatter a portion of the guided light as directional radiation using multi-beam elements of an array of multi-beam elements, each multi-beam element of the array of multi-beam elements comprising one or more of a diffraction grating, a micro-refractive element, and a micro-reflective element.
10. 1. A multi-view display system, comprising: a multi-view display configured to display a multi-view image; a processor; a memory storing a plurality of instructions, the plurality of instructions, when executed, causing the processor to: accessing a multi-view image comprising a plurality of tile-view images; generating a cross-correlation map from the multiview image by autocorrelating the multiview image with shifted copies of the multiview image; sampling the cross-correlation map at a plurality of predetermined locations in the cross-correlation map to identify a set of cross-correlation values; inputting a feature set of the multi-view images into a classifier, the feature set including the set of cross-correlation values; deriving a label from the classifier that indicates the multi-view format of the multi-view image; rendering the multiview image on a multiview display based on the multiview format; A multi-view display system, wherein at least one of the predetermined locations comprises a corner region of the cross-correlation map.
11. The instructions, when executed, cause the processor to: The multi-view display system of claim 10 , further comprising determining an aspect ratio of the multi-view image, the set of features further comprising the aspect ratio.
12. The multiple view display system of claim 10 , wherein at least one of the predetermined locations comprises a midpoint along one of a length and a width of the cross-correlation map.
13. The multiple view display system of claim 10 , wherein each cross-correlation value in the set of cross-correlation values is determined based on a single cross-correlation value.
14. The multiview display system of claim 10 , wherein the classifier is trained using training data, the training data including a training set of multiview images and corresponding training labels.
15. The multi-view display system of claim 14 , wherein the training labels include labels indicating a single tile format, a 1×2 tile format, and a 2×1 tile format.
16. 1. A computer program comprising executable instructions that, when executed by a processor of a computer system, perform an operation of multiview format detection, the operation comprising: accessing a multi-view image comprising a plurality of tile-view images; generating a cross-correlation map from the multiview image by autocorrelating the multiview image with shifted copies of the multiview image; sampling the cross-correlation map at a plurality of predetermined locations in the cross-correlation map to identify a set of cross-correlation values; detecting a multiview format of the multiview image by classifying a feature set of the multiview image comprising the set of cross-correlation values; upon detecting the multi-view format, interlacing pixels of the tile-view images of the multi-view image; At least one of the predetermined locations comprises a corner region of the cross-correlation map.
17. The operation is The computer program product of claim 16 , further comprising determining an aspect ratio of the multi-view image, wherein the set of features further comprises the aspect ratio.
18. The computer program product of claim 16 , wherein the multi-view format is a 1×2 tile format, a 2×1 tile format, or a 2×2 tile format.
Citation Information
Patent Citations
Display format identification method and device of video file as well as video player
CN102231829A
Electronic equipment, display control method and program
JP2013104963A
Multiview display and method having shifted color sub-pixels
WO2020222770A1