Machine-implemented method, machine-readable medium, and data processing system

By guiding the asymmetric cropping of the playback device using the original aspect ratio and vector metadata of roughly squares, the problem of multiple cropping adaptation in the prior art is solved, and efficient and quality-reserved content adaptation on different aspect ratio devices is achieved.

CN118556404BActive Publication Date: 2025-08-08DOLBY LABORATORIES LICENSING CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202180041911.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-20
Filing Date
2021-06-09
Publication Date
2025-08-08
Estimated Expiration
2041-06-09

AI Technical Summary

Technical Problem

The prior art requires multiple cropping or filling to adapt to display devices with different aspect ratios during content creation and playback, resulting in reduced image quality and destruction of creative intentions.

Method used

The original aspect ratio and associated metadata of roughly square are used to guide the playback device to crop and fill content asymmetrically so that it is adapted to display devices with different aspect ratios, and the crop direction is controlled through vector metadata to maintain the area of interest as the focus.

Benefits of technology

It enables efficient adaptation of content on different aspect ratio devices, maintaining image quality and creative frameworks, providing a smooth visual experience, and reducing unnecessary cropping and filling steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118556404B_ABST
    Figure CN118556404B_ABST
Patent Text Reader

Abstract

The described embodiments include systems and methods for producing and adapting images (such as video images) for presentation on display devices having various different aspect ratios (such as 4:3, 16:9, 9:16, etc.). In one embodiment, a method for producing content (such as video images) can begin by selecting an original aspect ratio and determining the position of an object in at least a first scene in the content. In one embodiment, the original aspect ratio can be approximately square (e.g., 1:1). Metadata can then be created based on the position of the object in the first scene to direct a playback device to asymmetrically crop the content relative to the position for display on a display device having an aspect ratio different from the original aspect ratio. Other methods and systems are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 068,201, filed on August 20, 2020; U.S. Provisional Patent Application No. 62 / 705,115, filed on June 11, 2020; and European Patent Application No. 20179451.8, filed on June 11, 2020, each of which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure relates generally to video processing. Background Art

[0004] Content creation (such as the creation of a film, television program, or animation) and content capture (such as the recording of a live sporting event or news event) require selecting an aspect ratio for the so-called image canvas. This selection can be deliberate (e.g., the content creator considers several possible options for the image canvas aspect ratio and selects one) or accidental (e.g., a photographer in a recording studio hastily grabs a specific camera with a predetermined aspect ratio without considering its capture aspect ratio). Once the image canvas aspect ratio is selected, the content is created or captured and then distributed for playback on devices, which may have many different aspect ratios. In many cases, the content creator will capture or create content with an image canvas having a first aspect ratio, and then crop or pad the content for a desired playback aspect ratio that differs from the first aspect ratio. The desired playback aspect ratio may be the aspect ratio that the content creator believes will be the most common aspect ratio used on playback devices. The content is then published and distributed to playback devices, which have many different aspect ratios that differ from the original canvas aspect ratio and the desired playback aspect ratio. All such playback devices will then be required to adapt the display of the content to match the display device coupled to the playback device by cropping or padding the content. In this example, the content is cropped and / or padded at least twice. This process of padding and cropping at least twice may result in unnecessary cropping or padding in the image, thereby hindering the preservation of the content creator's intent in the process of adapting the content multiple times to different aspect ratios. Summary of the Invention

[0005] Aspects and embodiments described in this disclosure may provide systems and methods that may use a generally square aspect ratio for an original image canvas and associated metadata that allows a variety of endpoint aspect ratios to be derived from original content using the original image canvas or a set of original image canvases.

[0006] In one embodiment, a method for producing content, such as a video image, may begin by selecting an original aspect ratio for an image canvas and determining the position of an object within at least a first scene within the content of the image canvas. In one embodiment, the original aspect ratio may be substantially square (e.g., 1:1). The object may be a region of interest in the content, such as an actor or other focal point of the scene. Metadata may then be created based on the position of the object in the first scene to direct a playback device to asymmetrically crop the content relative to the position for display on a display device having an aspect ratio different from the original aspect ratio. The metadata may direct the playback device to asymmetrically expand the view around the object based on the position of the object on the canvas and the aspect ratio of the playback device's display. In one embodiment, the metadata may also direct the asymmetrical expansion based on other factors, such as a desire to avoid partial inclusion of certain image elements, such as a human face, and may provide data for preventing partial inclusion of such elements (which may mean that such elements are completely excluded or completely included in the cropped view). Such image elements may be added to the region of interest to ensure that they are fully included in the view or completely excluded from the view. For example, such image elements can be added by defining the size of the region of interest for including such image elements. Content and metadata can be stored and then distributed to playback devices with different aspect ratios. The playback device can use the metadata to adapt the content (by cropping and padding as necessary) to the display used by the playback device. In one embodiment, metadata can be created on a scene-by-scene basis, and a scene can be as short as a single frame of video content, so that the metadata can be frame-by-frame, where a frame is an image presented during a single refresh interval on a display device. Therefore, in one embodiment, the metadata described herein can be created frame-by-frame to capture frame-by-frame changes over time. Moreover, the original aspect ratio may not be static during the content, and thus the original aspect ratio may change during the content, or even change on a scene-by-scene (or even frame-by-frame) basis for at least a portion of the content; changes in the original aspect ratio can be referred to as a variable aspect ratio that changes during the content.

[0007] In one embodiment, the native aspect ratio can be selected to be approximately square (e.g., 1:1) or closer to square than an aspect ratio of 16:9, such that the ratio of the length to the height of the native aspect ratio is less than the ratio of 16:9 (16 / 9=1.7778) and greater than or equal to 1:1. A substantially square native aspect ratio can ensure the widest range of choices for adapting content to most aspect ratios in the field of playback devices. In an alternative embodiment, an native aspect ratio can be selected to prioritize image quality in a vertical playback orientation (e.g., a portrait orientation), in which case the native aspect ratio can be approximately square (1:1), ranging from 1:1 to less than 1:1 down to 9:16. The canvas area varies for each frame or each scene or shot by shot to provide great flexibility in content creation. As described herein, the native aspect ratio can change over time for the content, so the content can include variable aspect ratios.

[0008] The metadata can be a vector that specifies a direction relative to an object (e.g., a region of interest) in the current scene. The playback device can use the metadata to construct cropping and / or padding of the image to render the scene based on the metadata. In effect, the metadata can direct the playback device to crop or expand the entire canvas in a direction away from the object while keeping the object as the focus of the cropped scene. The playback device can also use tone mapping and color volume mapping that are customized for the adapted aspect ratio so that the tone mapping and color volume mapping are based on the actual content (e.g., the region of interest) at the adapted aspect ratio rather than all the content in the original image canvas of the scene.

[0009] In one embodiment, the method of determining objects and object locations can be implemented for multiple different scenes across the same or different objects. In one embodiment, the method can be implemented for each scene or each frame or each group of frames or at least a subset of scenes in the content, such that for at least a subset of scenes, each scene or frame can be adapted to a different aspect ratio upon playback based on metadata created during content creation. In one embodiment, a scene (e.g., a group of one or more frames) can be a camera shot or take in a film production or other content production, and different scenes can have different objects, backgrounds, camera angles, etc.

[0010] During the content creation process, one or more previews displaying the content at different aspect ratios can be generated based on the generated metadata. After viewing the previews, the content creator can edit the metadata directly or by revising the position of an object, selecting a different object, etc., and then display the one or more previews to see whether the revisions have improved the appearance of the content in the different previews. In one embodiment, the one or more previews can be one or more rectangular overlays on the original canvas, with the content of the scene displayed in the overlays.

[0011] In one embodiment, a user interface on a playback device may allow a user to switch playback operation between cropping (e.g., cropping performed based on the metadata described herein) and filling (which reverts to the usual practice of filling pixels (typically black pixels) around a roughly square canvas to fill the entire aspect ratio of the display). In one embodiment, such switching may allow the user to see more of the original canvas than the focused view that could be provided by using cropping based on the metadata described herein. In one embodiment, the image may be further zoomed beyond what is necessary to match the aspect ratio of the playback device to provide an enhanced (closer) view of the object area. In one embodiment, the transition between the cropped, fill, and / or zoomed viewing states may be displayed as a smooth transition to provide the user / viewer of the playback device with the impression of a smooth or seamless transition between viewing states.

[0012] In one embodiment, the final composite image to be shown to the user may correspond to a superposition of multiple inputs in multiple windows or areas of the screen. For example, the primary input may be shown in a large window on the display, while the secondary input may be shown in a smaller window that is smaller or much smaller than the large window (similar to the picture-in-picture feature on a television). For example, the primary input window may display a view generated based on metadata to provide a cropped view of the original or full canvas based on the position of the object and the aspect ratio of the primary input window or display device, and the secondary window may show the full or original canvas without any cropping or padding. In one embodiment, one or both of these windows may be fully or partially zoomed. Furthermore, the methods and systems described herein may be used to optimize the playback of content for any size and aspect ratio of each window. If one of the windows is resized, these methods and systems may be used to use the metadata and the position of the object within the scene to adapt the output to the resized aspect ratio (of the resized window). Furthermore, the methods and systems described herein may be used with a single window (without secondary windows) such that content is cropped within the window based on the window's aspect ratio, and these methods and systems may optimize playback using metadata as described herein.

[0013] In another embodiment, the methods and systems described herein may be used to create interesting effects by focusing on an object of interest while applying slight zooms and pans while also optimizing tone mapping for selected areas to enhance transitions between playbacks of photos in a photo stream. This may be guided by metadata that may be thought of as an "expected motion path." In one embodiment, the "expected motion path" is used instead of tracking the viewer position in order to provide a "guided Ken Burns effect." The metadata is a series of position vector coordinates (X, Y, Z) relative to the screen that describe the expected motion path of the viewer over a specified time period. As used herein, the term "Ken Burns effect" refers to a type of pan and zoom effect used when showing still pictures in film and video production.

[0014] In one embodiment, the original or complete canvas may already contain some padding to fit the image to the aspect ratio of the canvas. In this case, an embodiment may use additional metadata to indicate the location of the active area within the canvas, and if present, the client device or playback device may use this additional metadata to adapt playback based only on the active area of the canvas (excluding the padding area).

[0015] The various aspects and embodiments described herein may include a non-transitory machine-readable medium that may store executable computer program instructions that, when executed, cause one or more data processing systems to perform the methods described herein when the computer program instructions are executed. The instructions may be stored in a non-transitory machine-readable medium, such as a non-volatile memory, such as flash memory, or volatile dynamic random access memory (DRAM) or other forms of memory.

[0016] The above summary does not include an exhaustive list of all embodiments and aspects of the present disclosure. All systems, media, and methods can be practiced according to all suitable combinations of the various aspects and embodiments summarized above and those disclosed in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention is illustrated by way of example and not limitation in the figures of the accompanying drawings in which like references indicate similar elements.

[0018] Figure 1A 、 Figure 1B and Figure 1C Examples of different aspect ratios of display devices that can be used with one or more embodiments described herein are shown.

[0019] Figure 2Ais a flow chart illustrating a method, according to one embodiment, that may be used to create content and metadata for adapting output to different aspect ratios.

[0020] Figure 2B is a flow chart illustrating a method, according to one embodiment, that may be used to adapt playback at a playback device based on the aspect ratio of the playback device and based on metadata associated with the content.

[0021] Figure 3A 、 Figure 3B 、 Figure 3C and Figure 3D Examples of object locations and associated metadata based on those locations are shown.

[0022] Figure 4A 、 Figure 4B 、 Figure 4C 、 Figure 4D 、 Figure 4E 、 Figure 4F 、 Figure 4G 、 Figure 4H 、 Figure 4I 、 Figure 4J 、 Figure 4K and Figure 4L An example is shown of how a playback device can use metadata and the aspect ratio of the playback device to asymmetrically crop an image on the original canvas based on the metadata and the aspect ratio.

[0023] Figure 5 is a flowchart illustrating a method for creating content according to one embodiment.

[0024] Figure 6A 、 Figure 6B 、 Figure 6C 、 Figure 6D 、 Figure 6E 、 Figure 6F and Figure 6G Depicted are examples of how a playback device may display an image based on image metadata and the relative position of the display to the viewer.

[0025] Figure 7 An example of displaying an image according to an embodiment of an image adaptation process is shown.

[0026] Figure 8 An example of a data processing system that can be used to create content and metadata as described herein is shown. The figure also shows an example of a data processing system that can be a playback device that uses the metadata to adapt playback based on the metadata and the aspect ratio of the playback device. DETAILED DESCRIPTION

[0027] Various embodiments and aspects will be described with reference to the details discussed below, and the accompanying drawings will illustrate the various embodiments. The following description and drawings are illustrative and should not be construed as limiting. Numerous specific details are described to provide a thorough understanding of the various embodiments. However, in some cases, well-known or conventional details have not been described in order to provide a concise discussion of the embodiments.

[0028] Throughout this specification, reference to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment. The various appearances of the phrase "in one embodiment" throughout this specification do not necessarily refer to the same embodiment. The processes depicted in the following figures are performed by processing logic that may include hardware (e.g., circuitry, dedicated logic, etc.), software, or a combination of the two. Although the processes are described below in terms of some sequential operations, it should be understood that some of the described operations may be performed in a different order. Furthermore, some operations may be performed in parallel rather than sequentially.

[0029] This specification includes material that is subject to copyright protection, such as computer program software. Copyright owners, including the assignee of the present invention, hereby reserve all rights, including copyright, in these materials. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the Patent and Trademark Office file or records, but otherwise reserves all copyright rights whatsoever. Copyright Dolby Laboratories, Inc.

[0030] Embodiments described herein can create metadata and use it to adapt content from an original canvas or a full canvas for output on different display devices with different aspect ratios. These display devices can be conventional LCD or LED displays that are part of a playback device (such as a tablet or smartphone or laptop or television), or they can be conventional displays that are separate from the playback device but coupled to the playback device to drive the display by providing output to the display. Figure 1A 、 Figure 1B and Figure 1C Three examples of 3 different aspect ratios are shown. In particular, Figure 1A An example of a display with an aspect ratio of 4:3 (the ratio of the length to the height of the viewable area of the display panel) is shown; thus, if the length of the display area on the display is 8 inches, the height of the display area on the display will be 6 inches if the display has an aspect ratio of 4:3. Most CRT televisions have this aspect ratio. Figure 1BAn example of a display panel is shown having a display area with an aspect ratio of 16:9; this aspect ratio is commonly used by display panels for laptop computers and televisions. Figure 1C An example of a display panel or image canvas having an aspect ratio of 1:1 is shown, so the image canvas is square (the length and height of the viewable area are the same). As will be described below, one embodiment uses an original image canvas having an aspect ratio of 1:1 or approximately a square image canvas during the content creation process. Reference will now be made to Figure 2A Provides an example of the content creation process.

[0031] Now refer to Figure 2A A method according to an embodiment may begin at operation 51, where an original aspect ratio of an image canvas is selected. Once selected, content is created using the image canvas. Content creation may involve creating images (such as computer-generated graphics or animations) or capturing content using a camera (e.g., a movie camera used with live actors) or other techniques known in the art for creating content, or a combination of such techniques. The same or different image canvases may be used during the content creation process; the canvas area may vary for each frame, each set of frames, or each scene to provide optimal flexibility in content creation. In one embodiment, the original image canvas may be a square canvas (having an aspect ratio of 1:1) or a substantially square canvas. In one embodiment, an image canvas is substantially square if it is closer to a square than a canvas having an aspect ratio of 16:9, such that the ratio of the length to the height of the image canvas is less than the ratio of 16:9 (approximately 1.778) and greater than or equal to 1:1. A substantially square original aspect ratio ensures the widest range of options for adapting content to most aspect ratios across a wide range of playback devices. Other embodiments may use image canvases that are not substantially square, but this may affect how well the scene adapts to different displays with different aspect ratios.

[0032] exist Figure 2AIn operation 53, the positions of objects within scenes within the content are determined. Operation 53 may be performed during content creation, after content creation (during the editing process of the created content), or both. Operation 53 may begin by identifying or determining objects or regions of interest within a particular scene. This identification or identification of objects or regions of interest may be performed on a scene-by-scene basis for all scenes within the content or for at least a subset of all scenes within the content. For example, a first scene may have a first identified object, and a second scene may have a second identified object that is different from the first identified object. Furthermore, different scenes may include the same object that has been identified in different scenes, but the object may be located differently in the different scenes. In one embodiment, the identification or determination of objects or regions of interest may be performed manually by the content creator or automatically by a data processing system. For example, the data processing system may use well-known face detection algorithms, image saliency analysis algorithms, or other well-known algorithms to automatically detect one or more faces within a scene. In one embodiment, the automatic detection of objects may be manually overridden by the content creator. In one embodiment, the content creator may instruct the data processing system used during the content creation process to automatically determine objects for a subset of scenes, while reserving another subset for manual identification by the content creator. Once the object has been identified, the position of the object in operation 53 may be determined manually or automatically based on the centroid of the object. For example, if the face of the subject has been identified as the object or region of interest, the centroid of the face may be used as the position of the object on the image canvas in operation 53. In one embodiment, if the position is automatically edited by the data processing system, the position may be manually edited by the content creator. In one embodiment, whether the position is determined manually or automatically, the content creation tool or editing tool may enable the user to select a center pixel (e.g., Sx, Sy) for the object, and optionally a width and height (e.g., Sw, Sh) for the object to use in adapting the image for playback at the aspect ratio of a particular display device. After operation 53 has been performed, processing may proceed to operation 55, as Figure 2A shown.

[0033] In operation 55, the data processing system can automatically determine metadata based on the position of the object in the specified scene. The metadata can specify how to adapt the playback on a display device with an aspect ratio different from the original aspect ratio of the original image canvas. For example, the metadata can specify how to expand the original aspect ratio from the position of the object in one or more directions within the original aspect ratio to the original aspect ratio, so as to crop the image of the original aspect ratio to adapt the image to the specific aspect ratio of the display device controlled by the playback device for playback. In one embodiment, the metadata can be represented as a vector specifying a direction away from the determined object. Figure 3A 、 Figure 3B 、 Figure 3Cand Figure 3D An example of such metadata is shown which may be in vector form. Figure 3A An example is shown where the object 105 is near the center of the scene 103 in the original image canvas 101 . Figure 3B An example of an object 111 being on the left side of the scene 109 in the original image canvas 101 is shown; Figure 3B As shown, object 111 is vertically centered along the left side of scene 109 . Figure 3C An example is shown where the object 117 is in the upper right corner of the original image canvas 101 in the scene 115 . Figure 3D An example of object 123 being in the lower right corner of scene 121 in original image canvas 101 is shown. Figure 3A In the example shown, the vector representing the metadata example is considered equal in all directions of the object and can therefore be considered to have a value of 0,0 in this case. Figure 3A In an example, the object is expanded by cropping the original image canvas based on the aspect ratio of the playback device and the orientation of the display device (e.g., landscape or portrait); Figures 4A to 4D Four examples are shown of how to adapt content for playback at different aspect ratios when the object is centered in the original image canvas (this will be described further below). Figure 3B In the example shown, vector 112 is an example of metadata that specifies the direction in which the original image canvas 101 is cropped so that the object 111 remains in focus regardless of the aspect ratio of the display device and the orientation of the display device. Figure 4E 、 Figure 4F 、 Figure 4G and Figure 4H Shows that when the object is Figure 3B Four examples of how to adapt content for playback at different aspect ratios when positioned as shown. Figure 3C In the example shown, vector 119 is an example of metadata that specifies the direction in which the original image canvas 101 is cropped so that the object 117 remains in focus regardless of the aspect ratio of the display device and the orientation of the display device. Figure 4I 、 Figure 4J 、 Figure 4K and Figure 4L Shows that when the object is Figure 3C Four examples of how content can be adapted for playback at different aspect ratios are shown. Figure 3D In the example shown, vector 125 may be metadata specifying a direction in which to crop original image canvas 101 so that object 123 remains in focus regardless of the aspect ratio of the display device and the orientation of the display device.

[0034] The vectors representing the metadata can guide the playback device on how to crop the original image canvas based on the metadata and the location of the object. In one embodiment, rather than cropping symmetrically around the object, vectors (such as vectors 112, 119, and 125) guide asymmetrical cropping relative to the object, as further explained below. Asymmetrical cropping in this manner can provide at least two advantages: (a) the aesthetic framing of the scene is better preserved; even after intermediate levels of zoom, the image with the object in the upper right corner (e.g., see Figure 3C ) still results in an image with an object in the upper right, which better preserves the creative intent of the framing than a symmetrical crop; and (b) an asymmetrical crop makes it possible to zoom in on the object without abruptly changing the zoom direction or zoom ratio. In one embodiment, a vector may include an x component (for the x-axis) and a y component (for the y-axis), and the vector may be represented by 2 values: Px and Py, where Px is the x component of the vector and Py is the y component of the vector. In one embodiment, Px may be defined as: Px = 2(0.5-Sx) and Py may be defined as Py = 2(0.5-Sy), where Sx and Sy are the center of the object relative to the original image canvas, with coordinates 0,0 being at the upper left corner of the original image canvas, and 1,1 being the coordinates of the lower right corner of the canvas, and 0.5,0.5 being the coordinates of the center of the original image canvas. In Figure 3B In the case of the example shown, Px=1 and Py=0, and the vector is considered to be positive in the horizontal direction. Figure 3C In the case of the example shown, Px = -1 and Py = 1, and the vector can be considered to be negative in the horizontal direction and positive in the vertical direction. Further details and examples of this metadata and how it is used during playback are provided below.

[0035] Return Reference Figure 2A After operation 55 is completed for a particular scene, processing may proceed to operation 57 where the content and metadata are stored for use during playback. Then, in operation 59, the data processing system used during content creation or during the editing process determines whether there is more content to process. For example, if there are further scenes that need to be processed according to Figure 2AIf the method shown in FIG. 5 is processed, the process returns to operation 51, in which a new original aspect ratio can be selected for the image canvas or the last used original aspect ratio of the image canvas can be continued for content creation and / or editing. In one embodiment, the decision in operation 59 can be a manual decision controlled by a human operator of the data processing system. If there is no further content to be processed, the saved content and metadata can be provided to one or more content distribution systems in operation 61. For example, the content and metadata can be provided to a cable network for distribution to a set-top box, or can be distributed to a content provider, such as a content provider that delivers content over the Internet. The content and metadata can be distributed using streaming media or by downloading an entire copy of the content and metadata.

[0036] Figure 2A The method shown in is a content creation method that can be performed in a movie studio or other facility where content is created. The method is typically performed separately from playback on a playback device, and Figure 2B An example of a method on a playback device is shown. However, in one embodiment, Figure 2A The method shown in can also be used with Figure 2B The methods shown in are performed together on the same device that both creates content and then displays the content on one or more display devices having an aspect ratio that is different from the aspect ratio of the original image canvas.

[0037] like Figure 2B The playback method shown may begin at operation 71. In operation 71, a playback device may receive content including image data and may also receive associated metadata. For example, metadata may be associated with a first scene, and the metadata may specify how to adapt playback to a display device having an aspect ratio different from the original aspect ratio used to create the scene, relative to the position of objects in the scene. In one embodiment, the metadata may be in the form of a vector, such as Figure 3B 、 Figure 3C and Figure 3D, vectors 112, 119, and 125 are shown in FIG. These vectors can be represented as Px, Py values that are provided along with the content of the scene associated with the metadata. Once the metadata is received in operation 71, the playback device can perform operation 73. In operation 73, the playback device adapts the content in the scene to the aspect ratio of a display device coupled to the playback device. For example, the display device can be a display panel of a television or a smartphone or a tablet computer, and the playback device adapts the content by cropping the content in the scene so that the content adapts to the aspect ratio of the display device. In one embodiment, this adaptation or cropping uses the metadata described herein to asymmetrically crop the content based on the position of the object and the metadata, and the metadata can be vectors as described herein. This adaptation or cropping can also include customizing tone mapping and color volume mapping based on the cropped content displayed within the aspect ratio of the display device (e.g., only containing the area of interest), rather than performing tone mapping and color volume mapping based on the entire image in the original image canvas.

[0038] Although detailed examples of implementations of operation 73 will be provided below, by reference to Figures 4A to 4L It is helpful to provide a general description of the adaptation process. Figure 4A 、 Figure 4B 、 Figure 4C and Figure 4D In the example shown, operation 73 crops the content symmetrically relative to object 105 in original image canvas 51 depending on the aspect ratio of the display device and the orientation of the display device. Figure 4A The aspect ratio 153 created by operation 73 , which has cropped the content in landscape mode symmetrically around the object 105 in the original image canvas 151 , is shown. Figure 4B An aspect ratio 155 resulting from the cropping in operation 73 , which crops the content in landscape mode to the aspect ratio 155 , is shown. Figure 4C An aspect ratio 157 resulting from the cropping in operation 73 , which crops the content in portrait mode to the aspect ratio 157 , is shown. Figure 4D The aspect ratio 159 resulting from the cropping in operation 73 is shown, which crops the content in portrait mode to the aspect ratio 159. Figure 4A 、 Figure 4B 、 Figure 4C and Figure 4D In each of these examples shown in , the vector metadata may be a vector value of Px=0, Py=0, at which the metadata causes the original image canvas 151 to be cropped symmetrically around the object 105 based on the aspect ratio of the display device used to display the content. Figure 4A 、 Figure 4B 、 Figure 4C and Figure 4DIn the example shown, the focus of the display output remains on the object 105 in all cases, regardless of the aspect ratio and display orientation (eg, portrait or landscape).

[0039] exist Figure 4E 、 Figure 4F 、 Figure 4G and Figure 4H In the example shown in , operation 73 crops the content in the original image canvas asymmetrically relative to object 111 based on vector 112 , which in this example specifies how to crop asymmetrically. Figure 4E An aspect ratio 161 resulting from the cropping in operation 73 , which crops the content in landscape mode to aspect ratio 161 based on vector 112 , is shown. Figure 4F An aspect ratio 163 resulting from the cropping in operation 73 , which crops the content in landscape mode to aspect ratio 163 based on vector 112 , is shown. Figure 4G An aspect ratio 165 resulting from the cropping in operation 73 , which crops the content in portrait mode to aspect ratio 165 based on vector 112 , is shown. Figure 4H The aspect ratio 167 resulting from the cropping in operation 73, which crops the content in portrait mode to the aspect ratio 167 based on vector 112, is shown. Figure 4E 、 Figure 4F 、 Figure 4G and Figure 4H As can be seen in the example shown in , cropping keeps the object 111 to the left of the cropped view within the original image canvas 151, whether in landscape or portrait mode and regardless of the aspect ratio of the display device.

[0040] exist Figure 4I 、 Figure 4J 、 Figure 4K and Figure 4L In the example shown in , operation 73 crops the content in the original image canvas asymmetrically relative to object 117 based on vector 119, which specifies how to crop asymmetrically in this example. Figure 4I An aspect ratio 171 resulting from the cropping in operation 73 , which crops the content in landscape mode to aspect ratio 171 based on vector 119 , is shown. Figure 4J An aspect ratio 173 resulting from the cropping in operation 73 , which crops the content in landscape mode to aspect ratio 173 based on vector 119 , is shown. Figure 4K An aspect ratio 175 resulting from the cropping in operation 73 , which crops the content in portrait mode to aspect ratio 175 based on vector 119 , is shown. Figure 4L The aspect ratio 177 resulting from the cropping in operation 73, which crops the content in portrait mode to the aspect ratio 177 based on vector 119, is shown. Figure 4I 、 Figure 4J 、 Figure 4K and Figure 4L As can be seen in the example shown, object 117 remains in the upper right corner of each cropped output regardless of the orientation of the display device and regardless of the aspect ratio of the display device.

[0041] Return Reference Figure 2B After operation 73, the method determines whether there is more content to be processed at operation 75. If there is more content to be processed, processing returns to operation 71 to continue receiving content and metadata and adapting the content to the display device as described herein. If there is no further content, processing proceeds to operation 77 and the method ends.

[0042] Figure 5 Another example of a method that can be performed during content creation or during editing of content that has already been created and stored and is ready for editing is shown. Figure 5 In operation 201, a content creator may select an original canvas aspect ratio, such as a 1:1 aspect ratio, and create content for the current scene. Then, in operation 203, the content creator or a data processing system may determine the position of a current object in the current scene. This determination of the current scene's position may be performed manually by the content creator or automatically by the data processing system (possibly employing manual adjustments or overwrites). Alternatively, an editing or content creation tool may allow the content creator or data processing system to set the size of the object and its center point, such as a point at coordinates Sx, Sy, as described herein. Then, in operation 205, the data processing system may calculate metadata based on the position that can be used to adapt the content to other aspect ratios. This metadata may also (optionally) include additional metadata, such as metadata describing padding around the image. The content creator may then cause previews of one or more other aspect ratios to be displayed by adapting the content in the current scene to such other aspect ratios. In other words, the content creator may cause the data processing system to generate a display of previews of the content such that the content creator may view each preview at different aspect ratios to determine whether fitting or cropping is desired or satisfactory. In one embodiment, the preview may be a rectangle overlaid on the image in the original image canvas, e.g. Figures 4A to 4L The rectangles shown in FIG show various aspect ratios (e.g., 161). These previews will show how the content will appear to the end user of the playback device. Cropping operations based on the use of metadata (e.g., Figure 2BThe position of the rectangle is obtained by cropping operation 73 in . Metadata provides the advantage of defined playback processing behavior to show an accurate preview of how the content will be presented on a variety of different playback devices with different display devices having different aspect ratios. Finding the playback behavior in this way allows the content creator to preview the exact end result and make any adjustments that may be desired or necessary. Therefore, in operation 209, the content creator can decide whether to adjust the adaptation for one or more aspect ratios. If the adaptation needs to be adjusted, the processing can return to operation 203, where the content creator can revise the position of the current object or select a different object using a content creation tool or editing tool to select a different position or a different object. If it is determined that no adjustment is necessary in operation 209, the content creator can move to the next scene and determine whether to process the next scene in operation 211. If all scenes that need to be processed have been processed, the processing can be completed and ended, as shown in FIG. Figure 5 On the other hand, if there are additional scenes to be processed, the process returns to operation 201, as shown Figure 5 shown.

[0043] The following subsections provide detailed examples of metadata and methods for using metadata to crop content for playback at different aspect ratios. In one embodiment, metadata may be specified in a compatible bitstream as follows: One or more rectangular regions define an object region:

[0044] The rectangle should be defined such that (Top) <= (1-Bottom) and (Left) <= (1-Right). The values of TopOffset, BottomOffset, LeftOffset, and RightOffset should be set accordingly.

[0045] Playback devices should enforce this behavior and tolerate non-compliant metadata.

[0046] In case the offset is 0, the entire image is indicated as the region of interest.

[0047] With a width and height of zero pixels, the upper left corner of the rectangle is indicated as the region of interest, corresponding to the center of the object.

[0048] Additional metadata, used as described below, may include the following:

[0049]

[0050] The coordinates may change per frame or per shot, or may be static for the entire content. Any changes between the image and the corresponding metadata can be fully frame synchronized.

[0051] If the canvas is resized before delivery, such as in an adaptive streaming environment, the offset coordinates are also updated accordingly.

[0052] The following subsections describe adapting content on a playback device and optimizing the color volume mapping of the adapted content on the playback device. This section assumes that the playback device performs all of these operations locally on the playback device, but in alternative embodiments, a centralized processing system may perform some of these operations for one or more playback devices coupled to the centralized processing system.

[0053] During playback, the playback device is responsible for adapting the canvas and associated metadata to the specific aspect ratio of the attached panel. This involves three operations, as described below. For example, in one embodiment:

[0054] 1: Calculate the region of interest and update the mapping curve:

[0055] The coordinates of the canvas region of interest (or the area to be displayed on the panel) are calculated by calculating the upper left and lower right pixels TLx, TLy, BRx, BRy and the width and height (CW, CH) of the canvas; for example, the method can perform the calculation based on the equation immediately below or the software implementation provided further below:

[0056] 1) TLx = (Sx - Px) * CW

[0057] 2) TLy=(Sy-Py)*CH

[0058] 3) BRx=(Sx+Px)*CW

[0059] 4) BRy=(Sy+Py)*CH

[0060] In addition to adaptively resizing the image based on the region of interest, the tone mapping algorithm can also be adjusted to achieve optimal tone mapping for the cropped region (rather than the entire original image in the original image canvas). This can be achieved by calculating additional metadata corresponding to the region of interest and using it to adjust the tone mapping curve, for example, as described in U.S. Patent No. 10,600,166 (which describes a display management process known in the art), where the tone mapping curve takes as an input a "smid" (mean luminance) parameter representing the average brightness of the source content. The adjustment using this new ROI luminance offset metadata (e.g., expressed as L12MidOffset) is calculated as follows:

[0061] SMid = (L1.Mid + L3MidOffset) / / Calculate the middle brightness of the entire frame

[0062] SMid'=SMid*(1-ZF)+(SMid+L12MidOffset)*ZF / / Adjust for ROI

[0063] Where ZF is the zoom fraction, so ZF=0 corresponds to full screen, and ZF=1 corresponds to fully zoomed into the object.

[0064] NOTE: L3MidOffset represents the offset of the L1.Mid value and may also be referred to as L3.Mid.

[0065] Another parameter adjusted in a similar manner is the optional global dimming algorithm used to optimize the mapping to the global dimming display. The global dimming algorithm takes two values, L4Mean and L4Power, as input. Before calculating the global dimming backlight, the L4Mean value is adjusted by the zoom factor as follows:

[0066] L4Mean'=L4Mean*(1-ZF)+(L4Mean+L12MidOffset)*ZF

[0067] 2: Crop and process the region of interest

[0068] To use memory efficiently and ensure consistent timing for playback devices, in a preferred embodiment, playback devices should follow these steps:

[0069] 1) Decode the bitstream encoded with Region of Interest (ROI) metadata (e.g., vectors described by Px, Py) and insert each frame into the decoded picture buffer.

[0070] 2) When it is time to display the current frame, only the portion of the image needed for the ROI is read back from memory, starting with the top left pixel (TLx,y), which is read some time t (the "latency") before it is presented to the panel. This latency t is determined by the time it takes the imaging pipeline to process the first pixel, and includes any spatial upsampling performed by the imaging pipeline.

[0071] 3) Once the entire region of interest has been read from memory, the decoded picture buffer can be overwritten with subsequent decoded pictures.

[0072] Once the cropped region of the image is read from memory, it is mapped to the dynamic range of the panel. This method may follow known techniques described in US Patent 10,600,166 using the adjusted mapping parameters from operation 1 above.

[0073] 3: Resize to output resolution

[0074] The final operation is to resize the image to the resolution of the panel. Obviously, the resolution or size of the final image may not match the resolution of the panel. Methods for resizing the image must be applied to achieve the desired resolution, which is well known in the art. Example methods can include bilinear or Lancsoz resampling or many methods including super-resolution or neural networks.

[0075] In an embodiment, without limitation, metadata used to signal the ROI and related parameters may be expressed as Level 12 (L12) metadata, as summarized below.

[0076] 1) Specify the coordinates of the ROI rectangle:

[0077] a. The rectangle is specified relative to the edge of the image, so that the default value of zero corresponds to the entire image

[0078] b. The offset is specified as a percentage of the image width and height with 16-bit precision. This method ensures that the metadata remains unchanged even when the image is resized.

[0079] c. If the offset results in a ROI with a width and / or height of zero pixels, the single pixel in the top left corner is considered to be the ROI.

[0080] 2) Average brightness of ROI

[0081] a. This value acts as an offset to the color volume metadata to optimize the rendering of the ROI. When the ROI is expanded to fill the screen, the color volume mapping will preserve more contrast in the ROI.

[0082] b. Calculated the same way as L1.Mid, but only using the pixels that make up the ROI. The value stored in the metadata is offset from the full screen value to ensure zero values are restored to use the L1.Mid value:

[0083] i.L12.MidOffset=ROI.Mid-L1.Mid-L3.Mid

[0084] Note: L1.Mid can be calculated as the average of the PQ-encoded maxRGB values of the image or as the average luminance. maxRGB is the maximum value of the pixel's color component values {R,G,B}. L3.Mid represents the offset to the 'Mid' PQ value present in the L1 metadata (L1.Mid).

[0085] c. The playback device smoothly interpolates this offset based on the relative size of the ROI being displayed. The value used by the display device can be generated as

[0086] L1.Mid+L3.Mid+f(L12.MidOffset), where f represents the interpolation function.

[0087] 3) Optionally, you can specify a mastering viewing distance

[0088] a. The optimal viewing distance is specified as a fraction of the reference viewing distance. This is used to ensure that the image does not scale when viewed from the same optimal viewing distance.

[0089] b. The default viewing distance is relative to a viewing angle of 17.7613 degrees (calculated from 2*atan(0.5 / 3.2), which is the ITU-R reference viewing angle for full HD content). A closer distance (e.g., 0.5) would correspond to a viewing angle of 17.7613 / 0.5 = 35.5226. Note that trigonometric functions have been omitted for simplicity and to ensure that different aspect ratios are calculated equally.

[0090] c. The range is 3 / 32 to 2, with increments of 1 / 128. The metadata is an 8-bit integer in the range 11 to 255, used to calculate the image height as follows:

[0091] Optimal viewing distance = (L12.MVD+1) / 128

[0092] d. If not specified or in the range of 0 to 10, the default value is 127, or the optimal viewing distance is equal to the reference viewing distance. For new content, this value may be smaller, such as a value of 63, indicating that the optimal distance is the same as the reference viewing distance.

[0093] 4) Optionally, you can specify the distance of the object (large part of the ROI) from the camera

[0094] a. This can be used to enhance the "look around" feature by allowing the image to pan and zoom at the correct speed in response to changes in the viewer's position. Objects that are farther away will pan and zoom at a slower rate than nearby objects.

[0095] 5) Optionally, you can specify an "expected viewer motion path"

[0096] a. This can be used to direct a "Ken Burns" effect during playback, even when viewer tracking is unavailable or not enabled. Examples would include picture frames. By directing the Ken Burns direction for panning and zooming, this feature allows the artist to specify the desired effect, whether zooming in or out on an object, without inadvertently cropping the primary subject out of the image.

[0097] 6) Optionally, you can specify separate layers for graphics or overlays

[0098] a. This allows scaling and compositing of graphics on top of an image independently of image scaling. Prevents important overlays or graphics from being cropped or scaled when the image is cropped or scaled. Preserves the "creative intent" of graphics and overlays

[0099] Preferably, only a single level 12 field is specified in the bitstream. In the event that multiple fields are specified, only the last one is considered valid. The metadata value can change from frame to frame, which is necessary for tracking ROIs within a video sequence. This field is extensible, allowing additional fields to be added for future versions.

[0100] The following provides software examples (eg, pseudo code) that can implement embodiments. Smart zoom: adapting content to playback devices

[0101] in:

[0102] [TDisplay_w TDisplay_h] = target display width and height

[0103] TargetAspectRatio=TDisplay_w / TDisplay_h

[0104] [Im_w Im_h] = source image width and height

[0105] [SwSh] = ROI width and height

[0106] [S TBLR ] = ROI top, bottom, left, and right pixels

[0107] [Rw Rh] = width and height of the ROI after reshaping

[0108] 1. Reshape the ROI to match the target display aspect ratio

[0109] ifSw>=Sh

[0110] Rw=Sw

[0111] Rh=Rw / TargetAspectRatio

[0112] if Sw <Sh

[0113] Rh=Sh

[0114] Rw=Sh*TargetAspectRatio

[0115] 2. Create an oversized source image canvas to match the target display aspect ratio

[0116] if Im_w>Im_h

[0117] if TDisplay_w>TDisplay_h

[0118] H=Im_h

[0119] W=H*TargetAspectRatio

[0120] if TDisplay_w <TDisplay_h

[0121] H=Im_w

[0122] W=W / TargetAspectRatio

[0123] ifIm_w <Im_h

[0124] if TDisplay_w>TDisplay_h

[0125] H=Im_h

[0126] W=H*TargetAspectRatio

[0127] If TDisplay_w <TDisplay_h

[0128] W=Im_w

[0129] H=W / TargetAspectRatio

[0130] overSizeSource=[HW]

[0131] padSize = [H-Im_h-W-Im_w]

[0132] 3. Transform the ROI to the size of the reshaped RO1

[0133] Im TBLR =[1 1 Im_h Im_w]

[0134] ROISizeDiff=[Rh-Sh-Rw-Sw]

[0135] % Calculate the percentage of pixels between the sides of the ROI and the source image border.

[0136] topRoom=(ST-Im T ) / Im_h

[0137] bottomRoom=(Im B -S B ) / Im_h

[0138] leftRoom=(SL-ImL) / Im_w

[0139] rightRoom=(ImR-SR) / Im_w

[0140] if topRoom <bottomRoom

[0141] % More available pixels below the ROI

[0142] topScalar=topRoom

[0143] bottomScalar=1-topScalar

[0144] otherwise

[0145] % More available pixels above the ROI

[0146] bottomScalar=bottomRoom

[0147] topScalar=1-bottomScalar

[0148] if leftRoom <rightRoom

[0149] % There are more available pixels on the left side of the ROI

[0150] leftScalar=leftRoom

[0151] rightScalar=1-leftScalar

[0152] otherwise

[0153] % There are more available pixels on the right side of the ROI

[0154] rightScalar=rightRoom

[0155] leftScalar=1-rightScalar

[0156] % Use ROISizeDiff and {tblr}Scalars to create the reshaped ROI coordinates

[0157] R T =ST-(ROISizeDiff*topScalar)

[0158] R B =S B +(ROISizeDiff*bottomScalar)

[0159] R L =S L -(ROISizeDiff*leftScalar)

[0160] R R =S R +(ROISizeDiff*rightScalar)

[0161] % Convert these source-relative coordinates to super-large source coordinates (specified as 'to-be-posted', or ')

[0162] R' T =RT+(padSize_h / 2)

[0163] R' B =R B +(padSize_h / 2)

[0164] R' L =R L +(padSize_w / 2)

[0165] R' R =R R +(padSize_w / 2)

[0166] % Get the coordinates of the enlarged ROI

[0167] Left=R' L -1

[0168] Right=oversizedSource_w-R' R

[0169] Top=R' T -1

[0170] Bottom=oversizedSource_h-R' B

[0171] RS TBLR =[Top Bottom Left Right]*zoomFactor

[0172] (For the calculation of zoomFactor, see the following subsection)

[0173] %Generate rescaled image

[0174] outputROI TL =R' TL -RS TL

[0175] outputROI BR =R' BR +RS BR

[0176] % Calculate zoomFactor

[0177] This section describes an example of how the zoom factor is calculated and used in the software examples provided above.

[0178] Default TDisplayDiagonal and ViewerDistance can be given as parameters in the configuration file (e.g. in inches)

[0179] % Calculate the display diagonal in pixels

[0180] Tdisplay_diagonalInPixels=hypotenuse(TDisplay_h,TDisplay_w)

[0181] % Calculate the number of pixels per inch

[0182] pixels_per_inch=TDisplay_diagonalInPixels / TDisplayDiagonal

[0183] % Calculate display height (inches)

[0184] TDisplay_heightInInches=TDisplay_h / pixels_per_inch

[0185] % Convert viewer distance from inches to image height

[0186] Viewer_distance_PH=viewerDistance / TDisplay_heightInInches

[0187] % zoom factor compares the optimal distance to the actual viewer distance

[0188] zoomFactor=Mastering_Distance_PH / Viewer_Distance_PH

[0189] Alternative embodiments of playback behavior are also provided in the Appendix.

[0190] Display adaptation based on the relative position of the observer and the display

[0191] When viewing a scene through a window, the scene's appearance varies depending on the observer's position relative to the window. For example, a person sees a larger portion of the outside scene when they are closer to the window than when they are farther away. Similarly, as the observer moves sideways, some parts of the image appear on one side of the window, while other parts are obscured on the other side.

[0192] If the window is replaced by a lens (zooming in or out), the outside scene will now appear larger (zooming in) or smaller (zooming out) compared to the real scene, but it will still provide the observer with the same experience as if they were moving relative to the window.

[0193] In contrast, when an observer views a digital image reproduced on a traditional display, the image does not change based on the observer's relative position to the display. In embodiments, this difference between the experience of viewing through a window and viewing a traditional display is addressed by adapting the image on the display based on the observer's relative position to the display, allowing them to view it through a window as if they were viewing a rendered scene. Such embodiments allow content creators (e.g., photographers, mobile users, or cinematographers) to better convey or share with their audiences the experience of being in a real scene.

[0194] In an embodiment, an example process of adapting image display according to the relative position of the observer and the display may include the following steps:

[0195] Obtain an image using a capture device (such as a camera), or load it from disk

[0196] Specify regions of interest (ROI) on captured images

[0197] Transmit image and ROI metadata to receiving device

[0198] On the receiving device, determine the viewer's position relative to the display

[0199] Displays images on the display based on ROI metadata, the screen’s aspect ratio, and the viewer’s position relative to the screen

[0200] Each of these steps is described in more detail below.

[0201] As non-limiting examples, the image may be obtained using a camera, or by loading it from disk or memory, or by capturing it from decoded video. The process may be applied to a single picture or frame or a sequence of pictures or frames.

[0202] A region of interest is an area within an image and generally corresponds to the most important portion of the image that should be retained in various display and viewing configurations. A region of interest, such as a rectangular region of an image, can be defined manually or interactively, for example, by allowing a user to draw a rectangle on the image using their finger, a mouse, a pointer, or some other user interface. In some embodiments, an ROI can be automatically generated by identifying a specific object in the image (e.g., a face, a car, a license plate, etc.). The ROI can also be automatically tracked across multiple frames of a video sequence.

[0203] There are many ways to estimate the distance and relative position of the viewer with respect to the screen. The following methods are provided by way of example only and not limitation. In an embodiment, the viewer position can be established using an imaging device near or integrated into the bezel of the display, such as an internal camera or an external webcam. The image from the camera can be analyzed to locate the head of a person in the image. This is done using traditional image processing techniques commonly used for "face detection", camera autofocus, autoexposure, or image annotation, etc. There is sufficient literature and technology to implement face detection for those skilled in the art to isolate the position of the viewer's head in the image. The return value of the face detection process is a rectangular bounding box of the viewer's head, or a single point corresponding to the center of the bounding box. In an embodiment, the viewer's position can be further refined by any of the following techniques:

[0204] a) Temporal filtering. This type of filtering can reduce measurement noise in the estimated head position, providing a smoother and more continuous experience. IIR filters can reduce noise, but the filtered position lags behind the actual position. Kalman filtering aims to reduce noise and predict the actual position based on a certain number of previous measurements. Both techniques are well known in the art.

[0205] b) Eye Position Tracking. Once the head position is identified, the estimated viewer position can be further refined by finding the positions of the viewer's eyes. This can involve further image processing, and the head tracking step can be skipped entirely. The viewer's position can then be updated to indicate the position directly between the eyes, or alternatively, to indicate the position of a single eye.

[0206] c) Faster update of measurement results: Faster (more frequent) measurement results are desired to obtain the most accurate current position of the viewer.

[0207] d) Depth Cameras. To improve the estimate of the distance between the observer and the camera, specialized cameras that directly measure distance can be employed. Examples include time-of-flight, stereo cameras, or structured light. Each of these is known in the art and is commonly used to estimate the distance from objects in the scene to the camera.

[0208] e) Infrared cameras. To improve performance in various ambient lighting conditions (i.e., dark rooms), infrared cameras can be used. These can directly measure the heat of a person's face or measure the IR light reflected from an IR emitter. Such devices are often used in security applications.

[0209] f) Distance calibration. The distance between the viewer and the camera can be estimated using image processing algorithms. This can then be calibrated to the distance from the screen to the viewer using the known displacement between the camera and the screen. This ensures that the displayed image is correct for the estimated viewer position.

[0210] g) Gyroscopes. These are widely available in mobile devices and can easily provide information about the orientation of the display (e.g., portrait mode vs. landscape mode) or the relative movement of the handheld display compared to the viewer.

[0211] It has been described previously herein how, given ROI metadata and the characteristics of the screen (aspect ratio), the rendered image can be adapted according to the region of interest and the assumed position of the viewer. In embodiments, if the assumed position of the viewer is replaced by its estimated position, as calculated by any of the techniques previously discussed, one or more of the following techniques may be used to adjust the display rendering, for example: Figures 6A to 6G Described in .

[0212] As an example, Figure 6A Depicted is an image 610 as originally captured and displayed on a display 605 (e.g., a mobile phone or tablet computer in portrait mode), without regard to ROI metadata. As an example, a rectangle 615 ("Hi") may represent a region of interest (e.g., a picture frame, a banner, a person's face, etc.).

[0213] Figure 6B An example of rendering an image 610 by considering a reference viewing position (eg, 3.2 picture heights, centered horizontally and vertically on the screen) is depicted, with appropriate scaling to magnify the ROI (615) while maintaining the original aspect ratio.

[0214] In an embodiment, Figure 6C As depicted, as the viewer moves away from the screen, or the display moves further away from the viewer, the image is further magnified, exhibiting the same effect as the viewer moving away from the window, thereby limiting his view of the outside scene. Similarly, as Figure 6D As depicted, as the viewer moves closer to the screen, or the display moves closer to the viewer, the image zooms out further, exhibiting the same effect as the viewer moving closer to a window, thereby magnifying his view of the outside scene.

[0215] In an embodiment, Figure 6EAs depicted, as the viewer moves to the right of the display, or the display moves to the viewer's left, the image shifts to the right, exhibiting the same effect as the viewer looking out a window to the left. Figure 6F As depicted, as the viewer moves to the left of the display, or the display moves to the right of the viewer, the image shifts to the left, exhibiting the same effect as a viewer looking out a window to the right.

[0216] Similar adjustments can be made when the viewer (or display) moves up or down, or a combination of these. Typically, the amount of image movement is based on the assumed or estimated depth of the scene in the image. For very shallow depths, the movement is less than the viewer's actual movement, while for very deep depths, the movement can be equal to the viewer's movement.

[0217] In an embodiment, all of the above operations can be adjusted according to the aspect ratio of the display. For example, in landscape mode, Figure 6G As depicted, the original image (610) is scaled and cropped so that the ROI 615 is centered in the viewer's field of view. As previously described, the displayed image (610-ROI-B) can then be further adjusted based on the viewer's relative position with respect to the screen.

[0218] In an embodiment, as the ROI approaches the edge of the image, it may move by smaller and smaller amounts to prevent it from abruptly reaching the edge and not moving any further. Thus, from near a reference position (e.g., 610-ROI-A), the image may be adjusted in a natural manner, like looking through a window, but the rate of movement may decrease as the edge of the captured image is approached. It is desirable to prevent an abrupt boundary between natural movement and no movement, and instead, it is preferable to smoothly scale the fraction of movement as the observer moves toward the maximum allowed amount.

[0219] Optionally, the image can be slowly re-centered to the viewer's actual viewing position over time, potentially allowing a greater range of movement and motion from the actual viewing position. For example, if the viewer starts viewing from a reference position and then moves towards the lower left corner of the screen, the image can be adjusted to pan up and to the right. From this new viewing position, the viewer will not be allowed to move further in the lower left direction. With this optional feature, the view can return to the center position over time, restoring the viewer's range of motion in all directions. Optionally, the amount to shift and / or scale the image based on the viewer's position can be determined in part by additional distance metadata, which describes the distance (depth) of the primary objects making up the ROI to the viewer. To simulate the experience of viewing through a window, the image should be adapted less for close distances than for farther distances.

[0220] In another embodiment, as previously described, an overlay image can optionally be composited with the adjusted image, where the overlay image's position remains static. This ensures that important information in the overlay image remains visible at all times and from all viewing positions. Furthermore, it enhances the immersiveness and realism of the experience, much like a semi-transparent overlay printed on a window.

[0221] In another embodiment, the color volume mapping can optionally be adjusted based on the actual area of the image being displayed, as previously described. For example, if the viewer moves to the right to better see bright objects in the scene, the metadata describing the dynamic range of the image can be adjusted to display a brighter image, and thus the tone mapping may cause the rendered image to be mapped slightly darker, thereby mimicking the adaptation effect experienced by a human observer looking at a scene through a window.

[0222] Referring to the "smart zoom" pseudo code described above (for a fixed distance between the observer and the screen), in an embodiment, the following changes need to be made to allow for observer position-adaptive smart zoom:

[0223] a) Instead of using an assumed reference viewing distance, the actual distance from the viewer to the screen (measured by any known technique) is used to calculate the previously described "viewerDistance" and "zoomFactor" parameters and generate the zoomed image.

[0224] b) Shift the scaled image at coordinates (x, y) based on the viewer's position on the screen. By way of example and not limitation, the viewer's position can be calculated with reference to their eye coordinates (x, y). In pseudocode, this can be expressed as:

[0225] %Generally query the eye position or face position

[0226] eye_left_xy0 / / Initial left eye position (origin)

[0227] eye_right_xy0 / / Initial right eye position (origin)

[0228] eye_left_xy i / / New left eye position

[0229] eye_right_xy i / / New right eye position

[0230] % Calculate the difference in eye position

[0231] %Im_d indicates how much to shift the generated image.

[0232] % It is contained in [0, 1]. Im_d = 0 results in

[0233] % No shift (function off)

[0234] delta_left_xy=-(eye_left_xy i +eye_left_xy0)*Im_d

[0235] delta_right_xy=-(eye_right_xyi+eye_right_xy0)*Im_d

[0236] %Average the differences

[0237] offset_x=(delta_left_x+delta_right_x) / 2

[0238] offset_y=(delta_left_y+delta_right_y) / 2

[0239] % Apply shift to scaled / zoomed image (e.g. image calculated using zoomFactor)

[0240] outputROI L =outputROI L -offset_x-(max(0, outputROI R -offset_x-Im_w));

[0241] outputROI R =outputROI R -offset_x+(max(0, -outputROI L ));

[0242] outputROI T =outputROIT-offset_y-(max(0,outputROI B -offset_y-Im_h));

[0243] outputROI B =outputROI B -offset_y+(max(0, -outputROI T ));

[0244] Figure 7An example process flow for displaying an image using a display adaptation process according to an embodiment is depicted. In step 705, the device receives an input image and parameters related to a region of interest within the image. If image adaptation is not enabled (e.g., "smart zoom" as described herein), then in step 715, the device generates an output image without considering the ROI metadata. However, if image adaptation is enabled, then in step 710, the device may use the ROI metadata and display parameters (e.g., aspect ratio) to generate an output version of the input image to emphasize the ROI of the input image. In addition, in some embodiments, the device may further adjust the output image based on the relative position and distance of the viewer from the display. The output image is displayed in step 720.

[0245] Figure 8 An example of a data processing system 800 that can be used with one embodiment is shown. For example, system 800 can be implemented to provide execution Figure 2A or Figure 5 Content creation or editing system or method performed Figure 2B Note that although Figure 8 The various components of the device are illustrated, but are not intended to represent any particular architecture or manner of interconnecting the components, as such details are not germane to the present disclosure. It will also be understood that network computers and other data processing systems or other consumer electronic devices having fewer or possibly more components may also be used with embodiments of the present disclosure.

[0246] like Figure 8 As shown, device 800, a form of data processing system, includes a bus 803 coupled to microprocessor(s) 805, ROM (read-only memory) 807, volatile RAM 809, and non-volatile memory 811. Microprocessor(s) 805 can retrieve instructions from memory 807, 809, and 811 and execute them to perform the operations described above. Microprocessor(s) 805 may include one or more processing cores. Bus 803 interconnects these various components and also interconnects components 805, 807, 809, and 811 to a display controller and display device 813, as well as peripheral devices such as input / output (I / O) devices 815 (which may be a touch screen, mouse, keyboard, modem, network interface, printer, and other devices well known in the art). Typically, I / O devices 815 are coupled to the system via I / O controller 810. Volatile RAM (random access memory) 809 is typically implemented as dynamic RAM (DRAM), which requires a continuous power supply to refresh or maintain data in memory.

[0247] The non-volatile memory 811 is typically a magnetic hard drive or a magnetic optical drive or an optical drive or a DVD RAM or flash memory or other type of memory system that can retain data (e.g., large amounts of data) even after the system is powered off. Typically, the non-volatile memory 811 will also be random access memory, but this is not required. Although Figure 8 The non-volatile memory 811 is shown as a local device directly coupled to the rest of the components in the data processing system, but it should be understood that embodiments of the present disclosure may utilize non-volatile memory that is remote from the system (e.g., a network storage device coupled to the data processing system via a network interface such as a modem, an Ethernet interface, or a wireless network). As is well known in the art, the bus 803 may include one or more buses connected to each other through various bridges, controllers, and / or adapters.

[0248] The part described above can be implemented with logic circuits such as dedicated logic circuits or with processing cores of other forms of microcontrollers or execution program code instructions. Therefore, the process of teaching discussed above can be performed with program codes such as machine executable instructions, and these machine executable instructions make the machine executing these instructions perform certain functions. In this context, "machine" can be a machine (for example, an abstract execution environment, such as a "virtual machine" (for example, a Java virtual machine), an interpreter, a common language runtime, a high-level language virtual machine, etc.) that converts intermediate form (or "abstract") instructions into processor-specific instructions and / or an electronic circuit (for example, a "logic circuit" implemented with a transistor) arranged on a semiconductor chip (designed to execute instructions, such as a general-purpose processor and / or a special-purpose processor). The process of teaching discussed above can also be performed by an electronic circuit (as a substitute for a machine or in combination with a machine) that is designed to execute a process (or a part thereof) without executing a program code.

[0249] The present disclosure also relates to an apparatus for performing the operations described herein. The apparatus may be specially constructed for the desired purpose, or the apparatus may comprise a general-purpose device that is selectively activated or reconfigured by a computer program stored in the apparatus. Such a computer program may be stored in a non-transitory computer-readable storage medium, such as, but not limited to, any type of disk, including a floppy disk, an optical disk, a CD-ROM, and a magneto-optical disk, a DRAM (volatile), a flash memory, a read-only memory (ROM), a RAM, an EPROM, an EEPROM, a magnetic card or an optical card, or any type of medium suitable for storing electronic instructions, and each coupled to a device bus.

[0250] A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). For example, non-transitory machine-readable media include read-only memory ("ROM"); random access memory ("RAM"); magnetic disk storage media; optical storage media; flash memory devices; and the like.

[0251] Articles of manufacture can be used to store program code. Articles of manufacture storing program code can be implemented as, but not limited to, one or more non-transitory memories (e.g., one or more flash memories, random access memories (static, dynamic or other)), optical disks, CD-ROMs, DVD ROMs, EPROMs, EEPROMs, magnetic or optical cards, or other types of machine-readable media suitable for storing electronic instructions. Program code can also be downloaded from a remote computer (e.g., a server) to a requesting computer (e.g., a client) via a data signal implemented in a propagation medium (e.g., via a communication link (e.g., a network connection)) and then stored in a non-transitory memory (e.g., DRAM or flash memory or both) of the client computer.

[0252] The foregoing detailed description has been presented in terms of algorithms and symbolic representations of operations on data bits within a device memory. These algorithmic descriptions and representations are the tools used by those skilled in the art of data processing to most effectively convey the substance of their work to others skilled in the art. An algorithm is generally considered to be a consistent sequence of operations that produces a desired result. The operations are those requiring physical manipulation of physical quantities. Typically, but not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, primarily for common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0253] It should be borne in mind, however, that all of these terms and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise expressly stated from the foregoing discussion, it should be understood that throughout this specification, discussions utilizing terms such as "receiving," "determining," "sending," "terminating," "waiting," "changing," and the like refer to the actions and processes of manipulating data represented as physical (electronic) quantities within registers and memories of a device, or a similar electronic computing device, and transforming that data into other data similarly represented as physical quantities within a device memory or register or other such information storage device, transmission device, or display device.

[0254] The processes and displays presented herein are not inherently associated with any particular device or other apparatus. Based on the teachings herein, various general-purpose systems can be used in conjunction with the program, or it may prove convenient to construct more specialized apparatuses to perform the described operations. The structures required for various of these systems will be apparent from the following description. In addition, the present disclosure is not described with reference to any particular programming language. It will be appreciated that the teachings of the present disclosure described herein may be implemented using various programming languages.

[0255] In the foregoing description, specific exemplary embodiments have been described. It will be apparent that various modifications may be made to these embodiments without departing from the broader spirit and scope set forth in the appended claims. Accordingly, the description and drawings are to be regarded in an illustrative rather than a restrictive sense.

[0256] Various aspects of the present invention may be understood from the following enumerated example embodiments (EEE):

[0257] EEE1. A machine-implemented method, the method comprising:

[0258] Select the original aspect ratio (AR) for the image canvas used for content creation;

[0259] determining, within at least a first scene of content on the image canvas, a first position of a first object in the at least first scene;

[0260] determining, based on the determined position of the first object, first metadata specifying how to adapt playback on a display device having an AR different from the original AR relative to the first position; and

[0261] The first metadata is stored, the first metadata and the content being to be used during playback or to be transmitted for use during playback.

[0262] EEE2. The method of EEE 1, wherein the original AR is approximately square.

[0263] EEE3. A method as described in EEE 2, wherein the substantially square conforms to one of the following: (1) closer to a square than a 16:9 AR, such that a ratio of the length to the height of the original AR is less than a ratio of 16:9 (16 / 9) but greater than or equal to 1:1; or (2) greater than a ratio of 9:16 but less than 1:1 when portrait mode is preferred; and wherein the original AR changes during the content.

[0264] EEE4. The method according to any one of EEEs 1 to 3, further comprising:

[0265] determining a plurality of objects for a plurality of scenes, the plurality of scenes including the first scene, and the plurality of objects including the first object;

[0266] A corresponding position within a corresponding scene is determined for each object in the plurality of scenes.

[0267] EEE5. The method of EEE 4, wherein the objects are determined scene by scene within the plurality of scenes; and wherein the method further comprises:

[0268] Shows a preview of how different aspect ratios will be cropped based on the metadata.

[0269] EEE6. The method of any one of EEEs 1 to 5, wherein the first metadata directs asymmetric cropping on a playback device to expand from the first object in the first scene for different ARs when adapting for playback. EEE7.

[0270] EEE7. A non-transitory machine-readable medium storing executable program instructions, which, when executed by a data processing system, cause the data processing system to perform the method of any one of EEE 1 to 6.

[0271] EEE8. A data processing system comprising a processing system and a memory, wherein the data processing system is configured to execute the method as described in any one of EEE 1 to 6.

[0272] EEE9. A machine-implemented method, the method comprising:

[0273] Receiving content including image data of at least a first scene and receiving first metadata associated with the first scene, the first metadata specifying how to adapt playback on a display device having an aspect ratio (AR) different than an original aspect ratio relative to a first position of a first object in the first scene, the first scene being created on an image canvas having the original aspect ratio; and

[0274] The output is adapted to an aspect ratio of the display device based on the first metadata.

[0275] EEE10. The method of EEE 9, wherein the original AR is approximately square.

[0276] EEE11. The method of EEE 10, wherein the substantially square is closer to a square than a 16:9 AR, such that a ratio of a length to a height of the original AR is less than a ratio of 16:9 (16 / 9), and wherein the original AR changes during the content. EEE12.

[0277] EEE12a. A method as described in any one of EEEs 9 to 11, wherein the content includes multiple scenes, the multiple scenes include the first scene, and each scene of the multiple scenes has a determined position of an object in the corresponding scene, wherein the object is determined scene by scene, and wherein adaptation for different AR is performed on a scene basis, and wherein, within each scene or frame, tone mapping is performed for the display device scene by scene or frame by frame based on the area of interest, and wherein each scene includes one or more frames.

[0278] EEE12b. A method as described in any one of EEEs 9 to 11, wherein the content includes multiple scenes, the multiple scenes include the first scene, and each scene of the multiple scenes has a determined position of an object in the corresponding scene, wherein the object is determined scene by scene, and wherein adaptation for different AR is performed on a scene basis, and wherein tone mapping is performed for the display device on a scene by scene or frame by frame based on which relative part of the adapted image is marked as the area of interest, and wherein each scene includes one or more frames.

[0279] EEE13. The method of any one of EEEs 9 to 12, wherein the first metadata directs asymmetric cropping on a playback device to expand from the first object in the first scene for different ARs when adapting for playback. EEE14. The method of any one of EEEs 9 to 12, wherein the first metadata directs asymmetric cropping on a playback device to expand from the first object in the first scene for different ARs when adapting for playback.

[0280] EEE14. The method according to EEE 9, further comprising:

[0281] receiving distance and position parameters related to a position of a viewer relative to the display device; and

[0282] Output of the first object is further adapted to the display device based on the distance and position parameters.

[0283] EEE15. A method as described in EEE 14, wherein further adapting the output of the first object to the display device includes amplifying the output of the first object when the viewing distance between the viewer and the display device increases, and reducing the output of the first object when the viewing distance between the viewer and the display device decreases.

[0284] EEE16. A method as described in EEE 14, wherein further adapting the output of the first object to the display device includes shifting the output of the first object to the left when the display device moves to the right relative to the viewer, and shifting the output of the first object to the right when the display device moves to the left relative to the viewer.

[0285] EEE17. The method according to any one of EEEs 9 to 16, further comprising:

[0286] receiving graphics data; and

[0287] A video output is generated comprising a composition of the graphics data and the adapted output.

[0288] EEE18. The method of any one of EEEs 9 to 17, wherein the first metadata further comprises syntax elements for defining an expected viewer motion path for directing a Ken Burns-related effect during playback. ...

[0289] EEE19. A non-transitory machine-readable medium storing executable program instructions, which, when executed by a data processing system, cause the data processing system to perform the method of any one of EEEs 9 to 18.

[0290] EEE20. A data processing system comprising a processing system and a memory, wherein the data processing system is configured to execute the method according to any one of EEEs 9 to 18.

[0291] appendix

[0292] Example playback behavior

[0293] The playback device is responsible for applying the specified reconstruction based on the image metadata, display configuration, and optional user configuration. In an example embodiment, the steps are as follows:

[0294] 1) Specify relative viewing distance as a fraction of the "default viewing distance". Depending on the complexity or version of the implementation, options include:

[0295] Use the default value of RelativeViewingDistance = 1.0.

[0296] Adjust it dynamically in one of two ways:

[0297] ο Automatically, when resizing the window or when entering a picture in picture mode:

[0298] RelativeViewingDistance=sqrt(WindowWidth 2 +WindowHeight 2 ) / sqrt(DisplayWidth 2 +DisplayHeight 2 )

[0299] o Manually, through user interaction (pinching, scrolling, sliding function bars, etc.)

[0300] Measure the viewer distance using a camera or other sensor, and divide the measured viewer distance (usually in meters) by the default viewing distance specified in the configuration file:

[0301] RelativeViewingDistance=ViewerDistance / DefaultViewingDistance

[0302] Note: In some embodiments, the value of the relative viewing distance may need to be bounded within a specific range (e.g., between 0.5 and 2.0). Two example bounding schemes are provided later in this section.

[0303] 2) Convert the relative viewing distance of the source to a relative angle

[0304]

[0305]

[0306] in,

[0307] (W, H) src is the width and height of the source image in pixels, and MasteringViewingDistance is the value provided by L12 metadata or other metadata, with a default value of 0.5

[0308] 3) Convert the relative viewing distance of the target into a relative angle

[0309]

[0310]

[0311] in,

[0312] (W, H) tgt is the width and height of the destination image in pixels, and

[0313] RelativeViewingDistance is calculated by step (1)

[0314] 4) Calculate the viewing angle (U, V) of the area of interest roi

[0315] U roi =U src ×W roi / W src

[0316] Vroi =V src ×H roi / H src

[0317] in,

[0318] (W, H) roi The width and height of the ROI in pixels, provided by L12 metadata or other metadata, the default value is (W, H) src ,and

[0319] (W, H) src are the width and height of the source image in pixels

[0320] 5) Rescale the target viewing angle to ensure the full ROI is displayed

[0321] S1=max(1,U roi / U tgt , V roi / V tgt )

[0322] 6) Rescale the target viewing angle to ensure that padding is applied in only one direction

[0323] S2=S1×min(1,max(U src / (U tgt ×S1), V src / (V tgt ×S1)))

[0324] U tgt =U tgt ×S2

[0325] V tgt =V tgt ×S2

[0326] 7) Find the corner coordinates (U, V)0 of the top left pixel of the ROI:

[0327] U0=U src ×X0 / (W src -1)

[0328] V0=V src ×Y0 / (H src -1)

[0329] in,

[0330] (X, Y)0 is the top left position of the ROI, from 0 to (W, H) src , provided by L12 metadata or other metadata, with a default value of (0, 0), and

[0331] (W, H) src are the width and height of the source image.

[0332] 8) Scale the top left corner of the ROI based on the distance to the edge and center the letterbox area when the target viewing angle is larger than the source viewing angle

[0333]

[0334]

[0335] 9) Convert corner coordinates to pixel coordinates

[0336]

[0337]

[0338]

[0339]

[0340] 10) Rescale the ROI (X, Y, W, H) calculated in the previous step to the output image resolution. Note: This can be done before or after applying tone mapping.

[0341] 11) Calculate the adjustment of tone mapping based on the relative size of ROI and source image

[0342]

[0343] in,

[0344] S mid is the value used as the midpoint of the tone curve for tone mapping

[0345] L1 mid and L3 midoffset Provided by L1 and L3 metadata

[0346] L12 midoffset Provided by L12 metadata, default value is 0.0

[0347] Constraining the Range of Relative Viewing Distance

[0348] For certain embodiments, two options are provided to constrain the RelativeViewingDistance from a potentially infinite range to a valid range (eg, between 0.5 and 2.0).

[0349] Hard-limited. The viewing distance is hard-limited (clipped) between the minimum and maximum viewing distances, with the ROI size remaining within the entire range. This method ensures optimal mapping at all viewing distances, but exhibits abrupt changes in behavior at the minimum and maximum viewing distances.

[0350] Soft bounding. To prevent abrupt changes in behavior at the minimum and maximum viewing distances, while also extending the range of viewing distances, a sigmoid function is applied to the viewing distances. This function has several key properties:

[0351] a) 1:1 mapping at default viewing distance to provide realistic and immersive response

[0352] b) The slope is 0 at minimum and maximum viewing distances to prevent abrupt changes in behavior

[0353] As an example, the function curve shown below maps a slightly larger measured viewing distance of 0.25 to 2.5 times the default viewing distance to a mapped viewing distance having a range of 0.5 to 2 times the default viewing distance.

[0354]

Claims

1. A machine-implemented method, the method comprising: receiving content including image data of at least a first scene and receiving first metadata associated with the first scene, the first metadata specifying how to adapt playback on a display device having an aspect ratio different than an original aspect ratio relative to a first position of a first object in the first scene, the first scene being created on an image canvas having the original aspect ratio; as well as adapting an output to an aspect ratio of the display device based on the first metadata; The method further comprises: receiving distance and position parameters related to a position of a viewer relative to the display device; and further adapting output of the first object to the display device based on the distance and position parameters, wherein the first metadata is in the form of a vector, the vector specifies a direction away from the first object, and wherein if the vector is equal in all directions about the first object, the first metadata causes the content to be symmetrically cropped around the first object in the image canvas based on the aspect ratio and orientation of the display device used to display the content, wherein the symmetrical cropping maintains the focus of the display output on the first object in all cases regardless of the aspect ratio and the orientation of the display device, and wherein, in the opposite case, if the vector is not equal in all directions of the first object, the first metadata causes an asymmetrical cropping of the content in the image canvas relative to the first object, the asymmetrical cropping being based on the aspect ratio, the orientation of the display device used to display the content, and the vector specifying how to asymmetrically crop the content relative to the first object, wherein the asymmetrical cropping maintains a position of a cropped view of the first object in the image canvas, the position corresponding to the first position of the first object in the first scene, and wherein the asymmetrical cropping maintains the position of the cropped view of the first object in the image canvas regardless of the orientation of the display device and regardless of the aspect ratio of the display device.

2. The method according to claim 1, wherein The native aspect ratio is square.

3. The method according to claim 1, wherein The original aspect ratio is closer to a square than an aspect ratio of 16:9, such that a ratio of a length to a height of the original aspect ratio is less than a ratio of 16:9, and wherein the original aspect ratio changes during the content.

4. The method according to claim 1, wherein The content comprises a plurality of scenes including the first scene, and each of the plurality of scenes has a determined position of an object in the corresponding scene, wherein the object is determined on a scene-by-scene basis, and wherein adaptation for different aspect ratios is performed on a scene-by-scene basis, and wherein, within each scene or frame, tone mapping is performed for the display device on a scene-by-scene or frame-by-frame basis based on a region of interest including the first object, and wherein each scene comprises one or more frames.

5. The method according to claim 1, wherein The first metadata directs asymmetric cropping on a playback device to expand from a first object in the first scene for different aspect ratios when adapted for playback.

6. The method of claim 1, wherein: Further adapting the output of the first object to the display device includes amplifying the output of the first object when a viewing distance between the viewer and the display device increases, and reducing the output of the first object when a viewing distance between the viewer and the display device decreases.

7. The method of claim 1, wherein: Further adapting the output of the first object to the display device includes shifting the output of the first object to the left when the display device moves right relative to the viewer, and shifting the output of the first object to the right when the display device moves left relative to the viewer.

8. The method of claim 1, further comprising: receiving graphic data; as well as A video output is generated comprising a composition of the graphics data and the adapted output.

9. The method of claim 1, wherein: The first metadata further includes syntax elements for defining an expected viewer motion path for directing a Ken Burns-related effect during playback.

10. A machine-implemented method, the method comprising: Select the original aspect ratio for the image canvas used for content creation; determining, within at least a first scene of content on the image canvas, a first position of a first object in the at least first scene; determining, based on the determined position of the first object and based on a distance between a viewer and a display device, first metadata specifying how to adapt playback on the display device having an aspect ratio different than the native aspect ratio relative to the first position; as well as storing the first metadata, wherein the first metadata and the content are to be used or transmitted for use during playback, wherein the first metadata is in the form of a vector, the vector specifies a direction away from the first object, and wherein if the vector is equal in all directions about the first object, the first metadata causes the content to be symmetrically cropped around the first object in the image canvas based on the aspect ratio and orientation of the display device used to display the content, wherein the symmetrical cropping maintains the focus of the display output on the first object in all cases regardless of the aspect ratio and the orientation of the display device, and wherein, in the opposite case, if the vector is not equal in all directions of the first object, the first metadata causes an asymmetrical cropping of the content in the image canvas relative to the first object, the asymmetrical cropping being based on the aspect ratio, the orientation of the display device used to display the content, and the vector specifying how to asymmetrically crop the content relative to the first object, wherein the asymmetrical cropping maintains a position of a cropped view of the first object in the image canvas, the position corresponding to the first position of the first object in the first scene, and wherein the asymmetrical cropping maintains the position of the cropped view of the first object in the image canvas regardless of the orientation of the display device and regardless of the aspect ratio of the display device.

11. The method according to claim 10, wherein: Different zoom factors for displaying the first object are provided for different distances between the viewer and the display device.

12. The method of claim 10, wherein: The native aspect ratio is square.

13. The method of claim 10, wherein: The native aspect ratio is one of: (1) closer to a square than an aspect ratio of 16:9, such that a ratio of length to height of the native aspect ratio is less than a ratio of 16:9 but greater than or equal to 1:1; or (2) greater than a ratio of 9:16 but less than 1:1 when portrait mode is preferred; and wherein the native aspect ratio varies during the content.

14. The method of claim 10, wherein: The content comprises a plurality of scenes, the plurality of scenes including the first scene, and each of the plurality of scenes having a determined position of an object in the corresponding scene, wherein the object is determined on a scene-by-scene basis, and wherein adaptation for different aspect ratios is performed on a scene-by-scene basis, and wherein tone mapping is performed for the display device on a scene-by-scene or frame-by-frame basis based on which relative portion of the adapted image is marked as a region of interest, wherein the region of interest includes the first object, and wherein each scene comprises one or more frames.

15. The method of claim 10, further comprising: determining a plurality of objects for a plurality of scenes, the plurality of scenes including the first scene, and the plurality of objects including the first object; as well as A corresponding position within a corresponding scene is determined for each object in the plurality of scenes.

16. The method of claim 15, wherein: determining objects within the plurality of scenes on a scene-by-scene basis; and wherein the method further comprises: Shows a preview of how different aspect ratios will be cropped based on the metadata.

17. The method of claim 10, wherein: The first metadata directs asymmetric cropping on a playback device to expand from a first object in the first scene for different aspect ratios when adapted for playback.

18. A non-transitory machine-readable medium storing executable program instructions, which, when executed by a data processing system, cause the data processing system to perform the method of any one of claims 1 to 17.

19. A data processing system comprising a processing system and a memory, wherein the data processing system is configured to execute the method according to any one of claims 1 to 17.

Citation Information

Patent Citations

  • Tone curve mapping for high dynamic range images

    US10600166B2

  • Display with automatic screen parameter adjustment based on the position of a detected viewer

    GB2467898A

  • Image pickup device and its operation control method

    JP2002027298A

  • Dynamically cropping digital content for display in any aspect ratio

    US20170249719A1

  • Video Motion Effect Generation Based On Content Analysis

    US20190297248A1