Producing and Adapting Video Images for Presentation on Displays with Different Aspect Ratios

By employing a substantially square aspect ratio and metadata-guided asymmetric cropping, the patent addresses the challenge of preserving content intent across diverse aspect ratios, ensuring optimal viewing experiences on different display devices.

JP7747667B2Active Publication Date: 2025-10-01DOLBY LABORATORIES LICENSING CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022575886
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-20
Filing Date
2021-06-09
Publication Date
2025-10-01
Estimated Expiration
2041-06-09

AI Technical Summary

Technical Problem

Content creation and distribution face challenges in preserving the creator's intent due to multiple adaptations of aspect ratios, leading to unnecessary cropping or padding when content is played on devices with different aspect ratios.

Method used

Using a substantially square aspect ratio for the original image canvas and creating metadata that guides playback devices to asymmetrically crop or expand content based on the object's position, ensuring the object remains in focus while adapting to various aspect ratios.

Benefits of technology

Preserves the creative intent by maintaining the object in focus and optimizing image quality across different display devices with varying aspect ratios, allowing for seamless transitions and improved viewing experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007747667000017
    Figure 0007747667000017
  • Figure 0007747667000018
    Figure 0007747667000018
  • Figure 0007747667000019
    Figure 0007747667000019
Patent Text Reader

Abstract

Described embodiments include systems and methods for producing images, such as video images, and adapting the images for presentation on playback devices having a variety of different aspect ratios, such as 4:3, 16:9, 9:16, etc. In one embodiment, a method for producing content, such as a video image, may begin by selecting an original aspect ratio and determining, within at least a first scene in the content, the position of an object within the first scene. In one embodiment, the original aspect ratio may be substantially square (e.g., 1:1). Metadata may then be created based on the object's position within the first scene, such that the metadata guides the playback device to asymmetrically crop the content relative to that position to display the content on a display device having an aspect ratio different from the original aspect ratio. Other methods and systems are also described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Related Applications) This application claims priority to U.S. Provisional Patent Application No. 63 / 068,201, filed August 20, 2020, U.S. Provisional Patent Application No. 62 / 705,115, filed June 11, 2020, and European Patent Application No. 20179451.8, filed June 11, 2020, each of which is incorporated by reference in its entirety. [Background technology]

[0002] Content creation, such as the creation of a movie or TV show or animation, and content capture, such as the recording of a live sporting event or news, requires the selection of an aspect ratio for what may be called an image canvas. This selection may be deliberate (e.g., a content creator considers several possible options for the image canvas aspect ratio and chooses one) or serendipitous (e.g., a camera operator in a recording studio picks up a particular camera with a predetermined aspect ratio without considering the aspect ratio for capture). Once the image canvas aspect ratio is selected, content is created or captured, and then distributed for playback on devices that may have many different aspect ratios. Often, content creators capture or create content in an image canvas with a first aspect ratio and then crop or pan the content to fit an expected playback aspect ratio that differs from the first aspect ratio. The expected playback aspect ratio may be the aspect ratio that the content creator believes to be the most common aspect ratio used by playback devices. The content is then released and distributed to playback devices with many different aspect ratios that differ from both the original (initial) canvas aspect ratio and the expected playback aspect ratio. These playback devices must then adapt the display of the content by cropping or padding the content to match the display device connected to the playback device. In this example, the content is cropped and / or padded at least twice. This process of padding and cropping at least twice may result in unnecessary cropping or padding of the image, thereby preventing the intent of the content creator from being preserved through the process of adapting the content multiple times to fit different aspect ratios. Summary of the Invention [Problem to be solved by the invention]

[0003] Aspects and embodiments described in this disclosure may provide systems and methods that may use a substantially square aspect ratio for an original image canvas and associated metadata that allows a variety of endpoint aspect ratios to be derived from original content that uses the original image canvas or set of original image canvases. [Means for solving the problem]

[0004] In one embodiment, a method for producing content, such as a video image, may begin by selecting an original aspect ratio for an image canvas and, within at least a first scene in the content on the image canvas, determining a position of an object within the first scene. In one embodiment, the original aspect ratio may be substantially square (e.g., 1:1). The object may be an area of ​​interest within the content, such as an actor or other focal point within the scene. Metadata may then be created based on the object's position within the first scene to guide a playback device to asymmetrically crop the content relative to the position and display the content on a display device having an aspect ratio different from the original aspect ratio. The metadata may guide how the playback device may asymmetrically expand the view around the object based on the object's position on the canvas and the aspect ratio of the playback device. In one embodiment, the metadata may also guide the asymmetric zoom based on other factors, such as a desire to avoid partial inclusion of certain image elements, such as a human face, and the metadata may provide data used to prevent partial inclusion of such elements (which may mean completely excluding such elements or completely including them in the cropped view). Adding such image elements to a region of interest may ensure that the image elements are completely included in the view or completely excluded from the view. For example, such elements may be added by specifying the size of the region of interest in which they should be included. The content and metadata may be stored and then distributed to playback devices with different aspect ratios. The metadata may be used by the playback device to adapt the content to the display used by the playback device by cropping or padding, if necessary. In one embodiment, metadata may be created per scene, where a scene may be as short as one frame of video content, in which case the metadata may be per frame, where a frame is one image presented on the playback device during a single refresh interval.Thus, in one embodiment, the metadata described herein may be created for each frame to capture frame-to-frame changes over time. Furthermore, the original aspect ratio may not be constant throughout the content, and thus may vary throughout the content, and even from scene to scene (or even frame to frame) for at least some of the content. Changes in the original aspect ratio may be referred to as a variable aspect ratio that changes during the content.

[0005] In one embodiment, the original aspect ratio may be selected to be substantially square, e.g., 1:1, or closer to square than a 16:9 aspect ratio, i.e., the ratio of length to height of the original aspect ratio is less than the 16:9 ratio (16 / 9 = 1.7778) but greater than or equal to 1:1. A substantially square original aspect ratio may ensure the greatest range of options for adapting content to most aspect ratios of many playback devices. In other embodiments, the original aspect ratio may be selected to prioritize image quality in a vertical playback orientation, e.g., portrait orientation. In this case, the original aspect ratio may be substantially square (1:1), ranging from 1:1 to less than 1:1, or even up to 9:16. To provide great flexibility in content creation, the canvas may vary from frame to frame, scene to scene, or shot to shot. It should be noted herein that the original aspect ratio may change over time, such that the content may include a variable aspect ratio.

[0006] The metadata may be a vector that specifies a direction to an object (e.g., a region of interest) in the current scene. The playback device may use the metadata to construct an image crop and / or padding to render the scene based on the metadata. In effect, the metadata may guide the playback device to keep the object in focus in the cropped scene while cropping or expanding away from the object to fill the entire canvas. The playback device may also use tone mapping and color volume mapping specifically tailored for the adapted aspect ratio to base the tone mapping and color volume mapping on the actual content (e.g., region of interest) within the adapted aspect ratio, rather than on the entire content in the original image canvas for the scene.

[0007] In one embodiment, the method for determining objects and object positions may be performed for multiple different scenes covering the same or different objects. In one embodiment, the method may be performed for each scene, frame, set of frames, or at least a subset of scenes in the content. As a result, for at least a subset of scenes, each scene or frame may be adapted during playback to fit different aspect ratios based on metadata created during content creation. In one embodiment, a scene (e.g., a set of one or more frames) may be a camera shot or a take within a film production or other content production. Different scenes may have different objects, backgrounds, camera angles, etc.

[0008] During the content creation process, one or more previews of the content displayed at different aspect ratios may be generated based on the generated metadata. After viewing the previews, the content creator may edit the metadata directly, such as by modifying the position of objects or selecting different objects. The content creator may then display one or more previews to check whether the modifications improve the appearance of the content in the different previews. In one embodiment, the one or more previews may be one or more rectangular overlays on the original canvas, with the content of the scene displayed in the overlays.

[0009] In one embodiment, the playback device's user interface allows a user to switch playback between cropping (cropping based on metadata as described herein) and padding. This reverts to the common practice of padding pixels (typically black pixels) around a substantially square canvas to fill the entire aspect ratio of the display. This allows the user to see more than half of the original canvas in one embodiment, rather than the focused view that may be provided using metadata-based cropping as described herein. In one embodiment, zooming the image beyond what is necessary to match the aspect ratio of the playback device may provide an improved (closer) view of the subject area. In one embodiment, transitions between cropping, padding, and / or zooming viewing states may appear smooth, allowing the user / viewer of the playback device to see a smooth or seamless transition between viewing states.

[0010] In one embodiment, the final composite image presented to the user may correspond to the superimposition of multiple inputs within multiple windows or regions of the screen. For example, the primary input may be shown within a large window on the display, while the secondary input may be shown within a smaller or significantly smaller window (similar to a picture-in-picture feature on a television). For example, the primary input window may display a view generated from metadata to provide a cropped view of the original canvas or entire canvas based on the position and aspect ratio of the subject on the primary input window or display device, while the secondary window may show the entire canvas or original canvas without cropping or padding. In one embodiment, one or both of these windows may be fully or partially zoomed. Furthermore, the methods and systems described herein may be used to optimize playback of content for any size and aspect ratio of each window. If one of the windows is resized, the methods and systems may be used to adapt the output to the resized aspect ratio (of the resized window) using metadata and the position of the subject in the scene. Furthermore, the methods and systems described herein may be used with a single window (without secondary windows) to crop content within the window based on the aspect ratio of the window. The methods and systems may use the metadata described herein to optimize playback.

[0011] In another embodiment, the methods and systems described herein may be used to enhance the transition between photo playbacks in a photostream by focusing on a subject of interest (object of interest), while applying subtle zooms and pans to create interesting effects, and simultaneously optimize tone mapping to selected regions. This may be guided by metadata, which may be thought of as an "intended path of movement." In one embodiment, an "intended path of movement" is used instead of tracking viewer position to provide a "guided Ken Burns effect." The metadata is a series of position vector (X, Y, Z) coordinates relative to the screen that describe the intended path of movement for the viewer over a specific period of time. As used herein, the term "Ken Burns effect" refers to a type of pan and zoom effect used when showing still images in film and video productions.

[0012] In one embodiment, the original canvas or entire canvas may already include some padding to fit the image to the aspect ratio of the canvas. In this case, an embodiment may use additional metadata to indicate the location of the active area within the canvas. If this metadata is present, the client or playback device may use this additional metadata to adapt playback based only on the active area of ​​the canvas (without including the padded area).

[0013] Aspects and embodiments described herein may include a non-transitory machine-readable medium that may store executable computer program instructions that, when executed, cause one or more data processing systems to perform the methods described herein. The instructions may be stored in a non-transitory machine-readable medium, such as a non-volatile memory, e.g., flash memory, or volatile dynamic random access memory (DRAM), or other form of memory.

[0014] The above summary is not an exhaustive list of all embodiments and aspects of the present disclosure, and all systems, media, and methods may be implemented from all appropriate combinations of the various aspects and embodiments summarized above, as well as all appropriate combinations of the various aspects and embodiments disclosed in the detailed description below.

[0015] The invention will now be described, by way of example only, with reference to the accompanying drawings, in which like reference symbols indicate like features and in which: [Brief explanation of the drawings]

[0016] [Figure 1A] FIG. 1A shows examples of different aspect ratios of a display device that may be used in one or more embodiments described herein. [Figure 1B] FIG. 1B shows examples of different aspect ratios of display devices that may be used in one or more embodiments described herein. [Figure 1C] FIG. 1C shows examples of different aspect ratios of display devices that may be used in one or more embodiments described herein.

[0017] [Figure 2A] FIG. 2A is a flow chart illustrating a method according to one embodiment that may be used to create content with metadata that adapts the output to different aspect ratios.

[0018] [Figure 2B] FIG. 2B is a flow chart illustrating a method according to one embodiment that may be used to adapt playback on a playback device based on the aspect ratio of the playback device and based on metadata associated with the content.

[0019] [Figure 3A] FIG. 3A shows examples of object locations and metadata associated with the objects based on those locations. [Figure 3B] FIG. 3B shows examples of object locations and metadata associated with the objects based on those locations. [Figure 3C] FIG. 3C shows examples of object locations and metadata associated with the objects based on those locations. [Figure 3D] FIG. 3D shows examples of object locations and metadata associated with the objects based on those locations.

[0020] [Figure 4A] FIG. 4A shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop an image on the original canvas based on the metadata and aspect ratio. [Figure 4B] FIG. 4B shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop the image on the original canvas based on the metadata and aspect ratio. [Figure 4C] FIG. 4C shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop the image on the original canvas based on the metadata and aspect ratio. [Figure 4D] FIG. 4D shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop the image on the original canvas based on the metadata and aspect ratio. [Figure 4E] FIG. 4E shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop the image on the original canvas based on the metadata and aspect ratio. [Figure 4F] FIG. 4F shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop an image on the original canvas based on the metadata and aspect ratio. [Figure 4G]FIG. 4G shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop the image on the original canvas based on the metadata and aspect ratio. [Figure 4H] FIG. 4H shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop the image on the original canvas based on the metadata and aspect ratio. [Figure 4I] FIG. 4I shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop an image on the original canvas based on the metadata and aspect ratio. [Figure 4J] FIG. 4J shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop an image on the original canvas based on the metadata and aspect ratio. [Figure 4K] FIG. 4K shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop an image on the original canvas based on the metadata and aspect ratio. [Figure 4L] FIG. 4L shows an example of how a playback device can use the metadata and the aspect ratio of the playback device to asymmetrically crop the image on the original canvas based on the metadata and aspect ratio.

[0021] [Figure 5] FIG. 5 is a flow chart illustrating a method for creating content according to one embodiment.

[0022] [Figure 6A] FIG. 6A shows an example of how a playback device can display an image based on image metadata and the relative position of the display to the viewer. [Figure 6B]FIG. 6B shows an example of how a playback device can display an image based on image metadata and the relative position of the display to the viewer. [Figure 6C] FIG. 6C shows an example of how a playback device can display an image based on image metadata and the relative position of the display to the viewer. [Figure 6D] FIG. 6D shows an example of how a playback device can display an image based on image metadata and the relative position of the display to the viewer. [Figure 6E] FIG. 6E shows an example of how a playback device can display an image based on image metadata and the relative position of the display to the viewer. [Figure 6F] FIG. 6F shows an example of how a playback device can display an image based on image metadata and the relative position of the display to the viewer. [Figure 6G] FIG. 6G shows an example of how a playback device can display an image based on image metadata and the relative position of the display to the viewer.

[0023] [Figure 7] FIG. 7 shows an example of displaying an image according to one embodiment of the image adaptation process.

[0024] [Figure 8] FIG. 8 illustrates an example data processing system that may be used to create the content and metadata described herein, and further illustrates an example data processing system that may be a playback device that uses the metadata to adapt playback, where the adaptation is based on the metadata and the aspect ratio of the playback device. DETAILED DESCRIPTION OF THE INVENTION

[0025] Various embodiments and aspects are described in detail below. The accompanying drawings illustrate these various embodiments. The following description and drawings illustrate examples of the invention and should not be construed as limiting the invention. Numerous specific details are set forth to provide a thorough understanding of the various embodiments. However, in some instances, well-known or conventional details are not described in order to concisely describe the embodiments.

[0026] References herein to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment. The appearance of the phrase "in one embodiment" in various places throughout this specification does not necessarily refer to the same embodiment. The processes illustrated in the figures described below are performed by processing logic, including hardware (e.g., circuits, dedicated logic, etc.), software, or a combination thereof. While these processes are described below as several sequential operations, it should be understood that some of the described operations may occur in a different order. Furthermore, some operations may occur in parallel rather than sequentially.

[0027] This description contains copyrighted material, including computer program software. The copyright holders, including the assignee of this invention, hereby reserve all rights, including copyright, in and to such material. The copyright holders have no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the U.S. Patent and Trademark Office files or records, but otherwise reserve all copyright rights. Copyright is Dolby Laboratories, Inc.

[0028] The embodiments described herein can create and use metadata that adapts the content of an original canvas or the entire canvas for output on different display devices with different aspect ratios. These display devices may be conventional LCD or LED displays that are part of a playback device, such as a tablet computer, smartphone, laptop computer, or television, or may be conventional displays that are connected to but not integral with the playback device that drives the display by outputting to the display. Figures 1A, 1B, and 1C show three examples of three different aspect ratios. Specifically, Figure 1A shows an example of a display with a 4:3 aspect ratio (the aspect ratio is the ratio of the length to the height of the viewable area of ​​the display panel). Thus, for a display with a 4:3 aspect ratio, if the display area is 8 inches long, the display area is 6 inches high. Most cathode ray tube televisions had this aspect ratio. Figure 1B shows an example of a display panel with a display area with an aspect ratio of 16:9. Display panels for laptop computers and televisions often use this aspect ratio. 1C shows an example of a display panel or image canvas with an aspect ratio of 1:1. The image canvas is square (the length and height of the viewable area are equal). As described below, in one embodiment, a source image canvas with an aspect ratio of 1:1, or a source image canvas that is substantially close to a square image canvas, is used in the content creation process. An example of the content creation process is described below with reference to FIG. 2A.

[0029] As shown in FIG. 2A , a method according to one embodiment may begin with operation 51. In operation 51, an original aspect ratio for the image canvas is selected. Once the original aspect ratio is selected, content is created using the image canvas. Content creation may involve creating an image (such as computer-generated graphics or animation), capturing content using a camera (such as a movie camera used by a live performer), other techniques known in the art for content creation, or a combination of these techniques. The content creation process may use the same image canvas or different image canvases. To provide the greatest flexibility in content creation, the canvas area may vary from frame to frame, from frame set to frame, or from scene to scene. In one embodiment, the original image canvas may be a square canvas (aspect ratio of 1:1) or a substantially square canvas. In one embodiment, an image canvas is considered substantially square if it is more square than a 16:9 aspect ratio canvas, i.e., the ratio of the length to height of the image canvas is less than 16:9 (i.e., approximately 1.778) and greater than or equal to 1:1. A substantially square original aspect ratio can ensure the widest range of options for adapting content to most aspect ratios used across many playback devices. Other embodiments may use a non-substantially square image canvas, but this can affect how well the scene adapts to different displays with different aspect ratios.

[0030] In operation 53 of FIG. 2A, the location of an object within a scene in the content is determined. Operation 53 may be performed during content creation, after content creation (during an editing process to edit the created content), or both during and after content creation. Operation 53 may begin by identifying or determining an object or region of interest in a particular scene. This determination or identification of an object or region of interest may be performed scene-by-scene for all scenes in the content, or for at least a subset of all scenes in the content. For example, a first scene may have a first identified object, and a second scene may have a second identified object that is different from the first identified object. Furthermore, different scenes may contain the same object identified in the different scenes, but the location of the object may differ between the different scenes. In one embodiment, the identification or determination of an object or region of interest may be performed manually by a content creator or automatically by a data processing system. For example, the data processing system may automatically detect one or more faces in a scene using a known face detection algorithm, image salience analysis algorithm, or other known algorithm. In one embodiment, the automatic detection of the subject may be manually overridden by the content creator. In one embodiment, the content creator may instruct a data processing system used in the content creation process to automatically determine the subject for one subset of scenes and to allow the content creator to manually determine the subject for another subset. Once the subject is determined, the position of the subject may be manually or automatically determined based on the subject's centroid in operation 53. For example, if the subject's face is identified as the subject or region of interest, the face's centroid may be used as the subject's position on the image canvas in operation 53. In one embodiment, if this position is automatically determined by the data processing system, the content creator may manually edit this position.In one embodiment, whether this position is determined manually or automatically, a user can use a content creation or editing tool to select the object's center pixel (e.g., Sx, Sy) and, optionally, the object's width and height (e.g., Sw, Wh). The object's center pixel and object's width and height are used in adapting the image for reproduction in the aspect ratio of a particular display device. After performing operation 53, processing continues to operation 55, as shown in FIG. 2A.

[0031] In operation 55, the data processing system may automatically determine metadata based on the position of the object within a particular scene. The metadata may specify how to accommodate playback on a display device having an aspect ratio different from the original aspect ratio of the original image canvas. For example, the metadata may specify how to scale an image in one or more directions within the original aspect ratio from the position of the object to the original aspect ratio in order to crop the image to accommodate playback in the particular aspect ratio of a display device controlled by the playback device. In one embodiment, the metadata may be represented as a vector specifying a direction from the determined object. Figures 3A, 3B, 3C, and 3D show examples of metadata that may be in the form of a vector. Figure 3A shows an example of an object 105 near the center of a scene 103 in the original image canvas 101. Figure 3B shows an example of an object 111 on the left side of a scene 109 in the original image canvas 101. As shown in Figure 3B, the object 111 is oriented vertically from the left edge of the scene 109 toward the center. FIG. 3C shows an example of object 117 in the upper right corner of original image canvas 101 within scene 115. FIG. 3D shows an example of object 123 in the lower right corner of scene 121 within original image canvas 101. In the example shown in FIG. 3A, the vector representing the example metadata is considered equal in all directions from the object and therefore can be considered to have the value 0,0 in this case. In the example of FIG. 3A, the magnification from the object occurs by cropping the original image canvas based on the aspect ratio of the playback device and the orientation of the display device (e.g., landscape or portrait). FIGS. 4A through 4D show four examples of how content can be adapted for playback at different aspect ratios when the object is centered in the original image canvas. These examples are described further below. In the example shown in FIG. 3B, vector 112 is example metadata that specifies the direction to crop the original image canvas 101 to focus on object 111, regardless of the aspect ratio and orientation of the display device. Figures 4E, 4F, 4G and 4H show four examples of how content can be adapted for playback in different aspect ratios when the subject is in the position shown in Figure 3B.In the example shown in Figure 3C, vector 119 is an example of metadata that specifies the direction to crop original image canvas 101 to focus on subject 117, regardless of the aspect ratio of the display device and the orientation of the display device. Figures 41, 4J, 4K, and 4L show four examples of how content can be adapted for playback at different aspect ratios when the subject is in the position shown in Figure 3C. In the example shown in Figure 3D, vector 125 can be metadata that specifies the direction to crop original image canvas 101 to focus on subject 123, regardless of the aspect ratio of the display device and the orientation of the display device.

[0032] The vector representing the metadata can guide the playback device on how to crop the original image canvas based on the metadata and the position of the object. In one embodiment, the vector (e.g., vectors 112, 119, and 125) guides asymmetric cropping around the object, as described further below, rather than symmetrically cropping around the object. Such asymmetric cropping can provide at least two advantages: (a) the aesthetic framing of the scene is better preserved; even after zooming an image with an object in the upper right corner (e.g., see FIG. 3C ) to an intermediate level, the object remains in the upper right portion of the image, better preserving the creative intent of the framing compared to symmetric cropping; and (b) asymmetric cropping allows zooming in on the object without abruptly changing the zoom direction or zoom rate. In one embodiment, the vector can include an x-component (for the x-axis) and a y-component (for the y-axis), and the vector can be represented by two values, Px and Py, where Px is the x-component of the vector and Py is the y-component of the vector. In one embodiment, Px may be defined as Px = 2(0.5 - Sx), and Py may be defined as Py = 2(0.5 - Sy), where Sx and Sy are the center of the subject relative to the original image canvas, coordinates 0,0 are the top left corner of the original image canvas, coordinates 1,1 are the bottom right corner of the canvas, and coordinates 0.5,0.5 are the center of the original image canvas. For the example shown in Figure 3B, Px = 1 and Py = 0, and therefore this vector may be considered positive horizontally. For the example shown in Figure 3C, Px = 1 and Py = 1, and therefore this vector may be considered negative horizontally and positive vertically. Further details and examples of this metadata and how it is used in playback are provided below.

[0033] As shown in FIG. 2A, after completing operation 55 for a particular scene, processing proceeds to operation 57. In operation 57, the content and metadata are saved for use during playback. Thereafter, in operation 59, a data processing system being used during the content creation or editing process determines whether there is more content to process. If there are more scenes to process, for example, using the method shown in FIG. 2A, processing returns to operation 51. In operation 51, a new original aspect ratio for the image canvas may be selected, or the original aspect ratio previously used for the image canvas may be used to continue creating and / or editing content. In one embodiment, the determination in operation 59 may be a manual determination controlled by a human operator operating the data processing system. If there is no content to process, then in operation 61, the saved content and metadata may be provided to one or more content distribution systems. For example, the content and metadata may be provided to a cable network for distribution to set-top boxes, or to a content provider, such as a content provider that distributes content over the Internet. The content and metadata may be distributed for use in streaming media, or the entire content and metadata may be downloaded for use.

[0034] The method illustrated in Figure 2A is a content creation method that may be performed by a movie studio or other facility that creates content. This method is typically performed separately from playback on a playback device. An example of a method performed on a playback device is shown in Figure 2B. However, in one embodiment, the method illustrated in Figure 2A may be performed on the same device as the method illustrated in Figure 2B, which creates the content and then displays the content on one or more display devices having an aspect ratio different from that of the original image canvas.

[0035] The playback method shown in FIG. 2B may begin with operation 71. In operation 71, a playback device may receive content including image data and may also receive associated metadata. For example, the metadata may be associated with a first scene, and the metadata may specify how to adapt playback for a display device having an aspect ratio different from the original aspect ratio in which the scene was created, relative to the location of objects in the scene. In one embodiment, the metadata may take the form of vectors, such as vectors 112, 119, and 125 shown in FIGS. 3B, 3C, and 3D. These vectors may be represented by Px and Py values ​​that provide the content for the scene associated with the metadata. After receiving the metadata in operation 71, the playback device may perform operation 73. In operation 73, the playback device adapts the content in the scene to the aspect ratio of a display device connected to the playback device. For example, the display device may be a display panel of a television, a smartphone, or a tablet computer. The playback device adapts the content in the scene by cropping the content to fit the aspect ratio of the display device. This adaptation or cropping uses metadata as described herein to crop the content, in one embodiment asymmetrically, based on the subject's position and the metadata, which may be vectors as described herein. This adaptation or cropping may also include tone mapping and color volume mapping specifically prepared for this purpose, which is based on the cropped content (e.g., including only the region of interest) displayed within the aspect ratio of the display device, as opposed to tone mapping and color volume mapping that is based on the entire image in the original image canvas.

[0036] While a detailed example of the performance of act 73 is provided below, it is useful to broadly describe the adaptation process with reference to Figures 4A through 4L. In the examples shown in Figures 4A, 4B, 4C, and 4D, act 73 crops content symmetrically about subject 105 in original image canvas 51 according to the aspect ratio of the display device and the orientation of the display device. Figure 4A shows aspect ratio 153 created by act 73. Act 73 cropped the content in original image canvas 151 symmetrically around subject 105 in landscape mode. Figure 4B shows aspect ratio 155 resulting from the cropping in act 73. Act 73 crops the content in landscape mode to achieve aspect ratio 155. Figure 4C shows aspect ratio 157 resulting from the cropping in act 73. Act 73 crops the content in portrait mode to achieve aspect ratio 157. Figure 4D shows aspect ratio 159 resulting from the cropping in act 73. In operation 73, the content is cropped in portrait mode to obtain aspect ratio 159. In each of these examples shown in Figures 4A, 4B, 4C, and 4D, the vector metadata may be the vector values ​​Px=0, Py=0, where the vector metadata crops the original image canvas 151 symmetrically around the subject 105 based on the aspect ratio of the display device used to display the content. In the examples shown in Figures 4A, 4B, 4C, and 4D, the display output remains in focus on the subject 105 in all cases, regardless of the aspect ratio of the display device and the display orientation (e.g., portrait or landscape).

[0037] In the examples shown in FIGS. 4E, 4F, 4G, and 4H, act 73 asymmetrically crops content in the original image canvas relative to subject 111 based on vector 112. Vector 112 specifies the asymmetric cropping method in this example. FIG. 4E shows aspect ratio 161 resulting from the cropping in act 73. Act 73 crops the content in landscape mode based on vector 112 to achieve aspect ratio 161. FIG. 4F shows aspect ratio 163 resulting from the cropping in act 73. Act 73 crops the content in landscape mode based on vector 112 to achieve aspect ratio 163. FIG. 4G shows aspect ratio 165 resulting from the cropping in act 73. Act 73 crops the content in portrait mode based on vector 112 to achieve aspect ratio 165. FIG. 4H shows aspect ratio 167 resulting from the cropping in act 73. In operation 73, the content is cropped in portrait mode based on vector 112 to obtain aspect ratio 167. In the examples shown in Figures 4E, 4F, 4G, and 4H, it can be seen that cropping maintains subject 111 in the left portion of the cropped view of original image canvas 151, regardless of landscape or portrait and the aspect ratio of the display device.

[0038] In the examples shown in FIGS. 4I, 4J, 4K, and 4L, act 73 asymmetrically crops content in the original image canvas relative to subject 117 based on vector 119. Vector 119 specifies the method of asymmetric cropping in this example. FIG. 4I shows aspect ratio 171 resulting from the cropping in act 73. Act 73 crops the content in landscape mode based on vector 119 to achieve aspect ratio 171. FIG. 4J shows aspect ratio 173 resulting from the cropping in act 73. Act 73 crops the content in landscape mode based on vector 119 to achieve aspect ratio 173. FIG. 4K shows aspect ratio 175 resulting from the cropping in act 73. Act 73 crops the content in portrait mode based on vector 119 to achieve aspect ratio 175. FIG. 4L shows aspect ratio 177 resulting from the cropping in act 73. In operation 73, the content is cropped in portrait mode based on vector 119 to obtain aspect ratio 177. In the examples shown in Figures 4I, 4J, 4K, and 4L, it can be seen that subject 117 remains in the upper right corner of each cropped output regardless of the orientation of the display device and the aspect ratio of the display device.

[0039] Returning to Figure 2B, after operation 73, the method determines whether there is more content to process in operation 75. If there is more content to process, processing returns to operation 71. Operation 71 continues to receive content and metadata and adapt the content for the display device as described above. If there is no content to process, processing continues to operation 77 and the method ends.

[0040] FIG. 5 illustrates another example of a method that may be performed during content creation or during editing of content that has been created, saved, and ready to be edited. In operation 201 shown in FIG. 5, a content creator may select an original canvas aspect ratio, such as a 1:1 aspect ratio, and create content within the current scene. Then, in operation 203, the content creator or the data processing system may determine the position of the current object for the current scene. The position may be determined manually by the content creator or automatically by the data processing system, leaving open the possibility of manually adjusting or overwriting the current scene. The content creator or the data processing system may optionally use editing or content creation tools to set the object's center point (e.g., the point with coordinates Sx, Sy described herein) as well as the object's size. Then, in operation 205, the data processing system may calculate metadata based on the position that can be used to adapt the content to other aspect ratios. This metadata may further include additional metadata, such as metadata describing padding around the image (if necessary). The content creator may then display previews tailored to one or more other aspect ratios. This may be done by adapting the content in the current scene to the other aspect ratio. In other words, the content creator may have the data processing system display previews of the content so that the content creator can view each preview at different aspect ratios and determine whether the adaptation or cropping is desirable or satisfactory. In one embodiment, the previews may be rectangles overlaid on the image in the original image canvas, such as rectangles showing the various aspect ratios (e.g., 161) shown in FIGS. 4A through 4L. These previews show how the content will appear to an end user of the playback device. The position of the rectangles is derived based on a metadata-based cropping operation, such as cropping operation 73 in FIG. 2B. The metadata provides the advantage of prescribed playback behavior, which provides a highly accurate preview of how the content will be rendered on a variety of different playback devices having different display devices with different aspect ratios.Knowing the playback behavior in this way allows the content creator to preview the exact final result and make any desired or necessary adjustments. The content creator may then decide whether to adjust the adaptation for one or more aspect ratios in operation 209. If adjusting the adaptation is desired, processing may return to operation 203. In operation 203, the content creator may modify the current object position, or perhaps select a different object using a content creation or editing tool to select a different position or a different object. If no adjustments are deemed necessary in operation 209, the content creator may proceed to the next scene and decide whether to process the next scene in operation 211. If all scenes desired to be processed have been processed, processing may be complete and may end, as shown in FIG. 5. On the other hand, if there are additional scenes to process, processing returns to operation 201, as shown in FIG. 5.

[0041] Below are detailed examples of metadata and how the metadata can be used to crop content for playback in different aspect ratios. In one embodiment, the metadata may be specified in a compliant bitstream as follows: One or more rectangular regions define the subject area. The rectangle should be defined such that (Top)<=(1-Bottom) and (Left)<=(1-Right). The TopOffset, BottomOffset, LeftOffset and RightOffset values ​​should be set accordingly. Playback devices should enforce this behavior and tolerate non-compliant metadata. A 0 offset indicates that the entire image is the region of interest. If the width and height are zero pixels, the top left corner of the rectangle is designated as the region of interest, which corresponds to the center of the subject.

[0042] Additional metadata used as described below may include: [Table 1]

[0043] The coordinates may vary from frame to frame or shot to shot, or may be the same for the entire content. Any changes may be fully frame-synchronous between the image and its corresponding metadata.

[0044] For example, if the canvas is resized before distribution in an adaptive streaming environment, the offset coordinates are updated accordingly.

[0045] In the following, we describe adapting content in a playback device and optimizing color volume mapping for the adapted content in the playback device. We assume that the playback device performs all of these operations locally on the playback device, although in other embodiments a centralized processing system may perform some of these operations on behalf of one or more playback devices connected to it.

[0046] During playback, the playback device is responsible for adapting the canvas and its associated metadata to the particular aspect ratio of the attached panel. This involves three operations, described below. For example, in one embodiment, the three operations are:

[0047] 1. Calculate the region of interest and update the mapping curve. The coordinates of the area of ​​interest on the canvas or the area to be displayed on the panel are calculated by calculating the top left and bottom right pixels TLx, TLy, BRx, and BRy, as well as the width and height (CW, CH) of the canvas. For example, the calculation may be based on the formula shown immediately below, or on the software implementation shown later. 1) TLx = (Sx - Px) * CW 2)TLy=(Sy-Py) * CH 3)BRx=(Sx+Px) * CW 4)BRy=(Sy+Py) * CH

[0048] In addition to adaptively resizing the image according to the region of interest, the tone mapping algorithm may also be adjusted to achieve optimal tone mapping for the cropped region (rather than the entire original image in the original image canvas). This may be achieved by calculating additional metadata corresponding to the region of interest and using this to adjust the tone mapping curve. This is described, for example, in U.S. Pat. No. 10,600,166, which describes a display management process known in the art. The tone mapping curve takes one input, a "smid" (mean luminance) parameter that represents the average brightness of the source content. The adjustment is calculated using this new ROI luminance offset metadata (denoted, for example, as L12MidOffset), as follows: SMid=(L1.Mid+L3MidOffset) / / Calculates the midpoint luminance for the entire frame. SMid'=SMid * (1-ZF)+(SMid+L12MidOffset) * ZF / / Adjust to fit ROI, where ZF is the zoom factor, e.g. ZF=0 corresponds to full screen, ZF=1 corresponds to fully zoomed in on the subject. Note: L3MidOffset means the offset beyond the L1.Mid value and may also be called L3.Mid.

[0049] Another parameter that is adjusted in a similar fashion is the optional global dimming algorithm, which is used to optimize the mapping for a global dimmed display. The global dimming algorithm takes two input values, L4Mean and L4Power. Before calculating the global dimming background, the L4Mean value is adjusted by the zoom ratio as follows: L4Mean'=L4Mean * (1-ZF)+(L4Mean+L12MidOffset) * ZF

[0050] 2. Cropping and processing of the region of interest. In a preferred embodiment, to use memory efficiently and ensure consistent timing of the playback device, the playback device should do the following: 1) Decode the bitstream encoded with region of interest (ROI) metadata (e.g., vectors represented by Px, Py) and insert individual frames into the decoded picture buffer. 2) When it is time to display the current frame, only the portion of the image needed by the ROI is read from memory, starting with the top left pixel (TLx,y). This pixel is read a time t (the "delay") before it is presented on the panel. This delay t is determined by the time it takes for the first pixel to be processed by the imaging pipeline, including any spatial upsampling performed by the imaging pipeline. 3) Once the entire region of interest has been read from memory, overwrite the decoded picture buffer with the next decoded picture.

[0051] Once the cropped region of the image is read from memory, it is mapped to the dynamic range of the panel, which may be done by known techniques such as those described in U.S. Patent No. 10,600,166, using the adjusted mapping parameters obtained in act 1 above.

[0052] 3. Resize to output resolution. The final operation is to resize the image to the resolution of the panel. Obviously, the resolution or size of the final image will not match the resolution of the panel. To achieve the desired resolution, methods of resizing the image must be applied, which are well known in the art. Example methods could include bilinear or Lanczos resampling, or many methods including super-resolution techniques or neural networks.

[0053] In one embodiment, the metadata used to signal the ROI and its associated parameters may be represented as, but is not limited to, Level 12 (L12) metadata, which is summarized below. 1) A rectangle that specifies the coordinates of the ROI a. This rectangle is specified at a relative offset from the edge of the image, so that the default value of zero corresponds to the entire image. b. The offset is specified as a percentage of the image width and height with 16-bit precision. This method ensures that the metadata remains constant even when the image is resized. c. If the offset results in a zero width and / or height of the ROI, then the single pixel in the top left corner is considered the ROI. 2) Average brightness of ROI This value acts as an offset for the color volume metadata to optimize the presentation of the ROI. As the ROI is expanded to fill the entire screen, the color volume mapping preserves the majority of the contrast in the ROI. b. Calculated the same as L1.Mid, but using only pixels containing the ROI. The value stored in the metadata is an offset from the full screen value, ensuring a value of zero will fall back to using the L1.Mid value. i.L12.MidOffset=ROI.Mid-L1.Mid-L3.Mid Note: L1.Mid may be calculated as the average of the image's PQ encoded maxRGB values, or as the average luminance. maxRGB is the maximum of a pixel's color component values ​​{R,G,B}. L3.Mid represents an offset to the "Mid" PQ value present in the L1 metadata (L1.Mid). c. The playback device smoothly interpolates this offset based on the relative size of the ROI being displayed. The value used by the display device is L1.Mid+L3.Mid+f(L12.MidOffset) where f denotes the interpolation function. 3) Mastering viewing distance can be specified if necessary. a. Specify the mastering viewing distance as a function of the reference viewing distance. This is used to ensure that the image is not scaled when viewed from the same mastering viewing distance. b. The default viewing distance is 17.7613 degrees (2°, the ITU-R reference viewing angle for Full HD content). * atan(0.5 / 3.2)). Closer distances (e.g., 0.5) correspond to a viewing angle of 17.7613 / 0.5=35.5226. However, for simplicity and to ensure that different aspect ratios are calculated equivalently, trigonometric functions have been omitted. c. The range is 3 / 32 to 2 in increments of 1 / 128. The metadata is an 8-bit integer in the range 11 to 255, which is Mastering viewing distance = (L12.MVD+1) / 128 This is used to calculate the picture height. d. If not specified or in the range 0 to 10, the default is 127, or a mastering viewing distance equal to the reference viewing distance. For new content, this value may be a smaller value, such as 63, to indicate that the mastering distance is equal to the reference viewing distance. 4) If necessary, the distance of the object from the camera (greater than half of the ROI) can be specified. This can be used to enhance the "look around" feature by panning and zooming the image at the correct speed in response to changes in viewing position, with distant objects panned and scaled at a slower rate than closer objects. 5) If necessary, it is possible to identify the "viewer's intended path of movement." a. This can be used to guide "Ken Burns" effects during playback even when viewer tracking is not available or can't be enabled. An example would be a photo frame. This feature allows the artist to specify the desired effect, whether zooming in to or out from a subject, by guiding the Ken Burns direction for panning and scaling, without the possibility of the main subject being accidentally cut out of the image due to cropping. 6) You can specify separate layers for graphics or overlays as needed. a. This allows graphics to be scaled and composited with images independently of image scaling. Prevents important overlays or graphics from being cropped or scaled when the image is cropped or scaled. Preserves the "creative intent" of graphics and overlays.

[0054] It is preferred to specify only a single level 12 field in a bitstream. If multiple fields are specified, only the last one is considered valid. The metadata value can change from frame to frame, which is necessary to track the ROI within the video sequence. The fields are extensible, allowing additional fields to be added for future versions.

[0055] An example of software (pseudocode) that may implement an embodiment is provided below.

[0056] [Table 2] JPEG0007747667000003.jpg181157JPEG0007747667000004.jpg189157JPEG0007747667000005.jpg164157

[0057] Further embodiments of the regeneration behavior are also described in the appendix.

[0058] Adapting the display based on the observer's relative position to the display When a scene is viewed through a window, it appears differently depending on the observer's position relative to the window. For example, if the observer is closer to the window, a larger portion of the scene outside is visible than if the observer is farther away. Similarly, as the viewer moves laterally, part of the image appears on one side of the window, while other parts of the image are hidden on the other side of the window.

[0059] If a lens (magnifying or reducing) is replaced by a window, the outside scene will appear larger (zoomed in) or smaller (zoomed out) than the actual scene, but the observer will still have the same experience as if they were moving relative to the window.

[0060] In contrast, when a viewer views a digital image reproduced on a conventional display, the image does not change depending on the viewer's position relative to the display. In one embodiment, to address the difference between the experience of viewing through a window and the experience of viewing a conventional display, the image on the display adapts depending on the viewer's position relative to the display, so that the viewer sees the rendered scene as if they were viewing it through a window. In such an embodiment, a content creator (e.g., a photographer, mobile user, or filmmaker) can better convey or share with an audience the experience of being in the actual scene.

[0061] In one embodiment, an example process for adapting image display depending on the position of a viewer relative to the display may include the following steps. Obtain an image using a capture device such as a camera, or load it from disk. Identify the region of interest (ROI) on the captured image. Sending images and ROIs to the receiving device. · Determine the viewer's position relative to the display on the receiving device. Displaying images on a display according to ROI metadata, screen aspect ratio, and viewer's position relative to the screen.

[0062] Each of these steps is described in more detail below. For example, but not limited to, the image may be obtained using a camera, or by loading from disk or memory, or by capturing from decoded video. This process may be applied to a single picture or frame, or to a sequence of pictures or frames.

[0063] A region of interest is a region within an image, typically corresponding to the most important portion of the image that should be preserved across a wide range of display and viewing structures. A region of interest, e.g., a rectangular region of an image, may be defined manually or interactively, e.g., by allowing a user to draw a rectangle on the image with a finger, mouse, pointer, or some other user interface. In some embodiments, an ROI may be automatically generated by identifying a particular object (e.g., a face, a car, a license plate, etc.) within an image. An ROI may also be automatically tracked across multiple frames in a video sequence.

[0064] There are many ways to estimate the viewer's distance and relative position with respect to the screen. The following methods are provided by way of example only and are not limiting: In one embodiment, an imaging device near or integrated into the display bezel, such as an internal camera or an external webcam, may be used to determine the viewer's position. Images from the camera may be analyzed to find the location of a person's head within the image. This is done using conventional image processing techniques commonly used for "face detection," camera autofocus, autoexposure, or image annotation. There is ample literature and technology available for skilled users to perform face detection and isolate the location of an observer's head within an image. The return value of the face detection process is a rectangular bounding box of the viewer's head or a single point corresponding to the center of that bounding box. In one embodiment, finding the viewer's position may be further improved by any of the following techniques:

[0065] a) Temporal filtering. This type of filtering can reduce measurement noise in the estimated head position, thus providing a smoother, more continuous experience. IIR filters can reduce noise, but the filtered position lags behind the actual position. Kalman filtering aims to both reduce noise and predict the actual position based on several previous measurements. Both of these techniques are well known in the art.

[0066] b) Eye position tracking. Once the head position has been identified, it is possible to improve the estimated viewer position by locating the viewer's eyes. This may involve further image processing, or may skip the head-locating step entirely. The viewer's position may then be updated to indicate a midpoint between the two eyes or a single eye position.

[0067] c) Faster updated measurements: To obtain the most accurate current location of a viewer, faster (more frequent) measurements are desirable.

[0068] d) Depth Cameras. To improve the estimation of the distance from the camera to the viewer, special cameras that measure distance directly can be used. Some examples are time-of-flight cameras, stereo cameras, or structured light. Each of these is known in the art and is commonly used to estimate the distance of objects in a scene relative to the camera.

[0069] e) Infrared cameras. Infrared cameras can be used to improve performance over a wide range of ambient light (e.g., dark rooms). These may measure facial heat directly or may measure reflected infrared light from an infrared transmitter. Such devices are commonly used in the security field.

[0070] f) Distance Calibration. The distance between the viewer and the camera can be estimated by image processing algorithms. This distance can then be calibrated to the screen-to-viewer distance using the known displacement between the camera and the screen. This ensures that the displayed image is correct for the estimated viewer position.

[0071] g) Gyroscopes, which are widely available in mobile devices and can easily provide information about the display orientation (e.g., portrait or landscape mode) or the relative movement of the handheld display with respect to the viewer.

[0072] It has already been described herein how the rendered image can be adapted to the region of interest and the assumed position of the viewer, taking into account the ROI metadata and the characteristics (aspect ratio) of the screen. In one embodiment, if the assumed position of the viewer is replaced by an estimated position computed by any of the techniques above, the rendering of the display can be adjusted using one or more of the following techniques. Examples are shown in Figures 6A to 6G.

[0073] 6A shows an original image 610 captured without regard to ROI metadata and displayed on a display 605 (e.g., a portrait-oriented mobile phone or tablet). By way of example, rectangle 615 ("Hi") may represent a region of interest (e.g., a picture frame, a banner, a face, etc.).

[0074] Figure 6B shows an example of rendering an image 610 by considering a reference viewing position (e.g., picture height 3.2, centered horizontally and vertically on the screen), where appropriate scaling enlarges the ROI (615) while preserving the original aspect ratio.

[0075] In one embodiment, as shown in Figure 6C, as the viewer moves further away from the screen or the display moves away from the viewer, the image zooms in more. This creates the same effect as if the viewer were moving away from a window and the outside scene appeared restricted. Similarly, as shown in Figure 6D, as the viewer moves closer to the screen or the display moves closer to the viewer, the image zooms out more. This creates the same effect as if the viewer were moving closer to a window and the outside scene appeared larger.

[0076] In one embodiment, as shown in Figure 6E, as the viewer moves to the right of the display, or as the display moves to the viewer's left, the image shifts to the right. This creates the same effect as if the viewer were looking through a window to the left. Similarly, as shown in Figure 6F, as the viewer moves to the left of the display, or as the display moves to the viewer's right, the image shifts to the left. This creates the same effect as if the viewer were looking through a window to the right.

[0077] Similar adjustments may be made if the viewer (or display) moves up or down, or a combination of these various movements. Generally, the image moves by an amount based on the assumed or estimated depth of the scene in the image. If the depth is very shallow, the movement will be less than the viewer's actual movement. If the depth is very deep, the movement can be equal to the viewer's movement.

[0078] In one embodiment, all of the above operations may be adjusted depending on the aspect ratio of the display. For example, in landscape mode, the original image (610) is scaled and cropped so that the ROI 615 is centered in the viewer's field of view, as shown in Figure 6G. The displayed image (610-ROI-B) may then be further adjusted depending on the viewer's position relative to the screen, as described above.

[0079] In one embodiment, the ROI may move by smaller and smaller amounts as it approaches the edge of the image. This prevents the ROI from suddenly reaching the edge and becoming stuck. Therefore, from near the reference position (e.g., 610-ROI-A), the image may adjust naturally as if viewed through a window, but the rate of movement may decrease as the edge of the captured image approaches. It is desirable to prevent an abrupt boundary between natural movement and no movement. It is preferable to smoothly scale the rate of movement while the viewer moves toward the maximum possible amount.

[0080] Optionally, the image may be slowly re-centered over time and moved to the viewer's actual viewing position. This may allow for greater range of movement and motion from the actual viewing position. For example, if a viewer begins viewing from a reference position and then moves toward the bottom left corner of the screen, the image may be adjusted to pan upward and to the right. The viewer is not permitted to move further toward the bottom left corner from this new viewing position. This optional feature allows the view to return to a centered position over time, thereby restoring the viewer's range of movement in all directions. Optionally, the amount of image shifting and / or scaling based on the viewer's position may be determined in part by additional distance metadata. The additional distance metadata describes the viewer's distance (depth) from the primary subject containing the ROI. To emulate the experience of viewing through a window, the image should adapt less for closer distances than for farther distances.

[0081] In another embodiment, the adjusted image may be used to create an overlay image, as desired, as described above, where the position of the overlay image remains fixed. This prevents important information in the overlay image from being visible at all times and from all viewing positions, and may further enhance the immersion and realism of the experience as a semi-transparent overlay printed on a window.

[0082] In another embodiment, as described above, the color volume mapping may be adjusted as needed depending on the actual area of ​​the displayed image. For example, if the viewer moves to the right to view a bright object in the scene, the metadata describing the dynamic range of the image may be adjusted to reflect that brighter image. The rendered image may therefore be mapped slightly darker by the tone mapping, thereby mimicking the adaptation effect a human observer experiences when viewing a scene through a window.

[0083] Referring to the pseudocode above for "intelligent zoom" (for a fixed distance between the viewer and the screen), the following modifications are required in one embodiment to allow intelligent zoom by adaptation of the viewer position: a) Instead of using a hypothetical reference viewing distance, the actual distance from the viewer to the screen (measured by any known technique) is used to calculate the above "viewerDistance" and "zoomFactor" parameters to generate the scaled image. b) Shifting the scaled image across (x,y) coordinates depending on the viewer's position on the screen. By way of example and not limitation, the viewer's position may be calculated with reference to the (x,y) coordinates of their eyes. In pseudo-code, this may be expressed as follows: [Table 3]

[0084] 7 illustrates an example process flow for displaying an image using a display adaptation process according to one embodiment. In step 705, a device receives an input image and parameters associated with a region of interest within the image. If image adaptation (e.g., "intelligent zoom" as described herein) is not enabled, the device generates an output image in step 715 without considering the ROI metadata. However, if image adaptation is enabled, the device may generate an output version of the input image in step 710 using the ROI metadata and display parameters (e.g., aspect ratio), thereby highlighting the ROI in the input image. Additionally, in some embodiments, the device may further adjust the output image depending on the viewer's relative position with respect to the display and the viewer's distance from the display. The output image is displayed in step 720.

[0085] FIG. 8 illustrates an example of a data processing system 800. Data processing system 800 may be used in an embodiment. For example, system 800 may be implemented to provide a content creation or editing system that performs the method of FIG. 2A or FIG. 5, or to provide a playback device that performs the method of FIG. 2B. It should be noted that while FIG. 8 illustrates various components of a device, it is not intended to represent a particular structure or manner in which the components are interconnected, as such details are not relevant to this disclosure. It should also be understood that network computers and other data processing systems or other consumer electronic devices having fewer components or perhaps more components may be used in embodiments of the present disclosure.

[0086] As shown in FIG. 8, device 800 is in the form of a data processing system and includes a bus 803. Bus 803 is connected to microprocessor(s) 805, ROM (Read Only Memory) 807, volatile RAM 809, and non-volatile memory 811. Microprocessor(s) 805 may retrieve instructions from memories 807, 809, and 811 and execute the instructions to perform the operations described above. Microprocessor(s) 805 may include one or more processing cores. Bus 803 interconnects these various components and further interconnects these components 805, 807, 809, and 811 with a display controller and display device 813 and peripheral devices. Peripheral devices include, for example, input / output (I / O) devices 815, which may be a touch screen, a mouse, a keyboard, a modem, a network interface, a printer, and other devices known in the art. Typically, input / output devices 815 are connected to the system through an input / output controller 810. Volatile RAM (Random Access Memory) 809 is typically implemented as dynamic RAM (DRAM), which requires continuous power to refresh or maintain the data in memory.

[0087] Non-volatile memory 811 is typically a magnetic hard drive, magneto-optical drive, optical drive, DVD RAM, flash memory, or other type of memory system that retains data (e.g., large amounts of data) even after power is removed from the system. Typically, non-volatile memory 811 is also random access memory, although this is not required. FIG. 8 illustrates non-volatile memory 811 as a local device directly connected to other components in the data processing system. However, it is understood that embodiments of the present disclosure may utilize non-volatile memory remote to the system, such as a network storage system connected to the data processing system via a network interface, such as a modem, Ethernet interface, or wireless network. Bus 803 may include one or more buses connected to each other through various bridges, controllers, and / or adapters, as is well known in the art.

[0088] Portions of the above may be implemented by logic circuitry, such as special-purpose logic circuitry, or by a microcontroller or other form of processing core that executes program code instructions. Thus, the processes taught above may be executed by program code, such as machine-executable instructions, which cause a machine that executes those instructions to perform a particular function. In this context, a "machine" may be a machine that translates intermediate-form (or abstract) instructions into processor-specific instructions (e.g., an abstract execution environment such as a "virtual machine" (e.g., a Java Virtual Machine), an interpreter, a Common Language Runtime, or a high-level language virtual machine), and / or electronic circuitry, such as a general-purpose processor and / or a special-purpose processor, provided on a semiconductor chip (e.g., a "logic circuit" implemented with transistors) designed to execute instructions. The processes taught above may also be executed by electronic circuitry (either as a substitute for a machine or in combination with a machine) designed to execute the process (or portions thereof), in which case program code is not executed.

[0089] The present disclosure further relates to apparatus for performing the operations described herein. The apparatus may be specially constructed for the required purposes or may include a general-purpose device selectively activated or reconfigured by a computer program stored on the device. Such computer program may be stored on a non-transitory computer-readable storage medium, such as, but not limited to, floppy disks, optical disks, disks of all types, including CD-ROMs and magneto-optical disks, DRAM (volatile), flash memory, read-only memory (ROM), RAM, EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each connected to the device's bus.

[0090] A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). For example, non-transitory machine-readable media include read-only memory (ROM), random-access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.

[0091] An article of manufacture may be used to store program code. The article of manufacture storing the program code may be embodied as one or more non-transitory memories (e.g., but not limited to, one or more flash memories, random access memories (static, dynamic, or other), optical disks, CD-ROMs, DVD-ROMs, EPROMs, EEPROMs, magnetic or optical cards, or other types of machine-readable media suitable for storing electronic instructions). The program code may also be downloaded from a remote computer (e.g., a server) to a requesting computer (e.g., a client) by a data signal embodied in a propagation medium (e.g., via a communications link (e.g., network connection)) and then stored in non-transitory memory (e.g., DRAM, flash memory, or both) within the client computer.

[0092] The above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a device memory. These algorithmic descriptions and representations are the tools used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. Algorithms are herein, and generally, conceived as self-consistent sequences of operations leading to a desired result. The operations require physical manipulations of physical quantities. These quantities usually, though not necessarily, take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0093] It should be remembered, however, that all these and similar terms are associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. As is clear from the above description, unless otherwise indicated, throughout this description, descriptions made using terms such as "receive," "determine," "send," "finish," "wait," and "modify" will be understood to refer to the acts and processes of devices or similar electronic computing devices. These devices or similar electronic computing devices manipulate data represented as physical (electronic) quantities in the device's registers and memory, and transform that data into other data also represented as physical quantities in the device's memory or registers, or other similar information storage, transmission, or display devices.

[0094] The processes and displays presented herein are not inherently related to any particular device or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the operations described herein. The required structure for a variety of these systems will appear from the description below. Further, the present disclosure is not described with reference to any particular programming language. It will be understood that a variety of programming languages ​​may be used to implement the teachings of the disclosure set forth herein.

[0095] While certain exemplary embodiments have been described hereinabove, it will be apparent that various modifications may be made to these embodiments without departing from the broader spirit and scope of the following claims. The specification and drawings are therefore to be considered illustrative rather than limiting of the disclosure.

[0096] Various aspects of the present invention can be understood from the following exemplary embodiments (EEE). EEE1. 1. A machine-implemented method comprising: Selecting the original aspect ratio (AR) for the image canvas used to create content; determining, within at least a first scene within the content on the image canvas, a first position of a first object within the at least first scene; determining, based on the determined position of the first object, first metadata specifying how to adapt the first portion for playback on the display device with AR that differs from the original AR; storing the first metadata when the first metadata and the content are used or transmitted for use during playback; A method comprising: EEE2. The method of EEE1, wherein the original AR is substantially square. EEE3. The method of EEE2, wherein the substantially square is either (1) closer to a square than AR16:9, i.e., the original AR has a length to height ratio less than a 16:9 ratio (16 / 9) but greater than or equal to 1:1, or (2) greater than a 9:16 ratio but less than 1:1 when portrait mode is preferred, and the original AR changes during content. EEE4. determining a plurality of objects including the first object for a plurality of scenes including the first scene; For each of the objects in the plurality of scenes, determining a corresponding position within a corresponding scene; The method of any one of EEE1 to 3, further comprising: EEE5. a subject is determined for each scene within the plurality of scenes, the method comprising: The method of claim 8, further comprising displaying a preview showing how cropping at different aspect ratios will be performed based on the metadata. EEE6. EEE1 to EEE5. The method of any one of EEE1 to EEE5, wherein the first metadata guides asymmetric cropping on a playback device to extend from the first subject in the first scene to accommodate different AR when adapted for playback. EEE7. A non-transitory machine-readable medium storing executable program instructions that, when executed by a data processing system, cause the data processing system to perform a method described in any one of EEE1 to 6. EEE8. A data processing system having a processing system and a memory, the data processing system being configured to perform a method according to any one of EEE1 to EEE6. EEE9. 1. A machine-implemented method comprising: receiving content including image data for at least a first scene and first metadata associated with the first scene, the first metadata specifying how to adapt a first position of a first subject in the first scene for playback on a display device having an aspect ratio (AR) different from a native aspect ratio, the first scene being created on an image canvas having the native aspect ratio; Adapting an output to the aspect ratio of the display device based on the first metadata; A method comprising: EEE10. The method of claim 8, wherein the original AR is substantially square. EEE11. The method of claim 30, wherein the substantially square shape is closer to a square than AR16:9, i.e., the ratio of length to height of the original AR is less than the ratio 16:9 (16 / 9), and the original AR changes during the content. EEE12a. 3. The method of claim 1, wherein the content comprises a plurality of scenes including the first scene, each scene of the plurality of scenes having a determined position for an object in the scene, the object being determined for each scene, adaptation for different AR being performed for each scene, and tone mapping being performed for the display device on a scene-by-scene or frame-by-frame basis and based on a region of interest within each scene or each frame, and each scene comprising one or more frames. EEE12b. 3. The method of claim 1, wherein the content comprises a plurality of scenes including the first scene, each scene of the plurality of scenes having a determined position for an object of the scene, the object being determined for each scene, adaptation for different AR being performed for each scene, tone mapping being performed for the display device on a scene-by-scene or frame-by-frame basis and based on which relative portions of the adapted images are referred to as regions of interest, and each scene comprising one or more frames. EEE13. 13. The method of claim 9, wherein the first metadata guides asymmetric cropping on a playback device to extend from the first subject in the first scene to accommodate different AR when adapted for playback. EEE14. receiving distance and position parameters relating to a viewer's position relative to the display device; further adapting the output of the first object to the display device based on the distance parameter and the position parameter; The method of EEE9, further comprising: EEE15. The method of claim EEE14, wherein further adapting the output of the first subject to the display device comprises upscaling the output of the first subject when a viewing distance between the viewer and the display device increases, and downscaling the output of the first subject when the viewing distance between the viewer and the display device decreases. EEE16. The method of claim EEE14, wherein further adapting the output of the first subject to the display device comprises shifting the output of the first subject to the left when the display device moves to the right relative to the viewer, and shifting the output of the first subject to the right when the display device moves to the left relative to the viewer. EEE17. receiving graphics data; generating a video output comprising a combination of the graphics data and the adapted output; The method of any one of EEE9 to 16, further comprising: EEE18. 8. The method of any one of EEE9 to 17, wherein the first metadata further comprises syntax elements that define a viewer's intended movement path to guide a Ken Burns related effect during playback. EEE19. A non-transitory machine-readable medium storing executable program instructions that, when executed by a data processing system, cause the data processing system to perform a method according to any one of EEE9 to 18. EEE20. A data processing system having a processing system and a memory, the data processing system being configured to perform a method according to any one of claims 8 to 18. appendix Playback behavior example

[0097] The playback device is responsible for applying a specific reframing depending on the image metadata, the display configuration and optionally user configuration. In an example embodiment, the steps are as follows: 1) Specify the relative viewing distance as a function of the "default viewing distance." Depending on the complexity or version of the implementation, options include: Use the default value of RelativeViewingDistance=1.0. Dynamically adjust it in one of two ways: -Automatic adjustment when resizing the window or putting the picture into picture mode. RelativeViewingDistance=sqrt(WindowWidth 2 +WindowHeight 2 ) / sqrt(DisplayWidth 2 +DisplayHeight 2 ) Manual adjustment through user interactions (such as pinching, scrolling, and sliding bars). Measure the viewing distance using a camera or other sensor and divide the viewer's measured distance (typically in meters) into the default viewing distance specified in a configuration file. RelativeViewingDistance=ViewerDistance / DefaultViewingDistance Note: In some embodiments, it may be necessary to bound the relative viewing distance value within a specific range (e.g., between 0.5 and 2.0). Two example bounding schemes are shown later in this section. 2) Convert the relative viewing distance of the source to a relative angle.

number

number

number

number

number

number

number

number

number

number

Claims

1. 1. A machine-implemented method comprising: receiving content including image data for at least a first scene and first metadata associated with the first scene, the first metadata specifying how to accommodate playback on a display device having an aspect ratio (AR) different from an original aspect ratio for a first position of a first subject in the first scene, the first scene being created on an image canvas having the original aspect ratio; adapting an output to the aspect ratio of the display device based on the first metadata; Including, receiving distance and position parameters relating to a viewer's position relative to the display device; further adapting the output of the first object to the display device based on the distance parameter and the position parameter; further comprising 10. The method of claim 1, wherein the content includes a plurality of scenes including the first scene, each scene of the plurality of scenes having a determined position for an object in the scene, the object being determined for each scene, and adaptation to different AR being performed for each scene, and tone mapping being performed for the display device on a scene-by-scene or frame-by-frame basis and based on which relative portion of the adapted image is referred to as an area of ​​interest, wherein the area of ​​interest includes the first object and each scene includes one or more frames.

2. A machine-implemented method comprising: receiving content including image data for at least a first scene and first metadata associated with the first scene, the first metadata specifying how to accommodate playback on a display device having an aspect ratio (AR) different from an original aspect ratio for a first position of a first subject in the first scene, the first scene being created on an image canvas having the original aspect ratio; adapting an output to the aspect ratio of the display device based on the first metadata; Including, receiving distance and position parameters relating to a viewer's position relative to the display device; further adapting the output of the first object to the display device based on the distance parameter and the position parameter; further comprising the first metadata is in the form of a vector specifying a direction from the first object; if the vectors are equal in all directions from the first object, the first metadata causes a symmetric cropping of the content around the first object in an image canvas based on the aspect ratio and orientation of the display device used to play the content, the symmetric cropping focusing display output on the first object in all cases regardless of the aspect ratio and orientation of the display device; if the vector is not equal in all directions from the first object, the first metadata causes an asymmetric cropping of the content in the image canvas relative to the first object, the asymmetric cropping based on the aspect ratio, orientation of the display device used to play the content, and the vector specifying how to asymmetrically crop the content relative to the first object, wherein the asymmetric cropping maintains the first object at a position of a cropped view in the image canvas that corresponds to the first position of the first object in the first scene, and maintains the first object at the position of the cropped view in the image canvas regardless of the orientation and the aspect ratio of the display device; 10. The method of claim 1, wherein the content includes a plurality of scenes including the first scene, each scene of the plurality of scenes having a determined position for an object in the scene, the object being determined for each scene, and adaptation to different AR being performed for each scene, and tone mapping being performed for the display device on a scene-by-scene or frame-by-frame basis and based on which relative portion of the adapted image is referred to as an area of ​​interest, wherein the area of ​​interest includes the first object and each scene includes one or more frames.

3. The method of claim 1 or 2, wherein the original AR is substantially square.

4. 4. The method of claim 3, wherein the substantially square shape is closer to a square than AR 16:9, i.e., the ratio of length to height of the original AR is less than the ratio of 16:9 (16 / 9), and the original AR changes during the content.

5. 5. The method of claim 1, wherein the first metadata guides asymmetric cropping on a playback device to extend from the first subject in the first scene to accommodate different AR when adapted for playback.

6. 6. The method of claim 1, wherein further adapting the output of the first subject to the display device comprises upscaling the output of the first subject when a viewing distance between the viewer and the display device increases, and downscaling the output of the first subject when the viewing distance between the viewer and the display device decreases.

7. 7. The method of claim 1, wherein further adapting the output of the first subject to the display device comprises shifting the output of the first subject to the left when the display device moves to the right relative to the viewer, and shifting the output of the first subject to the right when the display device moves to the left relative to the viewer.

8. receiving graphics data; generating a video output comprising a combination of the graphics data and the adapted output; The method of any one of claims 1 to 7, further comprising:

9. The method of claim 1 , wherein the first metadata further comprises syntax elements that define a viewer's intended movement path to guide a Ken Burns-related effect during playback.

10. A non-transitory machine-readable medium storing executable program instructions which, when executed by a data processing system, cause the data processing system to perform the method of any one of claims 1 to 9.

11. A data processing system having a processing system and a memory, the data processing system being configured to perform the method of any one of claims 1 to 9.

12. 1. A machine-implemented method comprising: Selecting an original aspect ratio (AR) for an image canvas used for content creation; determining, within at least a first scene within content on the image canvas, a first position of a first object within the at least first scene; determining first metadata specifying how to adapt playback on the display device with AR different from the original AR for the first location based on the determined position of the first object and based on a distance between a viewer and a display device, wherein the content includes a plurality of scenes including the first scene, each scene having a determined position for an object for that scene, the object being determined for each scene, and adaptation to different AR being adapted to occur on a scene-by-scene basis, and tone mapping being adapted for the display device on a scene-by-scene or frame-by-frame basis and based on which relative portion of the adapted image is referred to as a region of interest, wherein the region of interest includes the first object, and each scene includes one or more frames; storing the first metadata when the first metadata and the content are used or transmitted for use during playback; A method comprising:

13. A machine-implemented method comprising: Selecting an original aspect ratio (AR) for an image canvas used for content creation; determining, within at least a first scene within content on the image canvas, a first position of a first object within the at least first scene; determining first metadata specifying how to adapt playback on the display device with AR different from the original AR for the first location based on the determined position of the first object and based on a distance between a viewer and a display device, wherein the content includes a plurality of scenes including the first scene, each scene having a determined position for an object for that scene, the object being determined for each scene, and adaptation to different AR being adapted to occur on a scene-by-scene basis, and tone mapping being adapted for the display device on a scene-by-scene or frame-by-frame basis and based on which relative portion of the adapted image is referred to as a region of interest, wherein the region of interest includes the first object, and each scene includes one or more frames; storing the first metadata when the first metadata and the content are used or transmitted for use during playback; Including, the first metadata is in the form of a vector specifying a direction from the first object; if the vectors are equal in all directions from the first object, the first metadata causes the image canvas to symmetrically crop the content around the first object based on the aspect ratio and orientation of the display device used to play the content, the symmetric cropping focusing display output on the first object in all cases regardless of the aspect ratio and orientation of the display device; and if the vector is not equal in all directions from the first object, the first metadata causes an asymmetric cropping of the content in the image canvas relative to the first object, the asymmetric cropping based on the aspect ratio, orientation of the display device used to play the content, and the vector specifying how to asymmetrically crop the content relative to the first object, wherein the asymmetric cropping maintains the first object at a position of a cropped view in the image canvas that corresponds to the first position of the first object in the first scene, and maintains the first object at the position of the cropped view in the image canvas regardless of the orientation and the aspect ratio of the display device.

14. Displaying the first object for different distances between the viewer and the display device.

14. The method of claim 12 or 13, further comprising providing different zoom ratios for different focal lengths.

15. The method of any one of claims 12 to 14, wherein the original AR is substantially square.

16. 16. The method of claim 15, wherein the substantially square shape is either (1) closer to a square than AR 16:9, i.e., the original AR has a length to height ratio less than a 16:9 ratio (16 / 9) but greater than or equal to 1:1, or (2) greater than a 9:16 ratio but less than 1:1 when portrait mode is preferred, and the original AR changes during content.

17. determining a plurality of objects including the first object for a plurality of scenes including the first scene; For each of the objects in the plurality of scenes, determining a corresponding position within a corresponding scene; 17. The method of any one of claims 12 to 16, further comprising:

18. a subject is determined for each scene within the plurality of scenes, the method comprising:

18. The method of claim 17, further comprising displaying a preview showing how cropping at different aspect ratios will be performed based on the metadata.

19. 19. The method of claim 12, wherein the first metadata guides asymmetric cropping on a playback device to extend from the first subject in the first scene to accommodate different AR when adapted for playback.

20. A computer program comprising program instructions which, when executed by a data processing system, cause said data processing system to carry out a method according to any one of claims 12 to 19.

21. A data processing system having a processing system and a memory, the data processing system being configured to perform the method of any one of claims 12 to 19.

Citation Information

Patent Citations

  • Method and apparatus for providing information to a user observing a multi view content

    EP3416381A1

  • Display with automatic screen parameter adjustment based on the position of a detected viewer

    GB2467898A

  • Image processing device, image processing method, program, and storage medium

    JP2016063323A

  • Image processing device, control method for the same and program

    JP2016066892A

  • Dynamically cropping digital content for display in any aspect ratio

    US20170249719A1