Rendering generative media objects based on priority data

WO2026169813A1PCT designated stage Publication Date: 2026-08-13DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-08-13

Smart Images

  • Figure US2026014011_13082026_PF_FP_ABST
    Figure US2026014011_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for rendering generative media objects. One of the methods includes receiving input media; identifying, from the input media, a first set of media objects corresponding to objects generated by one or more AI generative models; classifying the first set of media objects into a plurality of object types, wherein the classifying is based on perceptual criteria; determining priority data for respective objects in the first set of media objects based on the respective determined object types, wherein a first object type associated with a first perceptual criterion is given a higher priority than a second object type associated with a second perceptual criterion that is determined to be less noticeable than the first perceptual criterion; and rendering generative media objects from respective objects in the first set of media objects based on their respective priority data.
Need to check novelty before this filing date? Find Prior Art

Description

D25016W001RENDERING GENERATIVE MEDIA OBJECTS BASED ON PRIORITY DATA CROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to U.S., Patent Application No. 63 / 754,382, filed February 5, 2025, which is incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] This invention relates to rendering generative media objects based on priority data.BACKGROUND

[0003] Generative artificial intelligence (Al) systems can generate media content that include one or more media objects. The generated media content can include anomalies, causing the output to deviate from accurately representing the physical world or the scene it is intended to depict. The generated media content can be displayed on one or more client devices for users to view, listen to, or otherwise experience. When these anomalies are noticeable, they can detract from the user’s overall experience with the generated media content.

[0004] Systems that generate media content can be configured to render that content by converting the underlying data into a displayable format. However, when the media content is produced by a generative Al system, rendering processes should be designed to reduce the visibility of artifacts or irregularities that may be present in the generated output. Accordingly, there is a need for rendering techniques that present Al-generated media in a manner that minimizes the perceptibility of such anomalies.SUMMARY

[0005] This specification relates to rendering generative media objects using a prioritization framework.

[0006] A system can be configured to process input media that includes generative media objects to render the generative media objects, e.g., for display on a device. Generative media objects can include media objects generated using a generative artificial intelligence (Al) system. Rendering generative media objects can consume large amounts of computational resources. Generative media objects that were generated using generative Al can include anomalies or artifacts that reduce the perceived authenticity orD25016W001realness of the generative media objects. For example, viewers or listeners of the generative media objects may be able to determine that the objects have been generated using Al by detecting the anomalies or artifacts included in the objects.

[0007] Implementations of the present disclosure include systems, apparatus, and methods to render generative media objects that are likely to appear authentic, e.g., to a consumer of the rendered generative media objects, in a way that conserves computational resources. As used throughout this specification, to “render” a generative media object can mean to convert raw data representing the generative media object into data that can be used directly to cause the generative media object to be displayed, e.g., on a device. As used throughout this specification, a “consumer” of rendered media objects is an entity, such as a person, that interacts with the rendered media objects in a manner intended according to a modality of the media objects. For example, a consumer can be an entity that views the rendered media objects, listens to the rendered media objects, or both.

[0008] The system classifies media objects from a received input into object types based on perceptual criteria that have different levels of noticeability, e.g., to a consumer of media objects. The system generates priority data based on the classification. For example, media objects that are related to perceptual criteria that are more likely to be noticeable can be given higher priority than those related to perceptual criteria that are less likely to be noticeable.

[0009] The system renders the generative media objects from the input media based on the priority data. For example, the system can allocate more computational resources to rendering objects assigned higher priority, so as to increase the likelihood that objects related to more noticeable perceptual criteria will be perceived as real. Increasing a likelihood that more noticeable objects are perceived as real can in turn increase the likelihood that the generative media objects overall are perceived as real. The system can limit its allocation of resources to rendering objects assigned lower priority to conserve computational resources while maintaining a high likelihood that the objects will be perceived as real.

[0010] In general, one aspect of the subject matter described in this specification can be embodied in methods that include the actions of receiving input media; identifying, from the input media, a first set of media objects corresponding to objects generated by one or more Al generative models; classifying the first set of media objects into a plurality of object types, wherein the classifying is based on one or more perceptual criteria;D25016W001determining priority data for respective objects in the first set of media objects based on the respective determined object types, wherein a first object type associated with a first perceptual criterion is given a higher priority than a second object type associated with a second perceptual criterion that is determined to be less noticeable than the first perceptual criterion; and rendering generative media objects from respective objects in the first set of media objects based on their respective priority data.

[0011] Other implementations of this aspect include corresponding computer systems, apparatus, computer program products, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods. A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.

[0012] The foregoing and other implementations can each optionally include one or more of the following features, alone or in combination. In some implementations, the rendering comprises allocating computational resources to render respective objects according to the priority data. In some implementations, allocating the computational resources comprises controlling one or more of following options based on the priority data: order of computation; temporal resolution; spatial resolution; latent space bit-depth; or quantity of neural network layers utilized.

[0013] In some implementations, allocating computational resources to rendering respective objects comprises allocating more computational resources to rendering respective objects with higher priority according to the respective priority data than are allocated to rendering respective objects with lower priority according to the respective priority data.

[0014] In some implementations, the methods further include identifying, from the input media, a second set of media objects that are pre-rendered and are different from the first set of media objects. In some implementations, the methods further include combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects to generate a media file; and causing storage or transmission of the media file. In some implementations, combining data corresponding to the rendered generative media objects and data corresponding to the second set ofD25016W001media objects comprises combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects with one or more alpha maps.

[0015] In some implementations, the method is executed at a client device. In some implementations, the methods further include identifying a second set of media objects that are pre-rendered and are different from the first set of media objects; combining the generative media objects with the second set of media objects to obtain combined media data; and causing playback of the combined media data using one or more media output devices of the client device.

[0016] In some implementations, the plurality of object types includes audio object types and / or video object types; and determining the priority data comprises determining audio priority data for the audio object types and / or determining video priority data for the video object types. In some implementations, the plurality of object types includes multimodal object types, and determining the priority data comprises determining multimodal priority data for the multimodal object types. In some implementations, the input media is received wirelessly from an audio / video media transmitter.

[0017] This specification uses the term “configured to” in connection with systems, apparatus, and computer program components. That a system of one or more computers is configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform those operations or actions. That one or more computer programs is configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform those operations or actions. That special-purpose logic circuitry is configured to perform particular operations or actions means that the circuitry has electronic logic that performs those operations or actions.

[0018] The subject matter described in this specification can be implemented in various implementations and may result in one or more of the following advantages.

[0019] The techniques described in this specification can reduce the number of computational resources used to render generative media objects that were received as part of an input media, e.g., from a remote transmitter. The computational resources are reduced by allocating computational resources to render respective objects according to their priority data. For example, rather than using large numbers of computational resources to render all media objects that are identified in the input media, a systemD25016W001employing the disclosed techniques can allocate substantial computational resources only to objects that are assigned a high priority according to the priority data. This allocation can result in an overall reduction in computational resources consumed.

[0020] Furthermore, because the priority data is derived based on determined object types that are associated with perceptual criteria, the likelihood that a consumer notices anomalies in the generated output is reduced, resulting in a more realistic depiction of the input media to the consumer. As used throughout this specification, a “consumer” of rendered media objects is an entity, such as a person, that can interact with the rendered media objects in a manner intended according to the modality of the media objects. For example, a consumer can be an entity that views the rendered media objects, listens to the rendered media objects, or both.

[0021] A first object type can be assigned a higher priority than a second object type if the first object type is associated with a first perceptual criterion that is more likely to be noticeable (c.g., to a consumer of the rendered generative media objects) than a second perceptual criterion associated with the second object type. The system can allocate more computational resources to rendering the first object type, e.g., to increase a likelihood that first object type appears realistic. Since the first perceptual criterion is more likely to be noticeable, allocating more computational resources to rendering the first object type can in turn increase a likelihood that the generative media objects rendered by the system overall appear realistic. For example, allocating more computational resources to rendering the first object type can decrease a likelihood that anomalies included in rendered generative media objects are noticeable.

[0022] The details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0023] FIG. 1 is a block diagram of an example media rendering system according to implementations of the present disclosure.

[0024] FIG. 2 is a flow diagram of an example media rendering system according to implementations of the present disclosure.D25016W001

[0025] FIG. 3A is an example video frame including video objects with anomalies that are likely to be noticeable.

[0026] FIG. 3B is an example ranking of object types, according to implementations of the present disclosure.

[0027] FIG. 4 is a flow chart of an example process for rendering generative media objects, according to implementations of the present disclosure.

[0028] FIG. 5 shows an example of a computing device and example of display devices that can be used to implement the techniques described here.

[0029] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0030] FIG. 1 is a block diagram of an example media rendering system 100. The media rendering system 100 is configured to render generative media objects based on input media that is received by the system 100. For example, the generative media objects can include one or more of: video objects, image objects, or audio objects. The system 100 can be configured to render the generative media objects in a way that conserves computational resources, improves perceived authenticity of the rendered media objects, or both.

[0031] The system 100 includes an identification and separation engine 102. The identification and separation engine 102 is configured to process input media 101 to identify different types of media objects included in the input media 101 and separate the different types of media objects. For example, the input media 101 can include data representing media objects. Media objects can have any of a variety of content types, such as image, video, audio, or text. A media object of the input media 101 can be a separable aspect of the input media 101. Examples of media objects can include objects that are represented visually in images or videos; backgrounds or other aspects of environments represented in images or videos; one or more images; one or more video frames of a video; one or more clips of audio content; portions of text; or any combination of these.

[0032] The identification and separation engine 102 separates prerendered media objects 103 that are included in the input media 101 from generative media objects 105 that are included in the input media 101. The prerendered media objects 103 can include data representing media objects that have already been rendered, e.g., such that the data representing the prerendered media objects 103 can be used directly to cause display ofD25016W001the media objects on a device. For example, an external rendering system can have previously rendered the prerendered media objects 103. An example prerendered media object 103 is an image or a video frame that is captured by a camera from a real-world scene.

[0033] The generative media objects 105 can include data representing media objects that have been generated by an external generative Al system and have not yet been rendered. For example, the generative media objects 105 can include data representing media objects that is not yet in a format that can be used directly to cause display of the media objects. An example generative media object 105 is an image or video frame generated based on two temporally consecutive prerendered media objects 103 — such as two sequential camera-captured images or frames — to represent a scene that occurs between them. For instance, if the camera captures images every tenth of a second, a generative media object 105 may be generated to depict the moment between two such captures.

[0034] In some implementations, the input media 101 can include other data in addition or instead of the prerendered objects. That other data can be used, for example, for rendering purposes, for example, to be used by one or both of the combination engines 108 and 112 described below. Examples of such data include, but is not limited to, objectbased signals such as abstract data objects (e.g. audio essence, frequency content of a discrete audio source contributed in generating the input media 101), associated positional metadata specifying how and / or where the audio essence should be rendered, associated positional metadata specifying how and / or where the generative media objects should be rendered as part of the overall media as to be combined with the prerendered media objects (e.g., at combination engine 112), etc.

[0035] In some implementations, the input media 101 can also include one or more of metadata or alpha maps, as described with reference to FIG. 2. The identification and separation engine 102 can identify the metadata, alpha maps, or both. The identification and separation engine 102 can separate the metadata, alpha maps, or both, from other components of the input media 101. The identification and separation engine 102 can use a data parser to parse the input media 101 to identify the different types of media objects, and optionally metadata and alpha maps, and separate them.

[0036] The system 100 includes a classification engine 104. The classification engine 104 is configured to process the generative media objects 105 to classify the generative media objects 105 into a plurality of object types. The plurality of object types can include object types that are each associated with different respective perceptual criteria.D25016W001

[0037] The perceptual criteria can each correspond to a respective probability of being noticeable. The respective probability of being noticeable for a perceptual criterion can be a probability that an anomaly included in an object of an object type associated with the perceptual criterion would be noticed by a consumer of the rendered object. Examples of perceptual criteria are described with reference to FIG. 2. The respective probability of being noticeable associated with each of the perceptual criteria can be determined by any of a variety of factors. In an embodiment, the respective probability of being noticeable can be determined by the amount of time elapsed before a consumer notices an anomaly included in a media object that satisfies the perceptual criterion, as described with reference to FIG. 3B.

[0038] The system 100 includes a priority data generation engine 106. The priority data generation engine 106 is configured to process classified object types 107 received from the classification engine 104 to generate priority data 109. The priority data 109 includes a respective priority assigned to each of the generative media objects 105 based on the object type into (which the media object has been classified,) and the perceptual criteria associated with the object type. The priority assigned to each generative media object can be used for determining an allocation of computational resources for rendering the generative media object, as described below.

[0039] The priority data 109 can have any of a variety of formats. For example, the priority data 109 can include a list of tuples, each tuple including an identification of a respective generative media object and an identification of a priority assigned to the generative media object. In some implementations, the priority can be a numerical value. For example, the priority can be a numerical value on a scale of numbers, where numbers at one end of the scale correspond to high priorities and numbers at the other end of the scale correspond to low priorities. In some implementations, the priority can be a descriptive label that describes the priority assigned to the generative media object (e.g., “high” or “low”).

[0040] The system 100 includes a rendering engine 108. The rendering engine 108 is configured to process the generative media objects 105 to render generative media objects from the generative media objects 105 and based on the priority data 109. As a result, the rendering engine 108 generates rendered generative media objects 111. The rendered generative media objects 111 can include data that represents the same objects represented by the data included in the generative media objects 105. The data included inD25016W001the rendered generative media objects 111 can be in a format that is able to be used directly to cause display of the objects.

[0041] The system 100 includes a combination engine 112. The combination engine 112 is configured to process the prerendered media objects 103 and the rendered generative media objects 111 to generate combined media data 113. The combined media data 113 can include data representing the same media objects represented by the data included in the input media 101. However, the data representing the generative media objects 105 that is included in the combined media data 113 can be in a format that can be used directly to cause display of the media objects (whereas the corresponding data in the input media 101 may not be in this format).

[0042] The combined media data 113 can be used by the system 100 to cause display of the media objects represented by the data on one or more devices. In some implementations, the system 100 can be implemented on a client device, such as a phone, tablet, laptop, computer, or any of a variety of computing devices. The system 100 can generate the combined media data 113 at the client device. The system 100 can be configured to cause playback of the combined media data 113 on one or more media output devices of the client device, such as display screens, speakers, or any of a variety of media output devices, and / or can be configured to transmit the combined media data 113 to another device, such as a remote display.

[0043] FIG. 2 is a flow diagram 200 illustrating an example media rendering system. The figure illustrates example inputs and outputs processed by, and example operations performed by a media rendering system. The media rendering system can be the media rendering system 100 of FIG. 1, or a similar system with the additional details described below for FIG. 2.

[0044] The media rendering system can receive input media data 202. For example, the input media 101 of FIG. 1 can include or represent the input media data 202. The input media data 202 can include media objects of any of a variety of content types. For example, the input media data 202 can include media objects that represent video content, e.g., video objects 210. The input media data 202 can include media objects that represent audio content, e.g., audio objects 208. The input media data 202 can include metadata 212 for the media objects. For each media object included in the input media data 202, the input media data 202 can include corresponding metadata 212 that describes the media object. The metadata 212 describing a media object can include data describing any of aD25016W001variety of features of the media object, such as type, size, duration, dimensions, provenance, qualitative features, or any combination of these.

[0045] The input media data 202 can include alpha maps 206 for the media objects. An alpha map for a media object can be a set of data elements that each correspond to a respective data element of the media object. The data element of the alpha map can indicate a level of visibility of the corresponding data element of the rendered media object. For example, the media object can be an image, and the data elements can be pixels. An alpha map for the media object can include, for each pixel of the image, a corresponding data element indicating the level of visibility of the pixel in the rendered image. For example, a first data element of the alpha map corresponding to a first pixel can indicate that the first pixel is fully visible in the rendered image, while a second data element of the alpha map corresponding to a second pixel can indicate that the second pixel is invisible or substantially opaque in the rendered image.

[0046] In response to receiving the input media data 202, the media rendering system identifies and separates various components of the input media data 202 in an operation 204. For example, the operation 204 performed by the media rendering system can include separating from one another the video objects 210, audio objects 208, metadata 212, and alpha maps 206 included in the input media data 202. The system can use an identification and separation engine, e.g., the identification and separation engine 102 of FIG. 1 or a separate engine (not shown) in connection with the identification and separation engine 102, to separate the various components. Each of the alpha maps 206, the audio objects 208, and the video objects 210 can processed in subsequent identification and separation operations, described below. The metadata 212 can be stored, e.g., in a memory of the system, while the system processes the alpha maps 206, audio objects 208, and video objects 210. After processing the alpha maps 206, audio objects 208, and video objects 210, the system can access the metadata 212 to use the metadata 212 for classifying the audio objects 208 and the video objects 210, as described below.

[0047] The system processes each of the audio objects 208 and / or the video objects 210, depending on what objects the input media data 202 included. In an operation 218, the system identifies prerendered audio objects 236 and generative audio objects 219 included in the audio objects 208. The operation 218 also includes separating the prerendered audio objects 236 from the generative audio objects 219. In an operation 220, the system identifies prerendered video objects 234 and generative video objects 221D25016W001included in the video objects 210. The operation 220 also includes separating the prerendered video objects 234 from the generative video objects 221.

[0048] In some implementations, the prerendered objects and the generative objects can have been labeled, e.g., by an external entity, prior to the system receiving the input media data 202. The system can then use the labels to identify and separate the prerendered objects and the generative objects. In some implementations, the system can identify and separate the prerendered objects and the generative objects by identifying artifacts (e.g., typical artifacts) present in the generative objects. For example, the system can identify the artifacts in one or more objects included in the input media 202 and, based on the identification, determine that the one or more objects are generative objects. The system can use the identified artifacts to distinguish the generative objects from the prerendered objects.

[0049] The prerendered and generative object identification and separation operations 218 / 220 can be performed by one or more corresponding engines in the system, e.g., the identification and separation engine 102 in FIG. 1.

[0050] The system can also process the alpha maps 206 in an operation 222 that includes separating alpha maps that are included in the alpha maps 206 according to a content type of the media object with each alpha map. For example, the alpha maps 206 can include alpha maps for media objects of a plurality of different content types, such as alpha maps for video objects, e.g., video alpha maps, and alpha maps for audio objects, e.g., audio alpha maps. The operation 222 can include separating the video alpha maps included in the alpha maps 206 from the audio alpha maps included in the alpha maps 206.

[0051] The system classifies each of the identified generative video objects 221 and / or the identified generative audio objects 219 in respective operations 214 and 216. The system can use the metadata 212 to classify the generative video objects 221 and / or generative audio objects 219.

[0052] The system classifies the generative video objects 221 and / or the generative audio objects 219 into a plurality of object types. In the depicted example, the plurality of object types includes both object types for video objects and object types for audio objects. The system classifies the generative video objects 221 into the object types for video objects, and the generative audio objects 219 into the object types for audio objects. For example, the system classifies each generative video object 221 into a respective object type for video objects, and each generative audio object 219 into a respective object type for audio objects.D25016W001

[0053] Each of the plurality of object types is associated with respective perceptual criteria. Each of the perceptual criteria is defined by whether or not a media object includes a feature that is likely to include anomalies that would be noticed by a consumer of the rendered media object. Object types for video objects are associated with perceptual criteria for video objects. Each of the perceptual criteria for video objects can be defined by whether or not a video object includes a visual feature that is likely to include anomalies that would be noticed by a viewer of the rendered video object.Examples of perceptual criteria for video objects include, but are not limited to: the video object includes biological motion; the video object includes articulated motion; the user blinked eyes when viewing the video object; the video object is not permanent (e.g., lasts less than a specific period of time, e.g., half a second); the video object includes physicsbased features; the video object includes merging objects or emerging objects; the video object includes distortions of perspective; the video object includes intermingling of real world textures with compression artifacts; the video object exhibits the Mannequin effect; the video object includes inconsistencies in depth of field focus: the video object includes inconsistencies in the progression of time; the video object includes small complex objects; the video object includes nonsensical text; the video object includes reflections; and the video object includes shadows. Each of these perceptual criteria for video objects corresponds to a respective object type for video objects that satisfy the corresponding perceptual criterion. For example, the object types for video objects that correspond to the perceptual criteria listed above are listed in the ranking of object types 350 of FIG. 3B.

[0054] Object types for audio objects are associated with perceptual criteria for audio objects. Each of the perceptual criteria for audio objects can be defined by whether or not an audio object includes an auditory feature that is likely to include anomalies that would be noticed by a listener of the rendered audio object. Examples of perceptual criteria for audio objects include: one or more mismatches in the synchronization between the audio object and motion of corresponding visual elements; one or more mismatches between the audio object and corresponding video objects; one or more contexts indicated by the audio object is incorrect; and the audio object includes nonsensical sound.

[0055] Each of the perceptual criteria corresponds to a respective probability of being noticeable, which can be the probability that an anomaly in a media object satisfying the perceptual criterion is noticeable to a consumer of the rendered media object, as described with reference to FIG. 1.D25016W001

[0056] In order to classify each of the generative video objects 221. e.g., in the operation 214, the system determines, for each of the perceptual criteria, whether the generative video object satisfies the perceptual criterion. In some examples, the system can use the metadata 212 to determine whether each generative video object satisfies each of the perceptual criteria.

[0057] If the system determines that the generative video object satisfies the perceptual criterion, the system can classify the generative video object into the object type associated with the perceptual criterion. For example, the system can determine that a given generative video object satisfies the perceptual criterion of including biological motion, e.g., if the given generative video object includes an image of a person waving her hand. In response, the system can classify the given generative video object into an object type that includes biological motion. The system can perform an analogous classification operation for each of the generative audio objects 219 in the operation 216.

[0058] The classification operations 214 / 216 can be performed by one or more classification engines of the system, e.g., the classification engine 104 in FIG. 1.

[0059] In response to classifying each of the generative video objects 221 and the generative audio objects 219 into the plurality of object types, the system (at operation 224) generates priority data for the generative video objects 221 and the generative audio objects 219 based on the determined object types. The priority data can indicate a prioritization of the rendering quality for the media objects based on their respective object types and associated perceptual criteria. The priority data can indicate, for each of the combined set of generative video objects 221 and the generative audio objects 219, a respective priority relative to the other generative media objects in the combined set. For example, media objects classified as object types associated with more noticeable perceptual criteria can be assigned higher priorities than media objects classified as object types associated with less noticeable perceptual criteria.

[0060] In some implementations, the system can generate the priority data using a prioritization model. The prioritization model can be a ranking of the object types into which the media objects (e.g., the generative video objects 221 and the generative audio objects 219) have been classified. The ranking of the object types indicated by the prioritization model can be based on the respective probability of being noticeable associated with the perceptual criteria associated with each object type. An example ranking of object types is illustrated in FIG. 3B.D25016W001

[0061] To generate the priority data for the media objects, the system can compare each pair of media objects with reference to the prioritization model. For each pair of media objects, the system can determine which media object of the pair has been classified into an object type with a higher ranking according to the prioritization model. The system can determine that the media object of the pair that has been classified into an object type with a higher ranking according to the prioritization model has a higher priority than the other media object of the pair.

[0062] By determining respective relative priorities for each pair of media objects in this way, the system can generate priority data that indicates, for each media object, the relative priority of the media object with respect to all other media objects. For example, the system can perform a sorting algorithm that sorts the media objects according to their relative priorities determined by comparing the pairs of media objects based on the prioritization model.

[0063] In some implementations, the system generates distinct priority data for each of the generative video objects 221 and the generative audio objects 219. The system can use distinct prioritization models for each of video objects and audio objects, respectively, to separately generate the priority data for each of the generative video objects 221 and the generative audio objects 219, respectively. Such separate priority data generation helps in reducing potential effects of audio artifacts on prioritizing and / or rendering video objects, and vice versa.

[0064] In some implementations, like the one shown in FIG. 2, the system generates the priority data based on both generative video objects 221 and generative audio objects 21 . Such configuration can provide a more realistic approach on prioritizing video objects based on their perceivability when both video and audio effect perceivability of a particular artifact, e.g., in a football game when a reporter described a ball movement that does not align with the corresponding video.

[0065] The prioritization model for video objects can indicate a ranking of the object types for video objects. The prioritization model for audio objects can indicate a ranking of the object types for audio objects. The priority data for the generative video objects 221 can indicate, for each generative video object, a respective priority relative only to other generative video objects 221 (e.g., and not taking into account generative audio objects 219). The priority data for the generative audio objects 219 can indicate, for each generative audio object, a respective priority relative only to other generative audio objects 219 (e.g., and not taking into account generative video objects 221).D25016W001

[0066] The priority data can have any of a variety of formats, as described with reference to FIG. 1. In the example of FIG. 2, the priority data includes a respective positive integer assigned to each media object. The positive integer indicates a ranking of the priority for the media object relative to the priorities for the other media objects of the combined set of generative video objects 221 and the generative audio objects 219. In some implementations, the positive integer assigned to each generative video object indicates a ranking of the priority for the generative video object relative to the priorities for the other generative video objects 221, and likewise for the audio objects 219.

[0067] For example, a first media object assigned the positive integer of “1” according to the priority data can have the highest priority out of all the media objects to which its priority is relative (e.g., whether the combined set of both generative video and generative audio objects, or only generative video objects, or only generative audio objects). A second media object assigned the positive integer of “2” according to the priority data can have a priority that is lower than that of the first media object, but higher than that of all other media objects to which its priority is relative. In this way, each media object assigned a respective positive integer can have a priority that is lower than that of media objects assigned positive integers that are less than the respective positive integer; but higher than that of media objects assigned positive integers that are greater than the respective positive integer.

[0068] Operation 224 can be performed by a prioritization engine of the system, e.g., by the priority data generation engine 106 in FIG. 1.

[0069] As a result of the operation 224, the system generates an ordered list 226 of the media objects according to their assigned priorities. The ordered list 226 can list the media objects such that media objects assigned higher priorities precede media objects assigned lower priorities. For example, if there are n media objects, the ordered list 226 can include n entries, with each entry including a media object and its assigned positive integer. The n entries can be listed such that the assigned positive integers included in the entries are in ascending consecutive order, as illustrated in FIG. 2.

[0070] In implementations in which the system generates distinct priority data for each of the generative video objects and the generative audio objects, the system can generate two separate ordered lists: one that lists the generative video objects in order of their priorities, and another that lists the generative audio objects in order of their priorities.

[0071] The system uses the priorities assigned to the generative media objects to determine an allocation of computational resources to rendering the generative mediaD25016W001objects. For example, a larger number of computational resources can be allocated to rendering generative media objects with higher priorities than are allocated to rendering generative media objects with lower priorities. In this way, more computational resources can be allocated to rendering generative media objects that are more likely to include anomalies that would be noticeable to a consumer of the rendered generative media object. The computational resources allocated to rendering a media object can be used to render the media object in a way that reduces a likelihood that anomalies of the media object are noticeable to a consumer. Therefore, allocating computational resources in this way can decrease a likelihood that media objects rendered by the system include anomalies that are noticeable to a consumer.

[0072] The system (at operation 228) can allocate computational resources based on the assigned priorities in any of a variety of ways. For example, the system can select one or more options from a specific (e.g., a predefined) list of prioritized rendering options (provided at 228) to determine how to allocate computational resources based on the assigned priorities. The list of prioritized rendering options includes multiple rendering operations that may be prioritized in a specific (e.g., predefined) order. An example list of prioritized rendering options can include (a) determining an order of computation based on the priorities, (b) determining the temporal resolution of the media objects based on the priorities, (c) determining the spatial resolution of the media objects based on the priorities, (d) determining the latent space bit depth with which the media objects are represented based on the priorities, and (e) determining the quantity of neural network layers used to render the generative media objects based on the priorities.

[0073] As noted, the system can allocate computational resources based on (a) determining an order of computation based on the priorities. For example, generative media objects assigned higher priorities can be rendered before generative media objects assigned lower priorities. For example, the generative media objects can be rendered in the order in which they are listed on the list 226.

[0074] The system can allocate computational resources based on priority by (b) determining the temporal resolution of the media objects based on the priorities. The temporal resolution of a media object can be the smallest time interval at which changes can be represented or perceived in a media object during rendering, with higher temporal resolutions corresponding to smaller time intervals. For example, the system can render generative media objects with higher priorities such that they have greater temporal resolution than that of rendered generative media objects with lower priorities.D25016W001

[0075] The system can allocate computational resources based on priority by (c) determining the spatial resolution of the media objects based on the priorities. The spatial resolution of a media object can be the smallest distance interval over which changes can be represented or perceived in a media object during rendering, with higher spatial resolutions corresponding to smaller distance intervals, e.g., depiction of finer detail in the rendered object. For example, the system can render generative media objects with higher priorities such that they have greater spatial resolution than that of rendered generative media objects with lower priorities.

[0076] The system can allocate computational resources based on priority by (d) determining the latent space bit depth with which the media objects are represented based on the priorities. For example, rendering a generative media object can include generating a representation of the media object in a latent space, which can be a lower-dimensional space than the space in which the media object was originally represented, e.g., in the input media data 202. The representation of the media object in the latent space can be a latent representation of the media object, and can be represented using bits. For example, the system can use a particular number of bits to represent each media object in the latent space. The latent space bit depth of a media object can be the number of bits used to represent the media object in the latent space. Larger bit depths can correspond to using more bits to represent the media object in the latent space. Using more bit can in turn result in the capacity to encode more detail in the latent representation of the media object, which can result in a rendered media object that is more detailed, of higher quality, or both. The system can represent media objects with higher priorities in the latent space using more bits than those used to represent media objects with lower priorities in the latent space.

[0077] The system can allocate computational resources based on priority by (e) determining the quantity of neural network layers used to render the generative media objects based on the priorities. For example, the system can use a neural network including a plurality of layers to render the generative media objects. Using more layers of the neural network to render a generative media object can increase a likelihood that the rendered generative media object is of high quality, includes anomalies that are less noticeable to a consumer, or both. The system can use more layers of the neural network to render generative media objects with higher priorities than it uses to render generative media objects with lower priorities.D25016W001

[0078] In response to using one or more options from the list of prioritized rendering options at 228, the system renders each of the generative video objects and each of the generative audio objects by allocating computational resources to rendering the object in the way indicated by the selected one or more options and based on the priority assigned to the object. For example, if the system uses option (d), the system can read the assigned priority of the media object, e.g., according to the list 226. The system can determine a number of bits to use to represent the media object in the latent space based on the priority for the media object that is read from the list 226. For example, the system can use an algorithm that maps an assigned priority to a number of bits. For example, the system can allocate a total number of bits to be used to represent all of the media objects, and then partition the total number of bits proportionally to each of the media objects based on its assigned priority.

[0079] In some implementations, depending on how long the list is, the system can use more than one of the options listed therein. In some implementations, the operations (i.c., options) in the list are not prioritized; rather the system selects one or more options from what is included in the list. The selection can be based on what objects are included in the prioritized objects received from 224.. The system can allocate computational resources to rendering the media objects based on the priorities assigned to the media objects in the ways indicated by each of the options that are selected.

[0080] Operation 228 can be performed by a rendering engine of the system, e.g., rendering engine 108 in FIG. 1.

[0081] As noted above, in some implementations, the system generates two respective ordered lists 226 of media objects, for each of the generative video objects and the generative audio objects. In such implementations, the system can determine allocations of computational resources for rendering the generative video objects separately from determining allocations of computational resources for rendering the generative audio objects. For example, the system can allocate a first set of computational resources to rendering the generative video objects, and use the respective ordered list 226 for the generative video objects to allocate the computational resources of the first set to rendering each of generative video objects. The system can allocate a second set of computational resources to rendering the generative audio objects, and use the respective ordered list 226 for the generative audio objects to allocate the computational resources of the second set to rendering each of generative audio objects.D25016W001

[0082] In such implementations, the system can allocate the computational resources for rendering the generative video and audio objects sequentially, e.g., video objects before audio objects or vice versa. For example, the system can determine to prioritize rendering all generative video objects above rendering all audio objects, e.g., by determining that anomalies included in the generative video objects are more likely to be noticeable than those included in the generative audio objects. In response, the system can allocate the computational resources for rendering the generative video objects before allocating the computational resources for rendering the generative audio objects.

[0083] In some implementations, the system can allocate the computational resources for rendering the generative video and audio objects in parallel. In some implementations, the system can render the generative video and audio objects sequentially, e.g., video objects before audio objects or vice versa. In some implementations, the system can render the generative video and audio objects in parallel.

[0084] In response to rendering the generative video objects and the generative audio objects by allocating computational resources according to the priority data, the system generates rendered generative audio objects 230 and rendered generative video objects 232.

[0085] The system combines the rendered generative video objects 232 with the prerendered video objects 234 in an operation 238. In some implementations, the system can combine the rendered generative video objects 232 with the prerendered video objects 234 using the alpha maps for video objects identified during the operation 222. For example, the system can combine the rendered generative video objects 232 with the prerendered video objects 234 using the alpha maps through alpha compositing. For example, the alpha maps can indicate a respective weight for each object. The system can combine the objects using linear interpolation according to the weights indicated by the alpha maps. For example, the combination of the objects can be determined by a weighted sum of the data representing the objects, where the weights in the weighted sum are the weights indicated by the alpha maps. Operation 238 can be performed by a combine engine of the system, e.g., combine engine 112 in FIG. 1.

[0086] As a result of combining the rendered generative video objects 232 with the prerendered video objects 234, the system generates combined video media. For example, the combined video media can be a video including one or more video frames that include the rendered generative video objects 232 and the prerendered video objects 234. In some implementations, the system can generate the combined video media by, for each videoD25016W001frame of the video and for each pixel of the video frame, using the alpha maps for video objects to determine one or more opacity values for a generative video object represented by the pixel, a prerendered video object represented by the pixel, or both. The system can generate pixel data for the pixel, to be included in the combined video media, using the determined one or more opacity values. In some implementations, the system can transmit the combined video media to a display system 242, e.g., to be viewed by a viewer.

[0087] The system combines the rendered generative audio objects 230 with the prerendered audio objects 236 in an operation 240. In some implementations, the system can combine the rendered generative audio objects 230 with the prerendered audio objects 236 using the alpha maps for audio objects identified during the operation 222. As a result of combining the rendered generative audio objects 230 with the prerendered audio objects 236, the system generates combined audio media. For example, the combined audio media can be an audio sound file that includes the rendered generative audio objects 230 and the prcrcndcrcd audio objects 236. In some implementations, the system can transmit the combined audio media to an audio system 244, e.g., to be listened to by a user of the audio system 244.

[0088] In some implementations, the system is configured to render generative media objects that only include video objects. For example, in such implementations, the input media data 202 can include media objects that only include video objects. As a result of the operation 204, the system can generate only the alpha maps 206 and the video objects 210. The system can skip the operations 218, 216, 222, and 240.

[0089] In some implementations, the system is configured to render generative media objects that only include audio objects. For example, in such implementations, the input media data 202 can include media objects that only include audio objects. As a result of the operation 204, the system can generate only the alpha maps 206 and the audio objects 208. The system can skip the operations 220, 214, 222, and 238.

[0090] FIG. 3A is an example video frame 300 that includes video objects that include anomalies that are likely to be noticeable. The video frame 300 can be represented by data included in the input media received by a media rendering system, e.g., the media rendering system 100 of FIG. 1. In response to receiving the input media including the data representing the video frame 300, the media rendering system can render the video frame 300, e.g., using the process 400 of FIG. 4 for rendering media objects. For example, the media rendering system can render the video frame 300 by allocating computational resources according to priorities assigned to video objects included in theD25016W001video frame 300. The priorities can be assigned to the video objects based on perceptual criteria associated with probabilities that anomalies included in the video objects are noticeable to viewers. The priorities can be assigned such that the video objects included in the video frame 300 that include anomalies that are more likely to be noticeable are assigned higher priorities, and therefore allocated more computational resources for rendering, than are allocated to other video objects included in the video frame 300.

[0091] The video frame 300 is partitioned according to a grid. For example, the grid into which the video frame 300 is partitioned includes sixty-four cells arranged in eight rows and eight columns. Each cell of the grid can include one or more video objects. As part of the process for rendering the video frame 300, the media rendering system can assign a priority to each of the video objects included in the each cell of the grid. As described above with reference to FIG. 2, the priority assigned to a video object can be based on an object type for the video object that is associated with perceptual criteria that indicate a likelihood that anomalies included in the video object arc noticeable. For example, video objects of object types that are associated with perceptual criteria that indicate a high likelihood that anomalies included in the video object are noticeable are assigned higher priorities than video objects of object types associated with perceptual criteria indicating a lower likelihood that anomalies included in the video object are noticeable.

[0092] Cells that include video objects that include anomalies that are likely to be noticeable are outlined in FIG. 3A. For example, the outlined cell 302 includes a video object 304 of a human hand waving. The media rendering system can classify the video object 304 into an object type that includes biological motion. The media rendering system can assign a probability to the video object 304 using a prioritization model that indicates a ranking of object types according to probabilities of being noticeable associated with the perceptual criteria associated with the object types, e.g., the ranking of object types 350 illustrated in FIG. 3B. For example, the media rendering system can determine the ranking of the object type into which the video object 304 has been classified (e.g., the object type that includes biological motion), and use the ranking 350 to determine a corresponding rank for the object type according to the ranking 350.

[0093] For example, the media rendering system can determine that the object type that includes biological motion has the highest rank according to the ranking 350. In response, the media rendering system can assign a high priority to the video object 304.

[0094] In some implementations, in response to determining that the object type that includes biological motion has the highest rank according to the ranking 350, the mediaD25016W001rendering system can store data indicating that the object type into which the video object 304 was classified has the highest rank according to the ranking 350, e.g., in a memory of the media rendering system. The system can store analogous data for each of the video objects of the video frame 300. The system can then assign a respective priority to each of the video objects of the video frame 300 using the stored data.

[0095] For example, the system can use a sorting algorithm that sorts the video objects into an ordered list such that the order of the video objects in the ordered list corresponds the ranks of the object types of the video objects according to the ranking 350. The system can assign priorities to the video objects according to their order in the ordered list (e.g., video objects that appear earlier in the ordered list can be assigned higher priorities than those assigned to video objects that appear later in the ordered list).

[0096] In some implementations, the system can assign a respective priority to each of the video objects of the video frame 300 using the stored data by applying a mapping between: i) the ranks of object types according to the ranking 350, and ii) priorities assigned to video objects that are classified into the object types. For example, the mapping can indicate that video objects of the object type with the highest rank according to the ranking 350, such as the video object 304, are to be assigned the higher priorities.

[0097] In response to assigning a high priority to the video object 304, the media rendering system can allocate a large number of computational resources to rendering the video object 304, as described above with reference to FIG. 2. The media rendering system can allocate computational resources to rendering each of the video objects included in the video frame 300 according to their assigned priorities, as described with reference to FIG. 2. In this way, media rendering system can render the video frame 300 such that larger number of computational resources are used to render video objects of object types the anomalies of which are more likely to be noticeable. Rendering the video frame 300 in this way can, in turn, decrease a likelihood that anomalies included in the rendered video frame 300 are noticeable.

[0098] FIG. 3B is an example ranking of object types 350. The ranking of object types 350 can be used by a media rendering system, e.g., the media rendering system 100 of FIG. 1, to assign priorities for rendering generative media objects. For example, the ranking 350 can represent or be included in a prioritization model used by the media rendering system to generate priority data for generative media objects that indicates the respective priorities assigned to the generative media objects.D25016W001

[0099] In the example of FIG. 3B, the ranking of object types 350 includes object types for video objects listed in the order corresponding to their respective ranks. The ranking 350 can be used in implementations in which a media rendering system generates distinct priority data for each of generative video objects and generative audio objects. In such implementations, the media rendering system can use the ranking 350 to generate the priority data for the generative video objects. The ranking 350 can be used in implementations in which a media rendering system is configured to render generative video data, e.g., separate from rendering generative audio data.

[0100] The ranking 350 is based on the violin plot 352. The violin plot 352 illustrates times elapsed prior to viewers of videos noticing anomalies in the video objects of the videos, for each of a plurality of different object types of the video objects. The values displayed on the violin plot 352 represent averages across multiple distinct videos and / or multiple distinct viewers. The thickness of each violin plot indicates how many vicwcr(s) looked at a respective artifact. The position of each violin plot indicates at what time instances the viewer(s) looked at the artifact. The averaging can occur across different videos. For simplicity, the description of the values below will refer to only a single viewer and a single video.

[0101] The violin plot 352 includes an x-axis 354 and a y-axis 356. The x-axis 354 represents an amount of time elapsed before a viewer notices an anomaly in a video object of a particular object type. The x-axis 354 represents the amounts of time in seconds on a logarithmic scale.

[0102] The y-axis 356 represents different object types for the video objects. For example, each discrete point on the y-axis 356, represented by a tick mark on the y-axis 356 in FIG. 3B, represents a respective object type for a video object.

[0103] Each object type on the y-axis 356 corresponds to a respective plotted “violin” (e.g., a violin-shaped group of points) at the same location on the y-axis as the object type. For example, the “biological motion” object type corresponds to the violin 360; the “articulated motion” object type corresponds to the violin 361; the “eye blinking” object type corresponds to the violin 362; the “object permanence” object type corresponds to the violin 363; the “physics” object type corresponds to the violin 364; the “merging objects / emergence” object type corresponds to the violin 365; the “perspective distortions” object type corresponds to the violin 366; the “intermingling of real world textures with compression artifacts” object type corresponds to the violin 367; the “Mannequin effect” object type corresponds to the violin 368; the “depth of field focusD25016W001inconsistencies” object type corresponds to the violin 369; the “arrow of time” corresponds to the violin 370; the “small complex objects” object type corresponds to the violin 371; the “nonsensical text” object type corresponds to the violin 372; the “reflections” object type corresponds to the violin 373; and the “shadows” object type corresponds to the violin 374.

[0104] Each violin includes a respective plurality of points, with each point representing a time elapsed before a viewer first noticed an anomaly in a video object of the corresponding object type after starting to view the video. For example, for each point, the position of the point along the x-axis 354 represents the amount of time elapsed prior to the viewer first noticing the anomaly in the video object of the object type corresponding to the violin including the point. For each violin, points located at the same positions along the x-axis 354 are stacked vertically on top of one another, such that the thickness of the violin at a given position along the x-axis 354 signifies the number of points in the violin located at the given position. Positions along the x-axis 354 where the violin is thicker are positions at which more points in the violin are located. Therefore, increased thickness of a violin at a given position along the x-axis 354 signifies that a larger number of viewers took the amount of time represented by the given position to notice an anomaly in a video object of the corresponding object type after starting to view the video.

[0105] The ranking of object types 350 is determined by the violin plot 352 in the following way: object types for which viewers, on average, noticed an anomaly in a video object the object type earlier are ranked higher according to the ranking 350. Object types for which viewers, on average, noticed an anomaly in a video object the object type later are ranked lower according to the ranking 350. For example, noticing an anomaly in a video object of a given object type earlier can indicate that anomalies in video objects of the given object type tend to be more noticeable than those in video objects of other object types. The given object type can be ranked higher according to the ranking 350, and, as a result, video objects of the given object type can be assigned higher priorities by the media rendering system. Consequentially, more computational resources can be allocated to rendering the video objects of the given object type, which can reduce a likelihood that anomalies in the rendered video objects are noticeable.

[0106] FIG. 4 is a flow chart of an example process 400 for rendering generative media objects. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, aD25016W001media rendering system, e.g., the media rendering system 100 depicted in FIG. 1, appropriately programmed in accordance with this specification, can perform the process 400.

[0107] The system receives input media (402). The input media received by the system can include one or more of the input media data 202 of FIG. 2. For example, the input media can include one or more of: media objects of any of a variety of object types; metadata: or alpha maps. For example, the media objects can include video objects, audio objects, or both. The media objects included in the input media can include prerendered media objects, which are media objects that have already been rendered, e.g., by an external system. The media objects included in the input media can include generative media objects, which are media objects that have been generated by one or more Al generative models and have not yet been rendered.

[0108] The input media can be received from any of a variety of external sources, e.g., an audio media transmitter, a video media transmitter, or a transmitter of both audio and video media. The input media can be received wirelessly.

[0109] The system identifies, from the input media, a first set of media objects corresponding to objects generated by one or more Al generative models (404). In some implementations, the system can first identify and separate media objects of different content types from the input media. For each of the content types, the system can then identify a respective first set of media objects corresponding to objects generated by one or more Al generative models.

[0110] For example, if the received input media includes media objects that include audio objects and video objects, the system can separate the audio objects from the video objects. The system can, from the separated audio objects, identify generative audio objects, e.g., using techniques described with respect to the operation 218 of FIG. 2. The generative audio objects can be the audio objects generated by one or more Al generative models included in the respective first set for the audio objects.

[0111] The system can, from the separated video objects, identify generative video objects, e.g., using techniques described with respect to the operation 220 of FIG. 2. The generative video objects can be the video objects generated by one or more Al generative models included in the respective first set for the video objects.

[0112] In some implementations, the system can also identify and separate alpha maps, metadata, or both, from the input media. In some implementations, the system canD25016W001also identify a second set of media objects that are pre-rendered and are different from the first set of media objects.

[0113] The system classifies the first set of media objects into a plurality of object types (406). The system classifies the first set of media objects based on one or more perceptual criteria. The perceptual criteria are defined by whether or not a media object includes a feature that is likely to include anomalies that would be noticed by a consumer of the rendered media object, as described with reference to FIG. 2. For example, a media object can satisfy a perceptual criterion if the media object includes the feature indicated by the perceptual criterion, the feature being a feature that is likely to include anomalies that would be noticed by a consumer of the rendered media object. Examples of perceptual criteria are described with reference to FIG. 2.

[0114] The object types into which the system classifies the first set of media objects are each associated with a respective perceptual criterion. A media object being classified into a particular object type can signify that the media object satisfies the associated perceptual criterion. For example, each object type can be defined to include all objects that satisfy the associated perceptual criterion.

[0115] For each media object in the first set of media objects, the system can classify the media object into one of the plurality of object types. The system classifies the media object into one of the object types by determining, for each of the perceptual criteria, whether the media object satisfies the perceptual criterion. The system can use metadata, e.g., metadata for the media objects that was separated out from the input media, to determine whether the media object satisfies the perceptual criterion. If the system determines that the media object satisfies the perceptual criterion, the system can classify the media object into the object type associated with the perceptual criterion.

[0116] In some examples, the system may determine that a given media object satisfies more than one perceptual criteria. The system can employ a tie-breaker algorithm to select one perceptual criterion from the more than one perceptual criteria satisfied by the media object. The system can then classify the media object into the object associated with the selected perceptual criterion.

[0117] In some implementations, the system can use a model to classify the first set of media objects into the plurality of object types. The model can be configured to process an input including metadata for the media objects in the first set and data indicating the perceptual criteria associated with the plurality of object types. In response to processing the input, the model can be configured to generate an output indicating aD25016W001classification of the first set of media objects into the plurality of object types, based on the perceptual criteria.

[0118] In some implementations in which the input media includes video objects, audio objects, or both, the plurality of object types can include video object types, audio object types, or both. The system can classify the video objects included in the input media into the video object types. The system can classify the audio objects included in the input media into the audio object types.

[0119] More generally, the input media can include multimodal media objects, e.g., media objects of multiple different modalities. The object types can then include multimodal object types. For example, for each of the multiple different modalities of the multimodal media objects, the object types can include a respective set of object types corresponding to the modality. The system can classify the media objects of each modality into its respective set of object types.

[0120] The system determines priority data for respective objects in the first set of media objects based on the respective determined object types (408). The priority data can include a respective priority assigned to each object. The respective priorities can have any of a variety of formats, as described with reference to FIG. 1.

[0121] The system can determine the priority data using a prioritization model, as described with reference to FIG. 2. For example, the prioritization model can include a ranking of the object types into which the first set of media objects have been classified, where the ranking is based on the respective probability of being noticeable associated with the perceptual criteria associated with each object type. The system can assign priorities to the respective objects using the ranking of the object types included in the prioritization model, such that objects of object types with higher rankings are assigned higher priorities than those assigned to objects of object types with lower rankings.

[0122] Assigning priorities in this way can result in the system assigning priorities to the respective objects such that objects of object types associated with perceptual criteria associated with higher probabilities of being noticeable are assigned higher priorities than those assigned to objects of object types associated with perceptual criteria associated with lower probabilities of being noticeable. For example, a first object type associated with a first perceptual criterion is given a higher priority than a second object type associated with a second perceptual criterion that is deteimined to be less noticeable than the first perceptual criterion. That the second perceptual criterion is determined to be less noticeable than the first perceptual criterion can mean that probability of beingD25016W001noticeable associated with the second perceptual criterion is determined to be lower than that associated with the first perceptual criterion.

[0123] In implementations in which the object types include video object types, audio object types, or both, the system can determine the priority data by determining audio priority data for the audio object types, determining video priority data for the video object types, or both.

[0124] For example, if the object types include both video object types and audio object types, the system can determine video priority data based on the video object types and the system can determine audio priority data that is different from the video priority data based on the audio object types. The system can determine the video priority data using a video prioritization model that includes a ranking of video object types into which the video objects have been classified. The system can assign priorities to the video objects using the ranking of the video object types included in the video prioritization model, as described above. The video priority data can include the priorities assigned to each of the video objects.

[0125] Separately from generation of the video priority data, the system can determine the audio priority data using an audio prioritization model, e.g., that is different from the video prioritization model, and includes a ranking of audio object types into which the audio objects have been classified. The system can assign priorities to the audio objects using the ranking of the audio object types included in the audio prioritization model, as described above. For example, the system can assign the priorities to the audio objects independently of assigning the priorities to the video objects. The audio priority data can include the priorities assigned to each of the audio objects.

[0126] Similarly, in implementations in which the object types include multimodal object types, the system can determine multimodal priority data for each of the multimodal object types, e.g., for each of the respective sets of object types corresponding to a modality of media objects.

[0127] The system renders generative media objects from respective objects in the first set of media objects based on the respective priority data (410). The system can render the generative media objects based on the respective priority data by allocating computational resources to render respective objects according to the priority data. For example, the system can allocate more computational resources to rendering respective objects with higher priority according to the priority data than are allocated to rendering respective objects with lower priority according to the priority data.D25016W001

[0128] In some implementations, the system allocates computational resources by controlling one or more of the following options based on the priority data: order of computation; temporal resolution; spatial resolution; latent space bit-depth; or quantity of neural network layers utilized. For example, the system can control each of the one or more options using techniques described with reference to FIG. 2.

[0129] In implementations in which the system determines both audio priority data and video priority data, the system can render generative audio objects based on the audio priority data and render generative video objects based on the video priority data. For example, the system can first allocate respective sets of computational resources to each of rendering the generative video objects and rendering the generative audio objects. The system can next allocate computational resources to rendering the generative video objects from the respective set for the video objects, and allocate computational resources to rendering the generative audio objects from the respective set for the audio objects. In some examples, the system can allocate computational resources to rendering the generative video and audio objects sequentially, e.g., first allocating computational resources to rendering the generative video objects based on the video priority data and then allocating the remaining computational resources to rendering the generative audio objects based on the audio priority data, or vice versa.

[0130] In some implementations, the system can render the generative media objects by, for each object, converting data representing the object into an appropriate format for rendering. For example, the system can convert data representing audio objects into waveform buffers. For example, the system can convert data representing video objects into one or more sequences of images including RGB pixels. After converting data representing each generative media object into an appropriate format, the system can render each generative media object by performing one or more post-processing operations on the generative media object. For example, for generative video objects, the post-processing operations can include one or more of: color correction, scaling, cropping, denoising, applying textures, calculating lighting, or aspect-ratio adjustment. For example, for generative audio objects, the post-processing operations can include one or more of: normalization, compression, noise removal, mixing effects, or spatialization.

[0131] In some implementations, the system can render the generative media objects using a model (e.g., at the rendering engine 108 in FIG. 1). For example, the model can be configured to process data indicating the generative media objects (e.g., 105) and the priority data (e.g., 109) to generate data indicating the rendered generativeD25016W001media objects (e.g., 111) in accordance with the priority data. The model can include one or more neural networks, such as generative neural networks.

[0132] The system (e.g., at the rendering engine 108) can allocate rendering resources to the model based on the priority data. In some implementations, the system adapts the inference process according to the priority date. For example, the system can use the priority data in the process of rendering the generative media objects (e.g., to produce 111 in FIG. 1) by allocating a latent space bit-depth, by specifying the number of neural network layers for the model, by determining an order of rendering computation, and / or by specifying at least one of the temporal or spatial resolutions for the rendered media objects.

[0133] After rendering the generative media objects, the system can cause storage or transmission of the rendered generative media objects. For example, the system can generate a media file including the rendered generative media objects. The system can cause storage or transmission of the media file.

[0134] In implementations in which the system also identifies a second set of media objects that are pre-rendered, the system can combine data corresponding to the rendered generative media objects and data corresponding to the second set of media objects to generate a media file. For example, the system can combine the data using one or more alpha maps, as described with reference to FIG. 2. The system can cause storage or transmission of the media file that is generated by combining the data.

[0135] In some implementations, the process 400 is executed at a client device. A client device can be any electronic device capable of requesting, receiving, and / or processing media data for presentation to a user. In some examples, the client device includes at least one processor and a memory storing instructions executable by the processor. Examples of client devices include smartphones, laptops, tablets, wearables, computers, top boxes, and gaming consoles, although this list is not limiting and the client device can be any of a variety of electronic devices.

[0136] For example, the system can identify a second set of media objects that are pre-rendered and are different from the first set of media objects, as described above. The system can combine the generative media objects with the second set of media objects to obtain combined media data. The system can then cause playback of the combined media data using one or more media output devices of the client device, such as speakers or display screens.D25016W001

[0137] FIG. 5 shows an example of a computing device 500 and example of a display devices 580 / 582 that can be used to implement the techniques described here. For example, the computing device 500 can be or include the media rendering system 100 of FIG. 1, and the display device 580 / 582 can be the client device described with reference to FIG. 4. The computing device 500 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The display device is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart-phones, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the inventions described and / or claimed in this document.

[0138] The computing device 500 includes a processor 502, a memory 504, a storage device 506, a high-speed interface 508 connecting to the memory 504 and multiple high-speed expansion ports 510, and a low-speed interface 512 connecting to a low-speed expansion port 514 and the storage device 506. Each of the processor 502, the memory 504, the storage device 506, the high-speed interface 508, the high-speed expansion ports 510, and the low-speed interface 512, are interconnected using various busses, and can be mounted on a common motherboard or in other manners as appropriate. The processor 502 can process instructions for execution within the computing device 500, including instructions stored in the memory 504 or on the storage device 506 to display graphical information for a GUI on an external input / output device, such as a display 516 coupled to the high-speed interface 508. In other implementations, multiple processors and / or multiple buses can be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices can be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).

[0139] The memory 504 stores information within the computing device 500. In some implementations, the memory 504 is a volatile memory unit or units. In some implementations, the memory 504 is a non-volatile memory unit or units. The memory 504 can also be another form of computer-readable medium, such as a magnetic or optical disk.

[0140] The storage device 506 is capable of providing mass storage for the computing device 500. In some implementations, the storage device 506 can be orD25016W001contain a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. A computer program product can be tangibly embodied in an information earner. The computer program product can also contain instructions that, when executed, perform one or more methods, such as those described above. The computer program product can also be tangibly embodied in a computer- or machine-readable medium, such as the memory 504, the storage device 506, or memory on the processor 502.

[0141] The high-speed interface 508 manages bandwidth-intensive operations for the computing device 500, while the low-speed interface 512 manages lower bandwidthintensive operations. Such allocation of functions is exemplary only. In some implementations, the high-speed interface 508 is coupled to the memory 504, the display 516 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 510, which can accept various expansion cards (not shown). In the implementation, the low-speed interface 512 is coupled to the storage device 506 and the low-speed expansion port 514. The low-speed expansion port 514, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.

[0142] The computing device 500 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a standard server 520, or multiple times in a group of such servers. In addition, it can be implemented in a personal computer such as a laptop computer 522. It can also be implemented as part of a rack server system 524. Alternatively, components from the computing device 500 can be combined with other components in a mobile device (not shown), such as a display device 550. Each of such devices can contain one or more of the computing device 500 and the display device 550, and an entire system can be made up of multiple computing devices communicating with each other.

[0143] The display device 550 includes a processor 552, a memory 564, an input / output device such as a display 554, a communication interface 566, and a transceiver 568, among other components. The display device 550 can also be provided with a storage device, such as a micro-drive or other device, to provide additional storage. Each of the processor 552, the memory 564, the display 554, the communication interface 566, and the transceiver 568, are interconnected using various buses, and several of theD25016W001components can be mounted on a common motherboard or in other manners as appropriate.

[0144] The processor 552 can execute instructions within the display device 550, including instructions stored in the memory 564. The processor 552 can be implemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 552 can provide, for example, for coordination of the other components of the display device 550, such as control of user interfaces, applications run by the display device 550, and wireless communication by the display device 550.

[0145] The processor 552 can communicate with a user through a control interface 558 and a display interface 556 coupled to the display 554. The display 554 can be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology. The display interface 556 can comprise appropriate circuitry for driving the display 554 to present graphical and other information to a user. The control interface 558 can receive commands from a user and convert them for submission to the processor 552. In addition, an external interface 562 can provide communication with the processor 552, so as to enable near area communication of the display device 550 with other devices. The external interface 562 can provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces can also be used.

[0146] The memory 564 stores information within the display device 550. The memory 564 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. An expansion memory 574 can also be provided and connected to the display device 550 through an expansion interface 572, which can include, for example, a SIMM (Single In Line Memory Module) card interface. The expansion memory 574 can provide extra storage space for the display device 550, or can also store applications or other information for the display device 550. Specifically, the expansion memory 574 can include instructions to carry out or supplement the processes described above, and can include secure information also. Thus, for example, the expansion memory 574 can be provide as a security module for the display device 550, and can be programmed with instructions that permit secure use of the display device 550. In addition, secure applications can be provided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non-hackable manner.D25016W001

[0147] The memory can include, for example, flash memory and / or NVRAM memory (non-volatile random access memory), as discussed below. In some implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The computer program product can be a computer- or machine-readable medium, such as the memory 564, the expansion memory 574, or memory on the processor 552. In some implementations, the computer program product can be received in a propagated signal, for example, over the transceiver 568 or the external interface 562.

[0148] The display device 550 can communicate wirelessly through the communication interface 566, which can include digital signal processing circuitry where necessary. The communication interface 566 can provide for communications under various modes or protocols, such as GSM voice calls (Global System for Mobile communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access), TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others. Such communication can occur, for example, through the transceiver 568 using a radio-frequency. In addition, short-range communication can occur, such as using a Bluetooth, WiFi, or other such transceiver (not shown). In addition, a GPS (Global Positioning System) receiver module 570 can provide additional navigation- and location-related wireless data to the display device 550, which can be used as appropriate by applications running on the display device 550.

[0149] The display device 550 can also communicate audibly using an audio codec 560, which can receive spoken information from a user and convert it to usable digital information. The audio codec 560 can likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of the display device 550. Such sound can include sound from voice telephone calls, can include recorded sound (e.g., voice messages, music files, etc.) and can also include sound generated by applications operating on the display device 550.

[0150] The display device 550 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a cellular telephone 580. It can also be implemented as part of a smart-phone 582, personal digital assistant, or other similar mobile device.D25016W001

[0151] In this specification, the term “database” can broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. A database can be implemented on any appropriate type of memory.

[0152] In this specification the term “engine” can broadly to refer to a softwarebased system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some instances, one or more computers will be dedicated to a particular engine. In some instances, multiple engines can be installed and running on the same computer or computers.

[0153] Operations can occur substantially concurrently in that the operations need not be exactly concurrent but can overlap at least in part. For instance, a first operation can begin and sometime after that a second operation can begin while the first operation is still occurring. Execution of the two operations, whether by the same system or different systems, can be substantially concurrently. In some examples, two operations can execute substantially concurrently when they have the same start time, same end time, or both.

[0154] In this specification, the term likely can mean that there is a likelihood that something might occur and that the likelihood satisfies a likelihood threshold. For instance, when determining that an anomaly is likely noticeable to a viewer, a system would determine a likelihood that the anomaly is noticeable to the viewer. The system would then determine whether the likelihood satisfies, e.g., is greater than or equal to, a likelihood threshold by comparing the two values. If so, the system determines that the anomaly is likely noticeable to the viewer. If not, the system determines that the anomaly is not likely noticeable to the viewer.

[0155] A number of implementations have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. For example, various forms of the flows shown above can be used, with operations re-ordered, added, or removed.

[0156] Implementations of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or inD25016W001combinations of one or more of them. Implementations of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, a data processing apparatus. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus. One or more computer storage media can include a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0157] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can be or include special purpose logic circuitry, e.g., a field programmable gate array (“FPGA”) or an application-specific integrated circuit (“ASIC”). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0158] A computer program, which may also be referred to or described as a program, software, a software application, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub-programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.D25016W001

[0159] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., a field programmable gate array (“FPGA”) or an application-specific integrated circuit (“ASIC”).

[0160] Computers suitable for the execution of a computer program include, by way of example, general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. A computer can be embedded in another device, e.g., a mobile telephone, a smart phone, a headset, a personal digital assistant (“PDA”), a mobile audio or video player, a game console, a Global Positioning System (“GPS”) receiver, or a portable storage device, e.g., a universal serial bus (“USB”) flash drive, to name just a few.

[0161] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0162] To provide for interaction with a user, implementations of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a liquid crystal display (“LCD”), an organic light emitting diode (“OLED”) or other monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball or a touchscreen, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user canD25016W001be received in any form, including acoustic, speech, or tactile input. In some examples, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser.

[0163] Implementations of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), e.g., the Internet.

[0164] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some implementations, a server transmits data, e.g., a Hypertext Markup Language (“HTML”) page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user device, which acts as a client. Data generated at the user device, e.g., a result of user interaction with the user device, can be received from the user device at the server.

[0165] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of what is being claimed, which is defined by the claims themselves, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claim may be directed to a subcombination or variation of a subcombination.D25016W001

[0166] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0167] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

[0168] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):EEE1. A method comprising:receiving input media;identifying, from the input media, a first set of media objects corresponding to objects generated by one or more Al generative models;classifying the first set of media objects into a plurality of object types, wherein the classifying is based on one or more perceptual criteria;determining priority data for respective objects in the first set of media objects based on the respective determined object types, wherein a first object type associated with a first perceptual criterion is given a higher priority than a second object type associated with a second perceptual criterion that is determined to be less noticeable than the first perceptual criterion; andrendering generative media objects from respective objects in the first set of media objects based on their respective priority data.D25016W001EEE2. The method of EEE1, wherein the rendering comprises allocating computational resources to render respective objects according to the priority data.EEE3. The method of EEE2, wherein allocating the computational resources comprises controlling one or more of following options based on the priority data:order of computation;temporal resolution;spatial resolution;latent space bit-depth; orquantity of neural network layers utilized.EEE4. The method of EEE2 or EEE3, wherein allocating computational resources to rendering respective objects comprises allocating more computational resources to rendering respective objects with higher priority according to the respective priority data than are allocated to rendering respective objects with lower priority according to the respective priority data.EEE5. The method of any preceding EEE, further comprising identifying, from the input media, a second set of media objects that are pre-rendered and are different from the first set of media objects.EEE6. The method of EEE5, further comprising:combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects to generate a media file; and causing storage or transmission of the media file.EEE7. The method of EEE6, wherein combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects comprises combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects with one or more alpha maps.EEE8. The method of any preceding EEE, wherein the method is executed at a client device.D25016W001EEE9. The method of EEE8, further comprising:identifying a second set of media objects that are pre-rendered and are different from the first set of media objects;combining the generative media objects with the second set of media objects to obtain combined media data; andcausing playback of the combined media data using one or more media output devices of the client device.EEE10. The method of any preceding EEE, wherein:the plurality of object types includes audio object types and / or video object types; anddetermining the priority data comprises determining audio priority data for the audio object types and / or determining video priority data for the video object types.EEE11. The method of any preceding EEE, wherein the plurality of object types includes multimodal object types, and determining the priority data comprises determining multimodal priority data for the multimodal object types.EEE12. The method of any preceding EEE, wherein the input media is received wirelessly from an audio / video media transmitter.EEE13. A system comprising:one or more computers and one or more storage devices on which are stored instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:receiving input media;identifying, from the input media, a first set of media objects corresponding to objects generated by one or more Al generative models;classifying the first set of media objects into a plurality of object types, wherein the classifying is based on one or more perceptual criteria;determining priority data for respective objects in the first set of media objects based on the respective determined object types, wherein a first object type associated with a first perceptual criterion is given a higher priority than a second object typeD25016W001associated with a second perceptual criterion that is determined to be less noticeable than the first perceptual criterion; andrendering generative media objects from respective objects in the first set of media objects based on their respective priority data.EEE14. The system of EEE13, wherein the rendering comprises allocating computational resources to render respective objects according to the priority data.EEE15. The system of EEE14, wherein allocating the computational resources comprises controlling one or more of following options based on the priority data:order of computation;temporal resolution;spatial resolution;latent space bit-depth; orquantity of neural network layers utilized.EEE16. The system of EEE 14, wherein allocating computational resources to rendering respective objects comprises allocating more computational resources to rendering respective objects with higher priority according to the respective priority data than are allocated to rendering respective objects with lower priority according to the respective priority data.EEE17. The system of any one of EEE13 to EEE16, wherein the operations further comprise identifying, from the input media, a second set of media objects that are prerendered and are different from the first set of media objects.EEE18. The system of EEE17, wherein the operations further comprise:combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects to generate a media file; and causing storage or transmission of the media file.EEE19. The system of EEE18, wherein combining data corresponding to the rendered generative media objects and data corresponding to the second set of mediaD25016W001objects comprises combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects with one or more alpha maps.EEE20. One or more non-transitory computer storage media encoded with computer program instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:receiving input media;identifying, from the input media, a first set of media objects corresponding to objects generated by one or more Al generative models;classifying the first set of media objects into a plurality of object types, wherein the classifying is based on one or more perceptual criteria;determining priority data for respective objects in the first set of media objects based on the respective determined object types, wherein a first object type associated with a first perceptual criterion is given a higher priority than a second object type associated with a second perceptual criterion that is determined to be less noticeable than the first perceptual criterion; andrendering generative media objects from respective objects in the first set of media objects based on their respective priority data.

Claims

1. D25016W001CLAIMS1. A method comprising:receiving input media;identifying, from the input media, a first set of media objects corresponding to objects generated by one or more Al generative models:classifying the first set of media objects into a plurality of object types, wherein the classifying is based on one or more perceptual criteria;determining priority data for respective objects in the first set of media objects based on the respective determined object types, wherein a first object type associated with a first perceptual criterion is given a higher priority than a second object type associated with a second perceptual criterion that is determined to be less noticeable than the first perceptual criterion; andrendering generative media objects from respective objects in the first set of media objects based on their respective priority data.

2. The method of claim 1, wherein the rendering comprises allocating computational resources to render respective objects according to the priority data.

3. The method of claim 2, wherein allocating the computational resources comprises controlling one or more of following options based on the priority data:order of computation;temporal resolution;spatial resolution;latent space bit-depth; orquantity of neural network layers utilized.

4. The method of claim 2 or 3, wherein allocating computational resources to rendering respective objects comprises allocating more computational resources to rendering respective objects with higher priority according to the respective priority data than are allocated to rendering respective objects with lower priority according to the respective priority data.D25016W0015. The method of any preceding claim, further comprising identifying, from the input media, a second set of media objects that are pre-rendered and are different from the first set of media objects.

6. The method of claim 5, further comprising:combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects to generate a media file: and causing storage or transmission of the media file.

7. The method of claim 6, wherein combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects comprises combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects with one or more alpha maps.

8. The method of any preceding claim, wherein the method is executed at a client device.

9. The method of claim 8, further comprising:identifying a second set of media objects that are pre-rendered and are different from the first set of media objects;combining the generative media objects with the second set of media objects to obtain combined media data; andcausing playback of the combined media data using one or more media output devices of the client device.

10. The method of any preceding claim, wherein:the plurality of object types includes audio object types and / or video object types; anddetermining the priority data comprises determining audio priority data for the audio object types and / or determining video priority data for the video object types.D25016W00111. The method of any preceding claim, wherein the plurality of object types includes multimodal object types, and determining the priority data comprises determining multimodal priority data for the multimodal object types.

12. The method of any preceding claim, wherein the input media is received wirelessly from an audio / video media transmitter.

13. A system comprising:one or more computers and one or more storage devices on which are stored instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:receiving input media;identifying, from the input media, a first set of media objects corresponding to objects generated by one or more Al generative models;classifying the first set of media objects into a plurality of object types, wherein the classifying is based on one or more perceptual criteria;determining priority data for respective objects in the first set of media objects based on the respective determined object types, wherein a first object type associated with a first perceptual criterion is given a higher priority than a second object type associated with a second perceptual criterion that is determined to be less noticeable than the first perceptual criterion; andrendering generative media objects from respective objects in the first set of media objects based on their respective priority data.

14. The system of claim 13, wherein the rendering comprises allocating computational resources to render respective objects according to the priority data.

15. The system of claim 14, wherein allocating the computational resources comprises controlling one or more of following options based on the priority data:order of computation;temporal resolution;spatial resolution;latent space bit-depth; orquantity of neural network layers utilized.D25016W00116. The system of claim 14, wherein allocating computational resources to rendering respective objects comprises allocating more computational resources to rendering respective objects with higher priority according to the respective priority data than are allocated to rendering respective objects with lower priority according to the respective priority data.

17. The system of any one of claims 13 to 16, wherein the operations further comprise identifying, from the input media, a second set of media objects that are pre-rendered and are different from the first set of media objects.

18. The system of claim 17, wherein the operations further comprise:combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects to generate a media file; and causing storage or transmission of the media file.

19. The system of claim 18, wherein combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects comprises combining data corresponding to the rendered generative media objects and data corresponding to the second set of media objects with one or more alpha maps.

20. One or more non-transitory computer storage media encoded with computer program instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:receiving input media;identifying, from the input media, a first set of media objects corresponding to objects generated by one or more Al generative models;classifying the first set of media objects into a plurality of object types, wherein the classifying is based on one or more perceptual criteria;determining priority data for respective objects in the first set of media objects based on the respective determined object types, wherein a first object type associated with a first perceptual criterion is given a higher priority than a second object type associated with a second perceptual criterion that is detemrined to be less noticeable than the first perceptual criterion; andD25016W001rendering generative media objects from respective objects in the first set of media objects based on their respective priority data.