Generating content with position control
The system addresses the lack of spatial control in generative AI by using a trained model to generate content with precise object placement and class maps, allowing for detailed control over object locations and sizes in generated content.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-01
- Publication Date
- 2026-04-09
AI Technical Summary
Generative artificial intelligence systems lack fine-grained control over the spatial arrangement and inclusion of objects in generated content, making it difficult for content creators to specify precise locations and sizes of objects within the output.
A method and system that utilize a trained generative machine learning model to generate content items with precise control over object locations and classes, using a content class map to indicate the spatial arrangement of objects, and iteratively refine the content and class map using confidence values and exclusion lists.
Enables content creators to precisely control the spatial arrangement of objects within generated content, such as images, videos, and audio, by aligning object locations with user-specified class information, enhancing the granularity of content generation.
Smart Images

Figure US2025048928_09042026_PF_FP_ABST
Abstract
Description
PATENTDolby Ref. D24126WO01GENERATING CONTENT WITH POSITION CONTROLCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 703481, filed on 4 October 2024, and European Patent Application No. 24215593.5, filed on 26 Nov. 2024, each of which is incorporated by reference herein in its entirety.TECHNICAL FIELD
[0002] This disclosure pertains to systems, methods, and media for generating content.BACKGROUND
[0003] Generative artificial intelligence (Al) is being increasingly used to create content (e.g., text, images, videos, sound, etc.) For example, generative Al may be used to create an image, e.g., based on a prompt such as “create an image of a dog on the beach.” However, it can be difficult for a content creator using generative Al to have more fine-grained control of the output of the generative Al.NOTATION AND NOMENCLATURE
[0004] Throughout this disclosure, including in the claims, the expression performing an operation “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to, the signal or data) is used in a broad sense to denote performing the operation directly on the signal or data, or on a processed version of the signal or data (e.g., on a version of the signal that has undergone preliminary filtering or pre-processing prior to performance of the operation thereon).
[0005] Throughout this disclosure including in the claims, the expression “system” is used in a broad sense to denote a device, system, or subsystem. For example, a subsystem that implements a decoder may be referred to as a decoder system, and a system including such a subsystem (e.g., a system that generates X output signals in response to multiple inputs, in which the subsystem generates M of the inputs and the other X - M inputs are received from an external source) may also be referred to as a decoder system.
[0006] Throughout this disclosure including in the claims, the term “processor” is used in a broad sense to denote a system or device programmable or otherwise configurable (e.g., withPATENTDolby Ref. D24126WO01 software or firmware) to perform operations on data (e.g., audio, or video or other image data).Examples of processors include a field-programmable gate array (or other configurable integrated circuit or chip set), a digital signal processor programmed and / or otherwise configured to perform pipelined processing on audio or other sound data, a programmable general purpose processor or computer, and a programmable microprocessor chip or chip set.PATENTDolby Ref. D24126WO01SUMMARY
[0007] Methods, systems, and media for generating content are provided herein. According to some embodiments, a method for generating content may involve receiving a request to generate a content item using a trained generative machine learning model, wherein the request comprises content class information indicating relative locations of one or more content classes within the content item. The method may further involve generating the content item using the trained generative machine learning model based at least in part on the request. The method may further involve concurrently with generating the content item, generating a content class map using the trained generative machine learning model, wherein the content class map indicates locations of one or more content classes within the content item that are aligned to the content class information included in the request.
[0008] In some examples, the content item is an image comprising a plurality of patches of pixels, and wherein the content class map indicates a content class of a plurality of content classes on a patch-by-patch basis.
[0009] In some examples, the content class information is generated using a large language model (LLM) based on a prompt provided to the LLM.
[0010] In some examples, the content class information of the request is a partially filled content class map, and wherein the content class map generated by the trained generative machine learning model is a filled in version of the partially filled content class map.
[0011] In some examples, the request comprises indications of one or more content classes not to be included in the generated content item.
[0012] In some examples, the request comprises a partially generated content item, and wherein the content item generated by the trained generative machine learning model is a completed version of the partially generated content item.
[0013] In some examples, the request comprises a previously generated content item and instructions to modify at least one content class included in the previously generated content item in the generated content item.
[0014] In some examples, the content class information is generated based on a spoken or text input from a user.PATENTDolby Ref. D24126WO01
[0015] In some examples, the content class information is received via a user interface that provides a blank content class map and receives user input specifying content classes to be placed in specified locations within the blank content class map.
[0016] In some examples, the content item comprises video content.
[0017] In some examples, the content item comprises audio content.
[0018] In some examples, the trained generative machine learning model iteratively generates portions of the content item, and wherein the trained generative machine learning model concurrently generates portions of the content class map. In some examples, iteratively generating the portions of the content item comprise generating a set of content item codes and selecting a subset of the generated set of content item codes based on confidence values associated with the set of content item codes, and wherein iteratively generating the portions of the content class map comprise generating a set of class map codes and selecting a subset of the set of class map codes based on confidence values associated with the set of class map codes.
[0019] In some examples, the portions of the content item are discrete. In some examples, the trained generative machine learning model utilizes a MaskGIT architecture.
[0020] In some examples, the trained generative machine learning is a generative adversarial network (GAN).
[0021] According to some embodiments, a method of training a generative machine learning model may involve obtaining a training set, the training set comprising a plurality of content items and a corresponding plurality of content class maps, wherein each content class map of the plurality of content class maps indicates locations of one or more content classes included in a corresponding content item. The method may further involve, for each content item of the plurality of content items: masking a portion of the content item and a corresponding portion of a content class map corresponding to the content item; providing the masked content item and the masked content class map to the generative machine learning model; obtaining, as an output of the generative machine learning model, a predicted content item and a predicted content class map; and updating weights associated with the generative machine learning model based on a difference between the predicted content item and the content item and a difference between the predicted content class map and the content class map corresponding to the content item. The method may further involve providing a trained generative machine learning model based on final weights associated with the generative machine learning model, wherein the trainedPATENTDolby Ref. D24126WO01 generative machine learning model is usable to generate a new content item not included in the training set and a corresponding content class map indicating locations of one or more content classes within the newly generated content item.
[0022] In some examples, the trained generative machine learning model utilizes a MaskGIT architecture.
[0023] In some examples, the new content item is one of: a still image, a video, or audio content.
[0024] In some examples, the new content item is a still image comprising a plurality of patches of pixels, and wherein the corresponding content class map indicates locations of the one or more content classes included in newly generated content item on a patch-by-patch basis.
[0025] In some examples, the weights associated with the generative machine learning model are updated using a cross-entropy loss function.
[0026] At least some aspects of the present disclosure may be implemented via an apparatus. For example, one or more devices may be capable of performing, at least in part, the methods disclosed herein. In some implementations, an apparatus is, or includes, an audio processing system having an interface system and a control system. The control system may include one or more general purpose single- or multi-chip processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or combinations thereof.
[0027] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Note that the relative dimensions of the following figures may not be drawn to scale.BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a block diagram of an example content generation system in accordance with some embodiments.
[0029] Figure 2A illustrates an example input to a content generation system in accordance with some embodiments.PATENTDolby Ref. D24126WO01
[0030] Figure 2B illustrates example output of a content generation system in accordance with some embodiments.
[0031] Figure 3 illustrates techniques for modifying content in accordance with some embodiments.
[0032] Figures 4A, 4B, and 4C are block diagrams illustrating example input configurations for content generation systems in accordance with some embodiments.
[0033] Figure 5 is a block diagram depicting an inference process of an example machine learning model in accordance with some embodiments.
[0034] Figure 6 is a block diagram depicting a training process for the machine learning model illustrated in Figure 5 in accordance with some embodiments.
[0035] Figure 7 is a flowchart of an example process for generating a content item in accordance with some embodiments.
[0036] Figure 8 is a flowchart of an example process for training a machine learning model to generate a content item in accordance with some embodiments.
[0037] Figure 9A shows a block diagram that illustrates examples of components of an apparatus capable of implementing various aspects of this disclosure.
[0038] Figure 9B illustrates a schematic block diagram of an example device architecture that may be used to implement various aspects of the present disclosure.
[0039] Figure 9C illustrates a schematic block diagram of an example CPU implemented in the device architecture of Figure 9B that may be used to implement various aspects of the present disclosure.
[0040] Like reference numbers and designations in the various drawings indicate like elements.PATENTDolby Ref. D24126WO01DETAILED DESCRIPTION OF EMBODIMENTS
[0041] Generative artificial intelligence (Al) is being increasingly used to create content (e.g., text, images, videos, sound, etc.) For example, generative Al may be used to create an image, e.g., based on a prompt such as “create an image of a dog on the beach.”However, it can be difficult for a content creator using generative Al to have more fine-grained control of the output of the generative Al. For example, a content creator may want to specify relative locations and / or relative sizes of objects within a generated image. As a more particular example, a content creator may want to create an image of a dog on the beach, where the dog is in the middle of the image, where water is in the foreground, etc. However, it can be difficult for content creators to specify content to be created with such granularity. For example, a model may have difficulty interpreting a detailed prompt and generating a content item according to spatial locations specified in the prompt.
[0042] Disclosed herein are techniques for providing control in content generation. In particular, the techniques disclosed herein may be used to create content using generative Al, where a user can specify locations of various objects (sometimes referred to herein as “classes” or “content classes”) within the content, objects (or classes) that are not to be included in the generated content, etc. For example, the techniques disclosed herein may allow a user to specify class information that indicates general locations of various objects to be included in a generated content item, relative sizing information, etc. A trained machine learning model may then generate content based on the specified class information where the objects included in the content adhere to the input class information.
[0043] Note that, as used herein, a “class” or “a content class” generally refers to an object included in a content item. For example, an image of a dog on the beach may include a first class corresponding to a dog, and a second class corresponding to a beach.
[0044] Figure 1 illustrates a block diagram of an example content generation system in accordance with some embodiments. As illustrated, content generation system 102 may take, as input, class information 105. Class information 105 may indicate types of content to be included in content 106, desired locations of particular types of content to be included in content 106, etc. For example, class information 105 may indicate that a first region is to have portions of a first object, a second region is to have portions of a second object, etc.PATENTDolby Ref. D24126WO01
[0045] As illustrated, content generation system 102 may generate, as an output, content 106. Content 106 may include image content, video content, audio content, text-based content, etc. Additionally, content generation system 102 may generate class map 108. Class map 108 may indicate classes of content at different locations of content 106. Note that examples of class information 105, content 106, and class map 108 are shown in and described below in connection with Figures 2A and 2B. Additionally, note that the particular inputs and outputs shown in Figure 1 are merely by way of example, and other possible input and output combinations are shown in and described below in connection with Figures 4 A, 4B, and 4C.
[0046] It should be understood that content generation system 102 may include a trained machine learning model having any suitable type of architecture capable of producing generative content. For example, in some implementations, the trained machine learning model may include a transformer network configured to utilize an attention mechanism. In some embodiments, portions of the content item may be generated in a discrete manner. In some such embodiments, the machine learning model may have an architecture similar to the Masked Generative Iterative Transformer (MaskGIT) architecture. Alternatively, in some embodiments, portions of the content item may be continuous. In such embodiments, the machine learning model may utilize a diffusion model. In some embodiments, the trained machine learning model may utilize a generative-adversarial network (GAN).
[0047] In some implementations, generated content may include a generated image. The image may include patches of pixels. By way of example, an image may be 256 x 256 pixels, which is comprised of patches of 16 x 16 pixels each. For a given generated image, a class map may be generated. The class map may indicate locations of classes of content within a given generated image. In some implementations, the class map may have the same resolution of the generated image. For example, given a 16 patch x 16 patch of generated image codes, the class map may also be 16 patches x 16 patches (e.g., where each patch comprises 16 x 16 pixels). In particular, each patch of the class map may indicate a class of content in a generated image. Accordingly, where the generated image includes one or more objects (e.g., a dog, a rock, a cookie, etc.), the class map may indicate locations within the generated image of each of the one or more objects.
[0048] Turning to Figure 2A, an example of class information 202 that may be provided as input to a content generation system is shown in accordance with some embodiments. For example, class information 202 of Figure 2A may correspond to class information 105 providedPATENTDolby Ref. D24126WO01 as input to content generation system 102 as shown in Figure 1. As illustrated, class information 202 includes regions 202a, 202b, 202c, 202d, and 202e each corresponding to a different class of object to be generated within an output generated content item. Each region of 202a-202e indicates the class of object as well as a relative location of the object. For example, region 202a indicates that sky is to be generated in an upper left region of the generated content item. Region 202d indicates that a rock is to be placed in the lower left of the generate content item.
[0049] Figure 2B illustrates two example generated images, 252 and 256, as well as corresponding generated class maps (254 and 258, respectively) that may be generated by a content generation system responsive to receiving class information 202 as an input. Note that each of generated images 252 and 256 include all of the classes of objects specified in class information 202 (e.g., sky, a tree, a rock, a dog, and grass), which are each in the general location specified in class information 202. Class maps 254 and 258 indicate for each patch of pixels in corresponding generated images 252 and 256 respectively, the object class the patch of pixels is depicting. For example, class map 254 depicts regions of generated image 252 corresponding to a dog, regions corresponding to sky, etc. Note that class information 202 is generally sparse and indicates general relative locations of objects in a to-be-generated content item, whereas class maps 254 and 258 indicate object classes at every location within generated images 252 and 256.
[0050] In some embodiments, a generated image may be modified, e.g., to modify a relative size and / or location of objects within the image. For example, a generated class map indicating object classes for objects within the generated image may be generated. A generated class map may be modified to generate a modified class map. The generated class map may be modified based on instructions which may be typed, spoken, etc. Alternatively, the generated class map may be directly modified, e.g., using a user interface in which a user can adjust portions of the generated class map. The modified class map may form the input class information to a model that is part of content generation system (e.g., the modified class map may correspond to class information 105 provided to content generation system 102 of Figure 1), which may then generate a new image and a corresponding new class map.
[0051] Figure 3 illustrates a diagram depicting modification of a generated image in accordance with some embodiments. As illustrated, a generated image 304 has an associated class map 306. Generated class map depicts object classes associated with each patch of pixels in the generated image 304. For example, generated class map 306 illustrates a patch of pixels corresponding to the bean in generated image 304. A modification instruction is used toPATENTDolby Ref. D24126WO01 generated modified class map 308. An example of the modification instruction is “make the tree taller.” As illustrated, the portion of generated class map 306 corresponding to the tree is made taller in modified class map 308. Note additionally that modified class map 308 has a region that is masked, or blanked relative to original class map 306, as that is a region that model 302 is to fill in to generate new image 310. The modified class map 308 is provided as input to model 302 (e.g., corresponding to class information 105 being provided as input to content generation system 102 in Figure 1). Model 302 is configured to generate, as output, new image 310 and associated new class map 312. Note that in new image 310, the portions of modified class map 308 that were not masked or blanked out (e.g., corresponding to the bear, the rock, etc.), new image 310 is the same relative to generated image 304. Additionally note that the tree in new image 310 is taller than the tree in generated image 304, which corresponds to the change in modified class map 308 relative to generated class map 306. Additionally, note the changes in new class map 312 relative to generated class map 306.
[0052] It should be noted that, in some embodiments, a large language model (LLM) agent may be utilized to parse an input text prompt (or spoken prompt) and generate a corresponding class map or modified class map. For example, in some embodiments, a text prompt may be provided to an LLM agent (e.g., “generate an image of a dog sitting on a rock with a tree in the background on the right”), and the LLM agent may generate a class map that corresponds to the prompt. For example, the LLM agent may generate a matrix with elements corresponding to objects identified in the prompt and with locations identified based on the prompt. Given the example prompt above, the LLM agent may generate a matrix with elements corresponding to a tree class in an upper right region of the matrix. Similarly, an LLM agent may be used to modify an input original class map based on an original image and a text prompt, e.g., as shown in and described above in connection with Figure 3. For example, the LLM agent may be receive original image 304, original class map 306, and the modification instruction, and may generate, as output, modified class map 308.
[0053] Turning to Figure 4A, an example configuration of a content generation system 422 similar to that shown in and described above in connection with Figure 1A is illustrated in accordance with some embodiments. Content generation system 422 receives a partial class map 425. Partial class map 425 may be a partially filled in version of the final class map generated by content generation system 422. An example is modified class map 308 of Figure 3 or class map 202 of Figure 2A. The partial class map 425 may be specified as a matrix of the same size as the final class map 428, where some or all the matrix elements are specified objects. The final classPATENTDolby Ref. D24126WO01 map 428 may align with the partial class map 425. For example, final class map 428 may include all the same objects in the same general locations as specified in a partial class map, with the remaining portions of the class map filled in by content generation system 422. Final class map 428 corresponds to content 426 generated by content generation system 422.
[0054] Turning to Figure 4B, an example configuration of a content generation system 452 similar to that shown in and described above in connection with Figure 4B is shown in accordance with some embodiments. Unlike what is shown in Figure 4B, in addition to taking, as input, partial class map 425, content generation system 452 takes, as input, partial image code 455. Partial image code 455 may be a portion of content 456. Content 456 may effectively “fill in” the missing portions of partial image code 455, in a process sometimes referred to as “inpainting.” Class map 458 may represent a filled in version of partial class map 425, similar to what is described above in connection with Figure 4A.
[0055] Turning to Figure 4C, an example configuration of a content generation system 472 similar to that shown in and described above in connection with Figure 4B is shown in accordance with some embodiments. Unlike what is shown in Figure 4B, rather than taking, as input, a partial image code, content generation system 472 takes, as input, an excluded class list 475. Excluded class list 475 may specify one or more objects that are not to be included in generated content 476. For example, excluded class list 475 may specify that generated content 476 is not to include a dog, is not to include a rock, is not to include a tree, etc. As illustrated, content generation system 472 may additionally take a partial class map 474. By way of example, partial class map 474 may indicate locations of one or more objects, such as a dog and a beach, and excluded class list 475 may specify that generated content 476 (which may include an image of a dog on a beach) is not to include, e.g., trees, people, etc. Class map 478 may locations of objects / classes within content 476, as described above.
[0056] It should be noted that Figures 1 and 4A-4C illustrate the same system which may be configured to take different inputs. In other words, a single content generation system may be configured to take, as input, a partial class map, a partial image map / code, and / or an excluded class list, and may receive none, some, or all of these inputs.
[0057] In some implementations, a content generation system may iteratively generate a map of image codes and a class map concurrently. For example, during each iteration, predicted image codes may be generated by the model, and the K image codes with the highest confidence values may be selected (e.g., frozen) for use in the subsequent iteration. This technique forPATENTDolby Ref. D24126WO01 generating a map of image codes is similar to that implemented in the Mask Generative Image Transformer (MaskGIT) algorithm. Using the techniques disclosed herein, a class map may be generated concurrently with the map of image codes in an iterative fashion. In particular, during each iteration, class map codes may be predicted for elements of a class map, and the J class map codes with the highest confidence values may be selected (e.g., frozen) for use in the subsequent iteration. In some implementations, the MaskGIT algorithm may be modified to generate a predicted class map at each iteration concurrent with prediction of the map of image codes.
[0058] Figure 5 is a block diagram depicting an inference process of an example machine learning model in accordance with some embodiments. The implementation illustrated in Figure 5 may be used at inference time, e.g., to generate new content items and class maps. As illustrated, a trained machine learning model 502 (which may be based on the MaskGIT algorithm) may receive a map of image codes 512, which may start empty or be partially filled. At each iteration, each element in the generated map of image codes 516 may have a corresponding confidence value indicating a confidence value or metric associated with the prediction. Selection block 514 may select the K elements of generated map of image codes 516 with the highest confidence values. The map of the image codes with the selected elements (e.g., selected by selection block 514) may then form the new input map of image codes 512 provided to trained machine learning model 502. After a predetermined number of iterations have been completed, the generated map of image codes 516 will be fully filled in. The predetermined number of iterations may be 8, 16, 32, 50, etc. The final, fully filled in generated map of image codes 516 may be used to generate content item 518 by providing the predicted map of image codes 516 to a trained decoder 534 configured to output content item 518.
[0059] In some implementations, for example, as shown in and described above in connection with Figure 4B, trained machine learning model 502 may perform “image inpainting” by filling in a partially filled in input image. In other words, generated content item 518 may represent a “filled in” version of a partially filled in initial image. In such embodiments, an initial image 504 may be provided to an encoder 506 configured to generate embeddings associated with initial image 504. The embeddings may be quantized by quantizer 508. The quantized embeddings may be provided to masking block 510 to generate the initial masked map of image codes 512 used in the first iteration of trained machine learning model 502. Note that the output of masking block 510 generally corresponds to partial image code 455 of Figure 4B, which then becomes the initial map of image codes 512 provided as input to trained machine learning model 502. In some implementations, the output of quantizer 508 may be “masked” using a user interface,PATENTDolby Ref. D24126WO01 which may allow a user to erase or “mask out” portions of the initial image that are to be inpainted.
[0060] Concurrently with the process described above to generate content item 518, trained machine learning model 502 may generate a class map 532. Trained machine learning model 522 may begin with an initial class map 522, which may be fully blank (e.g., fully masked), or may be initialized based on partial / full class map 520 (which may at least partially specify objects and / or locations of objects within generated content item 518). After generating map of image codes 516 in a given iteration, trained machine learning model 502 may generate a class map 530. Selection block 524 may be used to select the J class codes with the highest confidence values, where the selected class codes are frozen and form the basis for the class map 522 provided in the next iteration. Note that, in instances in which one or more excluded classes are provided in excluded class list 528, the predicted class map 530 may be filtered using filter block 526. In particular, filter block 526 may remove any predicted classes that correspond to classes on excluded class list 528 prior to the J class codes with the highest confidence values being selected.
[0061] Note that trained machine learning model 502 may be configured to generate predicted map of image codes 516 in a given iteration based on the class map 522 from the previous iteration. For example, filter block 526 may set the confidence values for class codes corresponding to objects on excluded class list 528 to a very small value (e.g., negative infinity). This may ensure that selection block 524 does not select class codes corresponding to an excluded class. If a fixed image code (e.g., selected by selection block 514) has a class code included on excluded class list 528, the image code may be masked (e.g., blanked, erased, removed, etc.) at the end of a given iteration such that the image code is then re-generated or repredicted in the next iteration. This may ensure that content item 518 is generated without including any classes on excluded class list 528.
[0062] In some implementations, partial / full class map 520 may be generated based on a prompt 511. Prompt 511 may be text-based, spoken, etc. Prompt 511 may be processed by large language model (LLM) 521. For example, LLM 521 may translate prompt 511 to partial / full class map 520. As a more particular example, LLM 521 may parse prompt 511 to identify one or more objects specified in prompt 511 and locations of the one or more objects specified in prompt 511, and may generate partial / full class map 520 based on the parsing.PATENTDolby Ref. D24126WO01
[0063] Figure 6 is a block diagram depicting a training process for the machine learning model illustrated in Figure 5 in accordance with some embodiments. In particular, machine learning model 602 may be a version of machine learning model 502 of Figure 5 prior to training. The training process may iterate through a set of training images, each associated with a ground truth class map indicating locations of objects in the training image. For each training image 604, an encoder block 606 may generate an embedding. The embedding may be quantized by quantizer block 608 to generate a map of image codes. A masking block 610 may mask the map of image codes to generate masked image codes 612. Image codes generated by quantizer block 608 may be randomly masked by masking block 610 to generate masked image codes 612. Note that masking block 610 may randomly select portions of image codes that are masked, which may be different in each training step. Machine learning model 602 may then generate predicted image codes 616 by predicting the image codes. Loss function block 650 may determine a loss based on a difference between predicted image codes 616 and the ground truth map of image codes generated by quantized block 608. In some embodiments, loss function block 650 may determine the loss using a cross entropy loss function. The loss may be used to update weights associated with machine learning model 602.
[0064] Concurrently with generating predicted image codes, machine learning model 602 may be trained to predict the class map for each training sample. In particular, for a given training image, the ground truth class map 620 may be masked by masking block 621 to generate a masked class map 622. For example masked class map 622 may be a blank matrix. The masked class map 622 may be provided as input to machine learning model 602, which may generate predicted class map 630. Loss function block 652 may determine a loss based on a difference between predicted class map 630 and ground truth class map 620. In some embodiments, loss function block 652 may determine the loss using a cross entropy loss function. The loss may be used to update weights associated with machine learning model 602.
[0065] It should be noted that a single image may be processed differently during inference (e.g., as shown in Figure 5) compared to during the training process (e.g., as shown in Figure 6). During the inference process, a single image is generated over a fixed number of steps (sometimes referred to herein as iterations), where, in each step, K image codes and J class codes are filled in. During the training process, each training image and ground truth class map is processed in a training step by masking L% of the image codes for the training image and M% of the class codes (the masked codes being chosen at random). The model is then penalized for incorrect predictions of the masked image codes and / or masked class codes (e.g., using a lossPATENTDolby Ref. D24126WO01 function as described above). It should be understood that L% of image codes masked during training may be unrelated to the K image codes selected at each step or iteration during the inference process, and similarly, the M% of class codes masked during training may be unrelated to the J class codes selected at each step or iteration during the inference process.
[0066] Note that cross entropy loss may generally be used to penalize a model and / or update weights when the model is trained to select from a discrete set of possibilities. In the example configuration shown in Figure 6, the model is trained to predict discrete image codes and discrete class codes. In some implementations, other types of loss functions may be used, e.g., utilizing mean square error (MSE) based on a probability of classification, or the like.
[0067] The training process described above may repeat for each training sample of a set of training samples. Note that multiple sources (e.g., two, three, four, ten, fifty, etc.) of training samples may be represented in the training set such that the machine learning model is trained to generate a class map indicating a class of a set of multiple candidate classes.
[0068] Figure 7 is a flowchart of an example process 700 for concurrently generating a content item and a class map in accordance with some embodiments. Blocks of process 700 may be executed by one or more processors and / or control systems of a computing device, such as a laptop computer, a server device, a remote cloud computing device, etc. An example of such a control system is control system 910 of Figure 9A. In some implementations, blocks of process 700 may be executed in an order other than what is shown in Figure 7. In some embodiments, two or more blocks of process 700 may be executed substantially in parallel. In some embodiments, one or more blocks of process 700 may be omitted.
[0069] Process 700 can begin at 702 by receiving a request to generate a content item using a trained generative machine learning model where the request comprises content class information indicating relative locations of one or more content classes within the content item. In some implementations, the generative machine learning model may generate discrete portions of the content item, e.g., using a masked transformer architecture (e.g., based on MaskGIT, or a similar type of algorithm). Alternatively, in some implementations, the generative machine learning model may generate continuous portions of the content item, e.g., using a diffusion model.
[0070] In some embodiments, the request may include content class information indicating relative locations of one or more content classes with the content item. For example, the contentPATENTDolby Ref. D24126WO01 class information may indicate relative locations of objects to be included in the generated content item, e.g., that a dog is to be on top of a rock, etc. In some embodiments, the request may include a partially filled in content item that is to be fully filled in by the trained generative machine learning mode (e.g., by performing “image inpainting” as described above). A partially filled in content item may be specified as a partial image. Note that, in some implementations, the content class information may be generated by an LLM agent based on a prompt (which may be text, spoken, etc.) provided to the LLM agent, as described above in connection with Figure 5.
[0071] At 704, process 700 can generate the content item using the trained generative machine learning model based at least in part on the request. For example, in instances in which the trained generative machine learning model iteratively generates discrete portions of the content item, process 700 can mask predicted representations of the content item, predict the masked regions, and select a subset of the predicted regions based on confidence values, as shown in and described above in connection with Figure 5. In instances in which a partially filled in content item is provided, process 700 may generate a quantized embedding of the partially filled in content item (e.g., as shown in and described above in connection with Figure 5) and utilize the quantized embedding an initial map of image codes at the start of a first iteration.
[0072] At 706, concurrently with generating the content item, process 700 can generate a content class map using the trained generative machine learning model, where the content class map indicates locations of one or more content classes within the content item that are aligned to the content class information included in the request. The content class map may indicate (e.g., for each patch of pixels of a generated image) a corresponding content class. For example, as shown in and described above in connection with Figure 2B, a content class map may indicate that a first patch of pixels of a generated image includes a dog, and a second patch of pixels of a generated image includes a rock.
[0073] Figure 8 is a flowchart of an example process 800 for training a machine learning model to concurrently generate a content item and a class map in accordance with some embodiments. Note that the process depicted in Figure 8 is associated with the model architecture depicted in Figures 5 and 6, and a different training process may be used for other architectures, such as a diffusion model. Blocks of process 800 may be executed by one or more processors and / or control systems of a computing device, such as a laptop computer, a server device, a remote cloud computing device, etc. An example of such a control system is control system 910 of Figure 9A. In some implementations, blocks of process 800 may be executed inPATENTDolby Ref. D24126WO01 an order other than what is shown in Figure 8. In some embodiments, two or more blocks of process 800 may be executed substantially in parallel. In some embodiments, one or more blocks of process 800 may be omitted.
[0074] Process 800 can begin at 802 by obtaining a training set, the training set comprising a plurality of content items and a corresponding plurality of content class maps, where each content class map of the plurality of content class maps indicates locations of one or more content classes included in the corresponding content item. Each content class map in the training set may be considered a ground truth class map (e.g., as shown in and described above in connection with Figure 6). The content items may include images, videos, audio content, etc.
[0075] At 804, process 800 can, for each content item in the plurality of content items, mask a portion of the content item and a portion of a content class map corresponding to the content item. The masked portion for the content item and the masked portion of the content class map may be randomly selected. Note that, prior to masking the portion of the content item, process 800 may generate a quantized embedding of the content item, which may serve as, e.g., the initial map of an image code, which may then be masked, as shown in and described above in connection with Figure 6. By masking the portion of the content item and the portion of the content class map, process 800 may effectively “hide” a portion of the content item and the portion of the content class map which is then to be predicted by the machine learning model.
[0076] At 806, process 800 can provide the masked content item and the masked content class map to the generative machine learning model. In some implementations, the generative machine learning model may additionally be provided with a prompt, which may be text-based, spoken, etc.
[0077] At 808, process 800 can obtain, as an output of the generative machine learning model, a predicted content item and a predicted content class map. Note that these may be generated concurrently by the generative machine learning model, e.g., as part of one iteration.
[0078] At 810, process 800 can update weights associated with the generative machine learning model based on a difference between the predicted content item and the content item and a difference between the predicted content class map and the ground truth content class map corresponding to the content item. The weights may be updated based on a loss. The loss may be a cross-entropy loss.PATENTDolby Ref. D24126WO01
[0079] At 812, process 800 can determine whether training is complete. For example, process 800 may determine whether all training samples of the training set have been iterated through. As another example, process 800 may determine whether the machine learning model has achieved a target accuracy. In some implementations, process 800 may determine whether training is complete based on whether a fixed number of training epochs have been completed. In some embodiments, process 800 may determine that training is to be completed early (e.g., prior to the fixed number of training epochs being completed). For example, process 800 may generate sample images after a predetermined number of training epochs (e.g., every 5 epochs, every 10 epochs, etc.). Responsive to determining that a target image quality has been achieved and / or that image quality is no longer improving by more than a predetermined threshold over successive training epochs, training may be determined to be complete even if the target number of epochs have not been completed.
[0080] Responsive to determining, at 812, that training is not complete (“no” at 812), process 800 can loop back to block 804 and can obtain another content item from the training set. Process 800 can then loop through blocks 804-812 until training is complete.
[0081] Conversely, if, at 812, process 800 determines that training is complete (“yes” at 812), process 800 can proceed to block 814 and can provide a trained generative machine learning model based on final weights assocaited with the generative machine learning model. As described above, the trained generative machine learning model is usable to generate a new content item not included in the training set and a corresponding content class map indicating locations of objects included in the new content item.
[0082] Figure 9A is a block diagram that shows examples of components of an apparatus capable of implementing various aspects of this disclosure. As with other figures provided herein, the types and numbers of elements shown in Figure 9A are merely provided by way of example. Other implementations may include more, fewer and / or different types and numbers of elements. According to some examples, the apparatus 900 may be configured for performing at least some of the methods disclosed herein. In some implementations, the apparatus 900 may be, or may include, a television, one or more components of an audio system, a mobile device (such as a cellular telephone), a laptop computer, a tablet device, a smart speaker, or another type of device.
[0083] According to some alternative implementations the apparatus 900 may be, or may include, a server. In some such examples, the apparatus 900 may be, or may include, an encoder.PATENTDolby Ref. D24126WO01Accordingly, in some instances the apparatus 900 may be a device that is configured for use within an audio environment, such as a home audio environment, whereas in other instances the apparatus 900 may be a device that is configured for use in “the cloud,” e.g., a server.
[0084] In this example, the apparatus 900 includes an interface system 905 and a control system 910. The interface system 905 may, in some implementations, be configured for communication with one or more other devices of an audio environment. The audio environment may, in some examples, be a home audio environment. In other examples, the audio environment may be another type of environment, such as an office environment, an automobile environment, a train environment, a street or sidewalk environment, a park environment, etc. The interface system 905 may, in some implementations, be configured for exchanging control information and associated data with audio devices of the audio environment. The control information and associated data may, in some examples, pertain to one or more software applications that the apparatus 900 is executing.
[0085] The interface system 905 may, in some implementations, be configured for receiving, or for providing, a content stream. The content stream may include audio data. The audio data may include, but may not be limited to, audio signals. In some instances, the audio data may include spatial data, such as channel data and / or spatial metadata. In some examples, the content stream may include video data and audio data corresponding to the video data.
[0086] The interface system 905 may include one or more network interfaces and / or one or more external device interfaces (such as one or more universal serial bus (USB) interfaces). According to some implementations, the interface system 905 may include one or more wireless interfaces. The interface system 905 may include one or more devices for implementing a user interface, such as one or more microphones, one or more speakers, a display system, a touch sensor system and / or a gesture sensor system. In some examples, the interface system 905 may include one or more interfaces between the control system 910 and a memory system, such as the optional memory system 915 shown in Figure 9A. However, the control system 910 may include a memory system in some instances. The interface system 905 may, in some implementations, be configured for receiving input from one or more microphones in an environment.
[0087] The control system 910 may, for example, include a general purpose single- or multichip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components.PATENTDolby Ref. D24126WO01
[0088] In some implementations, the control system 910 may reside in more than one device. For example, in some implementations a portion of the control system 910 may reside in a device within one of the environments depicted herein and another portion of the control system 910 may reside in a device that is outside the environment, such as a server, a mobile device (e.g., a smartphone or a tablet computer), etc. In other examples, a portion of the control system 910 may reside in a device within one environment and another portion of the control system 910 may reside in one or more other devices of the environment. For example, a portion of the control system 910 may reside in a device that is implementing a cloud-based service, such as a server, and another portion of the control system 910 may reside in another device that is implementing the cloud-based service, such as another server, a memory device, etc. The interface system 905 also may, in some examples, reside in more than one device. In some implementations, a portion of a control system may reside in or on an earbud.
[0089] In some implementations, the control system 910 may be configured for performing, at least in part, the methods disclosed herein. According to some examples, the control system 910 may be configured for implementing methods of generating content and class maps, training a model to generate such content, or the like.
[0090] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transitory media. Such non- transitory media may include memory devices such as those described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. The one or more non-transitory media may, for example, reside in the optional memory system 915 shown in Figure 9A and / or in the control system 910. Accordingly, various innovative aspects of the subject matter described in this disclosure can be implemented in one or more non-transitory media having software stored thereon. The software may, for example, generate content items and corresponding class maps, train a model to generate content items and corresponding class maps, etc. The software may, for example, be executable by one or more components of a control system such as the control system 910 of Figure 9A.
[0091] In some examples, the apparatus 900 may include the optional microphone system 920 shown in Figure 9A. The optional microphone system 920 may include one or more microphones. In some implementations, one or more of the microphones may be part of, or associated with, another device, such as a speaker of the speaker system, a smart audio device, etc. In some examples, the apparatus 900 may not include a microphone system 920. However,PATENTDolby Ref. D24126WO01 in some such implementations the apparatus 900 may nonetheless be configured to receive microphone data for one or more microphones in an audio environment via the interface system 910. In some such implementations, a cloud-based implementation of the apparatus 900 may be configured to receive microphone data, or a noise metric corresponding at least in part to the microphone data, from one or more microphones in an audio environment via the interface system 910.
[0092] According to some implementations, the apparatus 900 may include the optional loudspeaker system 925 shown in Figure 9A. The optional loudspeaker system 925 may include one or more loudspeakers, which also may be referred to herein as “speakers” or, more generally, as “audio reproduction transducers.” In some examples (e.g., cloud-based implementations), the apparatus 900 may not include a loudspeaker system 925. In some implementations, the apparatus 900 may include headphones. Headphones may be connected or coupled to the apparatus 900 via a headphone jack or via a wireless connection (e.g., BLUETOOTH).
[0093] Figure 9B illustrates a schematic block diagram of an example device architecture 901 (in this example, an apparatus 901) that may be used to implement various aspects of the present disclosure. The apparatus 901 of Figure 9B is an instance of the apparatus 900 of Figure 9A. Architecture 901 includes but is not limited to servers and client devices, systems, etc., which may be configured to perform the methods that are described with reference to any or all of Figures 7 and 8. As shown, the architecture 901 includes central processing unit (CPU) 941, which is capable of performing various processes in accordance with a program stored in, for example, read only memory (ROM) 942 or a program loaded from, for example, storage unit 948 to random access memory (RAM) 943. The CPU 941 may be, for example, an electronic processor 941. In these examples, the CPU 941 is an instance of the control system 910 of Figure 9 A and the ROM 942 and RAM 943 are instances of the memory system 915. In RAM 943, the data required when CPU 941 performs the various processes is also stored, as required. CPU 941, ROM 942, and RAM 943 are connected to one another via bus 944. Input / output ( I / O ) interface 945 is also connected to bus 944. The bus 944 and the RO) interface 945 are instances of the interface system 905 of Figure 9A.
[0094] The following components are connected to I / O interface 945: input unit 946, that may include a keyboard, a mouse, or the like; output unit 947 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 948 including a hard disk, orPATENTDolby Ref. D24126WO01 another suitable storage device; and communication unit 949 including a network interface card such as a network card (e.g., wired or wireless).
[0095] In some implementations, input unit 946 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0096] In some implementations, output unit 947 include systems with various number of speakers. Output unit 947 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0097] In some embodiments, communication unit 949 is configured to communicate with other devices (e.g., via a network). Drive 950 is also connected to I / O interface 945, as required. Removable medium 951 , such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 950, so that a computer program read therefrom is installed into storage unit 948, as required. A person skilled in the art would understand that although apparatus 901 is described as including the above-described components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alteration all fall within the scope of the present disclosure.
[0098] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 949, and / or installed from the removable medium 951, as shown in Figure 9B.
[0099] Figure 9C illustrates a schematic block diagram of an example CPU 941 implemented in the device architecture 901 of Figure 9B that may be used to implement various aspects of the present disclosure. The CPU 941 includes an electronic processor 960 and a memory 961. The electronic processor 960 is electrically and / or communicatively connected to the memory 961 for bidirectional communication. The memory 961 may store content generation software 962 and / or training software 563. The memory 961 may be, for example, a ROM, a RAM, or anotherPATENTDolby Ref. D24126WO01 non-transitory computer readable medium. The electronic processor 960 may implement the content generation software 962 stored in the memory 961 to perform, among other things, the method 700 of Figure 7. Additionally or alternatively, the electronic processor 960 may implement the training software 963 stored in the memory 961 to perform, among other things, the method 800 of Figure 8.
[0100] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 941 in combination with other components of Figure 9B), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor or other computing device (e.g., control circuitry). While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0101] Additionally, various blocks shown in the flowcharts may be viewed as method steps, and / or as operations that result from operation of computer program code, and / or as a plurality of coupled logic circuit elements constructed to carry out the associated function(s). For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the computer program containing program codes configured to carry out the methods as described above.
[0102] Some aspects of present disclosure include a system or device configured (e.g., programmed) to perform one or more examples of the disclosed methods, and a tangible computer readable medium (e.g., a disc) which stores code for implementing one or more examples of the disclosed methods or steps thereof. For example, some disclosed systems can be or include a programmable general purpose processor, digital signal processor, or microprocessor, programmed with software or firmware and / or otherwise configured to perform any of a variety of operations on data, including an embodiment of disclosed methods or steps thereof. Such a general purpose processor may be or include a computer system including anPATENTDolby Ref. D24126WO01 input device, a memory, and a processing subsystem that is programmed (and / or otherwise configured) to perform one or more examples of the disclosed methods (or steps thereof) in response to data asserted thereto.
[0103] Some embodiments may be implemented as a configurable (e.g., programmable) digital signal processor (DSP) that is configured (e.g., programmed and otherwise configured) to perform required processing on audio signal(s), including performance of one or more examples of the disclosed methods. Alternatively, embodiments of the disclosed systems (or elements thereof) may be implemented as a general purpose processor (e.g., a personal computer (PC) or other computer system or microprocessor, which may include an input device and a memory) which is programmed with software or firmware and / or otherwise configured to perform any of a variety of operations including one or more examples of the disclosed methods. Alternatively, elements of some embodiments of the inventive system are implemented as a general purpose processor or DSP configured (e.g., programmed) to perform one or more examples of the disclosed methods, and the system also includes other elements (e.g., one or more loudspeakers and / or one or more microphones). A general purpose processor configured to perform one or more examples of the disclosed methods may be coupled to an input device (e.g., a mouse and / or a keyboard), a memory, and a display device.
[0104] Example Embodiments:
[0105] Embodiment 1 : A method of generating content, the method comprising: receiving a request to generate a content item using a trained generative machine learning model, wherein the request comprises content class information indicating relative locations of one or more content classes within the content item; generating the content item using the trained generative machine learning model based at least in part on the request; and concurrently with generating the content item, generating a content class map using the trained generative machine learning model, wherein the content class map indicates locations of one or more content classes within the content item that are aligned to the content class information included in the request.
[0106] Embodiment 2: The method of embodiment 1, wherein the content item is an image comprising a plurality of patches of pixels, and wherein the content class map indicates a content class of a plurality of content classes on a patch-by-patch basis.PATENTDolby Ref. D24126WO01
[0107] Embodiment 3: The method of any one of embodiments 1 or 2, wherein the content class information is generated using a large language model (LLM) based on a prompt provided to the LLM.
[0108] Embodiment 4: The method of any one of embodiments 1-3, wherein the content class information of the request is a partially filled content class map, and wherein the content class map generated by the trained generative machine learning model is a filled in version of the partially filled content class map.
[0109] Embodiment 5: The method of any one of embodiments 1-4, wherein the request comprises indications of one or more content classes not to be included in the generated content item.
[0110] Embodiment 6: The method of any one of embodiments 1-5, wherein the request comprises a partially generated content item, and wherein the content item generated by the trained generative machine learning model is a completed version of the partially generated content item.
[0111] Embodiment 7: The method of any one of embodiments 1-6, wherein the request comprises a previously generated content item and instructions to modify at least one content class included in the previously generated content item in the generated content item.
[0112] Embodiment 8: The method of any one of embodiments 1-7, wherein the content class information is generated based on a spoken or text input from a user.
[0113] Embodiment 9: The method of any one of embodiments 1-8, wherein the content class information is received via a user interface that provides a blank content class map and receives user input specifying content classes to be placed in specified locations within the blank content class map.
[0114] Embodiment 10: The method of any one of embodiments 1-9, wherein the content item comprises video content.
[0115] Embodiment 11: The method of any one of embodiment 1-10, wherein the content item comprises audio content.
[0116] Embodiment 12: The method of any one of embodiments 1-11, wherein the trained generative machine learning model iteratively generates portions of the content item, andPATENTDolby Ref. D24126WO01 wherein the trained generative machine learning model concurrently generates portions of the content class map.
[0117] Embodiment 13: The method of embodiment 12, wherein iteratively generating the portions of the content item comprise generating a set of content item codes and selecting a subset of the generated set of content item codes based on confidence values associated with the set of content item codes, and wherein iteratively generating the portions of the content class map comprise generating a set of class map codes and selecting a subset of the set of class map codes based on confidence values associated with the set of class map codes.
[0118] Embodiment 14: The method of embodiment 12, wherein the portions of the content item are discrete.
[0119] Embodiment 15: The method of embodiment 14, wherein the trained generative machine learning model utilizes a MaskGIT architecture.
[0120] Embodiment 16: The method of any one of embodiments 1-15, wherein the trained generative machine learning is a generative adversarial network (GAN).
[0121] Embodiment 17: The method of any preceding embodiment, further comprising: modifying the generated content class map to generate a modified class map; providing the modified class map as input class information to the trained generative machine learning model; and generating, by the trained generative machine learning model, a new content item and a corresponding new class map.
[0122] Embodiment 18: The method of embodiment 17, wherein the generated content class map is modified based on instructions.
[0123] Embodiment 19: The method of embodiment 17, wherein the generated content class map is modified via a user interface configured to enable a user to adjust portions of the generated content class map.
[0124] Embodiment 20: A method of training a generative machine learning model, the method comprising: obtaining a training set, the training set comprising a plurality of content items and a corresponding plurality of content class maps, wherein each content class map of the plurality of content class maps indicates locations of one or more content classes included in a corresponding content item; for each content item of the plurality of content items: masking aPATENTDolby Ref. D24126WO01 portion of the content item and a corresponding portion of a content class map corresponding to the content item, providing the masked content item and the masked content class map to the generative machine learning model, obtaining, as an output of the generative machine learning model, a predicted content item and a predicted content class map, and updating weights associated with the generative machine learning model based on a difference between the predicted content item and the content item and a difference between the predicted content class map and the content class map corresponding to the content item; and providing a trained generative machine learning model based on final weights associated with the generative machine learning model, wherein the trained generative machine learning model is usable to generate a new content item not included in the training set and a corresponding content class map indicating locations of one or more content classes within the newly generated content item.
[0125] Embodiment 21: The method of embodiment 20, wherein the trained generative machine learning model utilizes a MaskGIT architecture.
[0126] Embodiment 22: The method of any one of embodiments 20 or 21, wherein the new content item is one of: a still image, a video, or audio content.
[0127] Embodiment 23: The method of any one of embodiments 20-22, wherein the new content item is a still image comprising a plurality of patches of pixels, and wherein the corresponding content class map indicates locations of the one or more content classes included in newly generated content item on a patch-by-patch basis.
[0128] Embodiment 24: The method of any one of embodiments 20-23, wherein the weights associated with the generative machine learning model are updated using a cross-entropy loss function.
[0129] Embodiment 25: A system comprising: one or more processors; and a non-transitory computer-readable medium storing instructions that, upon execution by the one or more processors, cause the one or more processors to perform any of the operations of embodiments 1- 21.
[0130] Embodiment 26: A non-transitory computer-readable medium storing instructions that, upon execution by one or more processors, cause the one or more processors to perform operations of any of embodiments 1 -24.PATENTDolby Ref. D24126WO01
[0131] While specific embodiments of the present disclosure and applications of the disclosure have been described herein, it will be apparent to those of ordinary skill in the art that many variations on the embodiments and applications described herein are possible without departing from the scope of the disclosure described and claimed herein. It should be understood that while certain forms of the disclosure have been shown and described, the disclosure is not to be limited to the specific embodiments described and shown or the specific methods described.
Claims
PATENTDolby Ref. D24126WO01CLAIMS1. A method of generating content, the method comprising: receiving a request to generate a content item using a trained generative machine learning model, wherein the request comprises content class information indicating relative locations of one or more content classes within the content item; generating the content item using the trained generative machine learning model based at least in part on the request; and concurrently with generating the content item, generating a content class map using the trained generative machine learning model, wherein the content class map indicates locations of one or more content classes within the content item that are aligned to the content class information included in the request.
2. The method of claim 1 , wherein the content item is an image comprising a plurality of patches of pixels, and wherein the content class map indicates a content class of a plurality of content classes on a patch-by-patch basis.
3. The method of any one of claims 1 or 2, wherein the content class information is generated using a large language model (LLM) based on a prompt provided to the LLM.
4. The method of any one of claims 1-3, wherein the content class information of the request is a partially filled content class map, and wherein the content class map generated by the trained generative machine learning model is a filled in version of the partially filled content class map.
5. The method of any one of claims 1-4, wherein the request comprises indications of one or more content classes not to be included in the generated content item.
6. The method of any one of claims 1-5, wherein the request comprises a partially generated content item, and wherein the content item generated by the trained generative machine learning model is a completed version of the partially generated content item.
7. The method of any one of claims 1-6, wherein the request comprises a previously generated content item and instructions to modify at least one content class included in the previously generated content item in the generated content item.PATENTDolby Ref. D24126WO018. The method of any one of claims 1-7, wherein the content class information is generated based on a spoken or text input from a user.
9. The method of any one of claims 1-8, wherein the content class information is received via a user interface that provides a blank content class map and receives user input specifying content classes to be placed in specified locations within the blank content class map.
10. The method of any one of claims 1-9, wherein the trained generative machine learning model iteratively generates portions of the content item, and wherein the trained generative machine learning model concurrently generates portions of the content class map.
11. The method of claim 10, wherein iteratively generating the portions of the content item comprise generating a set of content item codes and selecting a subset of the generated set of content item codes based on confidence values associated with the set of content item codes, and wherein iteratively generating the portions of the content class map comprise generating a set of class map codes and selecting a subset of the set of class map codes based on confidence values associated with the set of class map codes.
12. The method of claim 11, wherein the trained generative machine learning model utilizes a MaskGIT architecture.
13. A method of training a generative machine learning model, the method comprising: obtaining a training set, the training set comprising a plurality of content items and a corresponding plurality of content class maps, wherein each content class map of the plurality of content class maps indicates locations of one or more content classes included in a corresponding content item; for each content item of the plurality of content items: masking a portion of the content item and a corresponding portion of a content class map corresponding to the content item;PATENTDolby Ref. D24126WO01 providing the masked content item and the masked content class map to the generative machine learning model; obtaining, as an output of the generative machine learning model, a predicted content item and a predicted content class map; and updating weights associated with the generative machine learning model based on a difference between the predicted content item and the content item and a difference between the predicted content class map and the content class map corresponding to the content item; providing a trained generative machine learning model based on final weights associated with the generative machine learning model, wherein the trained generative machine learning model is usable to generate a new content item not included in the training set and a corresponding content class map indicating locations of one or more content classes within the newly generated content item.
14. The method of claim 13, wherein the trained generative machine learning model utilizes a MaskGIT architecture.
15. An apparatus configured to perform the method of any one of claims 1-14.