Three-dimensional (3D) data augmentation

WO2026102007A3PCT designated stage Publication Date: 2026-07-23QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
QUALCOMM INC
Filing Date
2025-11-05
Publication Date
2026-07-23

Smart Images

  • Figure US2025054169_23072026_PF_FP_ABST
    Figure US2025054169_23072026_PF_FP_ABST
Patent Text Reader

Abstract

A device includes a memory configured to store scene data that includes image data associated with an image of a scene, and includes placement space data that indicates a placement space associated with the image. The device also includes a processor coupled to the memory configured to identify a three-dimensional (3D) pose parameter associated with a 3D object model. The 3D pose parameter indicates a location within the placement space and an orientation. The processor is also configured to perform, based on the 3D pose parameter, a 3D transformation operation on the 3D object model to generate a first geometrically transformed 3D object model. The processor is also configured to generate, based on the image and the location, augmented image data that includes a modified image of the scene. The modified image includes at least a portion of the first geometrically transformed 3D object model.
Need to check novelty before this filing date? Find Prior Art

Description

QUALCOMM Ref. No. 2500012WO- 1 -THREE-DIMENSIONAL (3D) DATA AUGMENTATIONI. Cross-Reference to Related Applications

[0001] The present application claims priority from U.S. Provisional Patent Application No. 63 / 717,798, filed November 7, 2024, and U.S. Non-Provisional Patent Application No. 19 / 378,860, filed November 4, 2025, the contents of each of which are expressly incorporated herein by reference in their entirety.IL Field

[0002] The present disclosure is generally related to three-dimensional (3D) data augmentation.III. Description of Related Art

[0003] Advances in technology have resulted in smaller and more powerful computing devices. For example, there currently exist a variety of portable personal computing devices, including wireless telephones such as mobile and smart phones, tablets and laptop computers that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. Further, many such devices incorporate additional functionality such as a digital still camera, a digital video camera, a digital recorder, and an audio file player. Also, such devices can process executable instructions, including software applications, such as a web browser application, that can be used to access the Internet. As such, these devices can include significant computing capabilities.

[0004] A task model, e.g., a machine learning model such as a two-dimensional (2D) or three-dimensional (3D) detection or segmentation model, may be trained using training samples that include a first set of multiple images. After the task model is trained, the trained task model is tested or validated to determine an accuracy or identify a failure of the trained task model. For example, the trained task model may be tested or validated using testing samples that include a second set of multiple images that is different from the first set of multiple images. Training and testing / validating a task model can be a time consuming and compute intensive endeavor. Additionally, acquiring or generatingQUALCOMM Ref. No. 2500012WO- 2 - each of the testing data set and the training / validating data set can also be time consuming and compute intensive.

[0005] Generative models have been proposed to generate training and / or testing samples, such as by generating or augmenting one or more images. However, the conventional techniques proposed to generate training and / or testing samples have been limited to 2D perception tasks that require 2D semantic information, such as image classification, 2D image segmentation, or 2D object detection. These conventional 2D techniques do not introduce or generate 3D object content and therefore do not address a position, a pose, and a size of a 3D object while avoiding collisions with other objects. Accordingly, the conventional 2D techniques do not generate samples that support training and / or testing of 3D perception task. Additionally, the conventional 2D techniques operate in a static scenario that is independent to the 2D perception tasks, and, therefore, the samples are generated without any feedback from the 2D perception tasks that would ensure the usefulness of the generated samples for training and / or testing. Thus, the conventional 2D techniques are unable to generate useful and / or challenging samples that can be used to train or test / validate a 3D perception task.IV Summary

[0006] According to one implementation of the present disclosure, a device includes a memory configured to store scene data associated with a scene. The device also includes one or more processors coupled to the memory. The one or more processors are configured to obtain the scene data that includes image data associated with an image of the scene, and placement space data that indicates a placement space associated with the image. The one or more processors are configured to obtain a three- dimensional (3D) object model, and identify a 3D pose parameter associated with the 3D object model. The 3D pose parameter indicates a location within the placement space and an orientation. The one or more processors are also configured to perform, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model. The one or more processors are configured to generate, based on the image and the location, augmented image data that includes a modified image of the scene. The modified imageQUALCOMM Ref. No. 2500012WO- 3 - includes at least a portion of the first geometrically transformed 3D object model within the placement space.

[0007] According to another implementation of the present disclosure, a method includes obtaining scene data that includes image data associated with an image of a scene, and includes placement space data that indicates a placement space associated with the image. The method also includes obtaining a 3D object model, and identifying a 3D pose parameter associated with the 3D object model. The 3D pose parameter indicates a location within the placement space and an orientation. The method also includes performing, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model. The method includes generating, based on the image and the location, augmented image data that includes a modified image of the scene. The modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.

[0008] According to another implementation of the present disclosure, a non-transitory computer-readable medium stores instructions that are executable by one or more processors to cause the one or more processors to obtain scene data that includes image data associated with an image of a scene, and includes placement space data that indicates a placement space associated with the image. The instructions are also executable by the one or more processors to cause the one or more processors to obtain a 3D object model, and identify a 3D pose parameter associated with the 3D object model. The 3D pose parameter indicates a location within the placement space and an orientation. The instructions are also executable by the one or more processors to cause the one or more processors to perform, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model. The instructions are also executable by the one or more processors to cause the one or more processors to generate, based on the image and the location, augmented image data that includes a modified image of the scene. The modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.QUALCOMM Ref. No. 2500012WO- 4 -

[0009] According to another implementation of the present disclosure, an apparatus includes means for obtaining scene data that includes image data associated with an image of a scene, and includes placement space data that indicates a placement space associated with the image. The apparatus further includes means for obtaining a 3D object model, and means for identifying a 3D pose parameter associated with the 3D object model. The 3D pose parameter indicates a location within the placement space and an orientation. The apparatus further includes means for performing, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model. The apparatus further includes means for generating, based on the image and the location, augmented image data that includes a modified image of the scene. The modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.

[0010] Other aspects, advantages, and features of the present disclosure will become apparent after review of the entire application, including the following sections: Brief Description of the Drawings, Detailed Description, and the Claims.V. Brief Description of the Drawings

[0011] FIG. l is a block diagram of an example of a system operable to perform three- dimensional (3D) data augmentation, in accordance with one or more aspects of the present disclosure.

[0012] FIG. 2 is a diagram of an example of operations associated with the system of FIG. 1, in accordance with one or more aspects of the present disclosure.

[0013] FIG. 3 is a diagram of examples of placement spaces sampled based on a base distribution, in accordance with one or more aspects of the present disclosure.

[0014] FIG. 4 is a diagram of an example of a method of performing 3D data augmentation, in accordance with some aspects of the present disclosure.

[0015] FIG. 5 is a diagram of an example of images to illustrate 3D data augmentation, in accordance with one or more aspects of the present disclosure.QUALCOMM Ref. No. 2500012WO- 5 -

[0016] FIG. 6 is a diagram of an example of images to illustrate 3D data augmentation, in accordance with one or more aspects of the present disclosure.

[0017] FIG. 7 is a diagram of an example of an integrated circuit operable to perform 3D data augmentation, in accordance with some examples of the present disclosure.

[0018] FIG. 8 is a diagram of a mobile device operable to perform 3D data augmentation, in accordance with some examples of the present disclosure.

[0019] FIG. 9 is a diagram of a wearable electronic device operable to perform 3D data augmentation, in accordance with some examples of the present disclosure.

[0020] FIG. 10 is a diagram of a camera operable to perform 3D data augmentation, in accordance with some examples of the present disclosure.

[0021] FIG. 11 is a diagram of a headset, such as a virtual reality, mixed reality, or augmented reality headset, operable to perform 3D data augmentation, in accordance with some examples of the present disclosure.

[0022] FIG. 12 is a diagram of a mixed reality or augmented reality glasses device operable to perform 3D data augmentation, in accordance with some examples of the present disclosure.

[0023] FIG. 13 is a diagram of a second example of a vehicle operable to perform 3D data augmentation, in accordance with some examples of the present disclosure.

[0024] FIG. 14 is a diagram of an example of a method of performing 3D data augmentation, in accordance with some aspects of the present disclosure.

[0025] FIG. 15 is a block diagram of an illustrative example of a device that is operable to perform 3D data augmentation, in accordance with one or more aspects of the present disclosure.QUALCOMM Ref. No. 2500012WO- 6 -VI. Detailed Description

[0026] The above-described problems associated with generation of samples for training or testing / validating a task model, such as a three-dimensional (3D) task model, are solved using a 3D object model to perform 3D data augmentation as described herein. The present disclosure provides systems, devices, apparatus, methods, and computer-readable media for 3D data augmentation for media content systems. Some aspects more specifically relate to identification of a location and a pose of the 3D object model in a scene to generate a sample that is useful to train or test / validate a 3D perception model, such as a 3D detection model or a 3D segmentation model. For example, the 3D object model may be positioned in a 3D space and a rendering operation can be performed based on the 3D space to generate the sample. Other aspects more specifically relate to location and pose optimization of the 3D object model based on feedback from the 3D perception model.

[0027] In some embodiments, an illustrative system of the present disclosure is configured to obtain scene data that includes image data associated with an image of a scene, and includes placement space data that indicates a placement space associated with the image. The scene data may be provided to a 3D scene model, such as a 3D Gaussian Splatting model to generate a representation of the 3D geometry of the scene. In some implementations, the system also may use predicted monocular depth and semantic maps to generate the representation of the 3D geometry of the scene. The system is also configured to obtain a 3D object model, and identify a 3D pose parameter (e.g., a location within the placement space and an orientation) associated with the 3D object model. One or more 3D transformation operations (e.g., a rotation operation, a scaling operation, or a translation operation) may be performed on the 3D object model to generate a first geometrically transformed 3D object model that is positioned within the representation of the 3D geometry of the scene. The system may render an augmented image (e.g., a sample) of the scene that includes at least a portion of the first geometrically transformed 3D object model within the placement space.

[0028] In some embodiments, the system is configured to perform task-guided 3D augmentation based on feedback from a 3D perception task to enable generation of aQUALCOMM Ref. No. 2500012WO- 7 - sample in which the portion of the first geometrically transformed 3D object model included in the sample provides high uncertainty for or a failure of the 3D perception task. For example, the system may encode the placement space to generate a set of weights to initialize a flow network (e.g., a normalization network). The flow network may enable-back propagation of a gradient (associated with a loss function of the 3D perception task) so that an updated 3D pose for the 3D object model can be generated. To illustrate, the system may generate, based on the 3D pose applied to the 3D object model, a first sample that is provided to the 3D perception task. The 3D perception task performs a task perception operation on the first sample. Based on the task perception operation, a loss term associated with the task perception operation is evaluated to determine a parameter that is used to determine the updated 3D pose for the 3D object. The one or more 3D transformation operations may be performed on the 3D object model, based on the updated 3D pose, to generate a second geometrically transformed 3D object model. The system may render a second augmented image (e.g., a second sample) of the scene that includes at least a portion of the second geometrically transformed 3D object model within the placement space. The second augmented image may lead to a high uncertainty for the 3D perception tasks as compared to the first augmented image that includes the portion of the first geometrically transformed 3D object model.

[0029] Particular implementations of the subject matter described in this disclosure can be implemented to realize one or more of the following technical advantages. In some aspects, the techniques of the present disclosure related to 3D data augmentation provide a technical advantage of generating samples that have physical plausibility and accurate 3D object pose annotations, and result in high errors or failure for the 3D perception task. Accordingly, the generated samples enable improved training and / or testing / validation of the 3D perception task. Additionally, the techniques of the present disclosure improve the compute efficiency and reduces an amount of time to generate the samples as compared to conventional techniques.

[0030] Particular aspects of the present disclosure are described below with reference to the drawings. In the description, common features are designated by common reference numbers. As used herein, various terminology is used for the purpose of describingQUALCOMM Ref. No. 2500012WO- 8 - particular implementations only and is not intended to be limiting of implementations. For example, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Further, some features described herein are singular in some implementations and plural in other implementations. To illustrate, FIG. 1 depicts a device 102 including one or more processors (“processor(s)” 108 of FIG. 1), which indicates that in some implementations the device 102 includes a single processor 108 and in other implementations the device 102 includes multiple processors 108. For ease of reference herein, such features are generally introduced as “one or more” features and are subsequently referred to in the singular or optional plural (as indicated by “(s)”) unless aspects related to multiple of the features are being described.

[0031] In some drawings, multiple instances of a particular type of feature are used. Although these features are physically and / or logically distinct, the same reference number is used for each, and the different instances are distinguished by addition of a letter to the reference number. When the features as a group or a type are referred to herein - e.g., when no particular one of the features is being referenced, the reference number is used without a distinguishing letter. However, when one particular feature of multiple features of the same type is referred to herein, the reference number is used with the distinguishing letter.

[0032] As used herein, the terms “comprise,” “comprises,” and “comprising” may be used interchangeably with “include,” “includes,” or “including.” Additionally, the term “wherein” may be used interchangeably with “where.” As used herein, “exemplary” indicates an example, an implementation, and / or an aspect, and should not be construed as limiting or as indicating a preference or a preferred implementation. As used herein, an ordinal term (e.g., “first,” “second,” “third,” etc.) used to modify an element, such as a structure, a component, an operation, etc., does not by itself indicate any priority or order of the element with respect to another element, but rather merely distinguishes the element from another element having a same name (but for use of the ordinal term). As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to multiple (e.g., two or more) of a particular element.QUALCOMM Ref. No. 2500012WO- 9 -

[0033] As used herein, “coupled” may include “communicatively coupled,” “electrically coupled,” or “physically coupled,” and may also (or alternatively) include any combinations thereof. Two devices (or components) may be coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) directly or indirectly via one or more other devices, components, wires, buses, networks (e.g., a wired network, a wireless network, or a combination thereof), etc. Two devices (or components) that are electrically coupled may be included in the same device or in different devices and may be connected via electronics, one or more connectors, or inductive coupling, as illustrative, non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive signals (e.g., digital signals or analog signals) directly or indirectly, via one or more wires, buses, networks, etc. As used herein, “directly coupled” may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without intervening components.

[0034] In the present disclosure, terms such as “obtaining,” “determining,” “calculating,” “estimating,” “shifting,” “adjusting,” etc. may be used to describe how one or more operations are performed. It should be noted that such terms are not to be construed as limiting and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, “obtaining,” “generating,” “calculating,” “estimating,” “using,” “selecting,” “accessing,” and “determining” may be used interchangeably. For example, “obtaining,” “generating,” “calculating,” “estimating,” or “determining” a parameter (or a signal) may refer to actively generating, estimating, calculating, or determining the parameter (or the signal) or may refer to using, selecting, or accessing the parameter (or signal) that is already generated, such as by another component or device.

[0035] As used herein, the term “machine learning” should be understood to have any of its usual and customary meanings within the fields of computers science and data science, such meanings including, for example, processes or techniques by which one or more computers can learn to perform some operation or function without being explicitly programmed to do so. As a typical example, machine learning can be used to enable one or more computers to analyze data to identify patterns in data and generate aQUALCOMM Ref. No. 2500012WO- 10 - result based on the analysis. For certain types of machine learning, the results that are generated include data that indicates an underlying structure or pattern of the data itself. Such techniques, for example, include so called “clustering” techniques, which identify clusters (e.g., groupings of data elements of the data).

[0036] For certain types of machine learning, the results that are generated include a data model (also referred to as a “machine-learning model” or simply a “model”). Typically, a model is generated using a first data set to facilitate analysis of a second data set. For example, a first portion of a large body of data may be used to generate a model that can be used to analyze the remaining portion of the large body of data. As another example, a set of historical data can be used to generate a model that can be used to analyze future data.

[0037] Since a model can be used to evaluate a set of data that is distinct from the data used to generate the model, the model can be viewed as a type of software (e.g., instructions, parameters, or both) that is automatically generated by the computer(s) during the machine learning process. As such, the model can be portable (e.g., can be generated at a first computer, and subsequently moved to a second computer for further training, for use, or both). Additionally, a model can be used in combination with one or more other models to perform a desired analysis. To illustrate, first data can be provided as input to a first model to generate first model output data, which can be provided (alone, with the first data, or with other data) as input to a second model to generate second model output data indicating a result of a desired analysis. Depending on the analysis and data involved, different combinations of models may be used to generate such results. In some examples, multiple models may provide model output that is input to a single model. In some examples, a single model provides model output to multiple models as input.

[0038] Examples of machine-learning models include, without limitation, perceptrons, neural networks, support vector machines, regression models, decision trees, Bayesian models, Boltzmann machines, adaptive neuro-fuzzy inference systems, as well as combinations, ensembles and variants of these and other types of models. Variants of neural networks include, for example and without limitation, prototypical networks,QUALCOMM Ref. No. 2500012WO- 11 - autoencoders, transformers, self-attention networks, convolutional neural networks, deep neural networks, deep belief networks, etc. Variants of decision trees include, for example and without limitation, random forests, boosted decision trees, etc.

[0039] Since machine-learning models are generated by computer(s) based on input data, machine-learning models can be discussed in terms of at least two distinct time windows - a creation / training phase and a runtime phase. During the creation / training phase, a model is created, trained, adapted, validated, or otherwise configured by the computer based on the input data (which in the creation / training phase, is generally referred to as “training data”). Note that the trained model corresponds to software that has been generated and / or refined during the creation / training phase to perform particular operations, such as classification, prediction, encoding, or other data analysis or data synthesis operations. During the runtime phase (or “inference” phase), the model is used to analyze input data to generate model output. The content of the model output depends on the type of model. For example, a model can be trained to perform classification tasks or regression tasks, as non-limiting examples. In some implementations, a model may be continuously, periodically, or occasionally updated, in which case training time and runtime may be interleaved or one version of the model can be used for inference while a copy is updated, after which the updated copy may be deployed for inference.

[0040] In some implementations, a previously generated model is trained (or re-trained) using a machine-learning technique. In this context, “training” refers to adapting the model or parameters of the model to a particular data set. Unless otherwise clear from the specific context, the term “training” as used herein includes “re-training” or refining a model for a specific data set. For example, training may include so called “transfer learning.” In transfer learning a base model may be trained using a generic or typical data set, and the base model may be subsequently refined (e.g., re-trained or further trained) using a more specific data set.

[0041] A data set used during training is referred to as a “training data set” or simply “training data”. The data set may be labeled or unlabeled. “Labeled data” refers to data that has been assigned a categorical label indicating a group or category with which theQUALCOMM Ref. No. 2500012WO- 12 - data is associated, and “unlabeled data” refers to data that is not labeled. Typically, “supervised machine-learning processes” use labeled data to train a machine-learning model, and “unsupervised machine-learning processes” use unlabeled data to train a machine-learning model; however, it should be understood that a label associated with data is itself merely another data element that can be used in any appropriate machinelearning process. To illustrate, many clustering operations can operate using unlabeled data; however, such a clustering operation can use labeled data by ignoring labels assigned to data or by treating the labels the same as other data elements.

[0042] Training a model based on a training data set generally involves changing parameters of the model with a goal of causing the output of the model to have particular characteristics based on data input to the model. To distinguish from model generation operations, model training may be referred to herein as optimization or optimization training. In this context, “optimization” refers to improving a metric, and does not mean finding an ideal (e.g., global maximum or global minimum) value of the metric. Examples of optimization trainers include, without limitation, backpropagation trainers, derivative free optimizers (DFOs), and extreme learning machines (ELMs). As one example of training a model, during supervised training of a neural network, an input data sample is associated with a label. When the input data sample is provided to the model, the model generates output data, which is compared to the label associated with the input data sample to generate an error value. Parameters of the model are modified in an attempt to reduce (e.g., optimize) the error value. As another example of training a model, during unsupervised training of an autoencoder, a data sample is provided as input to the autoencoder, and the autoencoder reduces the dimensionality of the data sample (which is a lossy operation) and attempts to reconstruct the data sample as output data. In this example, the output data is compared to the input data sample to generate a reconstruction loss, and parameters of the autoencoder are modified in an attempt to reduce (e.g., optimize) the reconstruction loss.

[0043] FIG. 1 is a block diagram of an example of a system 100 operable to perform 3D data augmentation, in accordance with one or more aspects of the present disclosure. The system 100 includes a device 102 that is operable to perform 3D data augmentation as described herein.QUALCOMM Ref. No. 2500012WO- 13 -

[0044] The device 102 includes, or is coupled to, a memory 106, one or more processors 108 (collectively referred to herein as the “processor 108”), an image sensor 112, a display device 116, and a modem 118. The memory 106 may include one or more memories, such as a single memory or multiple different memories (of the same type or of different types).

[0045] The memory 106 is configured to store instructions 109, scene data 130, and 3D object models 138. To illustrate, the memory 106 includes or stores the instructions 109 that, when executed by the processor 108, cause the processor 108 to perform one or more operations as described herein. The 3D object models 138 include or indicate representations of objects, such as a vehicle (e.g., a car, truck, bus, etc.), an animal (e.g., a person, a dog, etc.), or other objects. In some embodiments, the 3D object models 138 include a 3D object model 150.

[0046] The scene data 130 includes image data 132, placement space data 134, and annotations 136. The image data 132 may include or be associated with an image of a scene. The placement space data 134 may include or indicate a placement space associated with the image. The annotations 136 include or indicate an annotation (e.g., data) associated with the image. For example, the annotations 136 may include or indicate a camera location, a camera pose, a foreground object, or a combination thereof. In some implementations, the scene data includes training data used to train a task model, such as a 3D task model (e.g., a 3D perception task model).

[0047] In some embodiments, the memory 106 also stores other information or data, such as one or more models. The one or more models may include a 3D task model (e.g., a 3D perception task model), a 3D scene model, a Gaussian splat model, a hypernetwork, a normalization model (e.g., a flow model), as illustrative, non-limiting examples.

[0048] The processor 108 includes an image generator 120, an objective operator 140, and a task operator 148. Each of the image generator 120, the objective operator 140, and the task operator 148, or a portion thereof, may be implemented by the processor 108 executing instructions (e.g., software), dedicated hardware (e.g., circuitry), aQUALCOMM Ref. No. 2500012WO- 14 - combination thereof. Although the processor 108 is described as including the objective operator 140 and the task operator 148, in other embodiments, the processor 108 may not include the objective operator 140, the task operator 148, or both. Additionally, or alternatively, although the objective operator 140 and the task operator 148 are described as being separate from the image generator 120, in other embodiments, the image generator 120 may include the objective operator 140, the task operator 148, or both. In some embodiments, the task operator may include the objective operator 140.

[0049] The image generator 120 is configured to obtain the scene data 130 and the 3D object model 150, and generate augmented image data 160 based on the scene data 130 and the 3D object model 150. To generate the augmented image data 160, the image generator 120 includes a pose determiner 122 and an image modifier unit 124.

[0050] The pose determiner 122 is configured to determine a 3D pose parameter 152 based on the scene data 130 (e.g., the placement space data 134, the annotations 136, or a combination thereof). The 3D pose parameter 152 may include or indicate a location 154 (associated with the placement space data 134), an orientation, a scale indicator (e.g., a scaling factor), or a combination thereof. As an illustrative example, the pose determiner 122 may determine the orientation, such as a yaw angle, based on a closest object (e.g., a foreground object or a background object) to the location. To illustrate, the annotations 136 may include an orientation (e.g., a yaw angle) of the closest object and the pose determiner 122 may assign the orientation of the closest object as the orientation included in or indicated by the 3D pose parameter 152. The pose determiner 122 may provide the 3D pose parameter 152 to the image modifier unit 124.

[0051] In some embodiments, the pose determiner 122 is also configured to determine the 3D pose parameter 152, or an updated 3D pose parameter that is based on the augmented image data 160. For example, the pose determiner 122 may include a hypernetwork / ?() and a normalization network^)- The hypernetwork / ?() is configured to determine a set of weights 6 to initialize the normalization network^)- In some embodiments, the hypemetwork A() is configured to determine the set of weights 6 based on the placement space data 134. The normalization network^) may include an invertible flow network that is configured to map a latent variable z to a 3D poseQUALCOMM Ref. No. 2500012WO- 15 - parameter T. For example, the inverse ( / z-1) of the flow network may be represented as / -1(z, 0) = T. In some embodiments, the latent variable is associated with a set of latent variables that have a Gaussian distribution such that ~ N(0, 1), where N() is a normalized set from zero to one. An example of the pose determiner 122 including the hypernetwork / ?() and the normalization network ) is described further herein at least with reference to FIG. 2.

[0052] The image modifier unit 124 is configured to obtain the scene data 130, the 3D object model 150, and the 3D pose parameter 152 and generate the augmented image data 160. In some embodiments, the image modifier unit 124 receives the 3D object model 150 and the 3D pose parameter 152 and performs, based on the 3D pose parameter 152 one or more transformation operations on the 3D object model 150 to generate the transformed 3D object model 156 (e.g., a first geometrically transformed 3D object model). The one or more transformation operations may include a rotation operation, a scaling operation, a translation operation, or a combination thereof. In some embodiments, the one or more transformation operations may be performed on the 3D object model 150 before the 3D object model is included in an image (e.g., the image data 132), after the 3D object model is included in the image (e.g., the image data 132), or a combination thereof.

[0053] The image modifier unit 124 may include a 3D scene model 126 and may be denoted herein as <b. In some embodiments, the 3D scene model 126 may include a Gaussian splatting model, such as a 3D Gaussian splatting model. Additionally, or alternatively, the 3D scene model 126 can be configured to generate a bird’s eye view based on the image data 132. The 3D scene model 126 can be configured to project a set of 3D Gaussian primitives X onto a 2D plane associated with a camera that generated the image data 132 - e.g., captured the image associated with the image data 132. The camera that generated the image data 132 may have had parameters co at the time of an image capture operation to capture the image. The parameters co may include or indicate a location of the camera, a pose of the camera, one or more image capture settings of the camera, or a combination thereof, as illustrative, non-limiting examples. In some examples, the parameters are included in or indicated by the annotations 136.QUALCOMM Ref. No. 2500012WO- 16 -

[0054] The 3D scene model 126 is configured to decompose the set of 3D Gaussian primitives A into a subset of static Gaussian points (Xb), representing the static parts of the scene such as the background, and a subset of dynamic parts (Xo), representing all movable objects like vehicles, humans, etc. The subset Xo= { X , X , . . ., X } can be further decomposed into foreground object instances in the scene. An image (x ) may be represented by the 3D scene model 126 as (Xb, Xo, co).

[0055] The image modifier unit 124 is configured to edit (e.g., modify) the scene by deleting, inserting, or manipulating the scene geometric or photometric information. For example, modifying (e.g., augmenting) the scene includes insertion of a new 3D object model (e.g., the 3D object model 150) with a set of Gaussian points (Xa) into the scene. To illustrate, the new 3D object model may be inserted in accordance with the one or more transformation operations, such as one or more transformation operations that have been determined to be physically plausible and to not cause the new 3D object model to collide with other parts, such as any of the subset of dynamic parts (Xo). The one or more transformation operations may be based on the 3D pose parameter 152, which is designated T. Additionally, the transformed 3D object model, such as the transformed 3D object model 156, is designated as T(xa).

[0056] After modification of the scene, the image modifier unit 124 is configured to render the scene to generate the augmented image data 160 - e.g., an augmented image x+may be represented as,

[0057] The image generator 120 is also configured to generate a ground truth 165, such as a ground truth that corresponds to the augmented image data 160 - e.g., the augmented image x+. The image generator 120 may provide the ground truth 165 to the objective operator 140.

[0058] The task operator 148 is configured to perform a 3D task operation, such as a 3D perception task operation. The 3D task operation may include a detection operation or a segmentation operation. In some embodiments, the task operator 148 includes a 3D task model that is configured to perform the 3D task operation. Although the task operation and the task model are described as being 3D, in other embodiments, the taskQUALCOMM Ref. No. 2500012WO- 17 - operation and the task model include a 2D task operation and a 2D task model, respectively. The task operator 148 is configured to perform the 3D task operation based on the augmented image data 160 to generate a task result 162. Additionally, the task operator 148 may generate task data 164 based on the 3D task operation. The task data 164 may include or indicate data, based on operation of the task operator 148 (e.g., the task operation) on the augmented image data 160, that enables an evaluation of a loss function of the task operation (e.g., a 3D task model). For example, the task data 164 may include or indicate the task result 162, an output probability or accuracy of the task operation, a gradient of the output probability or accuracy of the task operation, or a combination thereof.

[0059] The objective operator 140 is configured to determine a parameter 166, such as an updated latent variable, that is provided as feedback to the pose determiner 122 to generate an updated 3D pose parameter. The objective operator 140 includes an objective function 142 that is configured to calculate the parameter 166. The objective function 142 includes a loss function unit 144 associated with the task operator 148, and one or more regulator function unit 146 associated with the 3D pose parameter T 152, the 3D object model 150, the transformed 3D object model 156, a latent variable z, the set of weights 0, the augmented image data (e.g., the augmented image x+), or a combination thereof. In some implementations, the loss function unit 144 is represented as Z(). In some embodiments, the loss function unit 144 is / .(x+). Additionally, or alternatively, the regulator function unit 146 may include a visibility function r(r(xa)) and a density estimation function logpa)(z|T, 6 ), where z = f( T). The visibility function r(r(xa)) is configured to cause at least a portion of the transformed 3D object model to be included in a rendered image (e.g., an augmented image) and to avoid the transformed 3D object model from being occluded by another object and / or otherwise undetectable by barely having any portion of the transformed 3D object model within the rendered image (e.g., the augmented image). The density estimation function l°gP<u(zlT< 6 ), where z = f Tz), is configured to ensure that the transformed 3D object model is positioned inside the placement space and does not collide with another object.QUALCOMM Ref. No. 2500012WO- 18 -

[0060] Computing the exact likelihood function given a sample is in the placement space is used as a regularization term in optimizing the agent pose to ensure the agents are positioned inside the allowed region defined per individual frame / camera and do not collide with other objects. Thus, the optimization objective in Eq. (3) is modified as follows: In some embodiments, the objective function 142 to produce the parameter 166 is:where z = fi z), and is a scaling factor (e.g., a hyper-parameter). The objective function 142 may, for a particular 3D pose parameter, optimize the objective function 142 to maximize the loss associated with the task operator 148. In some embodiments, the objective function 142 may utilize quasi-Newton optimization technique, such as an Broyden-Fletcher-Goldfarb-Shanno (BFGS) technique, using line-searching to converge to a higher loss term.

[0061] The modem 118 is coupled to the processor 108 and is configured to transmit the augmented image data 160 (e.g., a training sample or a testing / validation sample) to a second device. In some embodiments, the modem 118 may be configured to receive data from another device. For example, the modem 118 may be configured to receive the scene data 130, or a portion thereof, the 3D object models 138, or a combination thereof, from another device.

[0062] The processor 108 is also coupled to an image sensor 112 and a display device 116. The image sensor 112 may include one or more cameras and may be configured to generate image data, such as the image data 132. The display device 116 is coupled to the processor 108 and is configured to output an image, such as an image that is associated with or corresponds to the image data 132, the augmented image data 160 (e.g., an augmented image or a modified image), or a combination thereof. In some examples, the display device 116 includes a display screen, a monitor or television, a projector, or a combination thereof.

[0063] The image sensor 112, the display device 116, or a combination thereof may be coupled to or integrated within the device 102. Although the device 102 is described asQUALCOMM Ref. No. 2500012WO- 19 - being coupled to or including the image sensor 112, the display device 116, and the modem 118, in other embodiments, one or more of these elements are optional, and in such embodiments, the device 102 may not include or be coupled to the image sensor 112, the display device 116, the modem 118, or a combination thereof.

[0064] During operation of the system 100, the processor 108 obtains the scene data 130, such as data that is associated with an image of a scene. The processor 108 may provide the scene data 130 to the image generator 120. For example, the pose determiner 122 may obtain the placement space data 134, and the image modifier unit 124 (e.g., the 3D scene model 126) may obtain the image data 132. Additionally, or alternatively, the image generator 120 (e.g., the pose determiner 122, the image modifier unit 124, or both) may obtain the annotations 136.

[0065] In some embodiments, the processor 108 (e.g., the image generator 120) identifies, based on the scene data 130 (e.g., the image data 132, the placement space data 134, and / or the annotations 136), one or more foreground objects included in the image, the placement space, one or more background objects, a camera pose, or a combination thereof. For example, the image modifier unit 124 may use the 3D scene model 126, such as a Gaussian splatting model (e.g., a 3D Gaussian splatting model), to identify the one or more foreground objects included in the image, the placement space, the one or more background objects, the camera pose, or a combination thereof. In some embodiments, the image modifier unit 124 (e.g., the 3D scene model 126) generates a 3D representation of the image, such as a 3D point cloud or other 3D representation.

[0066] The processor 108 also obtains the 3D object model 150. For example, the processor 108 may select the 3D object model 150 from the 3D object models 138. In some examples, the processor 108 randomly selects the 3D object model 150. Additionally, or alternatively, the processor 108 selects the 3D object model 150 based on a received input, such as an input from a user of the device 102. The processor 108 may provide the 3D object model 150 to the image generator 120 (e.g., the image modifier unit 124).QUALCOMM Ref. No. 2500012WO- 20 -

[0067] In some embodiments, the processor 108 (e.g., the image generator 120) determines, based on the scene data 130, a placement area of the placement space. For example, the pose determiner 122 may determine or identify the placement area based on the placement space data 134, the annotations 136, map data, or a combination thereof. In some embodiments, the image generator 120 (e.g., the pose determiner 122) is configured to perform one or more operations based on the placement area, as described further herein at least with reference to FIG. 2. For example, the image generator 120 (e.g., the pose determiner 122) may provide the placement area to a hypernetwork to generate a set of weights for a normalization network, and initiate the normalization network based on the set of weights. In some such examples, the image generator 120 (e.g., the pose determiner 122) may identify, based on the normalization network and a probability distribution, the location 154 associated with the placement space. Additionally, or alternatively, the image generator 120 (e.g., the pose determiner 122) can provide an indication of one or more foreground objects associated with the image to the hypemetwork. The one or more foreground objects may be associated with or indicated by the annotations 136 or determined by the image modifier unit 124. In some such examples, the set of weights is generated based on the placement area and the one or more foreground objects.

[0068] The processor 108 (e.g., the image generator 120) identifies a 3D pose parameter 152 associated with the 3D object model 150. The 3D pose parameter 152 may include or indicate a location 154 (within the placement space), an orientation, a scale value / factor, or a combination thereof. In some examples, the pose determiner 122 may randomly select the 3D pose parameter 152 based on the placement space data 134. To illustrate, the pose determiner 122 may identify or select the location 154 that is included in the placement space. The 3D pose parameter 152 may be determined based on the location. Additionally, or alternatively, the 3D pose parameter 152 can be determined based on the annotations 136, such as a camera pose. The pose determiner 122 may provide the 3D pose parameter 152 to the image modifier unit 124. In other embodiments, the pose determiner 122 may select a variable, such as a latent posevariable, that is used to determine the 3D pose parameter 152 based on theQUALCOMM Ref. No. 2500012WO- 21 - normalization network. For example, the processor 108 may randomly initialize (e.g., select) the variable based on a base distribution (e.g., a base Gaussian distribution.

[0069] The processor 108 (e.g., the image generator 120) performs, based on the 3D pose parameter 152, one or more 3D transformation operations on the 3D object model 150 to generate a transformed 3D object model 156 (e.g., a first geometrically transformed 3D object model). For example, the image modifier unit 124 may generate the transformed 3D object model 156 based on the 3D pose parameter 152. To perform the one or more 3D transformation operations, the processor (e.g., the image modifier unit 124) can scale the 3D object model 150, adjust an orientation of the 3D object model 150, perform a translation operation on the 3D object model 150 (or the scaled 3D object model), or a combination thereof.

[0070] The processor 108 (e.g., the image modifier unit 124) may provide the transformed 3D object model 156 to the 3D scene model 126. The 3D scene model 126 may position the transformed 3D object model 156 within the image (associated with the image data 132) or a 3D representation of the image. For example, the image modifier unit 124 (e.g., the 3D scene model 126) can position a center of the transformed 3D object model 156 at the location within the placement space. Alternatively, in some embodiments, the image modifier unit 124 provides the 3D object model 150 to the 3D scene model 126, and the 3D object model 150 is adjusted based on the 3D pose parameter 152 by the 3D scene model 126. To illustrate, the 3D scene model 126 may position the 3D object model 150 within the image (associated with the image data 132) or a 3D representation of the image and apply the 3D pose parameter 152 to the 3D object model 150 such that a center of the transformed 3D object model 156 is positioned at the location within the placement space. Accordingly, the image modifier unit 124 (e.g., the 3D scene model 126) modifies the image or the 3D representation to include at least a portion of the transformed 3D object model 156. Additionally, or alternatively, the image modifier unit 124 (e.g., the 3D scene model 126) can modify the image or the 3D representation to remove one or more foreground objects, one or more background objects, or a combination thereof.QUALCOMM Ref. No. 2500012WO- 22 -

[0071] The processor 108 (e.g., the image generator 120) generates augmented image data 160 that includes or is associated with a modified image of the scene (as compared to the image associated with the image data 132). For example, the image modifier unit 124 (e.g., the 3D scene model 126) can render the image or the 3D representation that includes at least a portion of the transformed 3D object model 156 to generate the augmented image data 160. The portion of the transformed 3D object model 156 may be positioned within the placement space. In some such examples, the augmented image data 160 is generated based on the image associated with the image data 132, the 3D object model 150, the 3D pose parameter 152, the location 154, or a combination thereof.

[0072] In some embodiments, the processor 108 (e.g., the image generator 120) provides the augmented image data 160 to the task operator 148. The task operator 148 is configured to perform a task operation, such as a 3D perception task operation, based on the augmented image data 160. The task operator 148 performs the task operation to generate a task result 162. In some such embodiments, the processor 108 may also generate one or more parameters 166 (hereinafter referred to as “the parameter 166”) based on the task operation performed by the task operator 148 on the augmented image data 160.

[0073] To generate the parameter 166, the processor 108 (e.g., the objective operator 140) is configured to obtain a ground truth 165 associated with or corresponding to the augmented image data 160, task data 164, or a combination thereof. For example, the processor 108 (e.g., the image generator 120) may be configured to generate the ground truth 165 associated with the 3D pose parameter 152, transformed 3D object model 156 (e.g., the first geometrically transformed 3D object model) within the placement space, the augmented image data 160, or a combination thereof. The ground truth 165 may include or indicate the 3D object model 150 (e.g., the transformed 3D object model 156), the 3D pose parameter 152, the variable (e.g., a latent pose-variable) used to determine the 3D pose parameter 152, the set of weights for the normalization network, or a combination thereof. The task data 164 may include or indicate data, based on operation of the task operator 148 (e.g., the task operation) on the augmented image data 160, that enables an evaluation of a loss function of the task operation (e.g., a 3D taskQUALCOMM Ref. No. 2500012WO- 23 - model). For example, the task data 164 may include or indicate the task result 162, an output probability or accuracy of the task operation, a gradient of the output probability or accuracy of the task operation, or a combination thereof.

[0074] The objective operator 140 (e.g., the objective function 142) generates the parameter 166 based on the ground truth 165 associated with or corresponding to the augmented image data 160, the task data 164, or a combination thereof. In some embodiments, the parameter 166 includes or indicates an updated variable (e.g., an updated latent pose-variable). The parameter 166 enables the image generator 120 (e.g., the pose determiner 122) to determine an updated 3D pose parameter that indicates another location within the placement space. Based on the updated 3D pose parameter, the image modifier unit 124 may perform the one or more 3D transformation operations on the 3D object model 150 (or on the transformed 3D object model 156) to generate a second geometrically transformed 3D object model. The image modifier unit 124 (e.g., the 3D scene model 126) can generate updated augmented image data that includes another modified image of the scene. The updated augmented image includes at least a portion of the second geometrically transformed 3D object model within the placement space. The updated augmented image data, when processed by the task operator 148, results in high errors or failure for the 3D perception task as compared to when the task operator 148 processes the augmented image data 160 using the 3D perception task.

[0075] In some embodiments, the processor 108 is configured to train, based on a training data set (e.g., one or more training samples) that includes at least a portion of the scene data 130, a task model to generate a trained task model. The trained task model, such as a trained detection model or a trained segmentation model, may include or correspond to a task function or task model that is used by the task operator 148 on the augmented image data 160. Additionally, or alternatively, the processor 108 is configured to generate a testing data set (e.g., one or more testing samples) that includes the augmented image data 160. The processor 108 may store the testing data set and / or use the testing data set to test or validate the trained task model.

[0076] In some examples, the device 102 corresponds to or is included in one of various types of devices, such that the processor 108 can be integrated in multiple types ofQUALCOMM Ref. No. 2500012WO- 24 - devices. In an illustrative example, the processor 108 is integrated in a wearable device, such as a wearable electronic device as depicted in FIG. 9, a virtual reality, mixed reality, or augmented reality headset as depicted in FIG. 11, a mixed reality or augmented reality glasses device as described with reference to FIG. 12, or another wearable device. In another illustrative example, the processor 108 is integrated in a mobile device (a mobile phone or a tablet) as depicted in FIG. 8, a camera as depicted in FIG. 10, a vehicle as depicted in FIG. 13, a computer or a server, or another system or device.

[0077] In some embodiments, a device (e.g., the device 102) includes a memory (e.g., the memory 106) configured to store scene data (e.g., the scene data 130) associated with a scene. The device also includes one or more processors (e.g., the processor 108) coupled to the memory. The one or more processors are configured to obtain the scene data that includes image data (e.g., the image data 132) associated with an image of the scene, and placement space data (e.g., the placement space data 134) that indicates a placement space associated with the image. The one or more processors are configured to obtain a 3D object model (e.g., the 3D object model 150), and identify a 3D pose parameter (e.g., the 3D pose parameter 152) associated with the 3D object model. The 3D pose parameter indicates a location (e.g., the location 154) within the placement space and an orientation. The one or more processors are also configured to perform, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model (e.g., the transformed 3D object model 156). The one or more processors are configured to generate, based on the image and the location, augmented image data (e.g., the augmented image data 160) that includes a modified image of the scene. The modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.

[0078] One technical advantage of the techniques described herein with reference to the device 102 includes an end-to-end differentiable approach for optimizing a position and a pose of the 3D object model 150 based on a gradient (e.g., the task data 164) from a 3D task model. For example, the techniques may generate the augmented image data 160 (e.g., a testing / validation sample) that includes at least a portion of the transformedQUALCOMM Ref. No. 2500012WO- 25 -3D object model 156 that has physical plausibility and accurate 3D object pose annotations, and results in high errors or failure for the 3D perception task.Accordingly, the augmented image data 160 (e.g., the generated sample) enables improved training and / or testing / validation of the 3D task model. Additionally, the techniques of the present disclosure improve the compute efficiency and reduce an amount of time to generate the augmented image data 160 as compared to conventional techniques that generate testing / validation samples.

[0079] FIG. 2 is a diagram of an example of operations associated with the system 100 of FIG. 1, in accordance with one or more aspects of the present disclosure. For example, the operations can be performed by or associated with the pose determiner 122 of the image generator 120 of the system 100. In addition to the pose determiner 122, FIG. 2 depicts a bird’s eye view (BEV) 200 of an example of the scene associated with the scene data 130.

[0080] The BEV 200 includes a placement space 240 (depicted as a roadway) and one or more foreground objects 234 (hereinafter referred to as “the foreground object 234”). Additionally, BEV 200 may include one or more background objects (not shown). The BEV 200 also indicates a camera pose 236 associated with an image, such as an image associated with the image data 132. In some embodiments, the foreground object 234, the camera pose 236, or a combination thereof, may include or correspond to the annotations 136.

[0081] The BEV 200 also includes or indicates a placement area 232 that is a portion of the placement space 240. The placement space 240, the placement area 232, or a combination thereof, may include or correspond to the placement space data 134. The placement area 232 may include or correspond to a portion of the placement space 240 that is included in the image associated with the camera pose 236. The placement area 232 may include a single continuous area or may include multiple discontinuous areas.

[0082] Referring to the pose determiner 122, the pose determining includes a hypernetwork 250 and a normalization network 260. The pose determiner 122 (e.g., the hypernetwork 250) is configured to receive an indication of the placement area 232, anQUALCOMM Ref. No. 2500012WO- 26 - indication of the foreground objects 234, or a combination thereof. In some embodiments, the hypernetwork 250 is configured to receive a combination of the placement area 232 and the foreground object 234 in which the foreground object 234 are removed from the placement area 232 to form the combination. Additionally, the pose determiner 122 (e.g., the normalization network 260) may receive a base distribution 262. The base distribution 262 may include a Gaussian distribution.

[0083] The hypernetwork 250 is configured to generate a set of weights 266 to initialize the normalization network 260. For example, the hypernetwork 250 may generate the set of weights 266 based on the placement area 232, the foreground objects 234, or a combination thereof. Additionally, or alternatively, the hypemetwork 250 may generate the set of weights 266 based on the annotations 136. The hypemetwork includes an encoder 252 and a decoder 256.

[0084] The encoder 252 is configured to encode the placement area 232, the foreground object 234, or a combination thereof, to generate a latent variable 254 in the latent space. In some examples, the encoder 252 may correspond to at least a portion of a variational autoencoder (VAE). In some other examples, the encoder 252 may include a point cloud encoder that is configured to receive and encode a 2D point cloud of the placement area 232, the foreground object 234, or a combination thereof. In yet other examples, the encoder 252 may include a vision transformer that is configured to encode a binary image of the placement area 232, the foreground objects 234, or a combination thereof. The latent variable 254 is provided to the decoder 256.

[0085] The decoder 256 may include a multi-layer perceptron (MLP) as the decoder 256. The decoder 256 is configured to generate the set of weights 266 based on the latent variable 254. The set of weights 266 may be configured to be used to initialize the normalization network 260.

[0086] The normalization network 260 is configured to receive the set of weights 266 and to initialize the normalization network 260 based on the set of weights 266. In some implementations, the normalization network 260 includes a flow network.QUALCOMM Ref. No. 2500012WO- 27 -

[0087] After initialization of the normalization network 260, the normalization network 260 is configured to generate a reconstruction 264. For example, to generate the reconstruction, the normalization network 260 is configured to map the base distribution 262 to the placement area 232. Examples of the mapping of a base distribution to a placement area are described further herein at least with reference to FIG. 3.

[0088] In some embodiments, the normalization network 260 is an invertible network. To illustrate, if a latent pose-variable is sampled from the base distribution 262 and provided to the initialized normalization network 260, the normalization network 260 is configured to output or identify a location, such as the location 154, within the placement area 232. Based on the camera pose 236 and the location, the pose determiner 122 may determine the 3D pose parameter 152. It is noted that the location identified by the normalization network 260 may not include or collide with the foreground object 234. Additionally, if a location (e.g., a pixel) within the placement area is selected and provided to the normalization network 260, the normalization network 260 is configured to output or identify a high probability portion of the base distribution 262. Accordingly, the initialized normalization network 260 is configured to favor positions that have a high likelihood of occurrence (with respect to the base distribution 262) and that will be within the placement area 232.

[0089] FIG. 3 is a diagram of examples of placement spaces sampled based on a base distribution, in accordance with one or more aspects of the present disclosure. FIG. 3 includes a first example 300, a second example 310, a third example 320, a fourth example 330, a fifth example 340, and a sixth example 350. Each of the examples 300, 310, 320, 330, 340, and 350 includes a base distribution (e.g., a Gaussian space) having a line that traverses the base distribution, a ground truth associated with a placement area, and a mapping of the base distribution (and the line) to the ground truth. For example, the first example 300 includes a base distribution 302, a ground truth 304, and a mapping 306. The second example 310 includes a base distribution 312, a ground truth 314, and a mapping 316. The third example 320 includes a base distribution 322, a ground truth 324, and a mapping 326. The fourth example 330 includes a base distribution 332, a ground truth 334, and a mapping 336. The fifth example 340QUALCOMM Ref. No. 2500012WO- 28 - includes a base distribution 342, a ground truth 344, and a mapping 346. The sixth example 350 includes a base distribution 352, a ground truth 354, and a mapping 356.

[0090] FIG. 4 is a diagram of an example of a method of performing 3D data augmentation, in accordance with some aspects of the present disclosure. In a particular aspect, one or more operations of the method 1400 are performed by the system 100, the device 102, the processor 108, the image generator 120, an integrated circuit, a mobile device (e.g., a mobile phone or a tablet computer device), a wearable electronic device, a camera, a virtual reality headset, a mixed reality headset, an augmented reality headset, a mixed reality or augmented reality glasses device, a vehicle, a server, or a combination thereof..

[0091] In some embodiments, the method 400 starts at block 1402. At block 404, the method 400 also includes selecting an image associated with a scene and a 3D object model. The image may include or correspond to the scene data 130, the image data 132, the placement space data 134, the annotations 136, or a combination thereof. The 3D object model may include or correspond to the 3D object models 138 or the 3D object model 150.

[0092] At block 406, the method 400 also includes determining a set of weights for a normalization network and initializing the normalization network. For example, the set of weights may be determined based on or using a hypemetwork, such as the hypernetwork 250. The set of weights and the normalization network may include or correspond to the set of weights 266 and the normalization network 260, respectively.

[0093] At block 408, the method 400 also includes selecting or determining a variable. For example, the variable may be randomly initialized based on a base distribution (e.g., a base Gaussian distribution, such as the base distribution 262, 302, 312, 322, 332, 342, or 352. In some implementations, the variable includes a latent pose-variable. At block 410, the method 400 also includes determining a 3D pose based on the normalization network and the variable. The 3D pose may include or correspond to the 3D pose parameter 152.QUALCOMM Ref. No. 2500012WO- 29 -

[0094] At block 412, the method 400 also includes applying the 3D pose to the 3D object model and positioning the geometrically transformed 3D object model in the scene. To illustrate, the 3D pose may be applied to the 3D object model to generate the geometrically transformed 3D object model. In some embodiments, applying the 3D model to the 3D object model includes scaling the 3D object model, rotating the 3D object model, or a combination thereof. Additionally, or alternatively, positioning the geometrically transformed 3D object model in the scene may include performing a translation operation or placing the geometrically transformed 3D object model at a location based on the 3D pose.

[0095] At block 413, the method 400 also includes rendering the scene including at least a portion of the geometrically transformed 3D object model to generate a rendered image. The rendered image may include or correspond to the augmented image data 160.

[0096] At block 414, the method 400 also includes performing a task operation based on the rendered image. For example, the task operator 148 may perform the task operation. Additionally, or alternatively, the task operation may be associated with or use a trained model, such as a trained detection model or a trained segmentation model, as illustrative, non-limiting examples.

[0097] At block 416, the method 400 also includes evaluating a loss term based on a ground truth associated with the geometrically transformed 3D object model to determine a parameter. For example, evaluating the loss term may be evaluated based on an objective function, such as the objective function 142. The ground truth may include or correspond to the transformed 3D object model 156 or the ground truth 165. The loss term may include or correspond to an output of the loss function unit 144, the task data 164, the task result 162, or a combination thereof. The parameter may include or correspond to the parameter 166. At block 418, the method 400 also includes backpropagating the parameter and updating the variable based on the parameter.

[0098] At block 422, the method 400 also includes determining whether a condition is satisfied. The condition can include satisfying a threshold number of iterations (e.g.,QUALCOMM Ref. No. 2500012WO- 30 - iterations of determining a 3D pose for the 3D object model), satisfying a threshold prediction accuracy (of the task operation and / or a task model), or identifying a failure of the task operation (or task model) on the rendered image in association with the geometrically transformed 3D object model. If the condition is not satisfied, the method 400 advances to block 410. If the condition is satisfied, the method advances to block 424.

[0099] At block 424, the method 400 also includes storing the rendered image. For example, the rendered image may be stored at the memory 106. Additionally, the rendered image may be included in test data that is configured to test or validate the task operation and / or the task model. At block 426, the method 400 ends. In some embodiments, after the method 400 ends at block 426, a next iteration of the method 400 may be started at block 402.

[0100] Although block 422 is described as being performed directly after block 418, in other embodiments, block 422 may be performed directly after a different block. For example, block 422 may additionally or alternatively be performed directly after block 413, block 414, or block 416, or block 424.

[0101] The method 400 of FIG. 4 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a controller, another hardware device, firmware device, or any combination thereof. As an example, the method 400 of FIG. 4 may be performed by a processor that executes instructions, such as described with reference to FIG. 15.

[0102] FIG. 5 is a diagram of images to illustrate 3D data augmentation, in accordance with one or more aspects of the present disclosure. For example, the images of FIG. 5 may be based on and / or illustrate one or more operations of the system 100, the device 102, the processor 108, the image generator 120, described with reference to the method 400 of FIG. 4, or a combination thereof.

[0103] FIG. 5 includes a first image 500 of a scene that indicates different 3D poses associated with a 3D object model, and a second image 520 of the scene that indicates aQUALCOMM Ref. No. 2500012WO- 31 - final 3D pose associated with the 3D object model. The 3D object model may include or correspond to the 3D object model 150.

[0104] Referring to the first image 500, the first image 500 indicates multiple 3D poses of the 3D object model. For example, each 3D pose is depicted as a 3D box having a position, an orientation, and a size. As shown, the first image 500 indicates a first 3D pose 502, a second 3D pose 504, a third 3D pose 506, a fourth 3D pose 508, a fifth 3D pose 510, a sixth 3D pose 512, and a seventh 3D pose 514. As illustrated in the first image 500, a pose optimization is performed to determine 3D pose parameters of the 3D object model to maximize a detection loss or produce a failure with respect to a 3D task model. The position of each of the 3D boxes may occur along a linear path or a nonlinear path.

[0105] Referring to the second image 520, the seventh 3D pose 514 is indicated as a final 3D pose of the 3D object model which has been transformed to a geometrically transformed 3D object model (e.g., a truck) included in the second image 520. The second image 520 may include or correspond to a modified / augmented image that is associated with the augmented image data 160. It is noted that the modified / augmented image would not include the 3D box positioned with respect to the geometrically transformed 3D object model (e.g., the truck).

[0106] FIG. 6 is a diagram of an example of images to illustrate 3D data augmentation, in accordance with one or more aspects of the present disclosure. For example, the images of FIG. 6 may be based on and / or illustrate one or more operations of the system 100, the device 102, the processor 108, the image generator 120, described with reference to the method 400 of FIG. 4, or a combination thereof.

[0107] FIG. 6 includes a first BEV 600 of a scene associated with the scene data 130, a second BEV 620, a third BEV 640, and an image 660 that indicates different 3D poses associated with a 3D object model. The 3D object model may include or correspond to the 3D object model 150.

[0108] The first BEV 600 includes a placement space 602 (depicted as a roadway) and one or more foreground objects 604 (hereinafter referred to as “the foreground objectQUALCOMM Ref. No. 2500012WO- 32 -604”). Additionally, the first BEV 600 may include one or more background objects (not shown). The first BEV 600 also indicates a camera pose 606 associated with an image, such as an image associated with the image data 132. In some embodiments, the foreground object 604, the camera pose 606, or a combination thereof, may include or correspond to the annotations 136.

[0109] The first BEV 600 also includes or indicates a placement area 232 that is a portion of the placement space 240. The placement space 240, the placement area 232, or a combination thereof, may include or correspond to the placement space data 134. The placement area 232 may include or correspond to a portion of the placement space 240 that is included in the image associated with the camera pose 236. The placement area 232 may include a single continuous area or may include multiple discontinuous areas.

[0110] The first BEV 600 also includes or indicates a set of candidate locations based on or associated with the placement space 602. The set of candidate locations may be selected, based on the placement space 602, such that each candidate location is within the placement space 602. Additionally, or alternatively, the set of candidate locations may be selected randomly, according to a pattern or a condition, or a combination thereof. As an illustrative example, an initial candidate location may be selected randomly and then additional candidate locations may be identified based on a pattern and / or to have a certain distance (e.g., minimum distance) between a closest neighboring candidate location. The set of candidate locations (e.g., circles) may include a first location 610, a second location 612, and a third location 614. In some embodiments, the first location 610 includes or corresponds to the location 154. As shown in the first BEV 600, the set of candidate locations includes ten locations; however, in other embodiments, the set of candidate locations may include fewer than ten locations or more than ten locations.[OHl] A search, such as a greedy search, may be performed to identify one or more candidate locations of the set of candidate locations to be used to generate 3D augmented data, such as the augmented image data 160. In some examples, for each candidate location, a 3D pose parameter (e.g., the 3D pose parameter 152) may beQUALCOMM Ref. No. 2500012WO- 33 - determined for the candidate location and a 3D object model (e.g., the 3D object model 150) may be transformed based on the 3D pose parameter and positioned at the candidate location. For each candidate location, augmented image data (e.g., the augmented image data 160) may be generated based on the transformed 3D object model positioned at the candidate location, and the generated augmented image data may be provided to a 3D task model (e.g., a 3D perception task model) to determine an uncertainty value associated with the candidate location. The augmented image data that causes a particular uncertainty (e.g., the highest uncertainty or an uncertainty that is greater than or equal to a threshold) may be selected to be stored or further processed. In some embodiments, the set of candidate locations may be sorted based on the uncertainty values and one or more candidate locations may be identified based on the uncertainty values. For example, two candidate locations may be identified to have their respective augmented image data stored or further processed.

[0112] In some embodiments, the augmented image data associated with the first location 610 is identified / selected based on a corresponding uncertainty. Referring to the second BEV 620, the second BEV 620 includes the identified / selected first location 610 (and does not include the other candidate locations of the set of candidate locations). The second BEV 620 also includes an indicator 622 of a transformed 3D object model associated with the first location 610.

[0113] In some embodiments, an updated location 650 and an updated 3D pose may be determined based on the first location 610 (e.g., the transformed 3D object model associated with the first location 610. The updated location 650 and the updated 3D pose may be determined as described herein at least with reference to FIGS. 1-3, 4, or 5. The updated location 650 may be associated with a higher uncertainty than the uncertainty of the first location 610. Referring to the third BEV 640, the third BEV 640 includes the updated location 650 and an indicator 652 of a transformed 3D object model generated based on the updated 3D pose and associated with the updated location 650.

[0114] The image 660 includes a first 3D box as the indicator 622 associated with the first location 610 and a second 3D box as the indicator 652 associated with the updatedQUALCOMM Ref. No. 2500012WO- 34 - location 650. It is noted that a transformed 3D object model based on the updated 3D posed is positioned within the indicator 652 at the updated location 650. The transformed 3D object model (and the second 3D box) is positioned in from of the foreground object 604 (e.g., an excavator) and does not overlap or collide with the foreground object 604. It is noted that the image 660 (e.g., a modified / augmented image) would not include the first and second 3D boxes. The image 660, or augmented image data corresponding thereto, may be stored as a testing or validation sample.

[0115] FIG. 7 depicts a diagram of an example of an integrated circuit 700 operable to perform 3D data augmentation, in accordance with some examples of the present disclosure. The integrated circuit 700 includes one or more processors 708 (herein after referred to as the “processor 708”) and a memory 706. The processor 708 and the memory 706 may include or correspond to the processor 108 and the memory 106, respectively. The processor 708 may include the image generator 720 and, optionally, the objective operator 140 (as indicated by the dashed box in FIG. 7). The image generator 720 may include or correspond to the image generator 120. The image generator 720 includes the pose determiner 122 and the image modifier unit 124. The memory 706 includes (e.g., stores) the scene data 130 and the 3D object models 138.

[0116] The integrated circuit 700 also includes an input interface 704, such as one or more bus interfaces, to enable the integrated circuit 700 to receive signals representing input data 770 for processing. For example, the input data 770 can correspond to or include the instructions 109, the scene data 130, the image data 132, the placement space data 134, the annotations 136, the 3D object models 138, the 3D object model, the 3D scene model 126, a 3D task model (e.g., a 3D perception task model), a training sample, the hypemetwork 250, the normalization network 260, or a combination thereof.

[0117] The integrated circuit 700 also includes an output interface 705, such as a bus interface, to enable the integrated circuit 700 to output signals representing output data 772. For example, the output data 772 can correspond to or include the 3D pose parameter 152, the location 154, the transformed 3D object model 156, the augmented image data 160, the task result 162, the task data 164, the ground truth 165, theQUALCOMM Ref. No. 2500012WO- 35 - parameter 166, a testing / validation sample, the set of weights 266, or a combination thereof.

[0118] The integrated circuit 700 includes the image generator 720, the objective operator 140, or a combination thereof enables implementation of 3D data augmentation in a system or a device. For example, the system or the device may include a mobile device (e.g., a mobile phone or tablet) as depicted in FIG. 8, a wearable electronic device as depicted in FIG. 9, a camera as depicted in FIG. 10, a virtual reality, mixed reality, or augmented reality headset as depicted in FIG. 11, a mixed reality or augmented reality glasses device, as described with reference to FIG. 12, or a vehicle as depicted in FIG. 13.

[0119] In some embodiments, the system or the device that includes the integrated circuit 700 also includes or is coupled to an image sensor (e.g., a camera), an input device (e.g., a microphone, a keyboard or touch screen, etc.), a display device, a speaker, a modem, or a combination thereof. For example, the image sensor, the display device, and the modem may include or correspond to the image sensor 112, the display device 116, and the modem 118, respectively.

[0120] In some embodiments, the system or the device that includes the integrated circuit 700 is operable to perform 3D data augmentation. For example, the image generator 720 may obtain the scene data 130 that includes image data (e.g., the image data 132) associated with an image of a scene, and includes placement space data (e.g., the placement space data 134) that indicates a placement space associated with the image. The image generator 720 obtains a 3D object model (from the 3D object models 138). The pose determiner 122 identifies a 3D pose parameter associated with the 3D object model. The 3D pose parameter indicates a location within the placement space and an orientation. The image modifier unit 124 performs, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model. The image modifier unit 124 also generates, based on the image and the location, augmented image data (e.g., the augmented image data 160) that includes a modified image of the scene. The modified image includes at least a portion of the first geometrically transformed 3D object modelQUALCOMM Ref. No. 2500012WO- 36 - within the placement space. The augmented image data may be stored at the memory 706 or output via the output interface 705. The image generator 720 (e.g., the pose determiner 122 and the image modifier unit 124) and, optionally, the objective operator 140 enables the system or the device to perform an end-to-end differentiable approach for optimizing a position and a pose of the 3D object model 150 based on a gradient (e.g., the task data 164) from a 3D task model.

[0121] FIG. 8 depicts a diagram of a mobile device 800 operable to perform 3D data augmentation, in accordance with some examples of the present disclosure. The mobile device 800 may include or correspond to a phone or a tablet, as illustrative, non-limiting examples. The mobile device 800 includes a camera 802 (e.g., an image sensor), a display 804 (e.g., a display screen), a microphone 806, a speaker 808, and the integrated circuit 700. Components of the integrated circuit 700, including the image generator 720 (e.g., the pose determiner 122 and / or the image modifier unit 124) and / or the objective operator 140, are integrated in the mobile device 800 and are illustrated using dashed lines to indicate internal components that are not generally visible to a user of the mobile device 800.

[0122] FIG. 9 depicts a diagram of a wearable electronic device 900 operable to perform 3D data augmentation, in accordance with some examples of the present disclosure. The wearable electronic device 900 may include or correspond to a “smart watch,” as an illustrative, non-limiting example. The wearable electronic device 900 includes a camera 902 (e.g., an image sensor), a display 904 (e.g., a display screen), a microphone 906, a speaker 908, and the integrated circuit 700. Components of the integrated circuit 700, including the image generator 720 (e.g., the pose determiner 122 and / or the image modifier unit 124) and / or the objective operator 140, are integrated in the wearable electronic device 900 and are illustrated using dashed lines to indicate internal components that are not generally visible to a user of the wearable electronic device 900.

[0123] FIG. 10 is a diagram of a camera device 1000 operable to perform 3D data augmentation, in accordance with some examples of the present disclosure. The camera device 1000 includes an image sensor 1002, a display 1004 (e.g., a display screen), aQUALCOMM Ref. No. 2500012WO- 37 - microphone 1006, a speaker 1008, and the integrated circuit 700. Components of the integrated circuit 700, including the image generator 720 (e.g., the pose determiner 122 and / or the image modifier unit 124) and / or the objective operator 140 are integrated in the camera device 1000 and are illustrated using dashed lines to indicate internal components that are not generally visible to a user of the camera device 1000.

[0124] FIG. 11 is a diagram of a headset 1100, such as a virtual reality, mixed reality, or augmented reality headset, operable to perform 3D data augmentation, in accordance with some examples of the present disclosure. A visual interface device is positioned in front of the user's eyes to enable display of augmented reality, mixed reality, or virtual reality images or scenes to the user while the headset 1100 is worn. The headset 1100 also includes a camera 1102 (e.g., an image sensor), a display 1104 (e.g., a display screen), a microphone 1106, a speaker 1108, and the integrated circuit 700. Components of the integrated circuit 700, including the image generator 720 (e.g., the pose determiner 122 and / or the image modifier unit 124) and / or the objective operator 140, are integrated in the headset 1100 and are illustrated using dashed lines to indicate internal components that are not generally visible to a user of the headset 1100.

[0125] FIG. 12 is a diagram of a mixed reality or augmented reality glasses device 1200 operable to perform 3D data augmentation, in accordance with some examples of the present disclosure. The glasses 1200 include a holographic projection unit 1204 configured to project visual data onto a surface of a lens 1205 or to reflect the visual data off of a surface of the lens 1205 and onto the wearer’s retina. The glasses 1200 also include a camera 1202 (e.g., an image sensor), a microphone 1206, a speaker 1208, and the integrated circuit 700. Components of the integrated circuit 700, including the image generator 720 (e.g., the pose determiner 122 and / or the image modifier unit 124) and / or the objective operator 140, are integrated in the glasses 1200 and are illustrated using dashed lines to indicate internal components that are not generally visible to a user of the glasses 1200.

[0126] FIG. 13 is a diagram of an example of a vehicle 1300 operable to perform 3D data augmentation, in accordance with some examples of the present disclosure. The vehicle 1300 may include or correspond to a car. The vehicle 1300 includes a cameraQUALCOMM Ref. No. 2500012WO- 38 -1302 (e.g., an image sensor), a display 1304 (e.g., a display screen), a microphone 1306, one or more speakers 1308, and the integrated circuit 700. Components of the integrated circuit 700, including the image generator 720 (e.g., the pose determiner 122 and / or the image modifier unit 124) and / or the objective operator 140, are integrated in the vehicle 1300 and are illustrated using dashed lines to indicate internal components that are not generally visible to a user of the vehicle 1300.

[0127] In some embodiments, a system or the device that includes the integrated circuit 700 may obtain the scene data (e.g., the scene data 130) that includes image data (e.g., the image data 132) associated with an image of a scene, and includes placement space data (e.g., the placement space data 134) that indicates a placement space associated with the image. The integrated circuit 700 obtains a 3D object model (from the 3D object models 150), and identifies a 3D pose parameter associated with the 3D object model. The 3D pose parameter indicates a location within the placement space and an orientation. The integrated circuit 700 performs, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model. The integrated circuit 700 also generates, based on the image and the location, augmented image data (e.g., the augmented image data 160) that includes a modified image of the scene. The modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space. The augmented image data may be stored at a memory, output via a modem, or displayed via a display (e.g., the display 804, 904, 1004, 1104, 1304 or the holographic projection unit 1204). Generating the augmented image data enables a device (e.g., the mobile device 800, the wearable electronic device 900, the camera device 1000, the headset 1100, the glasses 1200, the vehicle 1300, or a server) to perform an end-to-end differentiable approach for optimizing a position and a pose of an 3D object model based on a gradient (e.g., the task data 164) from a 3D task model. In some embodiments, the augmented image data (e.g., a testing / validation sample) may be configured to result in high errors or failure for a 3D perception task. Accordingly, the augmented image data (e.g., the generated sample) enables improved training and / or testing / validation of the 3D task model. Additionally, generation of the augmented image data may improve the compute efficiency and reduce an amount of time toQUALCOMM Ref. No. 2500012WO- 39 - generate the augmented image data as compared to conventional techniques that generate testing / validation samples.

[0128] The embodiments of the systems or devices as described with reference to FIGS. 7-13 are described, respectively, as including a display, a microphone, a speaker, a camera, or a combination thereof. As described with reference to FIGS. 7-13, the display and the camera may include or correspond to the display device 116 and the image sensor 112, respectively. It is noted that in other embodiments of the systems or devices of FIGS. 7-13, one or more of the systems or devices of FIGS. 7-13 may not include the display, the microphone, the speaker, the camera, or a combination thereof. Additionally, or alternatively, one or more of the systems or devices of FIGS. 7-13 may include an additional component. For example, the additional component may include a modem, such as the modem 118.

[0129] FIG. 14 is a diagram of an example of a method 1400 of performing 3D data augmentation, in accordance with some aspects of the present disclosure. In a particular aspect, one or more operations of the method 1400 are performed by the system 100, the device 102, the processor 108, the image generator 120, the integrated circuit 700, the processor 708, the image generator 720, the mobile device 800 (e.g., a mobile phone or a tablet computer device), the wearable electronic device 900, the camera 1000, the virtual reality, mixed reality headset 1100, the augmented reality headset, the mixed reality or augmented reality glasses device 1200, the vehicle 1300, a server, or a combination thereof.

[0130] In some embodiments, the method 1400 includes, at block 1402, obtaining scene data that includes image data associated with an image of a scene, and includes placement space data that indicates a placement space associated with the image. For example, the scene data, the image data, and the placement space data may include or correspond to the scene data 130, the image data 132, and the placement space data 134, respectively. In some embodiments, the scene data includes one or more annotations associated with the image. For example, the one or more annotations may include a camera location, a camera pose, a foreground object, or a combination thereof. The one or more annotations may include or correspond to the annotations 136. Additionally, orQUALCOMM Ref. No. 2500012WO- 40 - alternatively, the scene data may include training data used to train a task model. To illustrate, the training data may have been used to generate a task model used by the task operator 148.

[0131] At block 1404, the method 1400 also includes obtaining a 3D object model. For example, the 3D object model may include or correspond to the 3D object model 150. In some implementations, obtaining the 3D object model includes selecting the 3D object model from multiple 3D object models, such as the 3D object models 138.

[0132] At block 1406, the method 1400 further includes identifying a 3D pose parameter associated with the 3D object model. The 3D pose parameter indicates a location within the placement space and an orientation. The 3D pose parameter and the location may include or correspond to the 3D pose parameter 152 and the location 154, respectively.

[0133] At block 1408, the method 1400 includes performing, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model. For example, the first geometrically transformed 3D object model may include or correspond to the transformed 3D object model 156. In some embodiments, performing the one or more 3D transformation operations includes scaling the 3D object model, adjusting an orientation of the 3D object model, or a combination thereof.

[0134] At block 1410, the method 1400 includes generating, based on the image and the location, augmented image data that includes a modified image of the scene. For example, the augmented image data may include or correspond to the augmented image data 160. The modified image may include at least a portion of the first geometrically transformed 3D object model within the placement space. For example, generating the augmented image data may include positioning a center of the first geometrically transformed 3D object model at the location. In some embodiments, generating the augmented image data includes rendering the modified image.

[0135] In some embodiments, the method 1400 includes identifying, based on the scene data, one or more foreground objects included in the image. The one or moreQUALCOMM Ref. No. 2500012WO- 41 - foreground objects may include or correspond to the foreground object 234 or 604. Additionally, or alternatively, the method 1400 includes generating a 3D representation of the image, and modifying the 3D representation. For example, the 3D representation can be modified to include the first geometrically transformed 3D object model, to remove at least one foreground object of the one or more foreground objects, or a combination thereof. In some such examples, generating the augmented image data includes rendering the modified image based on the modified 3D representation.

[0136] In some embodiments, the method 1400 includes determining, based on the scene data, a placement area of the placement space. For example, the placement area may include or correspond to the placement area 232. Additionally, or alternatively, the method 1400 includes providing the placement area to a hypemetwork to generate a set of weights for a normalization network. The hypemetwork, the set of weights, and the normalization network may include or correspond to the hypemetwork 250, the set of weights 266, and the normalization network 260, respectively. In some embodiments, the method 1400 also includes providing an indication of one or more foreground objects associated with the image to the hypernetwork. In such embodiments, the set of weights can be generated based on the placement area and the one or more foreground objects. The indication of the one or more foreground object may include or correspond to the foreground object 234 or 604. The method 1400 may include initiating the normalization network based on the set of weights, and identifying, based on the normalization network and a probability distribution (e.g., the base distribution 262, 302, 312, 322, 332, 342, or 352), the location associated with the placement space.

[0137] In some embodiments, the method 1400 includes performing a task operation based on the augmented image data to generate a task result, and generating, based on the task operation, one or more parameters. For example, the task operation may be performed by or using the task operator 148. The one or more parameters and the task result may include or correspond to the parameter 166, and the task result 162 and / or the task data 164. In some examples, the method 1400 also includes determining an updated 3D pose parameter based on the one or more parameters, where the updated 3D pose parameter indicates another location within the placement space. In some embodiments, the method 1400 includes obtaining an output value of an objectiveQUALCOMM Ref. No. 2500012WO- 42 - function. For example, the objective function may include or correspond to the objective function 142. Additionally, the parameter may include or indicate the output value, such as a gradient value or a latent value. The output value may be based on a loss function associated with the task operation, one or more regularization operations, or a combination thereof. The updated 3D pose parameter may be determined based on the output value of the objective function.

[0138] In some embodiments, the method 1400 may include performing, based on the updated 3D pose parameter, the one or more 3D transformation operations to generate a second geometrically transformed 3D object model of the 3D object model. Additionally, the method 1400 may include generating, based on the image and the other location, updated augmented image data that includes another modified image of the scene. The other modified image includes at least a portion of the second geometrically transformed 3D object model within the placement space. For example, the other modified image may include or correspond to the image 520 or 660.

[0139] In some embodiments, the method 1400 includes determining a ground truth associated with the first geometrically transformed 3D object model within the placement space. For example, the ground truth may include or correspond to the transformed 3D object model 156 or the ground truth 165. The one or more parameters, such as the parameter 166, may be generated based on the ground truth and the task result (e.g., the task result 162, the task data 164, or a combination thereof).

[0140] In some embodiments, the method 1400 includes randomly selecting the location from the placement space. In some such embodiments, the 3D pose parameter is identified based on the location. Additionally, or alternatively, the method 1400 includes selecting a set of candidate locations based on the placement space, where the set of candidate locations includes the location. The set of candidate locations may include or correspond to the candidate locations 610, 612, and 614. For each candidate location of the set of candidate locations, the method 1400 may determine, based on a trained task model, a loss value associated with the candidate location. The trained task model may include or correspond to a model used or applied by the task operator 148. In some examples, the method 1400 includes identifying or selecting, based on the lossQUALCOMM Ref. No. 2500012WO- 43 - values, at least one candidate location of the set of candidate locations. For example, the first location 610 may be selected from the set of candidate locations that includes candidate locations 610, 612, and 614. The method 1400 can include determining, based on the loss value associated with the at least one candidate location, an updated 3D pose parameter associated with the at least one candidate location. For example, based on the loss value associated with the at least one candidate location, an updated location may be identified, such as the location 650, and the updated 3D pose parameter can be determined based on the updated location.

[0141] In some embodiments, the method 1400 includes training, based on a training data set that includes at least a portion of the scene data, a task model to generate a trained task model. The trained task model may include or correspond to a model used by the task operator 148. The trained task model may include a trained detection model or a trained segmentation model, as illustrative, non-limiting examples. The method 1400 may also include generating a testing data set that includes the augmented image data. Additionally, or alternatively, the method 1400 can include testing the trained task model based on the testing data set.

[0142] In some embodiments, the method 1400 includes generating the image data via one or more cameras, outputting the modified image via a display, or a combination thereof. For example, the one or more cameras and the display may include or correspond to the image sensor 112 and the display device 116, respectively. Additionally, or alternatively, the method 1400 includes transmitting the augmented image data to a second device via a modem. For example, the modem may include or correspond to the modem 118.

[0143] The method 1400 of FIG. 14 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a DSP, a controller, another hardware device, firmware device, or any combination thereof. As an example, the method 1400 of FIG. 14 may be performed by a processor that executes instructions, such as described with reference to FIG. 15.QUALCOMM Ref. No. 2500012WO- 44 -

[0144] It is noted that one or more blocks (or operations) described with reference to FIGS. 4 or 14 may be combined with one or more blocks (or operations) described with reference to another of the figures. For example, one or more blocks (or operations) of FIG. 14 may be combined with one or more blocks (or operations) of FIG. 4. As another example, one or more blocks associated with FIGS. 4 or 14 may be combined with one or more blocks (or operations) associated with FIGS. 1-13. Additionally, or alternatively, one or more operations described above with reference to FIGS. 1-14 may be combined with one or more operations described with reference to FIG. 15.

[0145] FIG. 15 is a block diagram of an illustrative example of a device 1500 that is operable to perform 3D data augmentation, in accordance with one or more aspects of the present disclosure. In various implementations, the device 1500 may have more or fewer components than illustrated in FIG. 15. In an illustrative implementation, the device 1500 may correspond to the device 102. In an illustrative implementation, the device 1500 may perform one or more operations described with reference to FIGS. 1- 14.

[0146] In a particular implementation, the device 1500 includes a processor 1506 (e.g., a central processing unit (CPU)). The device 1500 may include one or more additional processors 1510 (e.g., one or more DSPs). In a particular aspect, the processor 108 of FIG. 1 or the processor 708 of FIG. 7 corresponds to the processor 1506, the processors 1510, or a combination thereof. The processors 1510 may include a speech and music coder-decoder (CODEC) 1508 that includes a voice coder (“vocoder”) encoder 1536, a vocoder decoder 1538, or a combination thereof. Additionally, or alternatively, the processors 1510 may include an image generator 1580. The image generator 1580 may include or correspond to the image generator 120 or 720. The image generator 1580 may include a pose determiner 1598 and an image modifier 1599. The pose determiner 1598 and the image modifier 1599 may include or correspond to the pose determiner 122 and the image modifier unit 124, respectively. In some embodiments, the image generatorl580 may optionally include the objective operator 140.

[0147] In this context, the term “processor” refers to an integrated circuit consisting of logic cells, interconnects, input / output blocks, clock management components, memory,QUALCOMM Ref. No. 2500012WO- 45 - and optionally other special purpose hardware components, designed to execute instructions and perform various computational tasks. Examples of processors include, without limitation, central processing units (CPUs), digital signal processors (DSPs), neural processing units (NPU), graphics processing units (GPUs), field programmable gate arrays (FPGAs), microcontrollers, quantum processors, coprocessors, vector processors, other similar circuits, and variants and combinations thereof. In some cases, a processor can be integrated with other components, such as communication components, input / output components, etc. to form a system on a chip (SOC) device or a packaged electronic device.

[0148] Taking CPUs as a starting point, a CPU typically includes one or more processor cores, each of which includes a complex, interconnected network of transistors and other circuit components defining logic gates, memory elements, etc. A core is responsible for executing instructions to, for example, perform arithmetic and logical operations. Typically, a CPU includes an Arithmetic Logic Unit (ALU) that handles mathematical operations and a Control Unit that generates signals to coordinate the operation of other CPU components, such as to manage operations a fetch-decode- execute cycle.

[0149] CPUs and / or individual processor cores generally include local memory circuits, such as registers and cache to temporarily store data during operations. Registers include high-speed, small-sized memory units intimately connected to the logic cells of a CPU. Often registers include transistors arranged as groups of flip-flops, which are configured to store binary data. Caches include fast, on-chip memory circuits used to store frequently accessed data. Caches can be implemented, for example, using Static Random-Access Memory (SRAM) circuits.

[0150] Operations of a CPU (e.g., arithmetic operations, logic operations, and flow control operations) are directed by software and firmware. At the lowest level, the CPU includes an instruction set architecture (ISA) that specifies how individual operations are performed using hardware resources (e.g., registers, arithmetic units, etc.). Higher level software and firmware is translated into various combinations of ISA operations to cause the CPU to perform specific higher-level operations. For example, an ISAQUALCOMM Ref. No. 2500012WO- 46 - typically specifies how the hardware components of the CPU move and modify data to perform operations such as addition, multiplication, and subtraction, and high-level software is translated into sets of such operations to accomplish larger tasks, such as adding two columns in a spreadsheet. Generally, a CPU operates on various levels of software, including a kernel, an operating system, applications, and so forth, with each higher level of software generally being more abstracted from the ISA and usually more readily understandable by human users.

[0151] GPUs, NPUs, DSPs, microcontrollers, coprocessors, FPGAs, ASICS, and vector processors include components similar to those described above for CPUs. The differences among these various types of processors are generally related to the use of specialized interconnection schemes and ISAs to improve a processor’s ability to perform particular types of operations. For example, the logic gates, local memory circuits, and the interconnects therebetween of a GPU are specifically designed to improve parallel processing, sharing of data between processor cores, and vector operations, and the ISA of the GPU may define operations that take advantage of these structures. As another example, ASICs are highly specialized processors that include similar circuitry arranged and interconnected for a particular task, such as encryption or signal processing. As yet another example, FPGAs are programmable devices that include an array of configurable logic blocks (e.g., interconnected sets of transistors and memory elements) that can be configured (often on the fly) to perform customizable logic functions.

[0152] The device 1500 may include a memory 1586 and a CODEC 1534. The memory 1586 may include or correspond to the memory 106 or 706. The memory 1586 may include instructions 1556, that are executable by the one or more additional processors 1510 (or the processor 1506) to implement the functionality described with reference to the image generator 1580, or both. The instructions 1556 may include or correspond to the instructions 109. The memory 1586 also may include scene data 1582 and 3D object models 1583. The scene data 1582 and the 3D object models 1583 may include or correspond to the scene data 130 and the 3D object models 138, respectively. The device 1500 may include a modem 1570 coupled, via a transceiver 1550, to an antenna 1552. The modem 1570 may include or correspond to the modem 118.QUALCOMM Ref. No. 2500012WO- 47 -

[0153] The device 1500 may include a display 1528 coupled to a display controller 1526. One or more speakers 1592, one or more microphone(s) 1594, or both, may be coupled to the CODEC 1534. The CODEC 1534 may include a digital -to-analog converter (DAC) 1502, an analog-to-digital converter (ADC) 1504, or both. In a particular implementation, the CODEC 1534 may receive analog signals from the microphone(s) 1594, convert the analog signals to digital signals using the analog-to- digital converter 1504, and provide the digital signals to the speech and music codec 1508. In a particular implementation, the speech and music codec 1508 may provide digital signals to the CODEC 1534. The CODEC 1534 may convert the digital signals to analog signals using the digital -to-analog converter 1502 and may provide the analog signals to the speaker 1592.

[0154] In a particular implementation, the device 1500 may be included in a system-in- package or system-on-chip device 1522. In a particular implementation, the memory 1586, the processor 1506, the processors 1510, the display controller 1526, the CODEC 1534, and the modem 1570 are included in the system-in-package or system-on-chip device 1522. In a particular implementation, an input device 1530, a power supply 1544, and a camera 1545 are coupled to the system-in-package or the system-on-chip device 1522. For example, the camera 1545 may include or correspond to the image sensor 112. In some examples, the input device 1530 may include or be associated with the display device 116 or the display 1528. Moreover, in a particular implementation, as illustrated in FIG. 15, the display 1528, the input device 1530, the speaker(s) 1592, the microphone(s) 1594, the antenna 1552, the power supply 1544, and the camera 1545 are external to the system-in-package or the system-on-chip device 1522. In a particular implementation, each of the display 1528, the input device 1530, the speaker(s) 1592, the microphone(s) 1594, the antenna 1552, the power supply 1544, and the camera 1545 may be coupled to a component of the system-in-package or the system-on-chip device 1522, such as an interface or a controller. The display controller 1526, the display 1528, or both, may include or correspond to the display device 116.

[0155] The device 1500 may include a smart speaker, a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, aQUALCOMM Ref. No. 2500012WO- 48 - music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a car, a computing device, a communication device, an internet-of- things (loT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0156] In conjunction with the described implementations, an apparatus includes means for obtaining scene data that includes image data associated with an image of a scene, and includes placement space data that indicates a placement space associated with the image. For example, the means for obtaining the scene data can include the system 100, the device 102, the memory 106, the processor 108, the modem 118, the image generator 120, the pose determiner 122, the image modifier unit 124, the hypernetwork 250, the integrated circuit 700, the input interface 704, the memory 706, the processor 708, the image generator 720, the mobile device 800, the wearable electronic device 900, the camera device 1000, the headset 1100, the glasses 1200, the vehicle 1300, the device 1500, the processor 1506, the processor(s) 1510, the system-in-package or the system-on-chip device 1522, the transceiver 1550, the modem 1570, the image generator 1580, the memory 1586, the pose determiner 1598, the image modifier unit 1599, other circuitry configured to obtain scene data, a server, or a combination thereof.

[0157] The apparatus also includes means for obtaining a 3D object model. For example, the means for obtaining the 3D object model can include the system 100, the device 102, the memory 106, the processor 108, the modem 118, the image generator 120, the pose determiner 122, the image modifier unit 124, the hypernetwork 250, the integrated circuit 700, the input interface 704, the memory 706, the processor 708, the image generator 720, the mobile device 800, the wearable electronic device 900, the camera device 1000, the headset 1100, the glasses 1200, the vehicle 1300, the device 1500, the processor 1506, the processor(s) 1510, the system-in-package or the system- on-chip device 1522, the transceiver 1550, the modem 1570, the image generator 1580, the memory 1586, the pose determiner 1598, the image modifier unit 1599, other circuitry configured to obtain a 3D object model, a server, or a combination thereof.QUALCOMM Ref. No. 2500012WO- 49 -

[0158] The apparatus further includes means for identifying a 3D pose parameter associated with the 3D object model. For example, the means for identifying the 3D pose parameter can include the system 100, the device 102, the processor 108, the image generator 120, the pose determiner 122, the image modifier unit 124, the hypernetwork 250, the integrated circuit 700, the processor 708, the image generator 720, the mobile device 800, the wearable electronic device 900, the camera device 1000, the headset 1100, the glasses 1200, the vehicle 1300, the device 1500, the processor 1506, the processor(s) 1510, the system-in-package or the system-on-chip device 1522, the image generator 1580, the pose determiner 1598, the image modifier unit 1599, other circuitry configured to identify a 3D pose parameter, a server, or a combination thereof. The 3D pose parameter indicates a location within the placement space and an orientation.

[0159] The apparatus includes means for performing, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model. For example, the means for performing the one or more 3D transformation operations can include the system 100, the device 102, the processor 108, the image generator 120, the image modifier unit 124, the integrated circuit 700, the processor 708, the image generator 720, the mobile device 800, the wearable electronic device 900, the camera device 1000, the headset 1100, the glasses 1200, the vehicle 1300, the device 1500, the processor 1506, the processor(s) 1510, the system-in-package or the system-on-chip device 1522, the image generator 1580, the image modifier unit 1599, other circuitry configured to perform one or more 3D transformation operations, a server, or a combination thereof.

[0160] The apparatus includes means for generating, based on the image and the location, augmented image data that includes a modified image of the scene. For example, the means for generating the augmented image data can include the system 100, the device 102, the processor 108, the image generator 120, the image modifier unit 124, the integrated circuit 700, the processor 708, the image generator 720, the mobile device 800, the wearable electronic device 900, the camera device 1000, the headset 1100, the glasses 1200, the vehicle 1300, the device 1500, the processor 1506, the processor(s) 1510, the system-in-package or the system-on-chip device 1522, the image generator 1580, the image modifier unit 1599, other circuitry configured toQUALCOMM Ref. No. 2500012WO- 50 - generate augmented image data, a server, or a combination thereof. The modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.

[0161] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as the memory 1586) includes instructions (e.g., the instructions 1556) that, when executed by one or more processors (e.g., the one or more processors 1510 or the processor 1506), cause the one or more processors to obtain scene data (e.g., the scene data 1582) that includes image data associated with an image of a scene, and includes placement space data that indicates a placement space associated with the image. The instructions also cause the one or more processors to obtain a 3D object model (e.g., an object model of the 3D object models 1583), and identify a 3D pose parameter associated with the 3D object model. The 3D pose parameter indicates a location within the placement space and an orientation. The instructions further cause the one or more processors to perform, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model. The instructions also cause the one or more processors to generate, based on the image and the location, augmented image data that includes a modified image of the scene. The modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.

[0162] Particular aspects of the disclosure are described below in sets of interrelated Examples:

[0163] According to Example 1, a device includes a memory configured to store scene data associated with a scene; and one or more processors configured to obtain the scene data that includes: image data associated with an image of the scene; and placement space data that indicates a placement space associated with the image; obtain a three- dimensional (3D) object model; identify a 3D pose parameter associated with the 3D object model, the 3D pose parameter indicates a location within the placement space and an orientation; perform, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometricallyQUALCOMM Ref. No. 2500012WO- 51 - transformed 3D object model; and generate, based on the image and the location, augmented image data that includes a modified image of the scene, the modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.

[0164] Example 2 includes the device of Example 1, where the scene data includes one or more annotations associated with the image, the one or more annotations include a camera location, a camera pose, a foreground object, or a combination thereof.

[0165] Example 3 includes the device of Example 1 or Example 2, where the scene data includes training data used to train a task model.

[0166] Example 4 includes the device of any of Examples 1 to 3, where, to obtain the 3D object model, the one or more processors are configured to select the 3D object model from multiple 3D object models.

[0167] Example 5 includes the device of any of Examples 1 to 4, where, to perform the one or more 3D transformation operations, the one or more processors are configured to scale the 3D object model, adjust an orientation of the 3D object model, or a combination thereof.

[0168] Example 6 includes the device of any of Examples 1 to 5, where, to generate the modified image, the one or more processors are configured to position a center of the first geometrically transformed 3D object model at the location.

[0169] Example 7 includes the device of any of Examples 1 to 6, where the one or more processors are configured to determine, based on the scene data, a placement area of the placement space; and provide the placement area to a hypemetwork to generate a set of weights for a normalization network.

[0170] Example 8 includes the device of Example 7, where the one or more processors are configured to initiate the normalization network based on the set of weights; and identify, based on the normalization network and a probability distribution, the location associated with the placement space.QUALCOMM Ref. No. 2500012WO- 52 -

[0171] Example 9 includes the device of Example 7, where the one or more processors are configured to provide an indication of one or more foreground objects associated with the image to the hypernetwork.

[0172] Example 10 includes the device of Example 9, where the set of weights is generated based on the placement area and the one or more foreground objects.

[0173] Example 11 includes the device of any of Examples 1 to 10, where the one or more processors are configured to: perform a task operation based on the augmented image data to generate a task result; and generate, based on the task operation, one or more parameters.

[0174] Example 12 includes the device of Example 11, where the one or more processors are configured to determine an updated 3D pose parameter based on the one or more parameters, the updated 3D pose parameter indicates another location within the placement space; and perform, based on the updated 3D pose parameter, the one or more 3D transformation operations to generate a second geometrically transformed 3D object model of the 3D object model.

[0175] Example 13 includes the device of Example 12, where the one or more processors are configured to generate, based on the image and the other location, updated augmented image data that includes another modified image of the scene, and where the other modified image includes at least a portion of the second geometrically transformed 3D object model within the placement space.

[0176] Example 14 includes the device of Example 11, where the one or more processors are configured to obtain an output value of an objective function, the output value based on a loss function associated with the task operation, one or more regularization operations, or a combination thereof; and determine an updated 3D pose parameter based on the output value of the objective function.

[0177] Example 15 includes the device of any of Examples 11 to 14, where the one or more processors are configured to determine a ground truth associated with the firstQUALCOMM Ref. No. 2500012WO- 53 - geometrically transformed 3D object model within the placement space, and where the one or more parameters are generated based on the ground truth and the task result.

[0178] Example 16 includes the device of any of Examples 1 to 15, where the one or more processors are configured to randomly select the location from the placement space, and where the 3D pose parameter is identified based on the location.

[0179] Example 17 includes the device of any of Examples 1 to 15, where the one or more processors are configured to select a set of candidate locations based on the placement space, where the set of candidate locations includes the location; for each candidate location of the set of candidate locations, determine, based on a trained task model, a loss value associated with the candidate location; and identify, based on the loss values, at least one candidate location of the set of candidate locations; and determine, based on the loss value associated with the at least one candidate location, an updated 3D pose parameter associated with the at least one candidate location.

[0180] Example 18 includes the device of any of Examples 1 to 17, where the one or more processors are configured to identify, based on the scene data, one or more foreground objects included in the image; and generate a 3D representation of the image.

[0181] Example 19 includes the device of Example 18, where the one or more processors are configured to modify the 3D representation to include the first geometrically transformed 3D object model and remove at least one foreground object of the one or more foreground objects.

[0182] Example 20 includes the device of Example 19, where the one or more processors are configured to, to generate the augmented image data, render the modified image based on the modified 3D representation.

[0183] Example 21 includes the device of any of Examples 1 to 20, where the one or more processors are configured to train, based on a training data set that includes at least a portion of the scene data, a task model to generate a trained task model, the trained task model includes a trained detection model or a trained segmentation model; generateQUALCOMM Ref. No. 2500012WO- 54 - a testing data set that includes the augmented image data; and test the trained task model based on the testing data set.

[0184] Example 22 includes the device of any of Examples 1 to 21, further comprising one or more cameras coupled to the one or more processors and configured to generate the image data.

[0185] Example 23 includes the device of any of Examples 1 to 22, further comprising a display coupled to the one or more processors and configured to output the modified image.

[0186] Example 24 includes the device of any of Examples 1 to 23, further comprising a modem coupled to the one or more processors, the modem configured to transmit the augmented image data to a second device.

[0187] Example 25 includes the device of any of Examples 1 to 24, where the one or more processors are integrated in a mobile phone, a tablet computer device, a wearable electronic device, a virtual reality headset, a mixed reality headset, an augmented reality headset, or a camera.

[0188] Example 26 includes the device of any of Examples 1 to 24, where the one or more processors are integrated in a vehicle.

[0189] According to Example 27, a method includes obtaining scene data that includes: image data associated with an image of a scene; and placement space data that indicates a placement space associated with the image; obtaining a three-dimensional (3D) object model; identifying a 3D pose parameter associated with the 3D object model, the 3D pose parameter indicates a location within the placement space and an orientation; performing, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model; and generating, based on the image and the location, augmented image data that includes a modified image of the scene, the modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.QUALCOMM Ref. No. 2500012WO- 55 -

[0190] Example 28 includes the method of Example 27, where the scene data includes one or more annotations associated with the image, the one or more annotations include a camera location, a camera pose, a foreground object, or a combination thereof.

[0191] Example 29 includes the method of Example 27 or Example 28, where the scene data includes training data used to train a task model.

[0192] Example 30 includes the method of any of Examples 27 to 29, where obtaining the 3D object model includes selecting the 3D object model from multiple 3D object models.

[0193] Example 31 includes the method of any of Examples 27 to 30, where performing the one or more 3D transformation operations includes scaling the 3D object model, adjusting an orientation of the 3D object model, or a combination thereof.

[0194] Example 32 includes the method of any of Examples 27 to 31, where generating the augmented image data includes positioning a center of the first geometrically transformed 3D object model at the location.

[0195] Example 33 includes the method of any of Examples 27 to 32, and further includes: determining, based on the scene data, a placement area of the placement space; and providing the placement area to a hypemetwork to generate a set of weights for a normalization network.

[0196] Example 34 includes the method of Example 33, and further includes: initiating the normalization network based on the set of weights; and identifying, based on the normalization network and a probability distribution, the location associated with the placement space.

[0197] Example 35 includes the method of Example 33, and further includes providing an indication of one or more foreground objects associated with the image to the hypernetwork.

[0198] Example 36 includes the method of Example 35, where the set of weights is generated based on the placement area and the one or more foreground objects.QUALCOMM Ref. No. 2500012WO- 56 -

[0199] Example 37 includes the method of any of Examples 27 to 36, and further includes: performing a task operation based on the augmented image data to generate a task result; and generating, based on the task operation, one or more parameters.

[0200] Example 38 includes the method of Example 37, and further includes: determining an updated 3D pose parameter based on the one or more parameters, the updated 3D pose parameter indicates another location within the placement space; and performing, based on the updated 3D pose parameter, the one or more 3D transformation operations to generate a second geometrically transformed 3D object model of the 3D object model.

[0201] Example 39 includes the method of Example 38, and further includes generating, based on the image and the other location, updated augmented image data that includes another modified image of the scene, and where the other modified image includes at least a portion of the second geometrically transformed 3D object model within the placement space.

[0202] Example 40 includes the method of Example 37, and further includes: obtaining an output value of an objective function, the output value based on a loss function associated with the task operation, one or more regularization operations, or a combination thereof; and determining an updated 3D pose parameter based on the output value of the objective function.

[0203] Example 41 includes the method of any of Examples 37 to 40, and further includes determining a ground truth associated with the first geometrically transformed 3D object model within the placement space, and where the one or more parameters are generated based on the ground truth and the task result.

[0204] Example 42 includes the method of any of Examples 27 to 41, and further includes randomly selecting the location from the placement space, and where the 3D pose parameter is identified based on the location.

[0205] Example 43 includes the method of any of Examples 27 to 41, and further includes: selecting a set of candidate locations based on the placement space, where theQUALCOMM Ref. No. 2500012WO- 57 - set of candidate locations includes the location; for each candidate location of the set of candidate locations, determining, based on a trained task model, a loss value associated with the candidate location; and identifying, based on the loss values, at least one candidate location of the set of candidate locations; and determining, based on the loss value associated with the at least one candidate location, an updated 3D pose parameter associated with the at least one candidate location.

[0206] Example 44 includes the method of any of Examples 27 to 43, and further includes: identifying, based on the scene data, one or more foreground objects included in the image; and generating a 3D representation of the image.

[0207] Example 45 includes the method of Example 44, and further includes modifying the 3D representation to include the first geometrically transformed 3D object model and remove at least one foreground object of the one or more foreground objects.

[0208] Example 46 includes the method of Example 45, where generating the augmented image data includes rendering the modified image based on the modified 3D representation.

[0209] Example 47 includes the method of any of Examples 27 to 46, and further includes: training, based on a training data set that includes at least a portion of the scene data, a task model to generate a trained task model, the trained task model includes a trained detection model or a trained segmentation model; generating a testing data set that includes the augmented image data; and testing the trained task model based on the testing data set.

[0210] Example 48 includes the method of any of Examples 27 to 47, and further includes generating the image data via one or more cameras; or outputting the modified image via a display.

[0211] Example 49 includes the method of any of Examples 27 to 48, and further includes transmitting the augmented image data to a second device via a modem.

[0212] Example 50 includes the method of any of Examples 27 to 49, where the method is performed by a mobile phone, a tablet computer device, a wearable electronic device,QUALCOMM Ref. No. 2500012WO- 58 - a virtual reality headset, a mixed reality headset, an augmented reality headset, a camera, or a vehicle.

[0213] According to Example 51, a non-transitory computer-readable medium storing instructions that are executable by one or more processors to cause the one or more processors to obtain scene data that includes: image data associated with an image of a scene; and placement space data that indicates a placement space associated with the image; obtain a three-dimensional (3D) object model; identify a 3D pose parameter associated with the 3D object model, the 3D pose parameter indicates a location within the placement space and an orientation; perform, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model; and generate, based on the image and the location, augmented image data that includes a modified image of the scene, the modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.

[0214] Example 52 includes the non-transitory computer-readable medium of Example 51, where the scene data includes one or more annotations associated with the image, the one or more annotations include a camera location, a camera pose, a foreground object, or a combination thereof.

[0215] Example 53 includes the non-transitory computer-readable medium of Example 51 or Example 52, where the scene data includes training data used to train a task model.

[0216] Example 54 includes the non-transitory computer-readable medium of any of Examples 51 to 53, where, to obtain the 3D object model, the instructions are executable by the one or more processors to cause the one or more processors to select the 3D object model from multiple 3D object models.

[0217] Example 55 includes the non-transitory computer-readable medium of any of Examples 51 to 54, where, to perform the one or more 3D transformation operations, the instructions are executable by the one or more processors to cause the one or moreQUALCOMM Ref. No. 2500012WO- 59 - processors to scale the 3D object model, adjust an orientation of the 3D object model, or a combination thereof.

[0218] Example 56 includes the non-transitory computer-readable medium of any of Examples 51 to 55, where, to generate the modified image, the instructions are executable by the one or more processors to cause the one or more processors to position a center of the first geometrically transformed 3D object model at the location.

[0219] Example 57 includes the non-transitory computer-readable medium of any of Examples 51 to 56, where the instructions are executable by the one or more processors to cause the one or more processors to determine, based on the scene data, a placement area of the placement space; and provide the placement area to a hypemetwork to generate a set of weights for a normalization network.

[0220] Example 58 includes the non-transitory computer-readable medium of Example 57, where the instructions are executable by the one or more processors to cause the one or more processors to initiate the normalization network based on the set of weights; and identify, based on the normalization network and a probability distribution, the location associated with the placement space.

[0221] Example 59 includes the non-transitory computer-readable medium of Example 57, where the instructions are executable by the one or more processors to cause the one or more processors to provide an indication of one or more foreground objects associated with the image to the hypernetwork.

[0222] Example 60 includes the non-transitory computer-readable medium of Example 59, where the set of weights is generated based on the placement area and the one or more foreground objects.

[0223] Example 61 includes the non-transitory computer-readable medium of any of Examples 51 to 60, where the instructions are executable by the one or more processors to cause the one or more processors to: perform a task operation based on the augmented image data to generate a task result; and generate, based on the task operation, one or more parameters.QUALCOMM Ref. No. 2500012WO- 60 -

[0224] Example 62 includes the non-transitory computer-readable medium of Example61, where the instructions are executable by the one or more processors to cause the one or more processors to determine an updated 3D pose parameter based on the one or more parameters, the updated 3D pose parameter indicates another location within the placement space; and perform, based on the updated 3D pose parameter, the one or more 3D transformation operations to generate a second geometrically transformed 3D object model of the 3D object model.

[0225] Example 63 includes the non-transitory computer-readable medium of Example62, where the instructions are executable by the one or more processors to cause the one or more processors to generate, based on the image and the other location, updated augmented image data that includes another modified image of the scene, and where the other modified image includes at least a portion of the second geometrically transformed 3D object model within the placement space.

[0226] Example 64 includes the non-transitory computer-readable medium of Example 61, where the instructions are executable by the one or more processors to cause the one or more processors to obtain an output value of an objective function, the output value based on a loss function associated with the task operation, one or more regularization operations, or a combination thereof; and determine an updated 3D pose parameter based on the output value of the objective function.

[0227] Example 65 includes the non-transitory computer-readable medium of any of Examples 61 to 64, where the instructions are executable by the one or more processors to cause the one or more processors to determine a ground truth associated with the first geometrically transformed 3D object model within the placement space, and where the one or more parameters are generated based on the ground truth and the task result.

[0228] Example 66 includes the non-transitory computer-readable medium of any of Examples 51 to 65, where the instructions are executable by the one or more processors to cause the one or more processors to randomly select the location from the placement space, and where the 3D pose parameter is identified based on the location.QUALCOMM Ref. No. 2500012WO- 61 -

[0229] Example 67 includes the non-transitory computer-readable medium of any of Examples 51 to 65, where the instructions are executable by the one or more processors to cause the one or more processors to select a set of candidate locations based on the placement space, where the set of candidate locations includes the location; for each candidate location of the set of candidate locations, determine, based on a trained task model, a loss value associated with the candidate location; and identify, based on the loss values, at least one candidate location of the set of candidate locations; and determine, based on the loss value associated with the at least one candidate location, an updated 3D pose parameter associated with the at least one candidate location.

[0230] Example 68 includes the non-transitory computer-readable medium of any of Examples 51 to 67, where the instructions are executable by the one or more processors to cause the one or more processors to identify, based on the scene data, one or more foreground objects included in the image; and generate a 3D representation of the image.

[0231] Example 69 includes the non-transitory computer-readable medium of Example68, where the instructions are executable by the one or more processors to cause the one or more processors to modify the 3D representation to include the first geometrically transformed 3D object model and remove at least one foreground object of the one or more foreground objects.

[0232] Example 70 includes the non-transitory computer-readable medium of Example69, where, to generate the augmented image data, the instructions are executable by the one or more processors to cause the one or more processors to render the modified image based on the modified 3D representation.

[0233] Example 71 includes the non-transitory computer-readable medium of any of Examples 51 to 70, where the instructions are executable by the one or more processors to cause the one or more processors to train, based on a training data set that includes at least a portion of the scene data, a task model to generate a trained task model, the trained task model includes a trained detection model or a trained segmentation model;QUALCOMM Ref. No. 2500012WO- 62 - generate a testing data set that includes the augmented image data; and test the trained task model based on the testing data set.

[0234] Example 72 includes the non-transitory computer-readable medium of any of Examples 51 to 21, where the instructions are executable by the one or more processors to cause the one or more processors to receive the image data from one or more cameras.

[0235] Example 73 includes the non-transitory computer-readable medium of any of Examples 51 to 72, where the instructions are executable by the one or more processors to cause the one or more processors to initiate output the modified image via a display.

[0236] Example 74 includes the non-transitory computer-readable medium of any of Examples 51 to 73, where the instructions are executable by the one or more processors to cause the one or more processors to initiate transmission of the augmented image data to a second device via a modem.

[0237] Example 75 includes the non-transitory computer-readable medium of any of Examples 51 to 74, where the non-transitory computer-readable medium is integrated in a mobile phone, a tablet computer device, a wearable electronic device, a virtual reality headset, a mixed reality headset, an augmented reality headset, or a camera.

[0238] Example 76 includes the non-transitory computer-readable medium of any of Examples 51 to 74, where the non-transitory computer-readable medium is integrated in a vehicle.

[0239] According to Example 77, an apparatus includes means for obtaining scene data that includes image data associated with an image of a scene, and placement space data that indicates a placement space associated with the image; means for obtaining a three- dimensional (3D) object model; means for identifying a 3D pose parameter associated with the 3D object model, the 3D pose parameter indicates a location within the placement space and an orientation; means for performing, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model; and means for generating,QUALCOMM Ref. No. 2500012WO- 63 - based on the image and the location, augmented image data that includes a modified image of the scene, the modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.

[0240] Example 78 includes the apparatus of Example 77, where the scene data includes one or more annotations associated with the image, the one or more annotations include a camera location, a camera pose, a foreground object, or a combination thereof.

[0241] Example 79 includes the apparatus of Example 77 or Example 78, where the scene data includes training data used to train a task model.

[0242] Example 80 includes the apparatus of any of Examples 77 to 79, where the means for obtaining the 3D object model includes means for selecting the 3D object model from multiple 3D object models.

[0243] Example 81 includes the apparatus of any of Examples 77 to 80, where the means for performing the one or more 3D transformation operations includes means for scaling the 3D object model, adjusting an orientation of the 3D object model, or a combination thereof.

[0244] Example 82 includes the apparatus of any of Examples 77 to 81, where the means for generating the augmented image data includes means for positioning a center of the first geometrically transformed 3D object model at the location.

[0245] Example 83 includes the apparatus of any of Examples 77 to 82, and further includes: means for determining, based on the scene data, a placement area of the placement space; and means for providing the placement area to a hypemetwork to generate a set of weights for a normalization network.

[0246] Example 84 includes the apparatus of Example 83, and further includes: means for initiating the normalization network based on the set of weights; and means for identifying, based on the normalization network and a probability distribution, the location associated with the placement space.QUALCOMM Ref. No. 2500012WO- 64 -

[0247] Example 85 includes the apparatus of Example 83, and further includes means for providing an indication of one or more foreground objects associated with the image to the hypernetwork.

[0248] Example 86 includes the apparatus of Example 85, where the set of weights is generated based on the placement area and the one or more foreground objects.

[0249] Example 87 includes the apparatus of any of Examples 77 to 86, and further includes: means for performing a task operation based on the augmented image data to generate a task result; and means for generating, based on the task operation, one or more parameters.

[0250] Example 88 includes the apparatus of Example 87, and further includes: means for determining an updated 3D pose parameter based on the one or more parameters, the updated 3D pose parameter indicates another location within the placement space; and means for performing, based on the updated 3D pose parameter, the one or more 3D transformation operations to generate a second geometrically transformed 3D object model of the 3D object model.

[0251] Example 89 includes the apparatus of Example 88, and further includes means for generating, based on the image and the other location, updated augmented image data that includes another modified image of the scene, and where the other modified image includes at least a portion of the second geometrically transformed 3D object model within the placement space.

[0252] Example 90 includes the apparatus of Example 87, and further includes: means for obtaining an output value of an objective function, the output value based on a loss function associated with the task operation, one or more regularization operations, or a combination thereof; and means for determining an updated 3D pose parameter based on the output value of the objective function.

[0253] Example 91 includes the apparatus of any of Examples 87 to 90, and further includes means for determining a ground truth associated with the first geometricallyQUALCOMM Ref. No. 2500012WO- 65 - transformed 3D object model within the placement space, and where the one or more parameters are generated based on the ground truth and the task result.

[0254] Example 92 includes the apparatus of any of Examples 77 to 91, and further includes means for randomly selecting the location from the placement space, and where the 3D pose parameter is identified based on the location.

[0255] Example 93 includes the apparatus of any of Examples 77 to 91, and further includes: means for selecting a set of candidate locations based on the placement space, where the set of candidate locations includes the location; means for determining, for each candidate location of the set of candidate locations and based on a trained task model, a loss value associated with the candidate location; and means for identifying, based on the loss values, at least one candidate location of the set of candidate locations; and means for determining, based on the loss value associated with the at least one candidate location, an updated 3D pose parameter associated with the at least one candidate location.

[0256] Example 94 includes the apparatus of any of Examples 77 to 93, and further includes: means for identifying, based on the scene data, one or more foreground objects included in the image; and means for generating a 3D representation of the image.

[0257] Example 95 includes the apparatus of Example 94, and further includes means for modifying the 3D representation to include the first geometrically transformed 3D object model and remove at least one foreground object of the one or more foreground objects.

[0258] Example 96 includes the apparatus of Example 95, where the means for generating the augmented image data includes means for rendering the modified image based on the modified 3D representation.

[0259] Example 97 includes the apparatus of any of Examples 77 to 96, and further includes: means for training, based on a training data set that includes at least a portion of the scene data, a task model to generate a trained task model, the trained task model includes a trained detection model or a trained segmentation model; means forQUALCOMM Ref. No. 2500012WO- 66 - generating a testing data set that includes the augmented image data; and means for testing the trained task model based on the testing data set.

[0260] Example 98 includes the apparatus of any of Examples 77 to 97, and further includes means for generating the image data via one or more cameras; or outputting the modified image via a display.

[0261] Example 99 includes the apparatus of any of Examples 77 to 98, and further includes means for transmitting the augmented image data to a second device via a modem.

[0262] Example 100 includes the apparatus of any of Examples 77 to 99, where the apparatus is a mobile phone, a tablet computer device, a wearable electronic device, a virtual reality headset, a mixed reality headset, an augmented reality headset, a camera, or a vehicle.

[0263] Those of skill would further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or combinations of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or processor executable instructions depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, such implementation decisions are not to be interpreted as causing a departure from the scope of the present disclosure.

[0264] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM),QUALCOMM Ref. No. 2500012WO- 67 - registers, hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transient storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor may read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or user terminal.

[0265] The previous description of the disclosed aspects is provided to enable a person skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope possible consistent with the principles and novel features as defined by the following claims.

Claims

QUALCOMM Ref. No. 2500012WO- 68 -WHAT IS CLAIMED IS:

1. A device comprising: a memory configured to store scene data associated with a scene; and one or more processors configured to: obtain the scene data that includes: image data associated with an image of the scene; and placement space data that indicates a placement space associated with the image; obtain a three-dimensional (3D) object model; identify a 3D pose parameter associated with the 3D object model, the 3D pose parameter indicates a location within the placement space and an orientation; perform, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model; and generate, based on the image and the location, augmented image data that includes a modified image of the scene, the modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.

2. The device of claim 1, wherein, the scene data includes: one or more annotations associated with the image, the one or more annotations include a camera location, a camera pose, a foreground object, or a combination thereof; training data used to train a task model; or a combination thereof.

3. The device of claim 1, wherein: to obtain the 3D object model, the one or more processors are configured to select the 3D object model from multiple 3D object models; to perform the one or more 3D transformation operations, the one or moreQUALCOMM Ref. No. 2500012WO- 69 - processors are configured to scale the 3D object model, adjust an orientation of the 3D object model, or a combination thereof; and to generate the modified image, the one or more processors are configured to position a center of the first geometrically transformed 3D object model at the location.

4. The device of claim 1, wherein the one or more processors are configured to: determine, based on the scene data, a placement area of the placement space; provide the placement area to a hypernetwork to generate a set of weights for a normalization network; initiate the normalization network based on the set of weights; and identify, based on the normalization network and a probability distribution, the location associated with the placement space.

5. The device of claim 4, wherein: the one or more processors are configured to provide an indication of one or more foreground objects associated with the image to the hypernetwork; and the set of weights is generated based on the placement area and the one or more foreground objects.

6. The device of claim 1, wherein the one or more processors are configured to: perform a task operation based on the augmented image data to generate a task result; and generate, based on the task operation, one or more parameters.

7. The device of claim 6, wherein the one or more processors are configured to: determine an updated 3D pose parameter based on the one or more parameters, the updated 3D pose parameter indicates another location within the placement space; perform, based on the updated 3D pose parameter, the one or more 3D transformation operations to generate a second geometricallyQUALCOMM Ref. No. 2500012WO- 70 - transformed 3D object model of the 3D object model; and generate, based on the image and the other location, updated augmented image data that includes another modified image of the scene, the other modified image includes at least a portion of the second geometrically transformed 3D object model within the placement space.

8. The device of claim 6, wherein the one or more processors are configured to: obtain an output value of an objective function, the output value based on a loss function associated with the task operation, one or more regularization operations, or a combination thereof; and determine an updated 3D pose parameter based on the output value of the objective function.

9. The device of claim 6, wherein: the one or more processors are configured to determine a ground truth associated with the first geometrically transformed 3D object model within the placement space; and the one or more parameters are generated based on the ground truth and the task result.

10. The device of claim 1, wherein: the one or more processors are configured to randomly select the location from the placement space; and the 3D pose parameter is identified based on the location.

11. The device of claim 1, wherein the one or more processors are configured to: select a set of candidate locations based on the placement space, where the set of candidate locations includes the location; for each candidate location of the set of candidate locations, determine, based on a trained task model, a loss value associated with the candidate location; and identify, based on the loss values, at least one candidate location of the set ofQUALCOMM Ref. No. 2500012WO- 71 - candidate locations; and determine, based on the loss value associated with the at least one candidate location, an updated 3D pose parameter associated with the at least one candidate location.

12. The device of claim 1, wherein the one or more processors are configured to: identify, based on the scene data, one or more foreground objects included in the image; generate a 3D representation of the image; modify the 3D representation to include the first geometrically transformed 3D object model and remove at least one foreground object of the one or more foreground objects; and to generate the augmented image data, render the modified image based on the modified 3D representation.

13. The device of claim 1, wherein the one or more processors are configured to: train, based on a training data set that includes at least a portion of the scene data, a task model to generate a trained task model, the trained task model includes a trained detection model or a trained segmentation model; generate a testing data set that includes the augmented image data; and test the trained task model based on the testing data set.

14. The device of claim 1, further comprising one or more cameras coupled to the one or more processors and configured to generate the image data.

15. The device of claim 1, further comprising a display coupled to the one or more processors and configured to output the modified image.

16. The device of claim 1, further comprising a modem coupled to the one or more processors, the modem configured to transmit the augmented image data to a second device.QUALCOMM Ref. No. 2500012WO- 72 -17. The device of claim 1, wherein the one or more processors are integrated in a mobile phone, a tablet computer device, a wearable electronic device, a virtual reality headset, a mixed reality headset, an augmented reality headset, or a camera.

18. The device of claim 1, wherein the one or more processors are integrated in a vehicle.

19. A method of operating a processor of an audio device, the method comprising: obtaining scene data that includes: image data associated with an image of a scene; and placement space data that indicates a placement space associated with the image; obtaining a three-dimensional (3D) object model; identifying a 3D pose parameter associated with the 3D object model, the 3D pose parameter indicates a location within the placement space and an orientation; performing, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model; and generating, based on the image and the location, augmented image data that includes a modified image of the scene, the modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.

20. A non-transitory computer-readable medium storing instructions that are executable by one or more processors to cause the one or more processors to: obtain scene data that includes: image data associated with an image of a scene; and placement space data that indicates a placement space associated with the image; obtain a three-dimensional (3D) object model;QUALCOMM Ref. No. 2500012WO- 73 - identify a 3D pose parameter associated with the 3D object model, the 3D pose parameter indicates a location within the placement space and an orientation; perform, based on the 3D pose parameter, one or more 3D transformation operations on the 3D object model to generate a first geometrically transformed 3D object model; and generate, based on the image and the location, augmented image data that includes a modified image of the scene, the modified image includes at least a portion of the first geometrically transformed 3D object model within the placement space.