Modifying image data
By generating a three-dimensional bounding box and using a conditioning input to orient content, the system addresses the lack of orientation control in conventional diffusion models, achieving precise image modification.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2025-08-29
- Publication Date
- 2026-05-15
AI Technical Summary
Conventional diffusion models lack fine control over the orientation of content added to images, allowing users to indicate only what and where content should be added, but not how it should be oriented.
The system generates a three-dimensional bounding box based on an input image, using a conditioning input associated with the orientation of this box to process the image with a diffusion model, ensuring the output image includes an object oriented accordingly.
Enables fine control over the location and orientation of content insertion, aligning with user intent and resolving ambiguities in image modification tasks.
Smart Images

Figure US2025044334_15052026_PF_FP_ABST
Abstract
Description
PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO1MODIFYING IMAGE DATATECHNICAL FIELD
[0001] The present disclosure generally relates to processing image data. For example, aspects of the present disclosure include systems and techniques for modifying image data.BACKGROUND
[0002] Diffusion models include a family of algorithms for generative modelling that achieve high- quality performance in several tasks (e.g., generating and / or modifying images based on text). For example, a diffusion model may be used to add content to an input image based on a text instruction.SUMMARY
[0003] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.
[0004] Systems and techniques are described for modifying image data. According to at least one example, a method is provided for modifying image data. The method includes: generating a three- dimensional (3D) bounding box based on an input image; generating a conditioning input based on the 3D bounding box, wherein the conditioning input is associated with an orientation of the 3D bounding box; and generating an output image by processing at least a portion of the input image using a diffusion model based on the conditioning input, wherein the output image includes an object oriented based on an orientation of the 3D bounding box.
[0005] In another example, an apparatus for modifying image data is provided that includes at least one memory and at least one processor (e.g., configured in circuitry) coupled to the at least one memory. The at least one processor configured to: generate a three-dimensional (3D) bounding box based on an input image; generate a conditioning input based on the 3D bounding box, wherein the conditioning input is associated with an orientation of the 3D bounding box; and generate an outputPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO2 image by processing at least a portion of the input image using a diffusion model based on the conditioning input, wherein the output image includes an object oriented based on an orientation of the 3D bounding box.
[0006] In another example, a non-transitory computer-readable medium is provided that has stored thereon instructions that, when executed by one or more processors, cause the one or more processors to: generate a three-dimensional (3D) bounding box based on an input image; generate a conditioning input based on the 3D bounding box, wherein the conditioning input is associated with an orientation of the 3D bounding box; and generate an output image by processing at least a portion of the input image using a diffusion model based on the conditioning input, wherein the output image includes an object oriented based on an orientation of the 3D bounding box.
[0007] In another example, an apparatus for modifying image data is provided. The apparatus includes: means for generating a three-dimensional (3D) bounding box based on an input image; means for generating a conditioning input based on the 3D bounding box, wherein the conditioning input is associated with an orientation of the 3D bounding box; and means for generating an output image by processing at least a portion of the input image using a diffusion model based on the conditioning input, wherein the output image includes an object oriented based on an orientation of the 3D bounding box.
[0008] In some aspects, one or more of the apparatuses described herein is, can be part of, or can include an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a vehicle (or a computing device, system, or component of a vehicle), a mobile device (e.g., a mobile telephone or so-called “smart phone”, a tablet computer, or other type of mobile device), a smart or connected device (e.g., an Intemet-of-Things (loT) device), a wearable device, a personal computer, a laptop computer, a video server, a television (e.g., a network- connected television), a robotics device or system, or other device. In some aspects, each apparatus can include an image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each apparatus can include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each apparatus can include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each apparatus can include one or more sensors. In some cases, the one or more sensors can be used for determining a location of the apparatuses, a state of thePCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO3 apparatuses (e.g., a tracking state, an operating state, a temperature, a humidity level, and / or other state), and / or for other purposes.
[0009] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.
[0010] The foregoing, together with other features and aspects, will become more apparent upon referring to the following specification, claims, and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Illustrative examples of the present application are described in detail below with reference to the following figures:
[0012] FIG. 1 is an illustration of an example system for modifying image data;
[0013] FIG. 2A is an illustration of an example system for modifying image data, according to various aspects of the present disclosure;
[0014] FIG. 2B is an illustration of another example system for modifying image data, according to various aspects of the present disclosure;
[0015] FIG. 3 A is an illustration of yet another example system for modifying image data, according to various aspects of the present disclosure;
[0016] FIG. 3B is an illustration of yet another example system for modifying image data, according to various aspects of the present disclosure;
[0017] FIG. 3 C is an illustration of yet another example system for modifying image data, according to various aspects of the present disclosure;
[0018] FIG. 4A includes four example images including respective bounding boxes and objects as examples of images generated or modified according to various aspects of the present disclosure;
[0019] FIG. 4B includes four example images including respective bounding boxes and objects as examples of images generated or modified according to various aspects of the present disclosure;PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO4
[0020] FIG. 4C includes four example images including respective bounding boxes and objects as examples of images generated or modified according to various aspects of the present disclosure;
[0021] FIG. 5 includes three images overlaid with 2D bounding boxes and three image frames including 3D bounding boxes, according to various aspects of the present disclosure;
[0022] FIG. 6 is a flow diagram illustrating an example process for modifying image data, in accordance with aspects of the present disclosure;
[0023] FIG. 7 includes two sets of images that show the forward diffusion process (which is fixed) and the reverse diffusion process (which is learned) of a diffusion model, according to various aspects of the present disclosure;
[0024] FIG. 8 includes a diagram illustrating how diffusion data is distributed from initial data to noise using a diffusion model in the forward diffusion direction, according to various aspects of the present disclosure;
[0025] FIG. 9 is a diagram illustrating a U-Net architecture for a diffusion model, according to various aspects of the present disclosure;
[0026] FIG. 10 is a block diagram illustrating an example of a deep learning neural network that can be used to perform various tasks, according to some aspects of the disclosed technology;
[0027] FIG. 11 is a block diagram illustrating an example of a convolutional neural network (CNN), according to various aspects of the present disclosure; and
[0028] FIG. 12 is a block diagram illustrating an example computing-device architecture of an example computing device which can implement the various techniques described herein.DETAILED DESCRIPTION
[0029] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. However, it will be apparent that various aspects may be practiced without these specific details. The figures and description are not intended to be restrictive.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO5
[0030] The ensuing description provides example aspects only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the exemplary aspects will provide those skilled in the art with an enabling description for implementing an exemplary aspect. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the application as set forth in the appended claims.
[0031] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term “aspects of the disclosure” does not require that all aspects of the disclosure include the discussed feature, advantage, or mode of operation.
[0032] As mentioned previously, diffusion models include a family of algorithms for generative modelling that achieve high-quality performance in several tasks (e.g., generating and / or modifying images based on text). For example, a diffusion model may be used to add content to an input image based on a text instruction.
[0033] However, conventional techniques for using diffusion models to modify images lack fine controls. For example, conventional techniques may allow a user to indicate what content should be added to an input image, and in some cases, where the content should be added. However, conventional techniques do not allow a user to indicate an orientation of the content to be added.
[0034] Systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively referred to herein as “systems and techniques”) are described herein for modifying image data. For example, the systems and techniques described herein may allow for a diffusion model to be instructed regarding an orientation of content to insert into an image.
[0035] In a first example, a user may provide an input image to the systems and techniques. The input image may include an object, for example, a vehicle. The user may instruct the systems and techniques to replace the vehicle with another object (e.g., a different vehicle). The systems and techniques may determine an orientation of the vehicle in the image and replace the vehicle with the other vehicle. The other vehicle may have an orientation that matches the orientation of replaced vehicle.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO6
[0036] In a second example, a user may use a graphical user interface (GUI) to indicate a position and orientation of an object to add to an image. The systems and techniques may insert the object according to the indicated position and orientation.
[0037] In the present disclosure, the terms “content,” “image content,” and like terms may refer to image data (e.g., pixels) to add to an image (e.g., to replace pixels of the image). Similarly, references to adding “objects,” “vehicles,” and like terms to images may refer to adding pixels representative of the objects (e.g., image data) to an image.
[0038] The systems and techniques may be implemented in various fields and for various purposes. For example, the systems and techniques may be used to generate image data that may be used for training and / or testing purposes. For instances, the systems and techniques may generate testing and / or training data that may be used to test and / or train an automated driving system. For example, the systems and techniques may use existing images as input images and replace vehicles in the input images with other vehicles. Additionally or alternatively, the systems and techniques may insert vehicles into the input images. The generated images may be used to test and / or train image-based operations of an automated driving system.
[0039] As an example, the systems and techniques may be used to allow a user to interactively edit images using a device (e.g., a smartphone). For example, the systems and techniques my allow the user to capture an image and add content to the image. The systems and techniques may allow the user to orient the content according to the intent of the user.
[0040] As yet another example, the systems and techniques may be used to allow a user or developer to add virtual objects to a virtual scene, for example, for extended reality (XR) (which may include virtual reality (VR), augmented reality (AR), and / or mixed reality (MR)). For example, the systems and techniques may enable a user or developer to add image data to a user’s field of view of a scene. The systems and techniques may enable the user or developer to orient the added image according to the intent of the user or developer.
[0041] Various aspects of the application will be described with respect to the figures below.
[0042] FIG. 1 is an illustration of an example system 100 for modifying image data. For example, system 100 may provide an input image 102 and an instruction 104 to a diffusion model 106. Diffusion model 106 may generate an output image 108 based on input image 102 and instruction 104. DiffusionPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO7 model 106 may add object 112 to input image 102 to generate output image 108 (e.g., to replace object 110). However, the position and / or orientation of object 112 in output image 108 may, or may not, match the intent of the user that provided instruction 104 to diffusion model 106. For example, object 112 may be oriented at an angle that is the opposite of the intent of the user.
[0043] FIG. 2A is an illustration of an example system 200 for modifying image data, according to various aspects of the present disclosure. For example, system 200 may provide an input image 202 and a conditioning input 204 image modifier 206. Image modifier 206 may generate an output image 208 based on input image 202 and conditioning input 204. Image modifier 206 may add object 210 to input image 202 to generate output image 208. Image modifier 206 may generate output image 208 such that the position and / or orientation of object 210 in output image 208 is based on conditioning input 204. Conditioning input 204 may indicate the position and / or location for inserting object 210 into input image 202.
[0044] Input image 202 may be, or may include, any suitable image. Input image 202 may, for example, be an image captured by a vehicle, by a user device (e.g., a smartphone), or an XR device (e.g., an image captured by a scene-facing camera of the XR device).
[0045] Conditioning input 204 may be, or may include, an indication of a position and / or orientation at which image modifier 206 is to generate object 210 for insertion into input image 202 when generating output image 208. In some aspects, conditioning input 204 may be a colored two- dimensional (2D) bounding box. The coloring of conditioning input 204 may indicate the orientation at which to generate object 210.
[0046] Image modifier 206 may include a diffusion model that may be used to generate image data. Image modifier 206 may generate conditioning inputs and use the conditioning inputs to instruct the diffusion model regarding how to generate output image 208. For example, image modifier 206 may generate conditioning inputs based on conditioning input 204.
[0047] Output image 208 may be an image generated by image modifier 206 based on input image 202 and conditioning input 204. Output image 208 may include object 210. Object 210 may be, or may include, pixels representative of any suitable object. Object 210 (e.g., pixels representative of an object) may be generated by image modifier 206 (e.g., by a diffusion model of image modifier 206),PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO8
[0048] System 200 may allow for fine control of image modification . For example, the systems and techniques may allow for fine control of the location and orientation of content to insert into images. For example, system 200 may allow control over how an object inserted into an image should appear in the image to match a user’s expectations. System 200 may resolve ambiguities that may arise based on hypothetical system that allows a user to provide an inpainting mask. For example, an inpainting mask may indicate a position of an object but may leave open two possible orientations.
[0049] Image modifier 206 allows for modifying images based on inserting new obj ects into images or replacing objects in images. FIG. 2A is an illustration of an example of using image modifier 206 to insert object 210 into input image 202 to generate output image 208. FIG. 2B is an illustration of an example of using image modifier 206 to insert object 220 into input image 212 (e.g., to replace object 222) to generate output image 218.
[0050] FIG. 3A is an illustration of an example system 300a for modifying image data, according to various aspects of the present disclosure. In general, system 300a may modify input image 302 to insert an object to generate output image data 334.
[0051] Input image 302 may be, or may include, any suitable image. Input image 302 may, for example, be an image captured by a vehicle, by a user device (e.g., a smartphone), or an XR device (e.g., an image captured by a scene-facing camera of the XR device).
[0052] Object detector 304 may generate mask 306 based on input image 302. Mask 306 may be, or may include, an indication of pixels of input image 302 that represent an object. Object detector 304 may be, or may include, an object-detector machine-learning model. For example, object detector 304 may be trained (e.g., through an iterative back-propagation training process) to detect objects in images and to generate masks based on the detected objects.
[0053] Image modifier 308 may generate image data 310 based on input image 302 and mask 306. Image data 310 may be, or may include, a portion of input image 302. For example, image modifier 308 may crop a portion of input image 302 to generate image data 310. Additionally or alternatively, image modifier 308 may apply mask 306 to the cropped portion of input image 302. For example, image modifier 308 may change pixel values of pixels of input image 302 indicated by mask 306. For instance, image modifier 308 may set to [0, 0, 0] (e.g., black) pixels of the portion of input image 302PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO9 that represent the object detected by object detector 304. Image generator 326 may use image data 310 as an input image.
[0054] Object detector 312 may generate (3D) bounding box 316 based on input image 302 and camera intrinsic parameters 314. Camera intrinsic parameters 314 may be, or may include, intrinsic parameters of a camera that captured input image 302. Camera intrinsic parameters 314 may include, for example, information regarding lens distortion of a lens of the camera.
[0055] Object detector 312 may be, or may include, a 3D object-detector machine-learning model. For example, object detector 312 may be trained (e.g., through an iterative back-propagation training process) to detect objects in images and to generate 3D bounding boxes based on the detected objects. Object detector 312 may detect an object.
[0056] Three-dimensional 3D bounding box 316 may be, or may include, coordinates in a 3D space, size dimensions, and orientation information. The 3D space may relate to a scene depicted by input image 302. The coordinates may indicate a position of an object in the 3D space. The coordinates may include at least one coordinate in each of three orthogonal directions (e.g., an x-dimension, a y- dimension, and a z-dimension). The size dimensions may include a width, a length, and a height of the object. The orientation information may be, or may include, an indication of an orientation of the object in the 3D space in three orthogonal directions (e.g., roll, pitch, and yaw). In some aspects, 3D bounding box 316 may be simplified. For example, in some aspects, object detector 312 may determine that all objects are on a ground plane and may thus omit a z-dimension and / or a roll angle, a pitch angle, and / or a height dimension.
[0057] Renderer 318 may render a colored two-dimensional (2D) bounding box 320 based on 3D bounding box 316 and mask 306. For example, renderer 318 may render 3D bounding box 316 in an image plane. The position and size of 2D bounding box 320 may indicate the 3D position and the size of 3D bounding box 316.
[0058] Renderer 318 may color 2D bounding box 320 to indicate the orientation of 3D bounding box 316. For example, 3D bounding box 316 may be a box including six sides, renderer 318 may render 2D bounding box 320 such that each side of the box of 3D bounding box 316 corresponds to a different color or coloring. In the present disclosure, the term “coloring” may refer to one or more colors (e.g., pattern of colors) of something (e.g., a rendered bounding box). For example, rendererPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO10318 may render 3D bounding box 316 such that a top face of 3D bounding box 316 is colored according to a first coloring, a front face of 3D bounding box 316 is colored according to a second coloring, and a left face of 3D bounding box 316 is colored according to a third coloring.
[0059] In some aspects, the coloring of a given face of an example 3D bounding box 316 may be colored with two colors. Additionally, the given face of may be colored with according to a diagonal coloring pattern. For example, the given face may be divided into two sections from one corner to an opposite corner (e.g., a top-left corner to a bottom-right comer). The two sections may have different colors. As another example, the given face may be colored according to a coloring gradient between opposite comers. For example, a top-left comer of a face may be one color, and a bottom-right comer of the face may be another color. The remainder of the face may be colored according to a gradient between the two colors.
[0060] The diagonal coloring may allow an observer of an image to determine the size of the face even if part of the face is occluded or not included in the image. For example, if part of a face is in an image and part is outside the frame of the image, the diagonal line may allow observer to trace a line to a corner of the face that is not in the image and determine the size of the face.
[0061] Guide 322 may generate conditioning input 324 based on 2D bounding box 320. Conditioning input 324 may be, or may include, features (e.g., generated by an image encoder). Conditioning input 324 may be an implicit representation of 2D bounding box 320. Renderer 318 encodes the orientation of 3D bounding box 316 as color in 2D bounding box 320. Guide 322 generates conditioning input 324 based on 2D bounding box 320. As such, conditioning input 324 includes a feature-space representation of the orientation of 3D bounding box 316.
[0062] Guide 322 may be a machine-learning model trained to generate features based on images. For example, guide 322 may be, or may include, a machine-learning-model encoder trained to encode images as features.
[0063] Image generator 326 may generate output image data 330 based on image data 310, prompt 328, and conditioning input 324. Image generator 326 may be, or may include, a diffusion model trained to generate image data.
[0064] Prompt 328 may be, or may include, an instruction (e.g., formatted as text) for generating output image data 330. Image generator 326 may be a diffusion model trained to generate imagesPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO11 based on an input image and one or more conditioning inputs. Image generator 326 may process image data 310 based on prompt 328 and conditioning input 324 (e.g., using prompt 328 and conditioning input 324 as conditioning inputs).
[0065] Blender 332 may insert output image data 330 into input image 302 to generate output image data 334. For example, blender 332 may blend pixels of output image data 330 with pixels of input image 302 to generate output image data 334.
[0066] Object detector 304, image modifier 308, and blender 332 are optional in system 300a. For example, some aspects of the present disclosure may omit object detector 304, image modifier 308, and blender 332. In such aspects, image generator 326 may use input image 302 as the input image (rather than image data 310). Further, in such aspects, image generator 326 may generate output image data 334 (rather than generating output image data 330).
[0067] Object detector 304, image modifier 308, and blender 332 are optional in system 300a. For example, some aspects of the present disclosure may omit object detector 304, image modifier 308, and blender 332. In such aspects, image generator 326 may use input image 302 as the input image (rather than image data 310). Further, in such aspects, image generator 326 may generate output image data 334 (rather than generating output image data 330).
[0068] For example, FIG. 3B is an illustration of example system 300b for modifying image data, according to various aspects of the present disclosure. In general, system 300b may modify input image 302 to insert an object to generate output image data 334. System 300b may be substantially the same as system 300a, but system 300b may omit object detector 304, image modifier 308, and blender 332.
[0069] System 300a of FIG. 3 A and system 300b of FIG. 3B illustrate operations related to replacing an object in an image with another object (e.g., replacing pixels that represent the object with pixels that represent the other object).
[0070] FIG. 3C is an illustration of example system 300c for modifying image data, according to various aspects of the present disclosure. System 300c may modify input image 302 by inserting an object into input image 302.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO12
[0071] Guide 322 may be the same in FIG. 3A, FIG. 3B and FIG. 3C. Similarly, image generator 326 may be the same in FIG. 3A, FIG. 3B and FIG. 3C. Object inserter 336 may generate 2D bounding box 320. 2D bounding box 320 may be the same in FIG. 3A, FIG. 3B and FIG. 3C.
[0072] Object inserter 336 may generate 2D bounding box 320 based on a user input. For example, object inserter 336 may include a user interface that may allow a user to indicate a position, orientation, size, color, shape, type, etc. of an object to insert into input image 302.
[0073] FIG. 4A includes four example images including respective bounding boxes and objects as examples of images generated or modified according to various aspects of the present disclosure. For example, the user interface (e.g., of system 300c) may display an image of a scene to a user. The user may use the user interface to insert an object into the image and may position and orient the object within the image. For example, the user may be able to drag (e.g., using a mouse or touch screen) the object to a desired location within the image.
[0074] In operation a user may select an image (e.g., input image 302) an object, a position for the object, and an orientation of the object (e.g., using a user interface). System 300c may determine prompt 328 based on the object (e.g., a text description of the object). Object inserter 336 may determine an instance of a 3D bounding box based on the size of the object, the position of the object and the orientation of the object. Further, object inserter 336 may determine an instance of 2D bounding box 320 based on the 3D bounding box. object inserter 336 may provide 2D bounding box 320 to guide 322. Guide 322 may generate conditioning input 324 based on 2D bounding box 320 and provide 2D bounding box 320 to image generator 326. Image generator 326 may generate an instance of output image data 334 based on conditioning input 324 and prompt 328. System 300c may display output image data 334 to the user using the user interface.
[0075] The four images of FIG. 4A may be examples of images displayed to a user. Additionally, the four images of FIG. 4A include examples of bounding boxes, which may, or may not, be displayed to the user with the images. The bounding boxes may be for internal operations of system 300c and may, or may not, be displayed at the user interface.
[0076] In some aspects, the user may be able to adjust the object, the position of the object, and / or the orientation of the object based on the image displayed at the user interface. For example, the user interface may display an instance of output image data 334 (e.g., one of the images of FIG. 4A). ThePCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO13 user may drag the object to a new position, drag a portion of the object to rotate the object, or change the object into another object by providing another input. The user interface may accept the user input and object inserter 336 may generate another instance of a bounding box based on the user input. Further, object inserter 336 may generate another instance of 2D bounding box 320 based on the bounding box and provide 2D bounding box 320 to guide 322. Guide 322 may generate another instance of conditioning input 324 based on 2D bounding box 320 and provide the other instance of conditioning input 324 to image generator 326. Image generator 326 may generate another instance of output image data 334 and provide the other instance of output image data 334 to the user interface and the user interface may display the other instance of output image data 334.
[0077] In some aspects, object inserter 336 may use a determine a depth of the object in the scene based on a y position of the object within the image. Further, object inserter 336 may determine a size of the object in the output image based on the depth of the object in the scene.
[0078] FIG. 4B includes four example images including respective bounding boxes and objects as examples of images generated or modified according to various aspects of the present disclosure. For example, the user interface (e.g., of system 300c) may display an image of a scene to a user. The user may use the user interface to insert an object into the image and may position and orient the object within the image. For example, the user may be able to drag (e.g., using a mouse or touch screen) the object to a desired location within the image. The four images of FIG. 4B include an object oriented in four different orientations as an example.
[0079] FIG. 4C includes four example images including respective bounding boxes and objects as examples of images generated or modified according to various aspects of the present disclosure. For example, the user interface (e.g., of system 300c) may display an image of a scene to a user. The user may use the user interface to insert an object into the image and may position and orient the object within the image. For example, the user may be able to drag (e.g., using a mouse or touch screen) the object to a desired location within the image. The four images of FIG. 4C include an object positioned in four different positions as an example.
[0080] Image generator 326 may generate the four images of FIG. 4C to properly occlude the inserted object. For example, in some aspects, because image generator 326 generates output image data 334 based on features (e.g., conditioning input 324) which are based on 2D bounding box 320,PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO14 which are in turn based on 3D bounding box 316 (either based on a replaced object or object indicated by a user through a user interface), image generator 326 insert objects into the input images based on the 3D scene into which the 3D bounding box 316 is inserted. Additionally or alternatively, object detector 304, or another object-detection model, may generate object masks for other objects in the scene (e.g., objects present in the images). Then image generator 326 may resolve occlusions via the object masks. For example, image generator 326 may insert pixels into images to represent inserted objects based on the masks of the other objects, such that the inserted object is occluded by the other objects.
[0081] FIG. 5 includes three images overlaid with 2D bounding boxes and three image frames including 3D bounding boxes, according to various aspects of the present disclosure. Image 502, image 504, and image 506 are examples of input images. Additionally, image 502, image 504, and image 506 are overlaid with 2D bounding boxes 508, 2D bounding boxes 510, and 2D bounding boxes 512 respectively. For example, an object detector, such as object detector 312 of FIG. 3A may determine 3D bounding box 316 based on objects in each of image 502, image 504, and image 506. Additionally, image frame 518 may render 3D bounding box 316 as 2D bounding box 320 (e.g., 2D bounding boxes 508, 2D bounding boxes 510, and 2D bounding boxes 512).
[0082] Image frame 514, image frame 516, and image frame 518 are image frames with 2D bounding boxes 508, 2D bounding boxes 510, and 2D bounding boxes 512, respectively, positioned therein. For example, 2D bounding boxes 508, 2D bounding boxes 510, and 2D bounding boxes 512 may have image coordinates that correspond to positions in image frames.
[0083] Image frame 514, image frame 516, and image frame 518 may be respective examples of visual maps that may indicate positions and orientations of objects in an image to replace. Additionally or alternatively, image frame 514, image frame 516, and image frame 518 may be respective examples of visual maps that may indicate positions and orientations of object to add to an image (e.g., with or without replacing an object in the image).
[0084] FIG. 6 is a flow diagram illustrating an example process 600 for modifying image data, in accordance with aspects of the present disclosure. One or more operations of process 600 may be performed by a computing device (or apparatus) or a component (e.g., a chipset, codec, etc.) of the computing device. The computing device may be a mobile device (e.g., a mobile phone), a network-PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO15 connected wearable such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or augmented reality (AR) device, a vehicle or component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device with the resource capabilities to perform the one or more operations of process 600. The one or more operations of process 600 may be implemented as software components that are executed and run on one or more processors.
[0085] At block 602, a computing device (or one or more components thereof) may generate a three- dimensional (3D) bounding box based on an input image. For example, object detector 312 may generate 3D bounding box 316 based on input image 302.
[0086] In some aspects, the 3D bounding box may be, or may include, coordinates in a 3D space, wherein the 3D space is based on the input image; size dimensions; and orientation information. For example, 3D bounding box 316 may include 3D coordinates, size dimensions, and orientation information.
[0087] At block 604, the computing device (or one or more components thereof) may generate a conditioning input based on the 3D bounding box, wherein the conditioning input is associated with an orientation of the 3D bounding box. For example, guide 322 may generate conditioning input 324 based, at least in part, on an orientation of 3D bounding box 316. For instance, guide 322 may generate conditioning input 324 based on d bounding box 320. 2D bounding box 320 may be rendered by Tenderer 318 based on 3D bounding box 316, including based on an orientation of 3D bounding box 316.
[0088] In some aspects, the computing device (or one or more components thereof) may render the 3D bounding box as a two-dimensional (2D) bounding box. The conditioning input is generated based on the 2D bounding box. For example, Tenderer 318 may render 3D bounding box 316 as 2D bounding box 320 and guide 322 may generate conditioning input 324 based on 2D bounding box 320.
[0089] In some aspects, to generate the conditioning input, the computing device (or one or more components thereof) may process the 2D bounding box using a machine-learning model to generate the conditioning input. For example, guide 322 may be, or may include, a machine-learning model trained to generate conditioning inputs based on images including 2D bounding boxes (e.g., colored 2D bounding boxes, where the color indicates orientation).PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO16
[0090] In some aspects, the conditioning input may be, or may include, a feature-space representation of the 3D bounding box. For example, guide 322 may encode 2D bounding box 320 as a feature-space representation to generate conditioning input 324.
[0091] In some aspects, the 2D bounding box may be rendered with coloring that is different for each face of the 3D bounding box to indicate an orientation of the 3D bounding box. For example, Tenderer 318 may render 2D bounding box 320 with coloring that is different for each face of 3D bounding box 316.
[0092] In some aspects, the 2D bounding box may be rendered with two respective colors for each face of the 3D bounding box to indicate an orientation of the 3D bounding box. For example, Tenderer 318 may render 2D bounding box 320 with two respective colors for each face of 3D bounding box 316. For instance, Tenderer 318 may render 2D bounding box 320 with coloring similar to the coloring of d bounding boxes 508, d bounding boxes 510, and d bounding boxes 512 of FIG. 5, for example, with a gradient between opposite corners of each face of the bounding box.
[0093] In some aspects, the 2D bounding box may be rendered based on a gradient between two respective colors for each face of the 3D bounding box. For example, Tenderer 318 may render 2D bounding box 320 with a gradient between two respective colors for each face of 3D bounding box 316. For instance, Tenderer 318 may render 2D bounding box 320 with coloring similar to the coloring of d bounding boxes 508, d bounding boxes 510, and d bounding boxes 512 of FIG. 5, for example, with a gradient between opposite corners of each face of the bounding box.
[0094] In some aspects, the computing device (or one or more components thereof) may detect an object in the input image; and determine an orientation of the object in the input image. The orientation of the 3D bounding box is based on the orientation of the detected object. For example, object detector 312 may detect an object in input image 302 and generate 3D bounding box 316. 3D bounding box 316 may include orientation information. The orientation information may be determined based on the object detected in input image 302.
[0095] In some aspects, the detected obj ect in the input image is replaced by the obj ect in the output image. For example, image modifier 206 of FIG. 2B may replace object 222 with object 220 in output image 218.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO17
[0096] In some aspects, the computing device (or one or more components thereof) may generate a mask indicative of pixels of the input image that represent the detected object; and generate the at least a portion of the input image based on the mask. For example, object detector 304 may generate mask 306 indicative of pixels of input image 302 that are to be replaced.
[0097] In some aspects, the computing device (or one or more components thereof) may crop the input image to generate the at least a portion of the input image. For example, image modifier 308 may crop input image 302 based on mask 306 to generate image data 310.
[0098] In some aspects, the output image may be a first output image. The computing device (or one or more components thereof) may blend the output image with the input image to generate a second output image. For example, blender 332 may blend output image data 330 with input image 302 to generate output image data 334.
[0099] In some aspects, the 3D bounding box is generated based on a user input. For example, object inserter 336 may generate 2D bounding box 320 based on a user input.
[0100] In some aspects, the computing device (or one or more components thereof) may cause a display to display the input image to a user; receive, from a user interface, the user input relative to the input image; and generate the 3D bounding box based on the user input, wherein the 3D bounding box is simulated at a depth in a scene depicted by the input image based on the user input. For example, object inserter 336 may display input image 302 to a user (for example, FIG. 4A, FIG. 4B, and FIG. 4C include examples of images that may be displayed at a user interface). Object inserter 336 may , from the user interface, receive a user input from the user. Object inserter 336 may generate 3D bounding box 316 based on the user input. Object inserter 336 may generate 3D bounding box 316 based on a position (including depth) and orientation indicated by the user input.
[0101] In some aspects, the computing device (or one or more components thereof) may determine a mask based on a model of the object, wherein the input image is processed by the diffusion model based on the mask. For example, object inserter 336 may determine a mask based on the object to be inserted and image generator 326 may process input image 302 based on the mask.
[0102] In some aspects, the computing device (or one or more components thereof) may crop the input image to generate the at least a portion of the input image. For example, image modifier 308 may crop input image 302 based on mask 306 to generate image data 310.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO18
[0103] In some aspects, the output image may be a first output image. The computing device (or one or more components thereof) may blend the output image with the input image to generate a second output image. For example, blender 332 may blend output image data 330 with input image 302 to generate output image data 334.
[0104] At block 606, the computing device (or one or more components thereof) may generate an output image by processing at least a portion of the input image using a diffusion model based on the conditioning input, wherein the output image includes an object oriented based on an orientation of the 3D bounding box. For example, image generator 326 may generate output image data 330 (or output image data 334) by processing image data 310 (or input image 302) using a diffusion model based on conditioning input 324. Output image data 330 may include an object oriented based on an orientation of 3D bounding box 316.
[0105] In some aspects, the object may not be present in the input image. For example, image modifier 206 of FIG. 2A may insert object 210 into output image 208. As another example, image modifier 206 of FIG. 2B may replace object 222 with object 220 in output image 218.
[0106] In some aspects, the computing device (or one or more components thereof) may test an automated driving system based on the output image. For example, process 600 may be used to generate test data to test an automated driving system.
[0107] In some aspects, the computing device (or one or more components thereof) may train an automated driving system based on the output image. For example, process 600 may be used to generate training data to train an automated driving system.
[0108] In some aspects, the computing device (or one or more components thereof) may cause a display of an extended reality (XR) device to display the output image. For example, process 600 may be used to generate XR content for display by an XR system.
[0109] In some aspects, the computing device (or one or more components thereof) may test a perception model based on the output image. For example, process 600 may be used to generate test data to test a perception model.
[0110] In some aspects, the object in the output image is in a challenging orientation. For example, the object 210 may be inserted into output image 208 in an orientation that may be challenging for a perception model to detect the object.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO19[OHl] In some examples, as noted previously, the methods described herein (e.g., process 600 of FIG. 6, and / or other methods described herein) can be performed, in whole or in part, by a computing device or apparatus. In one example, one or more of the methods can be performed by system 200 of FIG. 2 A and FIG. 2B, image modifier 206 of FIG. 2A and FIG. 2B, system 300a of FIG. 3 A, system 300b of FIG. 3B, system 300c of FIG. 3C, or by another system or device. In another example, one or more of the methods (e.g., process 600, and / or other methods described herein) can be performed, in whole or in part, by the computing-device architecture 1200 shown in FIG. 12. For instance, a computing device with the computing-device architecture 1200 shown in FIG. 12 can include, or be included in, the components of the system 200, image modifier 206, system 300a, system 300b, system 300c, and can implement the operations of process 600, and / or other process described herein. In some cases, the computing device or apparatus can include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component(s) that are configured to carry out the steps of processes described herein. In some examples, the computing device can include a display, a network interface configured to communicate and / or receive the data, any combination thereof, and / or other component(s). The network interface can be configured to communicate and / or receive Internet Protocol (IP) based data or other type of data.
[0112] The components of the computing device can be implemented in circuitry. For example, the components can include and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or can include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.
[0113] Process 600, and / or other process described herein are illustrated as logical flow diagrams, the operation of which represents a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The orderPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO20 in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.
[0114] Additionally, process 600, and / or other process described herein can be performed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non- transitory.
[0115] FIG. 7 provides two sets of images 700 that show the forward diffusion process (which is fixed) and the reverse diffusion process (which is learned) of a diffusion model. As shown in the forward diffusion process of FIG. 7, noise 704 is gradually added to a first set of images 702 at different time steps for a total of T time steps (e.g., making up a Markov chain), producing a sequence of noisy samples Xxthrough XT.
[0116] Diffusion models from a training perspective will take an image and will slowly add noise to the image to obscure the information in the image. In some aspects, the noise 704 is Gaussian noise. Each time step can correspond to each consecutive image of the first set of images 702 shown in FIG. 7. The initial image Xoof FIG. 7 is of a vase of flowers. Addition of the noise 704 to each image (corresponding to noisy samples Xxto XT) results in gradual diffusion of the pixels in each image until the final image (corresponding to sample XT) essentially matches the noise distribution. For example, by adding the noise, each data sample Xtthrough XTgradually loses its distinguishable features as the time step becomes larger, eventually resulting in the final sample XTbeing equivalent to the target noise distribution, for instance a unit variance zero- Gaussian 7^(0, 1).
[0117] The second set of images 706 shows the reverse diffusion process in which XTis the starting point with a noisy image (e.g., one that has Gaussian noise). The diffusion model can be trained to reverse the diffusion process (e.g., by training a model Pe(xt-i I xt)) to generate new data. In some aspects, a diffusion model can be trained by finding the reverse Markov transitions that maximize the likelihood of the training data. By traversing backwards along the chain of time steps, the diffusionPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO21 model can generate the new data. For example, as shown in FIG. 7, the reverse diffusion process proceeds to generate Xoas the image of the vase of flowers. In other cases, the input data and output data can vary based on the task for which the diffusion model is trained.
[0118] As noted above, the diffusion model is trained to be able to denoise or recover the original image Xoin an incremental process as shown in the second set of images 706. In some aspects, the neural network of the diffusion model can be trained to recover Xtgiven Xt-1, such as provided in the below example equation:
[0119] A diffusion kernel can be defined as:Define
[0120] Sampling can be defined as follows:
[0121] In some cases, the / ?tvalues schedule (also referred to as a noise schedule) is designed such that
[0122] The diffusion model runs in an iterative manner to incrementally generate the input image Xo. In one example, the model may have twenty steps. However, in other examples, the number of steps can vary.
[0123] FIG. 8 is a diagram 800 illustrating how diffusion data is distributed from initial data to noise using a diffusion model in the forward diffusion direction, in accordance with some aspects. Note that the initial data q(X0) is detailed in the initial stage of the diffusion process. An illustrative example of the data q(X0) is the initial image of the flowers in a vase shown in FIG. 7. As the diffusion model iterates and iteratively adds sampled noise to the data from t = 0 to t = T, as shown in FIG. 8, the data becomes nosier and may ultimately result in pure noise (e.g., at q(XT)). The example of FIG. 8 illustrates the progression of the data and how it becomes diffused with noise in the forward diffusion process.
[0124] In some aspects, the diffused data distribution (e.g., as shown in FIG. 8) can be as follows:PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO22
[0125] In the above equation, q(xt) represents the diffused data distribution, q x0, xt) represents the joint distribution, q(x0) represents the input data distribution, and q(xt|x0) is the diffusion kernel. In this regard, the model can sample xt~ q(xt) by first sampling x0- q(xo) and then samplingxt ~ q(.xt\xo) (which may be referred to as ancestral sampling). The diffusion kernel takes the input and returns a vector or other data structure as output.
[0126] The following is a summary of a training algorithm and a sampling algorithm for a diffusion model. A training algorithm can include the following steps:1: repeat2: x0~ QOo)3: t ~ Uniform ({1, ■ ■ ■ > T })4: 6 ~ JV'(O, I)5: Take gradient descent step on6: until converged
[0127] A sampling algorithm can include the following steps:6: return x0
[0128] FIG. 9 is a diagram illustrating a U-Net architecture 900 for a diffusion model, in accordance with some aspects. The initial image 902 (e.g., a vase of flowers) is provided to the U-Net architecture 900 which includes a series of residual networks (ResNet) blocks and self-attention layers to represent the network ee(xt, t ). The U-Net architecture 900 also includes fully-connected layers 910. In some cases, time representation 912 can be sinusoidal positional embeddings or random Fourier features. Noisy output 908 from the forward diffusion process is also shown.
[0129] The U-Net architecture 900 includes a contracting path 904 and an expanding path 906 as shown in FIG. 9, which gives it the U-shaped architecture. The contracting path 904 can be a convolutional network that includes repeated convolutional layers (that apply convolutionalPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO23 operations), each followed by a rectified linear unit (ReLU) and a max pooling operation. When images are being processed (e.g., the image 902) during the contracting path 904, the spatial information of the image 902 is reduced as features are generated. The expanding path 906 combines the features and spatial information through a sequence of up-convolutions and concatenations with high-resolution features from the contracting path 904. Some of the layers can be self-attention layers, which leverage global interactions between semantic features at the end of the encoder to explicitly model full contextual information.
[0130] Latent diffusion models (also referred to as stable diffusion models) introduce a diffusion process in the latent space of a machine learning model (e.g., variational autoencoder (VAE) neural network), making the machine learning model more efficient while enabling high-resolution image synthesis. For example, an Encoder (E) - Decoder (£>) pair of a VAE can be trained to capture a lowdimensional latent distribution given by z = £(x) such that x D(z . The denoising process outlined above can be formulated in this latent space by training a U-Net (e.g., U-Net architecture 900 of FIG. 9), which may include ResNet blocks and attention modules in some cases, to predict the noise introduced in the forward diffusion process, which optimizes the objective given by the following:
[0131] Here, e is the total noise introduced to the noise-free latent z0~£’(x) by the scheduler in T steps, ztis the corresponding partially-noisy latent at diffusion timestep t, and c is conditioning (e.g., text prompt embedding provided as input). With the predicted noise E0, denoising diffusion implicit models (DDIM) sampling can be applied on zTover T steps iteratively to recover z0in the original latent data distribution, such as in the following:where atis a parameter for noise scheduler.
[0132] When adopting Stable Diffusion (SD) to video generation or video editing, a key factor is to ensure the temporal consistency of a generated frame relative to one or more previous frames in the video. In addition to modifications to the U-Net model (such as temporal attention and 2+1D convolutions), it helps to rely on control signals, and / or denoising diffusion implicit models (DDIM) inversion to start the denoising with a correlated set of noise latents.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO24
[0133] As noted above, various aspects of the present disclosure can use machine-learning models or systems.
[0134] FIG. 10 is an illustrative example of a neural network 1000 (e.g., a deep-learning neural network) that can be used to implement machine-learning based feature segmentation, implicit-neural- representation generation, rendering, classification, object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc.), feature extraction, authentication, gaze detection, gaze prediction, and / or automation. For example, neural network 1000 may be an example of, or can implement, object detector 304 or object detector 312.
[0135] An input layer 1002 includes input data. In one illustrative example, input layer 1002 can include data representing input image 302. Neural network 1000 includes multiple hidden layers, for example, hidden layers 1006a, 1006b, through 1006n. The hidden layers 1006a, 1006b, through hidden layer 1006n include “n” number of hidden layers, where “n” is an integer greater than or equal to one. The number of hidden layers can be made to include as many layers as needed for the given application. Neural network 1000 further includes an output layer 1004 that provides an output resulting from the processing performed by the hidden layers 1006a, 1006b, through 1006n. In one illustrative example, output layer 1004 can provide mask 306 or 3D bounding box 316.
[0136] Neural network 1000 may be, or may include, a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated with the nodes is shared among the different layers and each layer retains information as information is processed. In some cases, neural network 1000 can include a feed-forward network, in which case there are no feedback connections where outputs of the network are fed back into itself. In some cases, neural network 1000 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading in input.
[0137] Information can be exchanged between nodes through node-to-node interconnections between the various layers. Nodes of input layer 1002 can activate a set of nodes in the first hidden layer 1006a. For example, as shown, each of the input nodes of input layer 1002 is connected to each of the nodes of the first hidden layer 1006a. The nodes of first hidden layer 1006a can transform the information of each input node by applying activation functions to the input node information. The information derived from the transformation can then be passed to and can activate the nodes of thePCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO25 next hidden layer 1006b, which can perform their own designated functions. Example functions include convolutional, up-sampling, data transformation, and / or any other suitable functions. The output of the hidden layer 1006b can then activate nodes of the next hidden layer, and so on. The output of the last hidden layer 1006n can activate one or more nodes of the output layer 1004, at which an output is provided. In some cases, while nodes (e.g., node 1008) in neural network 1000 are shown as having multiple output lines, a node has a single output and all lines shown as being output from a node represent the same output value.
[0138] In some cases, each node or interconnection between nodes can have a weight that is a set of parameters derived from the training of neural network 1000. Once neural network 1000 is trained, it can be referred to as a trained neural network, which can be used to perform one or more operations. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a tunable numeric weight that can be tuned (e.g., based on a training dataset), allowing neural network 1000 to be adaptive to inputs and able to leam as more and more data is processed.
[0139] Neural network 1000 may be pre-trained to process the features from the data in the input layer 1002 using the different hidden layers 1006a, 1006b, through 1006n in order to provide the output through the output layer 1004. In an example in which neural network 1000 is used to identify features in images, neural network 1000 can be trained using training data that includes both images and labels, as described above. For instance, training images can be input into the network, with each training image having a label indicating the features in the images (for the feature-segmentation machine-learning system) or a label indicating classes of an activity in each image. In one example using object classification for illustrative purposes, a training image can include an image of a number 2, in which case the label for the image can be [0 0 1 0 0 0 0 0 0 0],
[0140] In some cases, neural network 1000 can adjust the weights of the nodes using a training process called backpropagation. As noted above, a backpropagation process can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update are performed for one training iteration. The process can be repeated for a certain number of iterations for each set of training images until neural network 1000 is trained well enough so that the weights of the layers are accurately tuned.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO26
[0141] For the example of identifying objects in images, the forward pass can include passing a training image through neural network 1000. The weights are initially randomized before neural network 1000 is trained. As an illustrative example, an image can include an array of numbers representing the pixels of the image. Each number in the array can include a value from 0 to 255 describing the pixel intensity at that position in the array. In one example, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or luma and two chroma components, or the like).
[0142] As noted above, for a first training iteration for neural network 1000, the output will likely include values that do not give preference to any particular class due to the weights being randomly selected at initialization. For example, if the output is a vector with probabilities that the object includes different classes, the probability value for each of the different classes can be equal or at least very similar (e.g., for ten possible classes, each class can have a probability value of 0.1). With the initial weights, neural network 1000 is unable to determine low-level features and thus cannot make an accurate determination of what the classification of the object might be. A loss function can be used to analyze error in the output. Any suitable loss function definition can be used, such as a crossentropy loss. Another example of a loss function includes the mean squared error (MSE), defined as ^totai=X “ (target — output')2. The loss can be set to be equal to the value of Etotal.
[0143] The loss (or error) will be high for the first training images since the actual values will be much different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output is the same as the training label. Neural network 1000 can perform a backward pass by determining which inputs (weights) most contributed to the loss of the network and can adjust the weights so that the loss decreases and is eventually minimized. A derivative of the loss with respect to the weights (denoted as dL / dW where W are the weights at a particular layer) can be computed to determine the weights that contributed most to the loss of the network. After the derivative is computed, a weight update can be performed by updating all the weights of the filters. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be denoted as w = wi - T] dL / dW, where w denotes a weight, ^denotes the initial weight, and p denotes a learning rate. The learning rate can be set to any suitable value, with a high learning rate including larger weight updates and a lower value indicating smaller weight updates.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO27
[0144] Neural network 1000 can include any suitable deep network. One example includes a convolutional neural network (CNN), which includes an input layer and an output layer, with multiple hidden layers between the input and out layers. The hidden layers of a CNN include a series of convolutional, nonlinear, pooling (for downsampling), and fully connected layers. Neural network 1000 can include any other deep network other than a CNN, such as an autoencoder, a deep belief nets (DBNs), a Recurrent Neural Networks (RNNs), among others.
[0145] FIG. 11 is an illustrative example of a convolutional neural network (CNN) 1100. The input layer 1102 of the CNN 1100 includes data representing an image or frame. For example, the data can include an array of numbers representing the pixels of the image, with each number in the array including a value from 0 to 255 describing the pixel intensity at that position in the array. Using the previous example from above, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luma and two chroma components, or the like). The image can be passed through a convolutional hidden layer 1104, an optional non-linear activation layer, a pooling hidden layer 1106, and fully connected layer 1108 (which fully connected layer 1108 can be hidden) to get an output at the output layer 1110. While only one of each hidden layer is shown in FIG. 11, one of ordinary skill will appreciate that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in the CNN 1100. As previously described, the output can indicate a single class of an object or can include a probability of classes that best describe the object in the image.
[0146] The first layer of the CNN 1100 can be the convolutional hidden layer 1104. The convolutional hidden layer 1104 can analyze image data of the input layer 1102. Each node of the convolutional hidden layer 1104 is connected to a region of nodes (pixels) of the input image called a receptive field. The convolutional hidden layer 1104 can be considered as one or more filters (each filter corresponding to a different activation or feature map), with each convolutional iteration of a filter being a node or neuron of the convolutional hidden layer 1104. For example, the region of the input image that a filter covers at each convolutional iteration would be the receptive field for the filter. In one illustrative example, if the input image includes a 28x28 array, and each filter (and corresponding receptive field) is a 5x5 array, then there will be 24x24 nodes in the convolutional hidden layer 1104. Each connection between a node and a receptive field for that node learns a weight and, in some cases, an overall bias such that each node learns to analyze its particular local receptivePCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO28 field in the input image. Each node of the convolutional hidden layer 1104 will have the same weights and bias (called a shared weight and a shared bias). For example, the filter has an array of weights (numbers) and the same depth as the input. A filter will have a depth of 3 for an image frame example (according to three color components of the input image). An illustrative example size of the filter array is 5 x 5 x 3, corresponding to a size of the receptive field of a node.
[0147] The convolutional nature of the convolutional hidden layer 1104 is due to each node of the convolutional layer being applied to its corresponding receptive field. For example, a filter of the convolutional hidden layer 1104 can begin in the top-left corner of the input image array and can convolve around the input image. As noted above, each convolutional iteration of the filter can be considered a node or neuron of the convolutional hidden layer 1104. At each convolutional iteration, the values of the filter are multiplied with a corresponding number of the original pixel values of the image (e.g., the 5x5 filter array is multiplied by a 5x5 array of input pixel values at the top-left corner of the input image array). The multiplications from each convolutional iteration can be summed together to obtain a total sum for that iteration or node. The process is next continued at a next location in the input image according to the receptive field of a next node in the convolutional hidden layer 1104. For example, a filter can be moved by a step amount (referred to as a stride) to the next receptive field. The stride can be set to 1 or any other suitable amount. For example, if the stride is set to 1, the filter will be moved to the right by 1 pixel at each convolutional iteration. Processing the filter at each unique location of the input volume produces a number representing the filter results for that location, resulting in a total sum value being determined for each node of the convolutional hidden layer 1104.
[0148] The mapping from the input layer to the convolutional hidden layer 1104 is referred to as an activation map (or feature map). The activation map includes a value for each node representing the filter results at each location of the input volume. The activation map can include an array that includes the various total sum values resulting from each iteration of the filter on the input volume. For example, the activation map will include a 24 x 24 array if a 5 x 5 filter is applied to each pixel (a stride of 1) of a 28 x 28 input image. The convolutional hidden layer 1104 can include several activation maps in order to identify multiple features in an image. The example shown in FIG. 11 includes three activation maps. Using three activation maps, the convolutional hidden layer 1104 can detect three different kinds of features, with each feature being detectable across the entire image.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO29
[0149] In some examples, a non-linear hidden layer can be applied after the convolutional hidden layer 1104. The non-linear layer can be used to introduce non-linearity to a system that has been computing linear operations. One illustrative example of a non-linear layer is a rectified linear unit (ReLU) layer. A ReLU layer can apply the function f(x) = max(0, x) to all of the values in the input volume, which changes all the negative activations to 0. The ReLU can thus increase the non-linear properties of the CNN 1100 without affecting the receptive fields of the convolutional hidden layer 1104.
[0150] The pooling hidden layer 1106 can be applied after the convolutional hidden layer 1104 (and after the non-linear hidden layer when used). The pooling hidden layer 1106 is used to simplify the information in the output from the convolutional hidden layer 1104. For example, the pooling hidden layer 1106 can take each activation map output from the convolutional hidden layer 1104 and generates a condensed activation map (or feature map) using a pooling function. Max-pooling is one example of a function performed by a pooling hidden layer. Other forms of pooling functions be used by the pooling hidden layer 1106, such as average pooling, L2-norm pooling, or other suitable pooling functions. A pooling function (e.g., a max-pooling filter, an L2-norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden layer 1104. In the example shown in FIG. 11, three pooling filters are used for the three activation maps in the convolutional hidden layer 1104.
[0151] In some examples, max-pooling can be used by applying a max-pooling filter (e.g., having a size of 2x2) with a stride (e.g., equal to a dimension of the filter, such as a stride of 2) to an activation map output from the convolutional hidden layer 1104. The output from a max-pooling filter includes the maximum number in every sub-region that the filter convolves around. Using a 2x2 filter as an example, each unit in the pooling layer can summarize a region of 2*2 nodes in the previous layer (with each node being a value in the activation map). For example, four values (nodes) in an activation map will be analyzed by a 2x2 max-pooling filter at each iteration of the filter, with the maximum value from the four values being output as the “max” value. If such a max-pooling filter is applied to an activation filter from the convolutional hidden layer 1104 having a dimension of 24x24 nodes, the output from the pooling hidden layer 1106 will be an array of 12x12 nodes.
[0152] In some examples, an L2-norm pooling filter could also be used. The L2-norm pooling filter includes computing the square root of the sum of the squares of the values in the 2x2 region (or otherPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO30 suitable region) of an activation map (instead of computing the maximum values as is done in maxpooling) and using the computed values as an output.
[0153] The pooling function (e.g., max-pooling, L2-norm pooling, or other pooling function) determines whether a given feature is found anywhere in a region of the image and discards the exact positional information. This can be done without affecting results of the feature detection because, once a feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Max-pooling (as well as other pooling methods) offer the benefit that there are many fewer pooled features, thus reducing the number of parameters needed in later layers of the CNN 1100.
[0154] The final layer of connections in the network is a fully-connected layer that connects every node from the pooling hidden layer 1106 to every one of the output nodes in the output layer 1110. Using the example above, the input layer includes 28 x 28 nodes encoding the pixel intensities of the input image, the convolutional hidden layer 1104 includes 3x24x24 hidden feature nodes based on application of a 5x5 local receptive field (for the filters) to three activation maps, and the pooling hidden layer 1106 includes a layer of 3x 12x 12 hidden feature nodes based on application of maxpooling filter to 2x2 regions across each of the three feature maps. Extending this example, the output layer 1110 can include ten output nodes. In such an example, every node of the 3x12x12 pooling hidden layer 1106 is connected to every node of the output layer 1110.
[0155] The fully connected layer 1108 can obtain the output of the previous pooling hidden layer 1106 (which should represent the activation maps of high-level features) and determines the features that most correlate to a particular class. For example, the fully connected layer 1108 can determine the high-level features that most strongly correlate to a particular class and can include weights (nodes) for the high-level features. A product can be computed between the weights of the fully connected layer 1108 and the pooling hidden layer 1106 to obtain probabilities for the different classes. For example, if the CNN 1100 is being used to predict that an object in an image is a person, high values will be present in the activation maps that represent high-level features of people (e.g., two legs are present, a face is present at the top of the object, two eyes are present at the top left and top right of the face, a nose is present in the middle of the face, a mouth is present at the bottom of the face, and / or other features common for a person).PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO31
[0156] In some examples, the output from the output layer 1110 can include an M-dimensional vector (in the prior example, M=10). M indicates the number of classes that the CNN 1100 has to choose from when classifying the object in the image. Other example outputs can also be provided. Each number in the M-dimensional vector can represent the probability the object is of a certain class. In one illustrative example, if a 10-dimensional output vector represents ten different classes of objects is [0 0 0.05 0.8 0 0.15 0 0 0 0], the vector indicates that there is a 5% probability that the image is the third class of object (e.g., a dog), an 80% probability that the image is the fourth class of object (e.g., a human), and a 15% probability that the image is the sixth class of object (e.g., a kangaroo). The probability for a class can be considered a confidence level that the object is part of that class.
[0157] FIG. 12 illustrates an example computing-device architecture 1200 of an example computing device which can implement the various techniques described herein. In some examples, the computing device can include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or computing device of a vehicle), or other device. For example, the computing-device architecture 1200 may include, implement, or be included in any or all of system 200 of FIG. 2A and FIG. 2B, image modifier 206 of FIG. 2A and FIG. 2B, system 300a of FIG. 3A, system 300b of FIG. 3B, system 300c of FIG. 3C, and / or other devices, modules, or systems described herein. Additionally or alternatively, computing-device architecture 1200 may be configured to perform process 600, and / or other process described herein.
[0158] The components of computing-device architecture 1200 are shown in electrical communication with each other using connection 1212, such as a bus. The example computing-device architecture 1200 includes a processing unit (CPU or processor) 1202 and computing device connection 1212 that couples various computing device components including computing device memory 1210, such as read only memory (ROM) 1208 and random-access memory (RAM) 1206, to processor 1202.
[0159] Computing-device architecture 1200 can include a cache of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1202. Computing-device architecture 1200 can copy data from memory 1210 and / or the storage device 1214 to cache 1204 for quick access by processor 1202. In this way, the cache can provide a performance boost that avoids processor 1202 delays while waiting for data. These and other modules can control or be configuredPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO32 to control processor 1202 to perform various actions. Other computing device memory 1210 may be available for use as well. Memory 1210 can include multiple different types of memory with different performance characteristics. Processor 1202 can include any general-purpose processor and a hardware or software service, such as service 1 1216, service 2 1218, and service 3 1220 stored in storage device 1214, configured to control processor 1202 as well as a special-purpose processor where software instructions are incorporated into the processor design. Processor 1202 may be a self- contained system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.
[0160] To enable user interaction with the computing-device architecture 1200, input device 1222 can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech and so forth. Output device 1224 can also be one or more of a number of output mechanisms known to those of skill in the art, such as a display, projector, television, speaker device, etc. In some instances, multimodal computing devices can enable a user to provide multiple types of input to communicate with computing-device architecture 1200. Communication interface 1226 can generally govern and manage the user input and computing device output. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0161] Storage device 1214 is a non-volatile memory and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile discs (DVDs), cartridges, random-access memories (RAMs) 1206, read only memory (ROM) 1208, and hybrids thereof. Storage device 1214 can include services 1216, 1218, and 1220 for controlling processor 1202. Other hardware or software modules are contemplated. Storage device 1214 can be connected to the computing device connection 1212. In one aspect, a hardware module that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1202, connection 1212, output device 1224, and so forth, to carry out the function.
[0162] The term “substantially,” in reference to a given parameter, property, or condition, may refer to a degree that one of ordinary skill in the art would understand that the given parameter, property,PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO33 or condition is met with a small degree of variance, such as, for example, within acceptable manufacturing tolerances. By way of example, depending on the particular parameter, property, or condition that is substantially met, the parameter, property, or condition may be at least 90% met, at least 95% met, or even at least 99% met.
[0163] Aspects of the present disclosure are applicable to any suitable electronic device (such as security systems, smartphones, tablets, laptop computers, vehicles, drones, or other devices) including or coupled to one or more active depth sensing systems. While described below with respect to a device having or coupled to one light projector, aspects of the present disclosure are applicable to devices having any number of light projectors and are therefore not limited to specific devices.
[0164] The term “device” is not limited to one or a specific number of physical objects (such as one smartphone, one controller, one processing system and so on). As used herein, a device may be any electronic device with one or more parts that may implement at least some portions of this disclosure. While the below description and examples use the term “device” to describe various aspects of this disclosure, the term “device” is not limited to a specific configuration, type, or number of objects. Additionally, the term “system” is not limited to multiple components or specific aspects. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. While the below description and examples use the term “system” to describe various aspects of this disclosure, the term “system” is not limited to a specific configuration, type, or number of objects.
[0165] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein. However, it will be understood by one of ordinary skill in the art that the aspects may be practiced without these specific details. For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks including functional blocks including devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO34
[0166] Individual aspects may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.
[0167] Processes and methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general-purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code, etc.
[0168] The term “computer-readable medium” includes, but is not limited to, portable or nonportable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and / or data. A computer-readable medium may include a non- transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, magnetic or optical disks, USB devices provided with non-volatile memory, networked storage devices, any suitable combination thereof, among others. A computer-readable medium may have stored thereon code and / or machineexecutable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via anyPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO35 suitable means including memory sharing, message passing, token passing, network transmission, or the like.
[0169] In some aspects the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non- transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0170] Devices implementing processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor(s) may perform the necessary tasks. Typical examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
[0171] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.
[0172] In the foregoing description, aspects of the application are described with reference to specific aspects thereof, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. ForPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO36 the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects, the methods may be performed in a different order than that described.
[0173] One of ordinary skill will appreciate that the less than (“<”) and greater than (“>”) symbols or terminology used herein can be replaced with less than or equal to (“<”) and greater than or equal to (“>”) symbols, respectively, without departing from the scope of this description.
[0174] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof.
[0175] The phrase “coupled to” refers to any component that is physically connected to another component either directly or indirectly, and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.
[0176] Claim language or other language reciting “at least one of’ a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any duplicate information or data (e.g., A and A, B and B, C and C, A and A and B, and so on), or any other ordering, duplication, or combination of A, B, and C. The language “at least one of’ a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases “at least one” and “one or more” are used interchangeably herein.
[0177] Claim language or other language reciting “at least one processor configured to,” “at least one processor being configured to,” “one or more processors configured to,” “one or more processors being configured to,” or the like indicates that one processor or multiple processors (in any combination) can perform the associated operation(s). For example, claim language reciting “at least one processor configured to: X, Y, and Z” means a single processor can be used to perform operationsPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO37X, Y, and Z; or that multiple processors are each tasked with a certain subset of operations X, Y, and Z such that together the multiple processors perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language reciting “at least one processor configured to: X, Y, and Z” can mean that any single processor may only perform at least a subset of operations X, Y, and Z.
[0178] Where reference is made to one or more elements performing functions (e.g., steps of a method), one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function). Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions.
[0179] Where reference is made to an entity (e.g., any entity or device described herein) performing functions or being configured to perform functions (e.g., steps of a method), the entity may be configured to cause one or more elements (individually or collectively) to perform the functions. The one or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more (or all) of the functions, and / or any combination thereof. Where reference to the entity performing functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to collectively perform the functions. When the entity is configured to cause more than one component to collectively perform the functions, each function need not be performed by each of those components (e.g., different functions may be performed by different components) and / or each function need not be performed in whole by only one component (e.g., different components may perform different sub-functions of a function).
[0180] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have beenPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO38 described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0181] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general-purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer- readable data storage medium including program code including instructions that, when executed, performs one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include memory or data storage media, such as random-access memory (RAM) such as synchronous dynamic random-access memory (SDRAM), read-only memory (ROM), non-volatile random-access memory (NVRAM), electrically erasable programmable readonly memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. The techniques additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.
[0182] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, an application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, suchPCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO39 as, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.
[0183] Illustrative aspects of the disclosure include:
[0184] Aspect 1. An apparatus for modifying image data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: generate a three-dimensional (3D) bounding box based on an input image; generate a conditioning input based on the 3D bounding box, wherein the conditioning input is associated with an orientation of the 3D bounding box; and generate an output image by processing at least a portion of the input image using a diffusion model based on the conditioning input, wherein the output image includes an object oriented based on the orientation of the 3D bounding box.
[0185] Aspect 2. The apparatus of aspect 1, wherein the 3D bounding box comprises: coordinates in a 3D space, wherein the 3D space is based on the input image; size dimensions; and orientation information.
[0186] Aspect 3. The apparatus of any one of aspects 1 or 2, wherein the object is not present in the input image.
[0187] Aspect 4. The apparatus of any one of aspects 1 to 3, wherein the at least one processor is configured to render the 3D bounding box as a two-dimensional (2D) bounding box, wherein the conditioning input is generated based on the 2D bounding box.
[0188] Aspect 5. The apparatus of aspect 4, wherein, to generate the conditioning input, the at least one processor is configured to process the 2D bounding box using a machine-learning model to generate the conditioning input.
[0189] Aspect 6. The apparatus of aspect 5, wherein the conditioning input comprises a featurespace representation of the 3D bounding box.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO40
[0190] Aspect 7. The apparatus of any one of aspects 4 to 6, wherein the 2D bounding box is rendered with coloring that is different for each face of the 3D bounding box to indicate an orientation of the 3D bounding box.
[0191] Aspect 8. The apparatus of any one of aspects 4 to 7, wherein the 2D bounding box is rendered with two respective colors for each face of the 3D bounding box to indicate an orientation of the 3D bounding box.
[0192] Aspect 9. The apparatus of any one of aspects 4 to 8, wherein the 2D bounding box is rendered based on a gradient between two respective colors for each face of the 3D bounding box.
[0193] Aspect 10. The apparatus of any one of aspects 1 to 9, wherein the at least one processor is configured to: detect an object in the input image; and determine an orientation of the object in the input image, wherein the orientation of the 3D bounding box is based on the orientation of the detected object.
[0194] Aspect 11. The apparatus of aspect 10, wherein the detected object in the input image is replaced by the object in the output image.
[0195] Aspect 12. The apparatus of any one of aspects 10 or 11, wherein the at least one processor is configured to: generate a mask indicative of pixels of the input image that represent the detected object; and generate the at least a portion of the input image based on the mask.
[0196] Aspect 13. The apparatus of any one of aspects 10 to 12, wherein the at least one processor is configured to crop the input image to generate the at least a portion of the input image.
[0197] Aspect 14. The apparatus of aspect 13, wherein the output image comprises a first output image, wherein the at least one processor is further configured to blend the output image with the input image to generate a second output image.
[0198] Aspect 15. The apparatus of any one of aspects 1 to 9, wherein the 3D bounding box is generated based on a user input.
[0199] Aspect 16. The apparatus of aspect 15, wherein the at least one processor is configured to: cause a display to display the input image to a user; receive, from a user interface, the user input relative to the input image; and generate the 3D bounding box based on the user input, wherein thePCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO413D bounding box is simulated at a depth in a scene depicted by the input image based on the user input.
[0200] Aspect 17. The apparatus of any one of aspects 15 or 16, wherein the at least one processor is configured to determine a mask based on a model of the object, wherein the input image is processed by the diffusion model based on the mask.
[0201] Aspect 18. The apparatus of any one of aspects 15 to 17, wherein the at least one processor is configured to crop the input image to generate the at least a portion of the input image.
[0202] Aspect 19. The apparatus of aspect 18, wherein the output image comprises a first output image, wherein the at least one processor is configured to blend the output image with the input image to generate a second output image.
[0203] Aspect 20. The apparatus of any one of aspects 1 to 19, wherein the at least one processor is configured to test an automated driving system based on the output image.
[0204] Aspect 21. The apparatus of any one of aspects 1 to 20, wherein the at least one processor is configured to train an automated driving system based on the output image.
[0205] Aspect 22. The apparatus of any one of aspects 1 to 21, wherein the at least one processor is configured to cause a display of an extended reality (XR) device to display the output image.
[0206] Aspect 23. The apparatus of any one of aspects 1 to 22, wherein the at least one processor is configured to test a perception model based on the output image.
[0207] Aspect 24. The apparatus of aspect 23, wherein the object in the output image is in a challenging orientation.
[0208] Aspect 25. A method for modifying image data, the method comprising: generating a three- dimensional (3D) bounding box based on an input image; generating a conditioning input based on the 3D bounding box, wherein the conditioning input is associated with an orientation of the 3D bounding box; and generating an output image by processing at least a portion of the input image using a diffusion model based on the conditioning input, wherein the output image includes an object oriented based on the orientation of the 3D bounding box.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO42
[0209] Aspect 26. The method of aspect 25, wherein the 3D bounding box comprises: coordinates in a 3D space, wherein the 3D space is based on the input image; size dimensions; and orientation information.
[0210] Aspect 27. The method of any one of aspects 25 or 26, wherein the object is not present in the input image.
[0211] Aspect 28. The method of any one of aspects 25 to 27, further comprising rendering the 3D bounding box as a two-dimensional (2D) bounding box, wherein the conditioning input is generated based on the 2D bounding box.
[0212] Aspect 29. The method of aspect 28, wherein generating the conditioning input comprises processing the 2D bounding box using a machine-learning model to generate the conditioning input.
[0213] Aspect 30. The method of aspect 29, wherein the conditioning input comprises a featurespace representation of the 3D bounding box.
[0214] Aspect 31. The method of any one of aspects 28 to 29, wherein the 2D bounding box is rendered with coloring that is different for each face of the 3D bounding box to indicate an orientation of the 3D bounding box.
[0215] Aspect 32. The method of any one of aspects 28 to 31, wherein the 2D bounding box is rendered with two respective colors for each face of the 3D bounding box to indicate an orientation of the 3D bounding box.
[0216] Aspect 33. The method of any one of aspects 28 to 32, wherein the 2D bounding box is rendered based on a gradient between two respective colors for each face of the 3D bounding box.
[0217] Aspect 34. The method of any one of aspects 25 to 33, further comprising: detecting an object in the input image; and determining an orientation of the object in the input image, wherein the orientation of the 3D bounding box is based on the orientation of the detected object.
[0218] Aspect 35. The method of aspect 34, wherein the detected object in the input image is replaced by the object in the output image.
[0219] Aspect 36. The method of any one of aspects 34 or 35, further comprising: generating a mask indicative of pixels of the input image that represent the detected object; and generating the at least a portion of the input image based on the mask.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO43
[0220] Aspect 37. The method of any one of aspects 34 to 36, further comprising cropping the input image to generate the at least a portion of the input image.
[0221] Aspect 38. The method of aspect 37, wherein the output image comprises a first output image, the method further comprising blending the output image with the input image to generate a second output image.
[0222] Aspect 39. The method of any one of aspects 25 to 33, wherein the 3D bounding box is generated based on a user input.
[0223] Aspect 40. The method of aspect 39, further comprising: displaying the input image to a user; receiving the user input relative to the input image; and generating the 3D bounding box based on the user input, wherein the 3D bounding box is simulated at a depth in a scene depicted by the input image based on the user input.
[0224] Aspect 41. The method of any one of aspects 39 or 40, further comprising determining a mask based on a model of the object, wherein the input image is processed by the diffusion model based on the mask.
[0225] Aspect 42. The method of any one of aspects 39 to 41, further comprising cropping the input image to generate the at least a portion of the input image.
[0226] Aspect 43. The method of aspect 42, wherein the output image comprises a first output image, the method further comprising blending the output image with the input image to generate a second output image.
[0227] Aspect 44. The method of any one of aspects 25 to 43, further comprising testing an automated driving system based on the output image.
[0228] Aspect 45. The method of any one of aspects 25 to 44, further comprising training an automated driving system based on the output image.
[0229] Aspect 46. The method of any one of aspects 25 to 45, further comprising displaying the output image at a display of an extended reality (XR) device.
[0230] Aspect 47. The method of any one of aspects 25 to 46, further comprising testing a perception model based on the output image.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO44
[0231] Aspect 48. The method of aspect 47, wherein the object in the output image is in a challenging orientation.
[0232] Aspect 49. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations according to any of aspects 25 to 48.
[0233] Aspect 50. An apparatus for providing virtual content for display, the apparatus comprising one or more means for perform operations according to any of aspects 25 to 48.
Claims
PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO45CLAIMSWHAT IS CLAIMED IS:
1. An apparatus for modifying image data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: generate a three-dimensional (3D) bounding box based on an input image; generate a conditioning input based on the 3D bounding box, wherein the conditioning input is associated with an orientation of the 3D bounding box; and generate an output image by processing at least a portion of the input image using a diffusion model based on the conditioning input, wherein the output image includes an object oriented based on the orientation of the 3D bounding box.
2. The apparatus of claim 1, wherein the 3D bounding box comprises: coordinates in a 3D space, wherein the 3D space is based on the input image; size dimensions; and orientation information.
3. The apparatus of claim 1, wherein the object is not present in the input image.
4. The apparatus of claim 1, wherein the at least one processor is configured to render the 3D bounding box as a two-dimensional (2D) bounding box, wherein the conditioning input is generated based on the 2D bounding box.
5. The apparatus of claim 4, wherein, to generate the conditioning input, the at least one processor is configured to process the 2D bounding box using a machine-learning model to generate the conditioning input.
6. The apparatus of claim 5, wherein the conditioning input comprises a feature-space representation of the 3D bounding box.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO467. The apparatus of claim 4, wherein the 2D bounding box is rendered with coloring that is different for each face of the 3D bounding box to indicate an orientation of the 3D bounding box.
8. The apparatus of claim 4, wherein the 2D bounding box is rendered with two respective colors for each face of the 3D bounding box to indicate an orientation of the 3D bounding box.
9. The apparatus of claim 4, wherein the 2D bounding box is rendered based on a gradient between two respective colors for each face of the 3D bounding box.
10. The apparatus of claim 1, wherein the at least one processor is configured to: detect an object in the input image; and determine an orientation of the object in the input image, wherein the orientation of the 3D bounding box is based on the orientation of the detected object.
11. The apparatus of claim 10, wherein the detected object in the input image is replaced by the object in the output image.
12. The apparatus of claim 10, wherein the at least one processor is configured to: generate a mask indicative of pixels of the input image that represent the detected object; and generate the at least a portion of the input image based on the mask.
13. The apparatus of claim 10, wherein the at least one processor is configured to crop the input image to generate the at least a portion of the input image.
14. The apparatus of claim 13, wherein the output image comprises a first output image, wherein the at least one processor is further configured to blend the output image with the input image to generate a second output image.PCT / US25 / 44334 29 August 2025 (29.08.2025)Qualcomm Ref. No. 2407683 WO4715. The apparatus of claim 1, wherein the 3D bounding box is generated based on a user input.
16. The apparatus of claim 15, wherein the at least one processor is configured to: cause a display to display the input image to a user; receive, from a user interface, the user input relative to the input image; and generate the 3D bounding box based on the user input, wherein the 3D bounding box is simulated at a depth in a scene depicted by the input image based on the user input.
17. The apparatus of claim 15, wherein the at least one processor is configured to determine a mask based on a model of the object, wherein the input image is processed by the diffusion model based on the mask.
18. The apparatus of claim 15, wherein the at least one processor is configured to crop the input image to generate the at least a portion of the input image.
19. The apparatus of claim 18, wherein the output image comprises a first output image, wherein the at least one processor is configured to blend the output image with the input image to generate a second output image.
20. The apparatus of claim 1, wherein the at least one processor is configured to test an automated driving system based on the output image.