Generating image content
By using a generative machine learning model to generate image data outside the field of view based on user input, the limitation of image data caused by the fixed field of view of the camera is solved, and flexible expansion and rotation effects of the image field of view are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-23
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, cameras cannot change the generated image data by altering the field of view of the captured image, resulting in limitations in the image data.
By using a generative machine learning model, new image data is generated based on user-input field-of-view change instructions, including image portions within the field of view and generated pixels outside the field of view, thus enabling image expansion and modification.
It allows users to easily and conveniently change the field of view of an image, generating a wider or rotating field of view effect, enhancing the richness and flexibility of image data.
Smart Images

Figure CN121816558A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure generally relates to image content. For example, aspects of the present disclosure include systems and techniques for generating image content. BACKGROUND
[0002] A camera can focus light from a scene (e.g., using a lens) onto an image sensor and use the image sensor to generate image data based on the sensed light. The camera can focus light from a field of view of the scene onto the image sensor based on the lens of the camera. The image data generated by the image sensor can represent the field of view of the camera and not the entire scene. SUMMARY
[0003] The following presents a simplified summary related to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary is presented in a simplified form to provide a basic understanding of some concepts relating to one or more aspects with regard to the mechanisms disclosed herein.
[0004] Systems and techniques for generating image content are described. According to at least one example, an apparatus for generating image content is provided. The apparatus includes a user interface configured to display an image of a field of view and receive user input indicating a desired change to the image, where the desired change to the image includes a change to the field of view, and at least one processor configured to provide at least a portion of the image and an indication of the desired change as input to a generative machine learning model, obtain an altered image from the generative machine learning model, where the altered image includes at least a portion of the image of the field of view and generated pixels outside of the field of view.
[0005] In another example, a method for generating image content is provided. The method includes displaying an image of a field of view at a user interface, receiving user input indicating a desired change to the image at the user interface, where the desired change to the image includes a change to the field of view, providing at least a portion of the image and an indication of the desired change as input to a generative machine learning model, and obtaining an altered image from the generative machine learning model, where the altered image includes at least a portion of the image of the field of view and generated pixels outside of the field of view.
[0006] In another example, an apparatus for generating image content is provided that includes at least one memory and at least one processor (e.g., configured in circuitry) coupled to the at least one memory. The at least one processor is configured to: display an image of a field of view at a user interface; receive user input indicating a desired change to the image at the user interface, wherein the desired change to the image includes a change to the field of view; provide at least a portion of the image and an indication of the desired change as input to a generative machine learning model; and obtain an altered image from the generative machine learning model, wherein the altered image includes at least a portion of the image of the field of view and generated pixels outside of the field of view.
[0007] In another example, a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: display an image of a field of view at a user interface; receive user input indicating a desired change to the image at the user interface, wherein the desired change to the image includes a change to the field of view; provide at least a portion of the image and an indication of the desired change as input to a generative machine learning model; and obtain an altered image from the generative machine learning model, wherein the altered image includes at least a portion of the image of the field of view and generated pixels outside of the field of view.
[0008] In another example, an apparatus for generating image content is provided. The apparatus includes means for displaying an image of a field of view at a user interface; means for receiving user input indicating a desired change to the image at the user interface, wherein the desired change to the image includes a change to the field of view; means for providing at least a portion of the image and an indication of the desired change as input to a generative machine learning model; and means for obtaining an altered image from the generative machine learning model, wherein the altered image includes at least a portion of the image of the field of view and generated pixels outside of the field of view.
[0009] In some aspects, one or more of the devices described herein are, can be part of, or can include a mobile device (e.g., a mobile telephone or so-called “smart phone,” a tablet computer, or other type of mobile device), an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a vehicle (or a computing device, component, or system of a vehicle), a smart or connected device (e.g., an Internet of Things (IoT) device), a wearable device, a personal computer, a laptop computer, a video server, a television (e.g., a network-connected television), a robotic device or system, or other device. In some aspects, each device can include one image sensor (e.g., one camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each device can include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each device can include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each device can include one or more sensors. In some cases, the one or more sensors can be used to determine a location of the device, a state of the device (e.g., a tracking state, an operational state, a temperature, a humidity level, and / or another state), and / or for other purposes.
[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in determining the scope of the claimed subject matter. The subject matter should be understood from readi ng the entire specification of the patent, including any claims, the
[0011] The foregoing and other features and aspects will become more apparent from reading the following specification in conjunction with the accompanying drawings in which: BRIEF DESCRIPTION OF DRAWINGS
[0012] An illustrative example of the present application is described below in detail with reference to the following figures:
[0013] Figure 1 is a block diagram illustrating an example device for generating image content in accordance with various aspects of the present disclosure;
[0014] Figure 2 includes an example representation of a device displaying an image (e.g., which can be captured by the device) in accordance with various aspects of the present disclosure and an example representation of a device displaying an image (e.g., which includes generated image content);
[0015] Figure 3example representation of a device displaying an image (e.g., which can be captured by the device) in accordance with various aspects of the disclosure and an example representation of a device displaying an image (e.g., which includes generated image content);
[0016] Figure 4 example representation of a device displaying an image (e.g., which can be captured by the device) in accordance with various aspects of the disclosure and an example representation of a device displaying an image (e.g., which includes generated image content);
[0017] Figure 5 example representation of a device displaying an image (e.g., which can be captured by the device) in accordance with various aspects of the disclosure and an example representation of a device displaying an image (e.g., which includes generated image content);
[0018] Figure 6 is a block diagram illustrating another example device for generating image content in accordance with various aspects of the disclosure;
[0019] Figure 7 example representation of a device displaying an image (e.g., which can be captured by a first camera of the device) in accordance with various aspects of the disclosure, an example representation of a device displaying another image (e.g., which can be captured by a second camera of the device), and an example representation of a device displaying yet another image (e.g., which includes generated image content);
[0020] Figure 8 is a block diagram illustrating another example device for generating image content in accordance with various aspects of the disclosure;
[0021] Figure 9 example representation of a device displaying an image (e.g., which can be captured by a camera of the device at a first time), an example representation of a device displaying another image (e.g., which can be captured by the camera of the device at a second time), and an example representation of a device displaying yet another image (e.g., which includes generated image content);
[0022] Figure 10 is a flow diagram illustrating another example process for generating image content in accordance with various aspects of the disclosure;
[0023] Figure 11 is a block diagram illustrating an example architecture of an image processing system in accordance with various aspects of the disclosure;
[0024] Figure 12 is a block diagram illustrating an example implementation of a system configured to perform one or more of the functions described herein, which can include a central processing unit (CPU);
[0025] Figure 13 includes two sets of images showing a forward diffusion process (which is fixed) and a reverse diffusion process (which is learned) of a diffusion model;
[0026] Figure 14 is a diagram illustrating how diffusion data is distributed from an initial data distribution to noise in a forward diffusion direction using a diffusion model in accordance with some aspects of the present disclosure;
[0027] Figure 15 is a diagram illustrating a U-Net architecture for a diffusion model in accordance with some aspects of the present disclosure.
[0028] Figure 16 is a block diagram illustrating an example of a deep learning neural network that can be used to implement a perception module and / or one or more validation modules in accordance with some aspects of the disclosed technology;
[0029] Figure 17 is a block diagram illustrating an example of a convolutional neural network (CNN) in accordance with various aspects of the present disclosure; and
[0030] Figure 18 is a block diagram illustrating an example computing device architecture of an example computing device that can implement various techniques described herein. DETAILED DESCRIPTION
[0031] Certain aspects of the present disclosure are provided below. Some of these aspects can be independently applied, and some of them can be combined, as will be apparent to those skilled in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects of the application. It is apparent, however, that various aspects can be practiced without
[0032] The following description provides examples of aspects only and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the following description of the exemplary aspects will provide those skilled in the art with an enabling description of how the exemplary aspects can be implemented. It is to be understood that various changes can be made in the function and arrangement of elements without departing from the scope of the application as set forth in the appended claims.
[0033] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as preferred or advantageous over other aspects. Likewise, the term “aspects of the disclosure” does not require that all aspects of the disclosure include the discussed feature, advantage or mode of operation.
[0034] As described above, a camera can focus light from a field of view onto an image sensor and generate image data representing the camera's field of view. After sensing the light and generating the image data, it can not be possible to change the image by changing the field of view used to capture the image data.
[0035] A generative machine learning model can be trained to generate image data based on a provided condition. Data generated by a generative machine learning model can be referred to as "hallucinations." An image can be provided to a generative machine learning model as a condition. Image data generated by a generative machine learning model can be new image data (e.g., based on the training of the generative machine learning model). The new image data can be conditioned on the provided image, but not copied from the provided image.
[0036] Directing a generative machine learning model to generate image data can be a cumbersome process. For example, the process of providing an image as a condition to a generative machine learning model, providing parameters to be used by the generative machine learning model, and / or directing the generative machine learning model to generate what can be complex. For example, a generative machine learning model can be provided with many different parameters that can influence the generation of generated image data. Further, a generative machine learning model can generate new image data conditioned on provided image data in a variety of different ways. Directing a generative machine learning model to generate which image data can be cumbersome.
[0037] Described herein are systems, apparatuses, methods (also referred to as procedures), and computer-readable media for generating image data (collectively referred to herein as "systems and techniques"). The systems and techniques described herein can include a user interface (UI) that can allow a user to generate image data using a generative machine learning model in a simple and convenient manner. For example, the systems and techniques can allow a user to quickly and easily alter an image, such as altering a field of view of an image, using a generative machine learning model (e.g., a diffusion neural network model). According to some aspects, the systems and techniques can include a user interface that can allow a user to observe a captured image, direct a desired change to a field of view of the image, and generate an altered image based on the desired change using a generative machine learning model.
[0038] For example, a user can capture an image of a field of view (e.g., using a camera of a device such as a mobile phone). The system and techniques can display the image to the user in a user interface (e.g., in an application on the mobile phone). While displaying the image, the user interface can receive input from the user indicating a desire to change the field of view of the image. For example, the user can perform a pinch-out gesture (e.g., on a touchscreen of a device such as a mobile phone). The pinch-out gesture can be a common gesture to indicate a desire to zoom out from a narrow field of view to a wider field of view. The user interface can interpret the pinch-out gesture as a desire to expand the field of view of the image. As another example, the user can perform a drag gesture (which can be a common gesture to indicate a desire to pan the field of view in a scene). The user interface can interpret the drag gesture as a desire to pan the field of view. As another example, the user can perform a rotate gesture (which can be a common gesture to indicate a desire to rotate the image). The user interface can interpret the rotate gesture as a desire to rotate the field of view within a frame of the image. As another example, the user can physically rotate a device that includes the user interface. For example, the device can include an orientation sensor that can detect rotation of the apparatus. The user interface can interpret the rotation of the apparatus as a desire to rotate a frame of the image (e.g., from a landscape frame to a portrait frame or vice versa).
[0039] The user interface can provide the image to the generative machine learning model along with one or more instructions and / or settings. For example, the user interface can provide the image as a condition to the generative machine learning model. Further, the user interface can translate the interpreted desire of the user into one or more instructions, settings, and / or conditions to provide to the generative machine learning model. For example, based on the interpreted desire to expand the field of view, the user interface can instruct the generative machine learning model to generate pixels outside of the field of view of the image (e.g., on all sides of the field of view of the image). As another example, based on the interpreted desire to pan the field of view, the user interface can instruct the generative machine learning model to generate pixels outside of the field of view of the image (e.g., on one side of the field of view of the image). As another example, based on the interpreted desire to rotate the field of view, the user interface can instruct the generative machine learning model to generate pixels outside of the field of view of the image (e.g., on all sides of the field of view of the image to fill in the frame of the image as the field of view is rotated within the frame). As another example, based on the interpreted desire to rotate the frame, the user interface can instruct the generative machine learning model to generate pixels outside of the field of view of the image (e.g., on both sides of the field of view of the image to fill in the rotated frame of the image).
[0040] Based on the received user input, the generative machine learning model can be used to generate image data outside of the field of view of the image. Generating image data outside of the field of view of the image can be referred to as “extrapolating out.” For example, the image can be provided as a condition to the generative machine learning model, and the generative machine learning model can be instructed to generate, based on the received user input, altered images that include the field of view image and generated pixels outside of the field of view. For example, the image can be a 1000 x 1000 pixel image of the field of view. The generative machine learning model can be instructed to extrapolate the image by generating image data to fill a 2000 x 2000 pixel image with the original image of the field of view in the center and newly generated pixel data outside of the field of view.
[0041] In some cases, such as where a user desires to rotate an image within a frame, the systems and techniques can rotate the image and crop the altered image to fit within the frame. In some cases, the systems and techniques can rotate the image prior to providing the image to the generative machine learning model. In other cases, the systems and techniques can rotate the altered image provided by the generative machine learning model.
[0042] In some cases, the systems and techniques can provide multiple images to the generative machine learning model as a condition for generating image data. For example, the systems and techniques can substantially simultaneously capture multiple images using multiple cameras (e.g., a wide-angle camera and a super-wide-angle camera of a device such as a mobile phone). The systems and techniques can provide the multiple captured images as a condition to the generative machine learning model along with instructions to alter one of the images according to the interpretation.
[0043] As another example, the systems and techniques can obtain multiple images of a scene and provide the multiple images of the scene as a condition to the generative machine learning model. The systems and techniques can determine that the multiple images belong to the same scene based on the time the images were captured (e.g., if the images were captured at approximately the same time, the systems and techniques can assume that the images are captures of the same scene), the location the images were captured (e.g., if the images were captured at approximately the same location, which can be determined based on the location of the device and / or a location service, the systems and techniques can assume that the images are captures of the same scene), and / or by comparing the images. As another example, the systems and techniques can provide instructions to a user to point a camera around a scene (e.g., while the systems and techniques capture multiple images) and / or capture multiple images of a scene so that the systems and techniques can use the multiple images to generate additional pixels.
[0044] After generating the modified image, the system and technology can display the modified image (or the final image rotated and / or cropped according to the interpretation expectations) to the user interface. In some cases, while the generative machine learning model is generating the modified image, the user interface may display an image surrounded by blurred pixels that fill the new frame.
[0045] Various aspects of this application will be described below with reference to the accompanying drawings.
[0046] Figure 1 This is a block diagram illustrating an example apparatus 100 for generating image content according to various aspects of this disclosure. Generally, the camera 112 of apparatus 100 can capture an image 114. In some aspects, apparatus 100 can obtain the image 114 from another source, for example, another computing device can obtain it via a communication interface (…). Figure 1 Image 114 (not illustrated) is sent to device 100. Image 114 may represent the field of view of a scene. User interface (UI 102) of device 100 may display image 114 on display 104 of UI 102. UI 102 may (e.g., at touch sensor 106 and / or using orientation sensor 108) receive user input 110 indicating a desired change 116 to image 114. The desired change 116 to image 114 may be or may include a change to the field of view of image 114. Device 100 may generate image 122 (which may include at least a portion of image 114 changed according to change 116) in response to user input 110 using generative machine learning models (e.g., generative machine learning model 120 and / or generative machine learning model 121). Device 100 may conditionally provide at least a portion of image 114 to generative machine learning model 120 (or generative machine learning model 121). Furthermore, device 100 can provide instructions to generative machine learning model 120 (or generative machine learning model 121) regarding the generation of image 122. These instructions may be based on change 116. Generative machine learning model 120 (or generative machine learning model 121) can generate image 122 based on the change 116 in the field of view, to include at least a portion of image 114 within the field of view, and generated pixels outside the field of view.
[0047] Device 100 can be or include any suitable device, including UI 102 and one or more processors 118. For example, device 100 can be a mobile device (e.g., a mobile phone), a camera, a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device having the resource capabilities to perform the operations described with respect to device 100. In some aspects, device 100 can include camera 112 for capturing image 114. In other aspects, device 100 can not include camera 112, and image 114 can be obtained from a camera external to device 100.
[0048] Camera 112 of device 100 can include an array of light sensors that convert light into image data. Camera 112 can include a lens for focusing light from a field of view of a scene onto the array of light sensors. Camera 112 can generate image 114, which can represent the field of view, e.g., image 114 can include data representing the color and intensity of light received by the light sensors from the field of view. Figure 11 Image processing system 1100 can be an example of camera 112. In some aspects, device 100 can receive image 114 from another source. For example, device 100 can receive image 114 via a communication interface.
[0049] The UI 102 can be or include any suitable components for a user to provide user input 110 to the device 100 and for the device 100 to provide output to the user (e.g., to display images 114 and / or images 122 to the user). The UI 102 can include hardware (e.g., the display 104, the touch sensor 106, and the orientation sensor 108 and / or the camera 109), as well as firmware and / or software to control and connect with the hardware. The UI 102 can use the display 104 to display images to the user, including the images 114 and / or the images 122. The UI 102 can use the touch sensor 106 to receive user input 110. The touch sensor 106 can be a capacitive touch screen. For example, the touch sensor 106 can be integrated with or layered with the display 104 so that a user can provide input related to the images 114. For example, the touch sensor 106 can be integrated with the display 104 so that a user can touch the touch sensor 106 at a location corresponding to a point of the images 114 displayed at the display 104. In this way, the UI 102 can receive input with respect to the images 114. Additionally or alternatively, the UI 102 can use the orientation sensor 108 of the UI 102 to receive user input 110 based on an orientation of the orientation sensor 108 and / or the display 104. The orientation sensor 108 can be or include any suitable components for determining an orientation of the device 100 or the display 104. For example, the orientation sensor 108 can include one or more inertial measurement units or gravity-based switches. Additionally or alternatively, the UI 102 can use the camera 109 (which can include one or more cameras on one or more surfaces of the device 100) to capture images (e.g., of a user). The camera 109 can be or include, for example, an active depth camera, an infrared (IR) camera, a red-green-blue (RGB) camera, a stereo camera, an eye-facing camera, or any combination thereof. The UI 102 can detect gestures in images of a user of the device 100. The UI 102 can interpret the gestures as user input. For example, the UI 102 can interpret gestures to receive user input; additionally or alternatively, the UI 102 can interpret eye movements to receive user input. Thus, the UI 102 can implement hand tracking technology and / or eye tracking technology. Additionally or alternatively, the UI 102 can include a keyboard, a keypad, a touchpad, any other output device, any other input device, or any combination thereof.
[0050] UI 102 can receive user input 110 and interpret the user input 110 relative to image 114. More specifically, UI 102 can interpret user input 110 as an instruction for a desired change 116 in the field of view of image 114. For example, a user can perform a pinch touch gesture on touch sensor 106 (e.g., touching touch sensor 106 at two separate points and bringing the touched points together). UI 102 can interpret the pinch touch gesture as a desire to expand the field of view of image 114. As another example, a user can perform a drag touch gesture on touch sensor 106 (e.g., touching one or more points on touch sensor 106 and moving the touched points one or more along a direction (such as along a substantially straight line). UI 102 can interpret the drag touch gesture as a desire to translate the field of view of image 114. As another example, a user can perform a rotational touch gesture on touch sensor 106 (e.g., touching one or more points on touch sensor 106 and moving the touched points in an arc). UI 102 can interpret the rotational touch gesture as an expectation to rotate the field of view within a frame of image 114. As another example, a user can rotate device 100. UI 102 can interpret the rotation of device 100 as an expectation to rotate a frame of image 114 (e.g., from a horizontal frame to a vertical frame or vice versa). As another example, a user can wave both hands to one side to indicate an expected translation. Camera 109 can capture an image of the user waving both hands, and UI 102 can track the hands in the image and interpret the waving hands as an expected translation. As another example, device 100 can be a head-mounted device or can be included in a head-mounted device. Camera 109 can be facing the user's eyes. The user can hold their gaze to one side of the image. UI 102 can interpret the gaze as an expectation to translate the image to the other side. UI 102 can generate change 116, which can be data indicating the desired change. For example, change 116 can encode the desired change as an instruction relative to image 114.
[0051] Processor 118 may be or may include one or more suitable processors configured to perform computational operations. Figure 12 System 1200 and / or Figure 18 The computing device architecture 1800 may be an example of processor 118. In addition, processor 118 may implement generative machine learning model 120. Additionally or alternatively, processor 118 may be coupled with another computing device (e.g., a server computer or laptop computer) that can implement generative machine learning model 121. Figure 1(Not illustrated) Communication. Generative machine learning model 121 may be the same as, substantially the same as, perform the same or substantially the same operations as generative machine learning model 120. All descriptions of the operations performed by generative machine learning model 120 and / or the training of generative machine learning model 120 also apply to generative machine learning model 121. All operations described as being performed by generative machine learning model 120 may be additionally or alternatively performed by generative machine learning model 121. Generative machine learning model 120 (and / or generative machine learning model 121) may be a trained generative machine learning model capable of generating new image data based on provided conditions and / or instructions. (Regarding...) Figure 12 , Figure 13 , Figure 14 and Figure 15 Additional details are provided regarding generative machine learning models and their training. (Based on information about...) Figure 12 , Figure 13 , Figure 14 and Figure 15 The provided description allows for the training of generative machine learning model 120 (and / or generative machine learning model 121).
[0052] Apparatus 100 may provide at least a portion of image 114 to generative machine learning model 120 (and / or generative machine learning model 121) as input or as a condition. For example, in some cases, such as when the desired change includes translation, apparatus 100 may select a portion of image 114 (e.g., a portion that will be included in image 122 after translation) and provide that portion of image 114 to generative machine learning model 120 (or generative machine learning model 121) as input or as a condition. Generative machine learning model 120 (and / or generative machine learning model 121) may generate image 122 based on at least a portion of image 114 (e.g., using at least a portion of image 114 as a condition). Further, apparatus 100 may instruct generative machine learning model 120 (and / or generative machine learning model 121) relative to the generated image 122 based on change 116. For example, processor 118 may determine operating parameters or operating instructions for generative machine learning model 120 (and / or generative machine learning model 121) based on change 116. Image 122 generated based on image 114 may include at least a portion of image 114 within the field of view of image 114, as well as generated pixels outside the field of view. For example, image 122 may include a portion or all of image 114 based on desired changes.
[0053] Additionally or alternatively, device 100 (e.g., using processor 118) can smooth the edges between the generated pixels and the pixels from image 114. For example, generative machine learning models 120 and / or 121 can be trained with images having a larger field of view (FOV) than a reference. For example, the input image can be or may include a cropped smaller centered FOV image, or several cropped smaller images that are centered (e.g., images captured simultaneously from multiple cameras) or dissimilarly centered (images captured sequentially from a single camera). As another example, device 100 can identify edges (e.g., based on the size of image 114), or be informed of the edges between the generated pixels and the pixels from image 114 (e.g., by generative machine learning model 120 or generative machine learning model 121), and smooth the edges using fusion, blending, and / or filtering techniques. For example, device 100 can generate new values for pixels (e.g., red, green, and / or blue values) based on the values of other pixels (e.g., surrounding pixels). For example, device 100 can determine the new value of a pixel based on the values of neighboring pixels, some of which are part of the original image 114, while others are part of the generated image data. As another example, the system and techniques can use filters (e.g., 3×3 pixel filters or 5×5 pixel filters) to, for example, average pixel values across edges.
[0054] In some respects, processor 118 may enable UI 102 to display image 122 at display 104. Additionally or alternatively, processor 118 may analyze image 122 or cause image 122 to be analyzed by one or more other processors. Additionally or alternatively, processor 118 may cause image 122 to be stored (e.g., in memory) or transmitted for later display and / or analysis at different locations.
[0055] By interpreting user input 110 into a change 116 related to image 114 and instructing generative machine learning model 120 (and / or generative machine learning model 121) based on change 116, device 100 can provide a simple and convenient way for the user to instruct generative machine learning model 120 (and / or generative machine learning model 121) to generate image content relative to image 122. For example, based on a pinch gesture (e.g., a touch pinch gesture, a pinch gesture made with both hands, or a pinch gesture based on the user crossing their eyes), device 100 can instruct generative machine learning model 120 (and / or generative machine learning model 121) to generate image content around image 114 in image 122 (e.g., as if image 122 were captured with a wider field of view). Figure 2Examples of generating image content based on an expanded field of view are provided. As another example, based on a drag gesture (e.g., a drag touch gesture, a drag gesture made with one or more of the user's hands, or a drag gesture based on the user's gaze at one side of the image), device 100 can instruct generative machine learning model 120 (and / or generative machine learning model 121) to generate image content on one side of image 114 in image 122 (e.g., as if image 122 were image 114 captured by a translational field of view) by changing 116. Figure 3 Examples of generating image content based on translational field of view are provided. As another example, based on rotational gestures (e.g., based on a rotational pinch gesture, a rotational gesture made by hand, or an eye-rolling gesture), device 100 can instruct generative machine learning model 120 (and / or generative machine learning model 121) to generate image content to fill the corners of image 122 around a rotated version of image 114 in image 122 (e.g., as if image 122 were captured from image 114 by a rotating camera). Figure 4 Examples of generating image content based on a rotated field of view are provided. As another example, based on the rotation of device 100, device 100 can instruct generative machine learning model 120 (and / or generative machine learning model 121) to generate content to the right and left of image 114, or above and below image 114, by changing 116 (e.g., as if image 122 were captured from image 114 by a camera rotated 90 degrees). Figure 5 The example provided is based on rotating the field of view by 90 degrees to generate image content.
[0056] Figure 2 Example representation 202 of apparatus 204 includes displaying image 208 (e.g., which may be captured by apparatus 204) according to various aspects of this disclosure, and example representation 214 of apparatus 204 includes displaying image 216 (e.g., including generated image content, such as generated pixels 218). Generally, apparatus 204 can implement systems and techniques by displaying image 208, receiving user input (e.g., pinch gesture 212), and generating image 216 (including generated pixels 218) based on image 208 and user input (e.g., pinch gesture 212).
[0057] For example, 202 indicates that it includes device 204 (which may be...) Figure 1 Example of device 100). Device 204 implements UI 206 (which can be provided by...). Figure 1 UI 102 in Figure 1 (Implemented on display 104). UI 206 displays image 208. Image 208 can be displayed by a camera of device 204 (e.g., ...). Figure 1 The camera 112) captures, for example, image 208 can be...Figure 1 An example of image 114. Image 208 may belong to field of view 210. In other words, image 208 may represent field of view 210 of the scene. For illustrative purposes, a dashed box within UI 206 (e.g., smaller than image 208) is used to illustrate field of view 210. In practice, field of view 210 may have the same extent as field of view 210 and / or may fill UI 206 because image 208 belongs to field of view 210. Other fields of view (e.g., fields of view 310, 410, 510, 710, 718, 910, and 918) are similarly illustrated with respect to UI 206 and their corresponding images (e.g., images 308, 408, 508, 708, 716, 908, and 916).
[0058] UI 206 can receive user input (e.g., pinch gesture 212). For example, UI 206 may include a touch sensor (e.g., Figure 1 (Touch sensor 106). The touch sensor can sense touch, and the UI 206 can interpret the touch as input. Furthermore, the UI 206 can interpret the touch as input relative to the image 208. For example, the UI 206 can interpret the pinch gesture 212 relative to the center of the image 208 and / or the degree of the pinch gesture 212 relative to the image 208. The UI 206 can determine changes based on the pinch gesture 212 (e.g., Figure 1 (Change 116). This change can be a change to the field of view 210. For example, UI 206 can interpret the pinch gesture 212 as an expectation of changing the field of view 210, for example, by expanding (e.g., by shrinking). UI 206 can determine the center of the pinch gesture 212 as the new center of the new image, and UI 206 can determine the extent of the pinch gesture 212 to determine the extent of the new field of view of the new image.
[0059] Device 204 may include one or more processors (e.g., Figure 1 The processor 118), and can implement generative machine learning models (e.g., Figure 1 Generative machine learning model 120). Additionally or alternatively, device 204 can be coupled with a generative machine learning model (e.g., Figure 1The device 204 communicates with another computing device (the generative machine learning model 121). The device 204 may provide an image 208 (e.g., as a condition) and instructions based on changes to the generative machine learning model. The generative machine learning model may generate an image 216 based on the image 208 and the instructions. The UI 206 may display the image 216 at the UI 206, as illustrated in representation 214. The image 216 may include the image 208 of the field of view 210 and pixels 218 generated outside the field of view 210. For example, based on a desired pinch gesture 212 instructing the expansion of the field of view 210, the image 216 may include the image 208 of the field of view 210 and pixels 218 generated on one or more sides of the field of view 210. Additionally or alternatively, the device 204 may smooth the edges between the generated pixels 218 and pixels from the image 208. For example, the generative machine learning model may be trained with an image having a larger field of view (FOV) than a reference. For example, the input image may be or may include a cropped smaller centered FOV image, or several cropped smaller images that are centered (e.g., images captured simultaneously from multiple cameras) or not centered (images captured sequentially from a single camera). Image 216 may appear as if it were image 208, as if image 208 were captured in a larger field of view.
[0060] Figure 3 Example representation 302 of apparatus 204 includes displaying image 308 (e.g., which may be captured by apparatus 204) according to various aspects of this disclosure, and example representation 314 of apparatus 204 includes displaying image 316 (e.g., including generated image content, such as generated pixels 318). Generally, apparatus 204 can implement systems and techniques by displaying image 308, receiving user input (e.g., drag gesture 312), and generating image 316 (including generated pixels 318) based on image 308 and user input (e.g., drag gesture 312).
[0061] For example, representation 302 includes a device 204 implementing UI 206, which displays image 308. Image 308 may be generated by a camera of device 204 (e.g., ...). Figure 1 The image 308 can be captured by camera 112. Figure 1 An example of image 114. Image 308 can be a field of view 310; for example, image 308 can represent the field of view 310 of a scene.
[0062] UI 206 can receive user input (e.g., drag gesture 312) and interpret the user input in relation to image 308. For example, UI 206 can determine the start point and end point of drag gesture 312 relative to image 308. Additionally or alternatively, UI 206 can determine the length and direction of drag gesture 312. UI 206 can determine changes (e.g., ...) based on drag gesture 312. Figure 1 (Change 116). This change can be a change to the field of view 310. For example, UI 206 can interpret drag gesture 312 as an expectation of changing the field of view 310, for example, by translating the field of view 310. UI 206 can determine the direction and length of the expected change to the field of view 310.
[0063] Device 204 can be based on a generative machine learning model of device 204 (e.g., Figure 1 The device 204 may provide at least a portion of the image 308 (e.g., as a condition) and instructions by modifying the generative machine learning model 120 and / or generative machine learning model 121. For example, based on a drag gesture 312, the device 204 may determine a portion of the image 308, such as a portion of the image 308 that can be retained in the frame after translation, and provide that portion to the generative machine learning model. Alternatively, the device 204 may provide the entire image 308 to the generative machine learning model. The generative machine learning model may generate an image 316 based on at least a portion of the image 308 and instructions, and the UI 206 may display the image 316 at the UI 206, as illustrated in representation 314. The image 316 may include at least a portion of the image 308 in the field of view 310 and generated pixels 318 outside the field of view 310. For example, based on a desired drag gesture 312 instructing a translation of the field of view 310, image 316 may include at least a portion of image 308 of the field of view 310 and generated pixels 318 on one or more sides of the field of view 310. Additionally or alternatively, device 204 may smooth the edges between the generated pixels 318 and the pixels from image 308. Image 316 may appear to resemble image 308 as if image 308 were captured within the translated field of view. Where image 316 includes a portion of image 308, the portion of image 308 included in image 316 may be identical to the portion of image 308 provided by device 204. Alternatively, the portion of image 308 included in image 316 may be a subset of the images provided by device 204. For example, device 204 may provide all (or most) of image 308, and image 316 may include a subset of the provided images.
[0064] Figure 4Example representation 402 of apparatus 204 includes displaying image 408 (e.g., which may be captured by apparatus 204) according to various aspects of this disclosure, and example representation 414 of apparatus 204 includes displaying image 416 (e.g., including generated image content, such as generated pixels 418). Generally, apparatus 204 can implement systems and techniques by displaying image 408, receiving user input (e.g., rotation gesture 412), and generating image 416 (including generated pixels 418) based on image 408 and user input (e.g., rotation gesture 412).
[0065] For example, representation 402 includes a device 204 implementing UI 206, which displays image 408. Image 408 may be generated by a camera of device 204 (e.g., ...). Figure 1 The image 408 can be captured by camera 112. Figure 1 An example of image 114. Image 408 can be a field of view 410; for example, image 408 can represent the field of view 410 of a scene.
[0066] UI 206 can receive user input (e.g., rotation gesture 412) and interpret the user input in relation to image 408. For example, UI 206 can interpret the start point of rotation gesture 412 relative to image 408 and the end point of rotation gesture 412 relative to the image. Additionally or alternatively, UI 206 can determine the rotation length and / or rotation direction of rotation gesture 412. UI 206 can determine changes (e.g., based on rotation gesture 412) Figure 1 (Change 116). This change can be a change to the field of view 410. For example, UI 206 can interpret the rotation gesture 412 as an expectation of changing the field of view 410, for example, by rotating the field of view 410. UI 206 can determine the rotation length and / or direction of the expected change to the field of view 410.
[0067] Device 204 can be based on a generative machine learning model of device 204 (e.g., Figure 1The device 204 may provide at least a portion of the image 408 (e.g., as a condition) and instructions by modifying generative machine learning models 120 and / or 121. For example, based on drag gesture 312, the device 204 may determine a portion of the image 408, such as excluding corners that can be cut off from the frame by rotating the image 408, and provide that portion to the generative machine learning model. Alternatively, the device 204 may provide the entire image 408 to the generative machine learning model. The generative machine learning model may generate an image 416 based on at least a portion of the image 408 and instructions, and the UI 206 may display the image 416 at the UI 206, as illustrated in representation 414. The image 416 may include at least a portion of the image 408 in the field of view 410 and generated pixels 418 outside the field of view 410. For example, based on a desired rotation gesture 412 instructing a rotational field of view 410, image 416 may include at least a portion of image 408 of field of view 410 and generated pixels 418 at one or more corners of image 416. Additionally or alternatively, device 204 may smooth the edges between generated pixels 418 and pixels from image 408. Image 416 may appear to resemble image 408 as if image 408 were captured within a rotating field of view. Where image 416 includes a portion of image 408, the portion of image 408 included in image 416 may be identical to the portion of image 408 provided by device 204. Alternatively, the portion of image 408 included in image 416 may be a subset of the images provided by device 204. For example, device 204 may provide all (or most) of image 408, and image 416 may include a subset of the provided images.
[0068] In some aspects, device 204 can provide image 408 (unrotated) to a generative machine learning model, and the generative machine learning model can generate an intermediate image that includes generated pixels outside of image 408 within a field of view 410. The intermediate image may not be rotated relative to image 408. For example, the intermediate image may include image 408 with an unrotated field of view 410 (e.g., at the center of the intermediate image). Device 204 can rotate and crop the intermediate image to generate image 416. In other aspects, device 204 can rotate image 408 and provide a rotated version of image 408 to the generative machine learning model. The generative machine learning model can generate image 416 based on the rotated version of image 408.
[0069] Figure 5Example representation 502 of apparatus 204 includes displaying an image 508 (e.g., which may be captured by apparatus 204) according to various aspects of this disclosure, and example representation 514 of apparatus 204 includes displaying an image 516 (e.g., including generated image content, such as generated pixels 518). Generally, apparatus 204 can implement systems and techniques by displaying image 508, receiving user input (e.g., apparatus 204 rotating from representation 502 to representation 514), and generating image 516 (which includes generated pixels 518) based on image 508 and user input (e.g., apparatus 204 rotating to representation 514).
[0070] For example, 502 may include a device 204 implementing UI 206, which displays image 508. Image 508 may be generated by a camera of device 204 (e.g., ...). Figure 1 The image 508 can be captured by camera 112. Figure 1 An example of image 114. Image 508 can be a field of view 510; for example, image 508 can represent the field of view 510 of a scene.
[0071] UI 206 can receive user input (e.g., device 204 rotates to display 514) and interpret user input related to image 508. UI 206 can determine changes based on the rotation of device 204 (e.g., Figure 1 (Change 116). This change can be a change to the field of view 510. For example, UI 206 can interpret the rotation of device 204 as an expectation of changing the field of view 510, for example, by rotating the field of view 510 by 90 degrees. For example, UI 206 can determine the rotation of device 204 as an expectation of rotating image 508 from a portrait format to a landscape format.
[0072] Device 204 can be based on a generative machine learning model of device 204 (e.g., Figure 1The generative machine learning model 120 and / or generative machine learning model 121 may be modified to provide at least a portion of image 508 (e.g., as a condition) and instructions. For example, based on the rotation of 204, the device 204 may determine a portion of image 508, such as excluding the top portion that can be cut off from the frame by the rotation of image 508, and provide that portion to the generative machine learning model. Alternatively, the device 204 may provide the entire image 508 to the generative machine learning model. The generative machine learning model may generate image 516 based on at least a portion of image 508 and instructions, and UI 206 may display image 516 at UI 206, as illustrated in representation 514. Image 516 may include at least a portion of image 508 in the field of view 510 and generated pixels 518 outside the field of view 510. For example, based on the rotation of the desired device 204 instructing the rotation of image 508, image 516 may include at least a portion of image 508 in field of view 510 and generated pixels 518 on one or more sides of at least a portion of image 508. Additionally or alternatively, device 204 may smooth the edges between the generated pixels 518 and the pixels from image 508. Image 516 may appear to be image 508 as if image 508 were captured from a camera rotated 90 degrees. When device 204 rotates from a vertical orientation to a horizontal orientation, a generative machine learning model may generate generated pixels 518 on both sides of at least a portion of image 508 in field of view 510. When device 204 rotates from a horizontal orientation to a vertical orientation, a generative machine learning model may generate generated pixels 518 above and below at least a portion of image 508 in field of view 510. When image 516 includes a portion of image 508, the portion of image 508 included in image 506 may be the same as the portion of image 508 provided by device 204. Alternatively, a portion of image 508 included in image 516 may be a subset of images provided by device 204. For example, device 204 may provide all (or most) of image 508, and image 516 may include a subset of the provided images.
[0073] In some respects, UI 206 can receive user input related to an image including generated content. In such cases, device 204 can generate additional content based on the user input and the image including the already generated content. For example, in generating... Figure 2 Following image 216 (e.g., by expanding the field of view of image 208), UI 206 can receive a drag gesture associated with image 216. Device 204 can generate another image based on the drag gesture and image 216. As another example, after generating... Figure 3Following image 316 (e.g., by translating the field of view of image 308), UI 206 can receive a rotation gesture associated with image 316. Device 204 can generate another image based on the rotation gesture and image 316.
[0074] In some aspects, UI 206 can receive multiple desired user inputs. For example, user input may include pinch gestures, drag gestures, and / or rotation gestures. UI 206 can interpret the user input and generate instructions based on the interpreted desired changes. In some cases, UI 206 can generate a single instruction based on all desired changes and instruct the generative machine learning model to generate an image once. In other cases, UI 206 can generate instructions based on each desired change and instruct the generative machine learning model to continuously generate multiple images to achieve the desired change.
[0075] Figure 6 This is a block diagram illustrating an example apparatus 600 for generating image content according to various aspects of this disclosure. Generally, camera 112 of apparatus 600 can capture image 114, and camera 602 can capture image 604. Alternatively, image 114 and / or image 604 can be obtained from another source (e.g., via a communication interface). Image 114 can represent a first field of view of a scene, and image 604 can represent a second field of view of a scene. Apparatus 600 can use a generative machine learning model to generate image 608 based on image 114 and image 604. In some aspects, processor 118 can use generative machine learning model 606 to generate image 608. In other aspects, apparatus 600 can use generative machine learning model 607 (which can be implemented on another computing device, such as a laptop computer, mobile phone, or server) to generate image 608. The apparatus 600 may provide at least a portion of image 114 and at least a portion of image 604 to generative machine learning model 606 (and / or generative machine learning model 607) as input or condition, and the generative machine learning model 606 (and / or generative machine learning model 607) may generate image 608 based on at least a portion of image 114 and at least a portion of image 604.
[0076] Additionally or alternatively, the UI 102 of device 100 may display image 114 (and / or image 604) at display 104 of UI 102. UI 102 may receive user input 110 (e.g., at touch sensor 106 and / or using orientation sensor 108) indicating a desired change 116 to image 114 (and / or image 604). The desired change 116 to image 114 (and / or image 604) may be or may include a change to a first field of view of image 114 and / or a change to a second field of view of image 604. Further, processor 118 may provide instructions to generative machine learning model 606 and / or generative machine learning model 607 regarding the generation of image 608. The instructions may be based on change 116. Generative machine learning model 606 and / or generative machine learning model 607 can generate image 608 based on change 116 to include image 114 of a first field of view and / or image 604 of a second field of view, as well as generated pixels outside the first and second fields of view.
[0077] Device 600 can be basically similar to Figure 1 The device 100 may perform many of the same operations as the device 100. For example, the UI 102 in the device 600 may be the same as that in the device 100, and the UI 102 may operate in the device 600 in the same way as the UI 102 in the device 100. In addition, in the device 600, the UI 102 may display an image 604 at the display 104 and receive user input 110 associated with the image 604. The UI 102 may generate a change 116 indicating a change in the field of view of the image 114 (or image 604).
[0078] Furthermore, camera 112 may be the same as, substantially similar to, or perform the same or substantially the same operation as camera 112 of device 100. Image 114 in device 100 may be the same as or substantially similar to image 114 in device 600. Camera 602 may be another instance of camera 112. For example, camera 602 may be... Figure 11 An example of an image processing system 1100. In some aspects, camera 602 may have a different focal length than camera 112. For example, camera 112 may include a wide-angle lens, and camera 602 may include an ultra-wide-angle lens. Additionally or alternatively, camera 112 may be located on device 600 at a different position than camera 602. Camera 112 may capture an image 114 with a field of view different from that of the image 604 captured by camera 602.
[0079] The processor 118 of device 600 may be the same as, substantially similar to, and / or perform the same or substantially the same operations as the processor 118 of device 100. However, the generative machine learning model 606 of device 600 may differ from the generative machine learning model 120 of device 100. For example, while generative machine learning model 120 may be trained to generate images that include content generated based on input images, generative machine learning model 606 may be trained to generate images that include content generated based on two or more input images. Similarly, generative machine learning model 607 may be trained to generate images that include content generated based on two or more input images. Regarding... Figure 12 , Figure 13 , Figure 14 and Figure 15 Additional details are provided regarding generative machine learning models and their training. (Based on information about...) Figure 12 , Figure 13 , Figure 14 and Figure 15 The provided description allows for the training of generative machine learning model 606 and / or generative machine learning model 607.
[0080] Apparatus 600 may provide at least a portion of image 114 and at least a portion of image 604 to generative machine learning model 606 (and / or generative machine learning model 607) as input or as conditions. Generative machine learning model 606 (and / or generative machine learning model 607) may generate image 608 based on at least a portion of image 114 and at least a portion of image 604 (e.g., using at least a portion of image 114 and at least a portion of image 604 as conditions). Further, apparatus 600 may instruct generative machine learning model 606 (and / or generative machine learning model 607) relative to the generated image 608 based on change 116. For example, processor 118 may determine operating parameters or operating instructions for generative machine learning model 606 (and / or generative machine learning model 607) based on change 116. Image 608 generated based on image 114 and image 604 may include at least a portion of image 114 within the field of view of image 114 and at least a portion of image 604 within the field of view of image 604, as well as generated pixels outside the field of view. For example, image 608 may include part or all of image 114 and / or part or all of image 604 based on desired changes. Additionally or alternatively, device 600 may smooth the edges between generated pixels and pixels from image 604.
[0081] In some respects, processor 118 may enable UI 102 to display image 608 at display 104. Additionally or alternatively, processor 118 may analyze image 608 or cause image 608 to be analyzed by one or more other processors. Additionally or alternatively, processor 118 may cause image 608 to be stored (e.g., in memory) or transmitted for later display and / or analysis at different locations.
[0082] By interpreting user input 110 as a change 116 related to image 114 and / or image 604 and instructing generative machine learning model 606 (and / or generative machine learning model 607) based on change 116, device 600 can provide a simple and convenient way for the user to instruct generative machine learning model 606 (and / or generative machine learning model 607) to generate image content relative to image 608. For example, based on a pinch gesture, device 600 can instruct generative machine learning model 606 (and / or generative machine learning model 607) to generate image content around image 114 and / or image 604 in image 608 (e.g., as if image 608 were captured with a wider field of view of image 114 or image 604). Figure 2 The example provided is based on expanding the field of view to generate image content. Figure 2 The example can be implemented by device 600 by adding a second input image. As another example, based on a drag gesture, device 600 can instruct generative machine learning model 606 (and / or generative machine learning model 607) to generate image content on the side of image 604 in image 114 and / or image 608 by changing 116 (e.g., as if image 608 were captured from image 114 and / or image 604 from a translated field of view). Figure 3 The example provided is an example of generating image content based on a translation field of view. Figure 3 The example can be implemented by device 600 by adding a second input image. As another example, based on a rotation gesture, device 600 can instruct generative machine learning model 606 (and / or generative machine learning model 607) to generate image content by changing 116 to fill the corners of image 608 around a rotated version of image 604 in image 114 and / or image 608 (e.g., as if image 608 were captured from image 114 and / or image 604 by a rotating camera). Figure 4 The example provided is an example of generating image content based on a rotated field of view. Figure 4The example can be implemented by device 600 by adding a second input image. As another example, based on the rotation of device 600, device 600 can instruct generative machine learning model 606 (and / or generative machine learning model 607) to generate content to the right and left of image 114 and / or image 604, or above and below image 114 and / or image 604, by changing 116 (e.g., as if image 608 were captured from image 114 and / or image 604 by a camera rotated 90 degrees). Figure 5 The example provided is based on rotating the field of view by 90 degrees to generate image content. Figure 5 An example can be implemented by device 600 by adding a second input image.
[0083] Figure 7 Example representation 702 of apparatus 204 includes displaying image 708 (e.g., which may be captured by a first camera of apparatus 204) according to various aspects of this disclosure, example representation 714 of apparatus 204 includes displaying image 716 (e.g., which may be captured by a second camera of apparatus 204), and example representation 720 of apparatus 204 includes displaying image 722 (e.g., including generated image content, such as generated pixel 724 and / or modified pixel 726). Generally, apparatus 204 can implement systems and techniques by displaying image 708 and / or image 716 (e.g., at separate times, in partitioned screens, or partially overlapping), receiving user input (e.g., pinch gestures, drag gestures, rotation gestures, or rotation of apparatus 204), and generating image 722 (including generated pixel 724 and / or modified pixel 726) based on image 708, image 716, and user input.
[0084] For example, 702 includes a device 204 implementing UI 206, which displays image 708. Image 708 may be generated by a first camera of device 204 (e.g., ...). Figure 6 The image 708 can be captured by camera 112. Figure 6 An example of image 114. Image 708 can be a first field of view 710, for example, image 708 can represent the field of view 710 of a scene.
[0085] Furthermore, 714 includes a device 204 implementing UI 206, which displays image 716. Image 716 may be generated by a second camera of device 204 (e.g., Figure 6 The image 716 can be captured by camera 602. Figure 6An example of image 604. Image 716 may belong to a second field of view 718; for example, image 716 may represent the field of view 718 of a scene (e.g., the same scene represented by image 708). The focal length of the camera used to capture image 708 may be different from the focal length of the camera used to capture image 716. For example, image 716 may represent a wider field of view of a scene than the same scene represented by image 716.
[0086] UI 206 can receive user input (e.g., pinch gesture, drag gesture, rotate gesture, or rotation from device 204) and interpret the user input in relation to image 708 and / or image 716. The change can be a change to the field of view 710 and / or field of view 718. For example, UI 206 can interpret the user input as an expectation to change the field of view 710 and / or field of view 718, such as by widening, panning, and / or rotating the field of view 710 and / or field of view 718.
[0087] Device 204 can be based on a generative machine learning model of device 204 (e.g., Figure 6The device 204 provides at least a portion of image 708 and at least a portion of image 716 (e.g., as conditions) and instructions by modifying generative machine learning models 606 and / or 607. For example, based on a desired change in the field of view, the device 204 can determine a portion of image 708 and / or a portion of image 716, for example, excluding a portion of image 716 that can be cut off from the frame by translation, and provide that portion to the generative machine learning model. Alternatively, the device 204 can provide the entire image 708 and / or the entire image 716 to the generative machine learning model. The generative machine learning model can generate image 722 based on at least a portion of image 708, at least a portion of image 716, and instructions, and UI 206 can display image 722 at UI 206, as illustrated in representation 720. Image 722 may include at least a portion of image 708 in a first field of view 710, at least a portion of image 716 in a second field of view 718, generated pixels 724 (which may be, for example, outside of fields of view 710 and 718), and / or modified pixels 726. For example, image 722 may include image 708 in field of view 710, image 716 in field of view 718, and generated pixels 724 on one or more sides of fields of view 710 and / or 718. Additionally or alternatively, image 722 may include modified pixels 726, which may be outside of field of view 710, inside field of view 718, and based on image 716. For example, a generative machine learning model may generate modified pixels 726 based on image 716. Modified pixels 726 may be the same as or different from pixels in image 716. For example, modified pixels 726 may have a different resolution than image 716. Alternatively, modified pixels 726 may be inside the field of view 710 of image 708. In such cases, the modified pixel 726 may be based on image 708 and / or image 716. The modified pixel 726 may be the same as or different from the pixels of image 708. For example, the modified pixel 726 may have a different resolution than image 708. Additionally or alternatively, device 204 may smooth the edges between the generated pixel 724 and the pixels from image 708 and / or image 716. Image 722 may appear to be image 708 of the scene as if image 708 were captured in a different field of view. Additionally or alternatively, image 722 may appear to be image 716 of the scene as if image 716 were captured in a different field of view. Where image 722 includes a portion of image 708 and / or a portion of image 716, the portions of image 708 and / or image 716 included in image 722 may be identical to the portions of image 708 and / or image 716 provided by device 204. Alternatively, portions of image 708 and / or image 716 included in image 722 may be subsets of the images provided by device 204.For example, device 204 may provide all (or most) of image 708, and image 722 may include a subset of the provided images.
[0088] Figure 8 This is a block diagram illustrating an example apparatus 800 for generating image content according to various aspects of this disclosure. Generally, a camera 112 of apparatus 100 can capture image 114 (e.g., at a first time) and image 802 (e.g., at a second time after the first time). Alternatively, images 114 and 802 can be obtained from another source; for example, apparatus 800 can receive images 114 and 802 from another device via a communication interface. Image 114 can represent a first field of view of the scene, and image 802 can represent a second field of view of the scene (e.g., based on the camera 112 having moved or changed orientation between the first and second times). Apparatus 100 can use generative machine learning models (e.g., generative machine learning model 806 and / or generative machine learning model 807) to generate image 808 based on at least a portion of image 114 and at least a portion of image 802. In some aspects, processor 118 can use generative machine learning model 806 to generate image 802. In other respects, the device 800 may use a generative machine learning model 807, which may be implemented on another computing device (e.g., a laptop computer, mobile phone, or server) to generate image 808. The processor 118 may provide at least a portion of image 114 and at least a portion of image 802 to generative machine learning models 806 and / or 807 as input or conditions, and generative machine learning models 806 and / or 807 may generate image 808 based on at least a portion of image 114 and at least a portion of image 802.
[0089] Additionally or alternatively, the UI 102 of device 100 may display image 114 (and / or image 802) at display 104 of UI 102. UI 102 may (e.g., at touch sensor 106 and / or using orientation sensor 108) receive user input 110 indicating a desired change 116 to image 114 (and / or image 802). The desired change 116 to image 114 (and / or image 802) may be or may include a change to a first field of view of image 114 and / or a change to a second field of view of image 802. Further, processor 118 may provide instructions to generative machine learning model 806 (and / or generative machine learning model 807) regarding the generation of image 808. These instructions may be based on change 116. Generative machine learning model 806 (and / or generative machine learning model 807) can generate image 808 based on change 116 to include at least a portion (e.g., pixels) of image 114 in the first field of view and / or at least a portion (e.g., pixels) of image 802 in the second field of view, as well as generated pixels outside the first and second fields of view.
[0090] Device 800 can be basically similar to Figure 1 The device 100 may perform many of the same operations as the device 100. For example, the UI 102 in the device 800 may be the same as the UI 102 in the device 100, and the UI 102 may operate in the device 800 in the same way as the UI 102 in the device 100. In addition, in the device 800, the UI 102 may display an image 802 at the display 104 and receive user input 110 related to the image 802. The UI 102 may generate a change 116 indicating a change in the field of view of the image 114 (and / or the image 802).
[0091] Furthermore, camera 112 may be the same as, substantially similar to, or perform the same or substantially the same operation as camera 112 of device 100. Image 114 in device 100 may be the same as or substantially similar to image 114 in device 600. Furthermore, camera 112 may capture image 802 (e.g., after camera 112 captures image 114). Camera 112 may capture image 114 with a different field of view than image 802 captured by camera 112. For example, when capturing image 802, camera 112 may be pointed in a different direction, moved, or zoomed to a different focal length compared to when capturing image 114.
[0092] The processor 118 of device 600 may be the same as, substantially similar to, and / or perform the same or substantially the same operations as the processor 118 of device 100. However, the generative machine learning model 806 of device 800 may differ from the generative machine learning model 120 of device 100. For example, while generative machine learning model 120 may be trained to generate images including content generated based on input images, generative machine learning model 806 may be trained to generate images including content generated based on two or more input images. Similarly, generative machine learning model 807 may be trained to generate images including content generated based on two or more input images. Regarding... Figure 12 , Figure 13 , Figure 14 and Figure 15 Additional details are provided regarding generative machine learning models and their training. (Based on information about...) Figure 12 , Figure 13 , Figure 14 and Figure 15 The provided description allows for the training of generative machine learning model 806 and / or generative machine learning model 807.
[0093] Furthermore, processor 118 may implement scene determiner 804. Scene determiner 804 may determine that image 114 and image 802 represent the same scene (e.g., although image 114 and image 802 represent different fields of view of the scene). Scene determiner 804 may determine that the scene is the same based on: the capture time of image 114 and image 802, the capture position of image 114 and image 802 (e.g., as recorded by the positioning service of device 800), the movement of camera 112 between the capture of image 114 and image 802 (e.g., as recorded by the inertial measurement unit), a comparison of image 114 and image 802, or any combination thereof. Processor 118 may determine, based on scene determiner 804, that image 114 and image 802 represent the same scene and / or based on change 116, determine that image 808 is generated based on image 114 and / or image 802. For example, if images 114 and 802 represent different fields of view of the same scene and change 116 involves two fields of view, processor 118 can determine to generate image 808 based on images 114 and 802.
[0094] Apparatus 800 may provide at least a portion of image 114 and at least a portion of image 802 to generative machine learning model 806 (and / or generative machine learning model 807) as input or as conditions. Generative machine learning model 806 (and / or generative machine learning model 807) may generate image 808 based on at least a portion of image 114 and at least a portion of image 802 (e.g., using image 114 and image 802 as conditions). Further, apparatus 800 may instruct generative machine learning model 806 (and / or generative machine learning model 807) relative to the generated image 808 based on change 116. For example, processor 118 may determine operating parameters or operating instructions for generative machine learning model 806 (and / or generative machine learning model 807) based on change 116. Image 808 generated based on image 114 and image 802 may include at least a portion of image 114 (field of view of image 114) and at least a portion of image 802 (field of view of image 802), as well as generated pixels outside the field of view. For example, based on desired changes, image 808 may include part or all of image 114 and / or part or all of image 802. Additionally or alternatively, device 800 may smooth the edges between generated pixels and pixels from image 114 and / or image 802.
[0095] In some respects, processor 118 may enable UI 102 to display image 808 at display 104. Additionally or alternatively, processor 118 may analyze image 808 or cause image 808 to be analyzed by one or more other processors. Additionally or alternatively, processor 118 may cause image 808 to be stored (e.g., in memory) or transmitted for later display and / or analysis at different locations.
[0096] By interpreting user input 110 as a change 116 related to image 114 and / or image 802 and instructing generative machine learning model 806 (and / or generative machine learning model 807) based on change 116, device 800 can provide users with a simple and convenient way to instruct generative machine learning model 806 (and / or generative machine learning model 807) relative to image 808.
[0097] For example, based on a pinch gesture, device 800 can instruct generative machine learning model 806 (and / or generative machine learning model 807) to generate image content around image 114 and / or image 802 in image 808 (e.g., as if image 808 were captured with a wider field of view of image 114 or image 802). Figure 2 The example provided is based on expanding the field of view to generate image content. Figure 2The example can be implemented by device 800 by adding a second input image. As another example, based on a drag gesture, device 800 can instruct generative machine learning model 806 (and / or generative machine learning model 807) to generate image content on one side of image 802 in image 114 and / or image 802 by changing 116 (e.g., as if image 802 were captured from image 114 and / or image 808 from a translated field of view). Figure 3 The example provided is an example of generating image content based on a translation field of view. Figure 3 The example can be implemented by device 800 by adding a second input image. As another example, based on a rotation gesture, device 800 can instruct generative machine learning model 806 (and / or generative machine learning model 807) to generate image content by changing 116 to fill the corners of image 808 around a rotated version of image 802 in image 114 and / or image 808 (e.g., as if image 808 were captured from image 114 or image 802 by a rotating camera). Figure 4 The example provided is an example of generating image content based on a rotated field of view. Figure 4 The example can be implemented by device 800 by adding a second input image. As another example, based on the rotation of device 800, device 800 can instruct generative machine learning model 806 (and / or generative machine learning model 807) to generate content to the right and left of image 114 and / or image 802, or above and below image 114 and / or image 802, by changing 116 (e.g., as if image 808 were captured from image 114 or image 802 by a camera rotated 90 degrees). Figure 5 The example provided is based on rotating the field of view by 90 degrees to generate image content. Figure 5 An example can be implemented by device 600 by adding a second input image.
[0098] Figure 9Example representation 902 of apparatus 204 includes displaying image 908 (e.g., which may be captured by a first camera of apparatus 204 in a first moment) according to various aspects of this disclosure, example representation 914 of apparatus 204 includes displaying image 916 (e.g., which may be captured by a camera of apparatus 204 in a second moment), and example representation 920 of apparatus 204 includes displaying image 922 (e.g., including generated image content, such as generated pixel 924 and / or modified pixel 926). Generally, apparatus 204 can implement systems and techniques by displaying image 908 and / or image 916 (e.g., at separate times, in a partitioned screen display, or partially overlapping), receiving user input (e.g., pinch gesture, drag gesture, rotation gesture, or rotation of apparatus 204), and generating image 922 (including generated pixel 924) based on image 908, image 916, and user input.
[0099] For example, 902 may include a device 204 implementing UI 206, which displays image 908. Image 908 may be generated by a camera of device 204 (e.g., ...). Figure 8 The camera 112) captures the image in the first instance. Image 908 can be... Figure 8 An example of image 114. Image 908 can be a first field of view 910, for example, image 908 can represent the field of view 910 of a scene.
[0100] Furthermore, 914 includes a device 204 implementing UI 206, which displays image 916. Image 916 can be captured by a camera (e.g., camera 112) of the capturing device 204, which captures image 908. Image 916 can be... Figure 8 An example of image 802. Image 916 may belong to a second field of view 918, for example, image 916 may represent the field of view 918 of a scene (e.g., the same scene represented by image 908).
[0101] UI 206 can receive user input (e.g., pinch gesture, drag gesture, rotate gesture, or rotation from device 204) and interpret the user input in relation to image 908 and / or image 916. The change can be a change to field of view 910 and / or field of view 918. For example, UI 206 can interpret the user input as a desire to change field of view 910 and / or field of view 918 (e.g., by expanding field of view 910 and / or field of view 918, panning field of view 910 and / or field of view 918, and / or rotating field of view 910 and / or field of view 918).
[0102] Device 204 can be based on a generative machine learning model of device 204 (e.g., Figure 8The device 204 provides at least a portion of image 908 and at least a portion of image 916 (e.g., as conditions) and instructions by modifying generative machine learning models 806 and / or 807. For example, based on a desired change in the field of view, the device 204 may determine a portion of image 908 and / or a portion of image 916 (e.g., excluding a portion of image 716 that can be cut off from the frame by translation) and provide that portion to the generative machine learning model. Alternatively, the device 204 may provide the entire image 708 and / or the entire image 716 to the generative machine learning model. The generative machine learning model may generate image 922 based on at least a portion of image 908, at least a portion of image 916, and instructions, and UI 206 may display image 922 at UI 206, as illustrated in representation 920. Image 922 may include at least a portion of image 908 in a first field of view 910, at least a portion of image 916 in a second field of view 918, generated pixels 924 (which may be outside fields of view 910 and 918), and / or modified pixels 926. For example, image 922 may include at least a portion of image 908 in field of view 910, at least a portion of image 916 in field of view 918, and generated pixels 924 on one or more sides of fields of view 910 and / or 918. Additionally or alternatively, image 922 may include modified pixels 926, which may be outside field of view 910, inside field of view 918, and based on image 916. For example, a generative machine learning model may generate modified pixels 926 based on image 916. Modified pixels 926 may be the same as or different from pixels in image 916. For example, modified pixels 926 may have a different resolution than image 916. Alternatively, the modified pixel 926 may be within the field of view 910 of image 908. In such cases, the modified pixel 926 may be based on image 908 and / or image 916. The modified pixel 926 may be the same as or different from the pixels of image 908. For example, the modified pixel 926 may have a different resolution than image 908. Additionally or alternatively, device 204 may smooth the edges between the generated pixel 924 and the pixels from image 908 and / or image 916. Image 922 may appear to be image 908 of the scene as if image 908 were captured in a different field of view. Additionally or alternatively, image 922 may appear to be image 916 of the scene as if image 916 were captured in a different field of view. Where image 922 includes a portion of image 908 and / or a portion of image 916, the portions of image 908 and / or image 916 included in image 922 may be identical to the portions of image 908 and / or image 916 provided by device 204. Alternatively, a portion of image 908 and / or a portion of image 916 included in image 922 may be a subset of the images provided by device 204.For example, device 204 may provide all (or most) of image 908, and image 922 may include a subset of the provided images.
[0103] Figure 10 This is a flowchart illustrating a process 1000 for generating image content according to various aspects of this disclosure. One or more operations of process 1000 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device (such as a watch), an extended reality (XR) device (such as a virtual reality (VR) device or an augmented reality (AR) device), a vehicle or a component or system of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, and / or any other computing device having the resource capabilities to perform process 1000. One or more operations of process 1000 may be implemented as software components that execute and run on one or more processors.
[0104] At box 1002, a computing device (or one or more components thereof) may display an image of the field of view at a user interface. In some aspects, the computing device (or one or more components thereof) may include a user interface and may display an image of the field of view at the user interface. In other aspects, the computing device (or one or more components thereof) may communicate with the user interface and may enable the user interface to display an image of the field of view. For example, Figure 1 The device 100 can display an image 208 of field of view 210 at the display 104 of UI 102.
[0105] At box 1004, the computing device (or one or more components thereof) may receive user input at a user interface indicating a desired change to an image, where the desired change to the image includes a change in the field of view. In some aspects, the computing device (or one or more components thereof) may include a user interface and may receive user input indicating a desired change to an image at the user interface. In other aspects, the computing device (or one or more components thereof) may communicate with the user interface, which may receive user input indicating a desired change to an image and provide that instruction to the computing device (or one or more components thereof). For example, device 100 may receive user input (e.g., pinch gesture 212, drag gesture 312, rotation gesture 412, or rotation of device 100) at UI 102 (e.g., at touch sensor 106, using orientation sensor 108, and / or using camera 109). The user input may indicate a desired change to the view 210 of image 208.
[0106] In some aspects, the user interface can interpret gestures to receive user input. In some aspects, the user interface can be or may include a touchscreen configured to sense touch. The user interface can be configured to interpret touch as gestures. For example, UI 102 may include touch sensor 106. In some aspects, the user interface can be configured to interpret pinch touch gestures as indications of desired changes, including expanding the field of view. For example, UI 102 may use touch sensor 106 to sense pinch gestures and interpret them as indications of desired expansion of the field of view 210 represented by image 208, for example, as per [reference to image 208]. Figure 2 As described and illustrated. In some aspects, the user interface can be configured to interpret drag-and-drop touch gestures as indications of desired changes, including panning, to alter the field of view. For example, UI 102 can use touch sensor 106 to sense drag-and-drop gestures and interpret them as indications of desired panning of the field of view 310 represented by image 308, for example, as per [reference to image 308]. Figure 3 As described and illustrated. In some aspects, the user interface can be configured to interpret rotational touch gestures as indications that desire a change, including rotation of the field of view. For example, UI 102 can use touch sensor 106 to sense rotational gestures and interpret them as indications that desire to rotate the field of view 410 represented by image 408, for example, as described in relation to Figure 4 The description and examples are as follows.
[0107] In some aspects, the user interface can interpret gestures to receive user input. In some aspects, the user interface includes a camera configured to capture images of the user. The user interface can be configured to interpret the user's posture as gestures. For example, UI 102 may include camera 109. Camera 109 may capture one or more images of the user. UI 102 may track various body parts of the user and interpret gestures based on the tracked body parts. In some aspects, the camera may be or may include at least one of the following: an active depth camera; an infrared (IR) camera; a red-green-blue (RGB) camera; a stereo camera; or an eye-facing camera. For example, camera 109 may include any or all of the following: an active depth camera; an infrared (IR) camera; a red-green-blue (RGB) camera; a stereo camera; or an eye-facing camera. In some aspects, the user interface can be configured to interpret at least one of the following: interpret a gesture to receive the user input; or interpret eye position as a gesture to receive the user input. For example, UI 102 can be configured to interpret gestures and / or eye positions (e.g., gaze, saccade, blink, etc.) as gestures.
[0108] In some aspects, a computing device (or one or more components thereof) may include or be coupled to an orientation sensor configured to sense the orientation of a setting, and wherein a user interface is configured to interpret rotation of the device as an indication of a desired change including rotation of the field of view. For example, the UI 102 of device 100 may include an orientation sensor 108 that can determine the orientation (and / or a change in orientation) of device 100. UI 102 may interpret the change in orientation as a desired change in the field of view of image 114. In some aspects, an image may be displayed by the user interface in a portrait format, and the changed image includes generated pixels on at least one side of at least a portion of the image in the field of view; or the image may be displayed by the user interface in a landscape format, and the changed image includes generated pixels at at least one of the at least portion of the image in the field of view, above or below. For example, the UI 206 of device 204 may display image 508 in a portrait format. A user may rotate device 204, and UI 206 may interpret the rotation as a desired change in the field of view of image 508. Therefore, device 204 can transmit image 508 along with an indication of the desired rotation of image 508 from a portrait format to a landscape format to a generative machine learning model. The generative machine learning model can generate image 516, which may include at least a portion of image 508. Device 204 can display image 516 at UI 206. As an example, UI 206 of device 204 can display image 516 in landscape format. A user can rotate device 204, and UI 206 can interpret the rotation as an indication of the desired rotation of image 516. Therefore, device 204 can transmit image 516 along with an indication of the desired rotation of image 516 from a landscape format to a portrait format to a generative machine learning model. The generative machine learning model can generate image 508, which may include at least a portion of image 516. Device 204 can display image 508 at UI 206.
[0109] At box 1006, a computing device (or one or more components thereof) can provide at least a portion of an image and an indication of a desired change as input to a generative machine learning model. In some aspects, the computing device (or one or more components thereof) can implement a generative machine learning model. In such aspects, the computing device (or one or more components thereof) can provide at least a portion of an image and an indication of a desired change as input to a local generative machine learning model. In other aspects, the computing device (or one or more components thereof) can communicate with another computing device (or one or more components thereof) that can implement a generative machine learning model. In such aspects, the computing device (or one or more components thereof) can provide at least a portion of an image and an indication of a desired change as input to a generative machine learning model at another computing device (or one or more components thereof). For example, apparatus 100 can provide generative machine learning models 120 and / or 121 with at least a portion of image 114 (e.g., a portion of image 208) and an indication of a desired change (e.g., change 116).
[0110] At box 1008, a computing device (or one or more components thereof) can obtain a modified image from a generative machine learning model, wherein the modified image includes at least a portion of the image within the field of view and generated pixels outside the field of view. In some aspects, the computing device (or one or more components thereof) can implement the generative machine learning model. In such aspects, the computing device (or one or more components thereof) can obtain the modified image directly from the generative machine learning model. In other aspects, the computing device (or one or more components thereof) can communicate with another computing device (or one or more components thereof) that can implement the generative machine learning model. In such aspects, the computing device (or one or more components thereof) can obtain the modified image from a generative machine learning model from another computing device (or one or more components thereof). For example, device 100 can obtain image 122 (from generative machine learning model 120 and / or generative machine learning model 121). Image 122 may include at least a portion of the image 114 within the field of view (e.g., at least a portion of the image 208 within the field of view 210) and generated pixels outside the field of view (e.g., generated pixel 218).
[0111] In some respects, a computing device (or one or more components thereof) can display a changed image. For example, a computing device (or one or more components thereof) can cause a user interface to display a changed image. For example, device 204 can display image 216 at UI 206. Additionally or alternatively, a computing device (or one or more components thereof) can cause a changed image to be displayed by the user interface, stored in memory, analyzed, or transmitted. For example, device 100 can cause image 122 to be displayed (e.g., at display 104), stored in memory (…Figure 1 (not illustrated) in, being analyzed (e.g., at processor 118), and / or being sent (e.g., via) Figure 1 (Communication interfaces not illustrated). Image 122 may be stored or transmitted for display or analysis at another device and / or at another time.
[0112] In some aspects, the altered image may be or may be included in pixels generated on at least both sides of at least a portion of the image in the field of view. For example, image 216 includes pixels 218 generated on both sides of image 208 in field of view 210. In some aspects, the altered image may be or may be included in pixels generated on at least one side of at least a portion of the image in the field of view. For example, image 316 includes pixels 318 generated on one side of image 308 in field of view 310. In some aspects, the altered image may be or may be included in pixels generated at at least one corner of the altered image. For example, image 416 includes pixels 418 generated at all four corners of image 416.
[0113] In some aspects, a computing device (or one or more components thereof) can generate a final image by rotating or trimming at least one of the altered images based on a change in the field of view. For example, device 100 can obtain image 122. Device 100 (using processor 118) can rotate and / or trim image 122 to generate a final image. For example, device 100 can obtain a version of image 408 that has been expanded. Device 100 can rotate and trim the expanded image to generate image 416.
[0114] In some aspects, a computing device (or one or more components thereof) can rotate or crop an image to obtain at least a portion of the image for provision to a generative machine learning model. For example, before providing image 114 to generative machine learning model 120 (or generative machine learning model 121), device 100 may obtain image 114 and rotate or crop image 114. Device 100 may provide the rotated and / or cropped image 114 to generative machine learning model 120 (or generative machine learning model 121). Generative machine learning model 120 (or generative machine learning model 121) may generate image 122 based on the rotated and / or cropped version of image 114. For example, device 100 may obtain image 408. Device 100 may rotate image 408, for example, as illustrated in the frame of image 416, and provide the rotated image 408 to the generative machine learning model. The generative machine learning model may generate image 416 based on the rotated image 408.
[0115] In some aspects, the image may be a first image. A computing device (or one or more components thereof) may acquire the first image from a first camera. A computing device (or one or more components thereof) may acquire a second image from a second camera. The computing device (or one or more components thereof) may provide at least a portion of the first image and at least a portion of the second image as input to a generative machine learning model. For example, apparatus 600 may include camera 112 and camera 602. Camera 112 may capture image 604, and camera 602 may capture image 604. Apparatus 600 may provide at least a portion of image 114 and at least a portion of image 604 to generative machine learning model 606 (and / or generative machine learning model 607). The generative machine learning model may generate a modified image based on at least a portion of the first image and at least a portion of the second image. For example, generative machine learning model 606 may generate image 608 based on at least a portion of image 114 and at least a portion of image 604. For example, apparatus 204 may acquire images 708 and 716, and provide at least a portion of image 708 and at least a portion of image 716 to a generative machine learning model. A generative machine learning model can generate image 722 based on at least a portion of image 708 and at least a portion of image 716. In some aspects, the first camera may have a first focal length, and the second camera may have a second focal length different from the first focal length. For example, camera 112 may have a different focal length than camera 602.
[0116] In some aspects, the image may be a first image of a scene captured in a first moment. A computing device (or one or more components thereof) may acquire a second image. The second image may represent the scene captured in a second moment. The computing device (or one or more components thereof) may provide at least a portion of the first image and at least a portion of the second image to a generative machine learning model as input. For example, camera 112 may capture image 114 in a first moment and image 802 in a second moment. Both image 114 and image 802 may represent the same scene. Apparatus 100 may provide at least a portion of image 114 and at least a portion of image 802 to a generative machine learning model 806 (or generative machine learning model 807) as input. The generative machine learning model may generate modified images based on at least a portion of image 114 and at least a portion of image 802. For example, apparatus 204 may acquire images 908 and 916 and provide at least a portion of image 908 and at least a portion of image 916 to a generative machine learning model. A generative machine learning model can generate image 922 based on at least a portion of image 908 and at least a portion of image 916. In some aspects, the field of view can be a first field of view of a scene; the first image belongs to the first field of view of the scene; and the second image belongs to a second field of view of the scene. For example, image 908 can represent a first field of view 910 of the scene. Image 916 can represent a second field of view 918 of the scene.
[0117] In some aspects, a computing device (or one or more components thereof) may determine the use of a second image based on a desired change to a first image. For example, the computing device (or one or more components thereof) may acquire a plurality of images. The computing device (or one or more components thereof) may determine the use of a second image based on a first image, a scene, and / or a desired change to the first image. For example, apparatus 800 includes a scene determiner 804 that can determine a relationship between the first image and the second image, a relationship between a scene captured by the first image and a scene captured by the second image, and / or a relationship between the first image and the second image based on a change to the first image.
[0118] In some aspects, the computing device (or one or more components thereof) may acquire additional pixels before acquiring the modified image and cause the user interface to display the image within the field of view and additional pixels outside the field of view; and in response to acquiring the modified image, cause the user interface to display the modified image. For example, device 204 may provide image 208 and a pinch gesture 212 to a generative machine learning model. Before receiving image 216, device 204 may cause UI 206 to display image 208 within the field of view 210 and pixels outside 208 (e.g., where generated pixel 218 is exemplified). In some aspects, pixels may be acquired from the generative machine learning model before the generated pixel 218 is acquired by the generative machine learning model. In some aspects, additional pixels may be blurred. For example, additional pixels displayed around image 208 within the field of view 210 may be blurred before device 204 acquires generated pixel 218 from the generative machine learning model.
[0119] In some respects, a computing device (or one or more components thereof) can smooth pixels at the edges between at least a portion of the image of the field of view and generated pixels outside the field of view. For example, device 204 can smooth pixels at the edges between the image 208 of the field of view 210 and generated pixels 218.
[0120] In some aspects, a computing device (or one or more components thereof) can implement a generative machine learning model. For example, apparatus 100 may include a generative machine learning model 120. In some aspects, the computing device (or one or more components thereof) may include a communication interface configured to: send at least a portion of the image and the indication of the desired change to the computing device implementing the generative machine learning model; and receive the changed image from the computing device. For example, apparatus 100 may include a communication interface ( Figure 1 (Not illustrated). Further, device 100 may send at least a portion of image 114 and change 116 to generative machine learning model 121. Generative machine learning model 121 may process image 114 according to change 116 to generate image 122 and provide image 122 to device 100.
[0121] In some examples, as previously noted, the methods described herein (e.g., Figure 10 The process 1000 and / or other methods described herein may be performed wholly or partially by a computing device or apparatus. In one example, one or more of the methods may be performed by... Figure 1 Device 100 Figure 6 Device 600, Figure 8 The device 800 or another system or device performs the operation. In another example, these methods (e.g., Figure 10One or more of the processes 1000 and / or other methods described herein may be used by Figure 18 The computing device architecture 1800 shown is implemented wholly or partially. For example, it has Figure 18 The computing device of the illustrated computing device architecture 1800 may include or be included in components of device 100, processor 118, device 600, and / or device 800, and may perform operation of process 1000 and / or other processes described herein. In some cases, the computing device or device may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.
[0122] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0123] Process 1000 and / or other processes described herein are illustrated as logic flowcharts, whose operations represent sequences of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.
[0124] Additionally, process 1000 and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer processes, or one or more applications) that executes jointly on one or more processors, implemented in hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0125] Figure 11 This is a block diagram illustrating an example architecture of an image processing system 1100 according to various aspects of the present disclosure. The image processing system 1100 includes various components for capturing and processing images, such as an image of scene 1106. The image processing system 1100 can capture image frames (e.g., still images or video frames). In some cases, a lens 1108 and an image sensor 1118 (which may include an analog-to-digital converter (ADC)) may be associated with an optical axis. In one exemplary example, both the photosensitive area of the image sensor 1118 (e.g., a photodiode) and the lens 1108 may be centered on the optical axis.
[0126] In some examples, the lens 1108 of the image processing system 1100 faces the scene 1106 and receives light from the scene 1106. The lens 1108 bends the incident light from the scene toward the image sensor 1118. The light received by the lens 1108 then passes through the aperture of the image processing system 1100. In some cases, the aperture (e.g., aperture size) is controlled by one or more control mechanisms 1110. In other cases, the aperture may have a fixed size.
[0127] One or more control mechanisms 1110 may control exposure, focus, and / or zoom based on information from image sensor 1118 and / or image processor 1124. In some cases, one or more control mechanisms 1110 may include multiple mechanisms and components. For example, control mechanism 1110 may include one or more exposure control mechanisms 1112, one or more focus control mechanisms 1114, and / or one or more zoom control mechanisms 1116. One or more control mechanisms 1110 may also include, in addition to Figure 11 Additional control mechanisms beyond those illustrated herein. For example, in some cases, one or more control mechanisms 1110 may include controls for controlling analog gain, flash, HDR, depth of field, and / or other image capture characteristics.
[0128] The focus control mechanism 1114 of the control mechanism 1110 can obtain focus settings. In some examples, the focus control mechanism 1114 stores the focus settings in a memory register. Based on the focus settings, the focus control mechanism 1114 can adjust the positioning of the lens 1108 relative to the positioning of the image sensor 1118. For example, based on the focus settings, the focus control mechanism 1114 can adjust the focus by moving the lens 1108 closer to or further away from the image sensor 1118 by actuating a motor or servo system (or other lens mechanism). In some cases, the image processing system 1100 may include additional lenses. For example, the image processing system 1100 may include one or more microlenses on each photodiode of the image sensor 1118. These microlenses can each bend light received from the lens 1108 toward the corresponding photodiode before the light reaches the photodiode.
[0129] In some examples, focus settings may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. Focus settings may be determined using control mechanism 1110, image sensor 1118, and / or image processor 1124. Focus settings may be referred to as image capture settings and / or image processing settings. In some cases, lens 1108 may be fixed relative to the image sensor and focus control mechanism 1114.
[0130] Exposure control mechanism 1112 of control mechanism 1110 can obtain exposure settings. In some cases, exposure control mechanism 1112 stores exposure settings in a memory register. Based on the exposure settings, exposure control mechanism 1112 can control the aperture size (e.g., aperture size or aperture value), the duration of aperture opening (e.g., exposure time or shutter speed), the duration of light collection by the sensor (e.g., exposure time or electronic shutter speed), the sensitivity of image sensor 1118 (e.g., ISO speed or film speed), the analog gain applied by image sensor 1118, or any combination thereof. Exposure settings may be referred to as image capture settings and / or image processing settings.
[0131] The zoom control mechanism 1116 of the control mechanism 1110 can obtain zoom settings. In some examples, the zoom control mechanism 1116 stores the zoom settings in a memory register. Based on the zoom settings, the zoom control mechanism 1116 can control the focal length of an assembly (lens assembly) including lens elements such as lens 1108 and one or more additional lenses. For example, the zoom control mechanism 1116 can control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more lenses relative to each other. The zoom settings may be referred to as image capture settings and / or image processing settings. In some examples, the lens assembly may include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly may include a focusing lens (in some cases, this focusing lens may be lens 1108) that first receives light from scene 1106, where the light then passes through the focusing zoom system between the focusing lens (e.g., lens 1108) and image sensor 1118 before reaching image sensor 1118. In some cases, a focused zoom system may include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference from each other), with a negative (e.g., diverging, concave) lens between the two positive lenses. In some cases, zoom control mechanism 1116 moves one or more lenses in the focused zoom system, such as a negative lens and one or both positive lenses. In some cases, zoom control mechanism 1116 can control zoom by capturing images from an image sensor (e.g., including image sensor 1118) among a plurality of image sensors at a zoom setting corresponding to the zoom setting. For example, image processing system 1100 may include a wide-angle image sensor with a relatively low zoom and a telephoto image sensor with a greater zoom. In some cases, zoom control mechanism 1116 may capture images from the corresponding sensor based on the selected zoom setting.
[0132] Image sensor 1118 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image produced by image sensor 1118. In some cases, different photodiodes may be covered by different filters. In some cases, different photodiodes may be covered in different color filters, and light that matches the color of the filter covering the photodiode can thus be measured. Various color filter arrays can be used, such as, for example, and not limited to, Bayer color filter arrays, four-color filter arrays (QCFA), and / or any other color filter array.
[0133] In some cases, image sensor 1118 may optionally or additionally include opaque and / or reflective masks that block light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles. In some cases, opaque and / or reflective masks may be used for phase detection autofocus (PDAF). In some cases, opaque and / or reflective masks may be used to block portions of the electromagnetic spectrum from reaching the photodiodes of the image sensor (e.g., IR cutoff filters, UV cutoff filters, bandpass filters, low-pass filters, high-pass filters, etc.). Image sensor 1118 may also include an analog gain amplifier for amplifying the analog signal output from the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signal output from the photodiodes (and / or the analog signal amplified by the analog gain amplifier) into a digital signal. In some cases, certain components or functions discussed with respect to one or more control mechanisms in control mechanism 1110 may alternatively or additionally be included in image sensor 1118. Image sensor 1118 may be a charge-coupled device (CCD) sensor, an electron multiplication CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide semiconductor (CMOS), an N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.
[0134] Image processor 1124 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 1128), one or more host processors (including host processor 1126), and / or related to Figure 18 The computing device architecture 1800 may include one or more processors of any other type discussed. The host processor 1126 may be a digital signal processor (DSP) and / or other types of processor. In some specific implementations, the image processor 1124 is a single integrated circuit or chip (e.g., referred to as a system-on-a-chip or SoC) that includes the host processor 1126 and the ISP 1128. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 1130), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, and connectivity components (e.g., Bluetooth). ™This includes components such as the Global Positioning System (GPS), any combination thereof, and / or other components. I / O port 1130 may include any suitable input / output port or interface according to one or more protocols or specifications, such as Inter-Integrated Circuit 2 (I2C) interface, Inter-Integrated Circuit 3 (I3C) interface, Serial Peripheral Interface (SPI) interface, Serial General Purpose Input / Output (GPIO) interface, Mobile Industrial Processor Interface (MIPI) (such as MIPI CSI-2 physical (PHY) layer port or interface), Advanced High Performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In an exemplary example, host processor 1126 may communicate with image sensor 1118 using the I2C port, and ISP 1128 may communicate with image sensor 1118 using the MIPI port.
[0135] Image processor 1124 can perform multiple tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving input, managing output, managing memory, or some combination thereof. Image processor 1124 can store image frames and / or processed images in random access memory (RAM) 1120, read-only memory (ROM) 1122, cache, memory unit, another storage device, or some combination thereof.
[0136] Various input / output (I / O) devices 1132 may be connected to the image processor 1124. I / O devices 1132 may include a display screen, keyboard, keypad, touchscreen, touchpad, touch-sensitive surface, printer, any other output device, any other input device, or any combination thereof. In some cases, text may be input into the image processing device 1104 via the physical keyboard or keypad of the I / O device 1132, or via a virtual keyboard or keypad on the touchscreen of the I / O device 1132. I / O devices 1132 may include one or more ports, jacks, or other connectors that enable a wired connection between the image processing system 1100 and one or more peripheral devices, through which the image processing system 1100 may receive data from and / or send data to one or more peripheral devices. I / O devices 1132 may include one or more wireless transceivers that enable a wireless connection between the image processing system 1100 and one or more peripheral devices, through which the image processing system 1100 may receive data from and / or send data to one or more peripheral devices. Peripheral devices may include any type of I / O device 1132 discussed earlier, and they can be considered I / O devices 1132 in themselves once they are coupled to ports, jacks, wireless transceivers or other wired and / or wireless connectors.
[0137] In some cases, the image processing system 1100 may be a single device. In other cases, the image processing system 1100 may be two or more independent devices, including an image capture device 1102 (e.g., a camera) and an image processing device 1104 (e.g., a computing device coupled to the camera). In some embodiments, the image capture device 1102 and the image processing device 1104 may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly coupled together via one or more wireless transceivers. In some embodiments, the image capture device 1102 and the image processing device 1104 may be disconnected from each other.
[0138] like Figure 11 As shown, the vertical dashed line will Figure 11The image processing system 1100 is divided into two parts, namely image capture device 1102 and image processing device 1104. Image capture device 1102 includes a lens 1108, a control mechanism 1110, and an image sensor 1118. Image processing device 1104 includes an image processor 1124 (including an ISP 1128 and a host processor 1126), RAM 1120, ROM 1122, and I / O devices 1132. In some cases, certain components exemplified in image capture device 1102 (such as ISP 1128 and / or host processor 1126) may be included in image capture device 1102. In some examples, image processing system 1100 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof.
[0139] Image processing system 1100 may be part of or implemented by a single computing device or multiple computing devices. In some examples, image processing system 1100 may be part of electronic devices (or multiple electronic devices), such as camera systems (e.g., digital cameras, IP cameras, video cameras, security cameras, etc.), telephone systems (e.g., smartphones, cellular phones, conferencing systems, etc.), laptops or notebook computers, tablet computers, set-top boxes, smart TVs, display devices, game consoles, XR devices (e.g., HMDs, smart glasses, etc.), IoT (Internet of Things) devices, smart wearable devices, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices.
[0140] Although the image processing system 1100 is shown as including certain components, those skilled in the art will understand that the image processing system 1100 may include more than [other components]. Figure 11 The components shown are those of the majority of the components. The components of the image processing system 1100 may include software, hardware, or one or more combinations of software and hardware. For example, in some embodiments, the components of the image processing system 1100 may include electronic circuitry or other electronic hardware and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, GPU, DSP, CPU, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing the image processing system 1100.
[0141] In some examples, Figure 18 The computing device architecture 1800 shown and further described below may include an image processing system 1100, an image capture device 1102, an image processing device 1104, or a combination thereof.
[0142] As noted above, various aspects of this disclosure may utilize machine learning models or systems.
[0143] Figure 12 An example implementation of system 1200 is illustrated, which may include a central processing unit (CPU 1202) (which may be a multi-core CPU) configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with computing devices (e.g., a weighted neural network), task information, and other information may be stored in a memory block associated with the neural processing unit (NPU 1208), a memory block associated with the CPU 1202, a memory block associated with the graphics processing unit (GPU 1204), a memory block associated with the digital signal processor (DSP 1206), memory 1216, and / or may be distributed across multiple blocks. Instructions executed at the CPU 1202 may be loaded from the process memory associated with the CPU 1202 or may be loaded from memory 1216.
[0144] System 1200 may also include additional processing blocks tailored for specific functions, such as GPU 1204, DSP 1206, connectivity engine 1218 (which may include fifth-generation (5G) connectivity, fourth-generation LTE (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and multimedia processor 1212 capable of, for example, detecting and recognizing gestures. In one specific implementation, the NPU is implemented in CPU 1202, DSP 1206, and / or GPU 1204. System 1200 may also include one or more sensor processors 1214, one or more image signal processors (ISP 1210), and / or navigation engine 1220, which may include a global positioning system. In some examples, sensor processor 1214 may be associated with or connected to one or more sensors for providing sensor input to sensor processor 1214. For example, the one or more sensors and sensor processor 1214 may be provided in the same computing device, coupled to the same computing device, or otherwise associated with the same computing device.
[0145] System 1200 may be implemented as a system-on-a-chip (SoC). System 1200 may be based on an Advanced Reduced Instruction Set Computer (RISC) machine (ARM) instruction set. System 1200 and / or its components may be configured to perform machine learning techniques according to various aspects of this disclosure discussed herein. For example, system 1200 and / or its components may be configured to implement machine learning models (e.g., quantized trained machine learning models) as described herein and / or according to various aspects of this disclosure.
[0146] Machine learning (ML) can be considered a subset of artificial intelligence (AI). ML systems can include algorithms and statistical models that computer systems can use to perform various tasks through pattern-dependent inference and speculation, without the use of explicit instructions. An example of an ML system is a neural network (also known as an artificial neural network), which can include groups of interconnected artificial neurons (e.g., neuron models). Neural networks can be used in a variety of applications and / or devices, such as image and / or video decoding, image analysis and / or computer vision applications, Internet Protocol (IP) cameras, Internet of Things (IoT) devices, autonomous vehicles, service robots, and more.
[0147] Individual nodes in a neural network mimic biological neurons by taking input data and performing simple operations on that data. The results of these simple operations on the input data are selectively passed to other neurons. Weights are associated with each vector and node in the network, and these values constrain how the input data relates to the output data. For example, the input data for each node can be multiplied by its corresponding weight, and the products can be summed. The sum of the products can be adjusted with optional biases, and activation functions can be applied to the results to produce the node's output signal or "output activation" (sometimes called a feature map or activation map). The weights can initially be determined by an iterative stream of training data through the network (e.g., weights are established during training phases where the network learns how to identify a particular category based on the characteristics of its typical input data).
[0148] There are different types of neural networks, such as Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Generative Adversarial Networks (GANs), Multilayer Perceptron (MLP) neural networks, Transformer Neural Networks, and Diffusion-based Neural Networks. For example, a Convolutional Neural Network (CNN) is a feedforward artificial neural network. A CNN may comprise a collection of artificial neurons, each possessing a receptive field (e.g., a localized region of the input space) and collectively tiling the input space. RNNs work on the principle of storing the layer's output and feeding that output back to the input to help predict the layer's outcome. A GAN is a generative neural network that learns patterns in the input data so that the neural network model can generate new synthetic outputs, which may reasonably come from the original dataset. A GAN may comprise two neural networks operating together: a generative neural network that generates the synthetic output and a discriminative neural network that evaluates the authenticity of the output. In an MLP neural network, data is fed into the input layer, and one or more hidden layers provide an abstraction level to the data. The output layer can then be predicted based on this abstract data.
[0149] Deep learning (DL) is an example of machine learning techniques and can be considered a subset of ML. Many DL methods are based on neural networks, such as RNNs or CNNs, and utilize multiple layers. Using multiple layers in a deep neural network allows for the progressive extraction of higher-level features from a given raw data input. For example, the output of the first layer of artificial neurons becomes the input of the second layer, the output of the second layer becomes the input of the third layer, and so on. The layers located between the input and output of the entire deep neural network are often called hidden layers. Hidden layers learn (e.g., are trained) by transforming intermediate inputs from previous layers into slightly more abstract and complex representations that can be provided to subsequent layers until the final or desired representation is obtained as the final output of the deep neural network.
[0150] As noted above, neural networks are examples of machine learning systems and can include an input layer, one or more hidden layers, and an output layer. Data is provided from input nodes in the input layer, processed by hidden nodes in one or more hidden layers, and output is produced by output nodes in the output layer. Deep learning networks typically include multiple hidden layers. Each layer of a neural network can include a feature map or activation map, which can include artificial neurons (or nodes). Feature maps can include filters, kernels, etc. Nodes can include one or more weights used to indicate the importance of nodes in one or more layers. In some cases, deep learning networks may have a series of many hidden layers, where earlier layers are used to determine simple and low-level properties of the input, and later layers build a hierarchy of more complex and abstract properties.
[0151] Deep learning architectures can learn hierarchical structures of features. For example, if presented with visual data, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to recognize spectral power at specific frequencies. The second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to recognize common visual objects or spoken phrases.
[0152] Deep learning architectures perform particularly well when applied to problems with a natural hierarchical structure. For example, the classification of motorized vehicles can benefit from first learning to identify features such as wheels, windshields, and others. These features can then be combined in different ways at higher levels to identify cars, trucks, and airplanes.
[0153] Neural networks can be designed to have multiple connectivity patterns. In feedforward networks, information is passed from lower layers to higher layers, where each neuron in a given layer communicates with neurons in higher layers. As described above, hierarchical representations can be built in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be passed to another neuron in the same layer. Recurrent architectures can help identify patterns across more than one block of input data that is sequentially delivered to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when the recognition of higher-level concepts can aid in discerning specific lower-level features of the input.
[0154] Figure 13 Two sets of images (1300) are provided, illustrating the forward diffusion process (which is fixed) and the reverse diffusion process (which is learned) of the diffusion model. For example... Figure 13 As shown in the forward diffusion process, noise 1303 is gradually added to the first set of images 1302 at different time steps in a total of T time steps (e.g., forming a Markov chain), thereby generating a series of noisy samples X1 to X2. T .
[0155] From a training perspective, the diffusion model acquires an image and slowly adds noise to it to destroy the information within the image. In some respects, the noise is Gaussian noise. Each time step can correspond to... Figure 13 Each consecutive image in the first set of images 1302 shown. Figure 13The initial image X0 is an image of a cat. Noise 1303 is added to each image (corresponding to noisy samples X1 to X). T This causes the pixels in each image to gradually spread out until the final image (corresponding to sample X) is reached. T This essentially matches the noise distribution. For example, by adding noise, as the time step increases, each data sample X1 to X... T Gradually losing its distinguishable features, it eventually leads to the final sample X T Equivalent to the target noise distribution, such as a unit variance, zero Gaussian distribution. .
[0156] The second set of images, 1304, illustrates the reverse diffusion process, where X... T It is the starting point for noisy images (e.g., images with Gaussian noise). The diffusion model can be trained to reverse the diffusion process (e.g., by training model p). θ- (x t-1 | x t This generates new data. In some respects, the diffusion model can be trained by finding the inverse Markov transformation that maximizes the likelihood of the training data. By traversing backward along the time-step chain, the diffusion model can generate new data. For example, as... Figure 13 As shown, inverse diffusion is then performed to generate X0 as an image of the cat. In other cases, the input and output data may vary depending on the task for which the diffusion model was trained.
[0157] As noted above, the diffusion model is trained to denoise or restore the original image X0 in a progressive process, as shown in the second set of images 1304. In some respects, the neural network of the diffusion model can be trained to denoise or restore the original image X0 in a progressive process. t-1 Restore X in the case t As shown in the following example equation:
[0158]
[0159] The diffusion nucleus can be defined as follows:
[0160] definition
[0161] Sampling can be defined as follows:
[0162]
[0163] In some cases, Value scheduling (also known as noise scheduling) is designed to make and .
[0164] The diffusion model runs iteratively to progressively generate the input image X0. In one example, the model may have twenty steps. However, in other examples, the number of steps can vary.
[0165] Figure 14 This is an illustration of how a diffusion model can be used to distribute diffusion data from initial data to noise in the forward diffusion direction, based on some aspects. Note that the initial data q(X0) is detailed in the initial stage of the diffusion process. An illustrative example of the data q(X0) is... Figure 13 The initial image of the cat is shown. As the diffusion model iterates and sampling noise is added to the data iteratively from t=0 to t=T, as... Figure 14 As shown, the data becomes noisier and may eventually result in pure noise (e.g., in q(X)). T ) place). Figure 14 The example illustrates the progress of the data and how the data spreads along with noise during the forward diffusion process.
[0166] In some respects, the distribution of diffuse data (e.g., such as...) Figure 14 (As shown) can be as follows:
[0167] .
[0168] In the above equation, Indicates the distribution of diffusion data. Indicates the joint distribution. This represents the distribution of the input data, and It is a diffusion kernel. In this respect, the model can be improved by first sampling... And then sample To sample (This can be called ancestor sampling). The diffusion kernel takes input and returns a vector or other data structure as output.
[0169] The following is an overview of the training and sampling algorithms for the diffusion model. The training algorithm may include the following steps:
[0170] repeat
[0171]
[0172]
[0173]
[0174] Perform gradient descent steps on the following expression
[0175]
[0176] Until convergence
[0177] The sampling algorithm may include the following steps:
[0178]
[0179] for
[0180]
[0181]
[0182] Finish
[0183] return
[0184] Figure 15 This is a diagram illustrating a U-Net architecture 1500 for a diffusion model, based on some aspects. An initial image 1502 (e.g., of a cat) is provided to the U-Net architecture 1500, which includes a series of Residual Network (ResNet) blocks and self-attention layers to represent the network. The U-Net architecture 1500 also includes a fully connected layer 1508. In some cases, the time representation 1510 can be a sinusoidal position embedding or a random Fourier feature. Noise output 1506 from the forward diffusion process is also shown.
[0185] U-Net architecture 1500 includes, for example, Figure 15 The contraction path 1504 and expansion path 1505 are shown, giving it a U-shaped architecture. The contraction path 1504 can be a convolutional network comprising repeated convolutional layers (which apply convolutional operations), each followed by a rectified linear unit (ReLU) and max-pooling operation. When processing an image (e.g., image 1502) during the contraction path 1504, the spatial information of image 1502 is reduced as features are generated. The expansion path 1505 combines features and spatial information through a series of up-convolutions and cascading with high-resolution features from the contraction path 1504. Some layers can be self-attention layers, which explicitly model complete contextual information by leveraging global interactions between semantic features at the encoder ends.
[0186] Figure 16This is an exemplary example of a neural network 1600 (e.g., a deep learning neural network) that can be used to implement machine learning-based image generation, feature segmentation, implicit neural representation generation, rendering, classification, object detection, image recognition (e.g., face recognition, object recognition, scene recognition, etc.), feature extraction, authentication, gaze detection, gaze prediction, and / or automation. For example, neural network 1600 can be an example of or can implement generative machine learning models 120, 121, 606, 607, 806, and / or 807.
[0187] Input layer 1602 includes input data. In an exemplary example, input layer 1602 may include data representing images 114, 208, 308, 408, 508, 604, 708, 716, 802, 908, and / or 916. Neural network 1600 includes multiple hidden layers 1606a, 1606b through 1606n. Hidden layers 1606a, 1606b through 1606n comprise “n” hidden layers, where “n” is an integer greater than or equal to one. Multiple hidden layers can be made to include as many layers as needed for a given application. Neural network 1600 also includes an output layer 1604, which provides the output produced by the processing performed by hidden layers 1606a, 1606b through 1606n. In an exemplary example, output layer 1604 may provide images 122, 216, 316, 416, 516, 608, 722, 808 and / or 922.
[0188] The neural network 1600 may be or may include a multi-layer neural network with interconnected nodes. Each node may represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains the information while processing it. In some cases, the neural network 1600 may include a feedforward network, in which case there are no feedback connections in which the network's output is fed back into itself. In some cases, the neural network 1600 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read.
[0189] Information can be exchanged between nodes through node-to-node interconnects between layers. Nodes in input layer 1602 can activate the node set in the first hidden layer 1606a. For example, as shown, each input node in input layer 1602 is connected to each node in the first hidden layer 1606a. Nodes in the first hidden layer 1606a can transform the information of each input node by applying an activation function to the input node information. The information derived from this transformation can then be passed to nodes in the next hidden layer 1606b, activating those nodes, which can then perform their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of hidden layer 1606b can then activate nodes in the next hidden layer, and so on. Finally, the output of hidden layer 1606n can activate one or more nodes in output layer 1604, providing the output at those nodes. In some cases, although a node in neural network 1600 (e.g., node 1608) is shown as having multiple output lines, the node has a single output and is shown as all lines output from the node representing the same output value.
[0190] In some cases, each node or the interconnection between nodes may have weights, which are a set of parameters derived from the training of the neural network 1600. Once the neural network 1600 is trained, it can be called a trained neural network, which can be used to perform one or more operations. For example, the interconnection between nodes may represent a piece of information about the nodes learned in the interconnection. The interconnection may have tunable numerical weights that can be tuned (e.g., based on the training dataset), thereby allowing the neural network 1600 to adapt to the input and learn as more and more data is processed.
[0191] The neural network 1600 can be pre-trained to process features from the data in the input layer 1602 using different hidden layers 1606a, 1606b to 1606n, so as to provide an output through the output layer 1604. In an example where the neural network 1600 is used to identify features in an image, the neural network 1600 can be trained using training data that includes both images and labels, as described above. For example, training images can be input into the network, where each training image has a label indicating features in the image (for feature segmentation machine learning systems) or a label indicating the category of activity in each image. In an example where object classification is used for illustrative purposes, the training images may include images of the number 2, in which case the label of the image could be [0 0 1 0 0 0 0 0 0 0].
[0192] In some cases, the Neural Network 1600 can use a training process called backpropagation to adjust the weights of its nodes. As noted above, the backpropagation process can include forward pass, loss function, back pass, and weight update. For each training iteration, forward pass, loss function, back pass, and parameter update are performed. For each set of training images, this process can be repeated up to a certain number of iterations until the Neural Network 1600 is trained well enough to accurately tune the weights of each layer.
[0193] For an example of identifying objects in an image, the forward pass may include passing a training image through a neural network 1600. The weights are initially randomized before training the neural network 1600. As an illustrative example, the image may include a numerical array representing the pixels of the image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or lightness and two chroma components, etc.).
[0194] As noted above, for the first training iteration of a neural network 1600, the output may include values due to the weights being randomly selected during initialization without prioritizing any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values for each class may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). Using the initial weights, the neural network 1600 cannot determine low-level features and therefore cannot make an accurate determination of what the object's classification might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined, such as cross-entropy loss. Another example of a loss function is mean squared error (MSE), which is defined as... The loss can be set to equal E. total The value of .
[0195] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training label. The Neural Network 1600 performs backpropagation by determining which inputs (weights) contribute most to the network's loss and can adjust the weights to reduce and eventually minimize the loss. The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) can be calculated to determine the weights that contribute most to the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. A weight update can be represented as... Where w represents the weight, wi Let represent the initial weights, and η represent the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, while a lower value indicates smaller weight updates.
[0196] Neural Network 1600 can include any suitable deep network. An example includes a Convolutional Neural Network (CNN), which includes an input layer and an output layer, with multiple hidden layers between them. The hidden layers of a CNN include a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. Neural Network 1600 can include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), etc.
[0197] Figure 17 This is an exemplary example of a Convolutional Neural Network (CNN) 1700. The input layer 1702 of the CNN 1700 includes data representing an image or frame. For example, the data could include a numerical array representing pixels of an image, where each number in the array includes a value from 0 to 255 describing the pixel intensity at that location in the array. Using the previous example from above, the array could include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or lightness and two chroma components, etc.). The image can be passed through a convolutional hidden layer 1704, an optional non-linear activation layer, a pooling hidden layer 1706, and a fully connected layer 1708 (which may be hidden) to obtain an output at the output layer 1710. Although... Figure 17 The diagram shows only one hidden layer in each hidden layer, but those skilled in the art will understand that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in the CNN 1700. As previously described, the output may indicate a single category of an object, or may include probabilities that best describe the category of an object in an image.
[0198] The first layer of CNN 1700 can be a convolutional hidden layer 1704. The convolutional hidden layer 1704 analyzes the image data input to layer 1702. Each node in the convolutional hidden layer 1704 is connected to a region of the input image called a receptive field (pixel). The convolutional hidden layer 1704 can be thought of as one or more filters (each filter corresponding to a different start or feature map), where each convolutional iteration of the filter is a node or neuron in the convolutional hidden layer 1704. For example, the region of the input image covered by the filter at each convolutional iteration will be the receptive field of the filter. In an exemplary example, if the input image consists of a 28×28 array and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in the convolutional hidden layer 1704. Each connection between a node and its receptive field learns weights and, in some cases, learns an overall bias, such that each node learns to analyze its specific local receptive field in the input image. Each node in the convolutional hidden layer 1704 will have the same weights and biases (called shared weights and shared biases). For example, the filter has a weight (digital) array and the same depth as the input. For the image frame example, the filter would have a depth of 3 (based on the three color components of the input image). An exemplary example of the filter array size is 5×5×3, corresponding to the size of the receptive field of a node.
[0199] The convolutional property of the convolutional hidden layer 1704 is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filter of the convolutional hidden layer 1704 may begin at the top left corner of the input image array and may convolve around the input image. As noted above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 1704. In each convolutional iteration, the value of the filter is multiplied by the corresponding number of original pixel values of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values at the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain the sum of that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 1704. For example, the filter may move a step size (called stride) to the next receptive field. The stride may be set to 1 or any other suitable amount. For example, if the stride is set to 1, the filter will move 1 pixel to the right in each convolutional iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result at that location, thus determining a sum value for each node of the convolutional hidden layer 1704.
[0200] The map construction from the input layer to the convolutional hidden layer 1704 is called an activation map (or feature map). An activation map includes values for each node representing the filter results at each location in the input volume. Activation maps may include arrays containing various sums of values produced by the filter for each iteration of the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 1704 may include several activation maps to identify multiple features in the image. Figure 17 The example shown includes three activation maps. Using these three activation maps, the convolutional hidden layer 1704 can detect three different types of features, each of which is detectable across the entire image.
[0201] In some examples, nonlinear hidden layers can be applied after the convolutional hidden layer 1704. Nonlinear layers can be used to introduce nonlinearity into a system that has already computed linear operations. An exemplary example of a nonlinear layer is the Corrected Linear Unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0, x) to all values in the input volume, which changes all negative activations to 0. Therefore, ReLU can increase the nonlinearity of the CNN 1700 without affecting the receptive field of the convolutional hidden layer 1704.
[0202] A pooling hidden layer 1706 can be applied after the convolutional hidden layer 1704 (and, in use, after the non-linear hidden layer). The pooling hidden layer 1706 is used to simplify the information in the output of the convolutional hidden layer 1704. For example, the pooling hidden layer 1706 can take each activation map output from the convolutional hidden layer 1704 and use a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. The pooling hidden layer 1706 uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. Pooling functions (e.g., max pooling filters, L2 norm filters, or other suitable pooling filters) are applied to each activation map included in the convolutional hidden layer 1704. Figure 17 In the example shown, three pooling filters are used to convolve the three activation maps in the hidden layer 1704.
[0203] In some examples, max pooling can be used by applying a max pooling filter (e.g., of size 2×2) with a stride (e.g., equal to the dimension of the filter, such as a stride of 2) to the activation map output from convolutional hidden layer 1704. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer summarizes a region of 2×2 nodes from the previous layer (each node being a value in the activation map). For example, four values (nodes) in the activation map will be analyzed by the 2×2 max pooling filter at each iteration of the filter, with the maximum of the four values being output as the "maximum" value. If such a max pooling filter is applied to an activation filter of 24×24 nodes from convolutional hidden layer 1704, the output from pooling hidden layer 1706 will be an array of 12×12 nodes.
[0204] In some examples, L2 norm pooling filters may also be used. L2 norm pooling filters involve calculating the square root of the sum of squares of the values in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as done in max pooling), and using the calculated value as the output.
[0205] Pooling functions (e.g., max pooling, L2 norm pooling, or other pooling functions) determine whether a given feature is found anywhere in a region of the image and discard the exact location information. This can be done without affecting the results of feature detection because once a feature has been found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) offers the benefit of having far fewer pooling features, thus reducing the number of parameters required in subsequent layers of the CNN 1700.
[0206] The final connection in the network is a fully connected layer, which connects each node from the pooling hidden layer 1706 to each output node in the output layer 1710. Using the example above, the input layer comprises 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 1704 comprises 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for filtering) to three activation maps, and the pooling hidden layer 1706 comprises 3×12×12 hidden feature nodes based on applying a max-pooling filter to a 2×2 region in each of the three feature maps. Extending this example, the output layer 1710 may comprise ten output nodes. In such an example, each node of the 3×12×12 pooling hidden layer 1706 is connected to each node of the output layer 1710.
[0207] The fully connected layer 1708 takes the output of the previous pooling hidden layer 1706 (which should represent an activation map of high-level features) and determines the features most relevant to a particular class. For example, the fully connected layer 1708 can determine the high-level features most relevant to a particular class and may include weights (nodes) for those high-level features. The product between the weights of the fully connected layer 1708 and the pooling hidden layer 1706 can be computed to obtain the probabilities for different classes. For example, if the CNN 1700 is used to predict that the object in an image is a person, there will be high values in the activation map representing the high-level features of a person (e.g., two legs, a face at the top of the object, two eyes at the top left and top right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).
[0208] In some examples, the output from output layer 1710 may include an M-dimensional vector (M=10 in the previous example). M indicates the number of classes the CNN 1700 must choose from when classifying objects in an image. Other example outputs are also available. Each number in the M-dimensional vector represents the probability that an object belongs to a certain class. In an exemplary example, if the 10-dimensional output vector represents objects of ten different classes as [0 0 0.05 0.8 0 0.15 0 0 0 0], then the vector indicates a 5% probability that the image is an object of the third class (e.g., a dog), an 80% probability that the image is an object of the fourth class (e.g., a person), and a 15% probability that the image is an object of the sixth class (e.g., a kangaroo). The probability of a class can be considered as the confidence level that an object is part of that class.
[0209] Figure 18 An example computing device architecture 1800 is illustrated, illustrating example computing devices capable of implementing the various technologies described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device within a vehicle), or other devices. For example, computing device architecture 1800 may include, implement, or be included in any or all of device 100, device 600, device 800, and / or image processing system 1100. Additionally or alternatively, computing device architecture 1800 may be configured to perform process 1000 and / or other processes described herein.
[0210] The components of computing device architecture 1800 are shown to communicate electrically with each other using a connection 1812 (such as a bus). The example computing device architecture 1800 includes a processing unit (CPU or processor) 1802 and a computing device connection 1812 that couples various computing device components, including computing device memories 1810 (such as read-only memory (ROM) 1808 and random access memory (RAM) 1806), to the processor 1802.
[0211] The computing device architecture 1800 may include a cache of high-speed memory that is directly connected to, very close to, or integrated as part of the processor 1802. The computing device architecture 1800 may copy data from memory 1810 and / or storage device 1814 to cache 1804 for fast access by the processor 1802. In this way, the cache can provide performance improvements by avoiding latency for the processor 1802 while waiting for data. These and other modules may control or be configured to control the processor 1802 to perform various actions. Additional computing device memory 1810 may also be used. Memory 1810 may include various different types of memory with different performance characteristics. The processor 1802 may include any general-purpose processor and hardware or software services configured to control the processor 1802 (such as services 11816, 1818, and 31820 stored in storage device 1814), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 1802 may be a self-contained system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors can be symmetric or asymmetric.
[0212] To enable user interaction with the computing device architecture 1800, input device 1822 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. Output device 1824 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, projector, television, speaker equipment, etc. In some instances, multi-mode computing devices allow users to provide multiple types of input to communicate with computing device architecture 1800. Communication interface 1826 typically controls and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0213] Storage device 1814 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as magnetic tape cassettes, flash memory cards, solid-state memory devices, digital versatile optical discs, magnetic tape cartridges, random access memory (RAM) 1806, read-only memory (ROM) 1808, and hybrid forms thereof. Storage device 1814 may include services 1816, 1818, and 1820 for controlling processor 1802. Other hardware or software modules are envisioned. Storage device 1814 may be connected to computing device connection 1812. In one aspect, a hardware module performing a specific function may include software components stored in a computer-readable medium connected to necessary hardware components, such as processor 1802, connection 1812, output device 1824, etc., to perform that function.
[0214] With reference to a given parameter, property, or condition, the term "substantially" may mean that a person skilled in the art would understand that a given parameter, property, or condition is satisfied with a small degree of variance (such as, for example, within acceptable manufacturing tolerances). For example, depending on the specific parameter, property, or condition that is substantially satisfied, the parameter, property, or condition may be satisfied at least 90%, at least 95%, or even at least 99%.
[0215] Various aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet, laptop, vehicle, drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although devices having or coupled to a light projector are described below, various aspects of this disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.
[0216] The term "device" is not limited to one or a specific number of physical objects (such as a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that implement at least some parts of this disclosure. Although the following description and examples use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a specific configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or specific aspects. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. Although the following description and examples use the term "system" to describe various aspects of this disclosure, the term "system" is not limited to a specific configuration, type, or number of objects.
[0217] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks comprising devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring aspects.
[0218] Various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but it may have additional steps not included in the diagrams. Processes can correspond to methods, functions, procedures, subroutines, subroutines, etc. When a process corresponds to a function, its termination may correspond to the function returning to its calling function or the main function.
[0219] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, cause or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.
[0220] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media can include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, magnetic disks or optical disks, USB devices equipped with non-volatile memory, network storage devices, any suitable combinations thereof, etc. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments can be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, independent variables, parameters, data, etc., can be transmitted, forwarded, or sent through any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0221] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as power consumption, carrier signals, electromagnetic waves, and the signals themselves.
[0222] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.
[0223] Instructions, media for delivering such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.
[0224] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that the inventive concepts can be implemented and employed in a variety of other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. The various features and aspects of the applications described above can be used individually or in combination. Furthermore, without departing from the broader scope of this specification, aspects can be used in any number of environments and applications beyond those described herein. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.
[0225] Those skilled in the art will understand that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“≤”) and greater than or equal to (“≥”) symbols without departing from the scope of this description.
[0226] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0227] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0228] Claim language or other languages that state "at least one of" and / or "one or more of" in a set indicate that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language stating "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the language of a claim stating "at least one of A and B" or "at least one of A or B" may refer to A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.
[0229] Claim language or other languages that state "at least one processor, the at least one processor being configured to," "at least one processor being configured to," "one or more processors, the one or more processors being configured to," "one or more processors being configured to," etc., indicate that one or more processors (in any combination) are capable of performing associated operations. For example, claim language stating "at least one processor, the at least one processor being configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each assigned a specific subset of tasks of operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, claim language stating "at least one processor, the at least one processor being configured to: X, Y, and Z" may mean that any single processor can perform only a subset of operations X, Y, and Z.
[0230] When referring to one or more elements that perform functions (e.g., steps of a method), one element may perform all functions, or more than one element may jointly perform these functions. When more than one element jointly performs these functions, each function does not need to be performed by every single element (e.g., different functions may be performed by different elements), and / or each function does not need to be performed by only one element as a whole (e.g., different elements may perform different sub-functions of a function). Similarly, when referring to one or more elements configured to cause another element (e.g., a device) to perform functions, one element may be configured to cause another element to perform all functions, or more than one element may be jointly configured to cause another element to perform these functions.
[0231] When referring to an entity that performs or is configured to perform functions (e.g., steps of a method) (e.g., any entity or device described herein), the entity may be configured to cause one or more elements (individually or collectively) to perform those functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more of those functions, and / or any combination thereof. When referring to an entity that performs functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform those functions collectively. When the entity is configured to cause more than one component to perform those functions collectively, each function does not need to be performed by every single component (e.g., different functions may be performed by different components), and / or each function does not need to be performed by only one component as a whole (e.g., different components may perform different sub-functions of a function).
[0232] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.
[0233] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.
[0234] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.
[0235] The exemplary aspects of this disclosure include:
[0236] Aspect 1. An apparatus for generating image content, the apparatus comprising: a user interface configured to: display an image of a field of view; and receive user input indicating a desired change to the image, wherein the desired change to the image includes a change to the field of view; and at least one processor configured to: provide at least a portion of the image and the indication of the desired change as input to a generative machine learning model; and obtain a modified image from the generative machine learning model, wherein the modified image includes at least a portion of the image of the field of view and generated pixels outside the field of view.
[0237] Aspect 2. The apparatus according to aspect 1, wherein the user interface is configured to interpret gestures to receive the user input.
[0238] Aspect 3. The apparatus according to aspect 2, wherein the user interface includes a touchscreen configured to sense touch, and wherein the user interface is configured to interpret the touch as the gesture.
[0239] Aspect 4. The apparatus according to aspect 3, wherein the user interface is configured to interpret pinch touch gestures as indications that the desired change includes expanding the field of view.
[0240] Aspect 5. The apparatus according to any one of Aspects 3 or 4, wherein the user interface is configured to interpret drag-and-drop touch gestures as indications that the desired change includes panning to change the field of view.
[0241] Aspect 6. The apparatus according to any one of Aspects 3 to 5, wherein the user interface is configured to interpret a rotational touch gesture as an instruction that the desired change includes rotating the field of view.
[0242] Aspect 7. The apparatus according to any one of Aspects 2 to 6, wherein the user interface includes a camera configured to capture an image of a user, and wherein the user interface is configured to interpret the user's posture as the gesture.
[0243] Aspect 8. The apparatus according to aspect 7, wherein the camera includes at least one of: an active depth camera; an infrared (IR) camera; a red-green-blue (RGB) camera; a stereo camera; or an eye-facing camera.
[0244] Aspect 9. The apparatus according to any one of Aspects 7 or 8, wherein the user interface is configured to interpret at least one of: interpreting a gesture to receive the user input; or interpreting an eye position as a gesture to receive the user input.
[0245] Aspect 10. The apparatus according to any one of Aspects 1 to 9, wherein the altered image comprises the generated pixels on at least both sides of at least a portion of the image in the field of view.
[0246] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein the altered image comprises the generated pixels on at least one side of at least a portion of the image in the field of view.
[0247] Aspect 12. The apparatus according to any one of Aspects 1 to 11, wherein the at least one processor is further configured to generate a final image by rotating or trimming at least one of the altered image based on the change in the field of view.
[0248] Aspect 13. The apparatus according to any one of Aspects 1 to 12, wherein the at least one processor is further configured to rotate or crop the image to obtain the at least portion of the image, for provision to the generative machine learning model.
[0249] Aspect 14. The apparatus according to any one of Aspects 1 to 13, wherein: the image includes a first image; the apparatus further includes a first camera configured to capture the first image; the apparatus further includes a second camera configured to capture a second image; and the at least one processor is configured to provide at least a portion of the first image and at least a portion of the second image as input to the generative machine learning model.
[0250] Aspect 15. The apparatus according to aspect 14, wherein the first camera has a first focal length and the second camera has a second focal length different from the first focal length.
[0251] Aspect 16. The apparatus according to any one of Aspects 1 to 15, wherein: the image comprises a first image of a scene captured at a first time; and the at least one processor is further configured to provide at least a portion of the first image and at least a portion of the second image of the scene captured at a second time as input to the generative machine learning model.
[0252] Aspect 17. The apparatus according to aspect 16, wherein: the field of view includes a first field of view of the scene; the first image belongs to the first field of view of the scene; and the second image belongs to a second field of view of the scene.
[0253] Aspect 18. The apparatus of any one of claims 16 or 17, wherein the at least one processor is further configured to determine the use of the second image based on the desired change to the first image.
[0254] Aspect 19. The apparatus according to any one of Aspects 1 to 18, wherein: before the at least one processor obtains the altered image, the at least one processor is configured to obtain additional pixels and cause the user interface to display the image of the field of view and the additional pixels outside the field of view; and in response to the at least one processor obtaining the altered image, the at least one processor is configured to cause the user interface to display the altered image.
[0255] Aspect 20. The apparatus according to aspect 19, wherein the additional pixel is blurred.
[0256] Aspect 21. The apparatus according to any one of aspects 1 to 20, wherein the user interface is further configured to display the changed image.
[0257] Aspect 22. The apparatus according to any one of aspects 1 to 21, wherein the at least one processor is further configured to cause the changed image to be displayed by the user interface, stored in a memory, analyzed, or transmitted.
[0258] Aspect 23. The apparatus according to any one of aspects 1 to 22, wherein the at least one processor is further configured to smooth pixels at the edges between at least a portion of the image of the field of view and the generated pixels outside the field of view.
[0259] Aspect 24. The apparatus according to any one of aspects 1 to 23, wherein the at least one processor implements the generative machine learning model.
[0260] Aspect 25. The apparatus according to any one of aspects 1 to 24, the apparatus further comprising a communication interface configured to: transmit at least a portion of the image and the indication of the desired change to a computing device implementing the generative machine learning model; and receive the changed image from the computing device.
[0261] Aspect 26. The apparatus according to any one of aspects 1 to 25, wherein the apparatus further comprises an orientation sensor configured to sense the orientation of the apparatus, and wherein the user interface is configured to interpret rotation of the apparatus as an indication that the desired change includes rotating the field of view.
[0262] Aspect 27. The apparatus according to any one of Aspects 1 to 26, wherein one of the following conditions exists: the image is displayed by the user interface in a vertical format, and the altered image includes the generated pixels on at least one side of at least a portion of the image in the field of view; or the image is displayed by the user interface in a horizontal format, and the altered image includes the generated pixels at at least one above or below at least a portion of the image in the field of view.
[0263] Aspect 28. The apparatus according to any one of aspects 1 to 27, wherein the altered image includes the generated pixel at at least one corner of the altered image.
[0264] Aspect 29. A method for generating image content, the method comprising: displaying an image of a field of view at a user interface; receiving user input at the user interface indicating a desired change to the image, wherein the desired change to the image includes a change to the field of view; providing at least a portion of the image and the indication of the desired change as input to a generative machine learning model; and obtaining a modified image from the generative machine learning model, wherein the modified image includes at least a portion of the image of the field of view and generated pixels outside the field of view.
[0265] Aspect 30. The method according to aspect 29, wherein receiving the user input includes interpreting at least one of the following as the user input: touch gesture; gesture; or eye position.
[0266] Aspect 31. The apparatus according to any one of Aspects 1 to 28, wherein: the user input indicates two or more desired changes to the field of view of the image; and the at least one processor is configured to provide the instructions for the two or more desired changes to the generative machine learning model.
[0267] Aspect 32. The apparatus according to any one of Aspects 1 to 28 or 31, wherein: the user input indicates two or more desired changes to the field of view of the image; and the at least one processor is configured to provide two or more instructions of the two or more corresponding desired changes to the generative machine learning model.
[0268] Aspect 33. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing the at least one processor, when executed, to perform any one of aspects 29 to 30.
[0269] Aspect 34. An apparatus for providing virtual content for display, the apparatus comprising one or more components for performing operations according to any one of Aspects 29 to 30.
Claims
1. An apparatus for generating image content, the apparatus comprising: A user interface configured to display an image of the field of view; as well as Receive user input indicating a desired change to the image, wherein the desired change to the image includes a change to the field of view; and At least one processor, said at least one processor being configured to: Provide at least a portion of the image and an indication of the desired change as input to the generative machine learning model; An altered image is obtained from the generative machine learning model, wherein the altered image includes at least a portion of the image within the field of view and generated pixels outside the field of view.
2. The apparatus of claim 1, wherein the user interface is configured to interpret gestures to receive the user input.
3. The apparatus of claim 2, wherein the user interface includes a touchscreen configured to sense touch, and wherein the user interface is configured to interpret the touch as the gesture.
4. The apparatus of claim 3, wherein the user interface is configured to interpret pinch touch gestures as indications that the desired change includes expanding the field of view.
5. The apparatus of claim 3, wherein the user interface is configured to interpret drag-and-drop touch gestures as indications that the desired change includes panning to change the field of view.
6. The apparatus of claim 3, wherein the user interface is configured to interpret rotational touch gestures as instructions to change the desired field of view.
7. The apparatus of claim 2, wherein the user interface includes a camera configured to capture an image of a user, and wherein the user interface is configured to interpret the user's posture as the gesture.
8. The apparatus of claim 7, wherein the camera comprises at least one of the following: Active depth camera; Infrared (IR) camera; Red, green, and blue (RGB) camera; 3D camera; or A camera facing the eyes.
9. The apparatus of claim 7, wherein the user interface is configured to interpret at least one of the following: Interpreting gestures to receive the user input; or The eye position is interpreted as a gesture to receive the user input.
10. The apparatus of claim 1, wherein the altered image comprises the generated pixels on at least both sides of at least a portion of the image in the field of view.
11. The apparatus of claim 1, wherein the altered image comprises the generated pixels on at least one side of at least a portion of the image in the field of view.
12. The apparatus of claim 1, wherein the at least one processor is further configured to generate a final image by rotating or trimming at least one of the altered image based on the change in the field of view.
13. The apparatus of claim 1, wherein the at least one processor is further configured to rotate or crop the image to obtain the at least portion of the image, for provision to the generative machine learning model.
14. The apparatus according to claim 1, wherein: The image includes a first image; The device also includes a first camera configured to capture the first image; The device also includes a second camera configured to capture a second image; as well as The at least one processor is configured to provide at least a portion of the first image and at least a portion of the second image as input to the generative machine learning model.
15. The apparatus of claim 14, wherein the first camera has a first focal length, and the second camera has a second focal length different from the first focal length.
16. The apparatus according to claim 1, wherein: The images include a first image of the scene captured in a first-time event; and The at least one processor is further configured to provide at least a portion of the first image and at least a portion of the second image of the scene captured at a second time as input to the generative machine learning model.
17. The apparatus according to claim 16, wherein: The field of view includes the first field of view of the scene; The first image belongs to the first field of view of the scene; and The second image belongs to the second field of view of the scene.
18. The apparatus of claim 16, wherein the at least one processor is further configured to determine the use of the second image based on the desired change to the first image.
19. The apparatus according to claim 1, wherein: Before the at least one processor obtains the modified image, the at least one processor is configured to obtain additional pixels and cause the user interface to display the image within the field of view and the additional pixels outside the field of view; as well as In response to the at least one processor obtaining the changed image, the at least one processor is configured to cause the user interface to display the changed image.
20. The apparatus of claim 19, wherein the additional pixel is blurred.
21. The apparatus of claim 1, wherein the user interface is further configured to display the changed image.
22. The apparatus of claim 1, wherein the at least one processor is further configured to cause the altered image to be displayed by the user interface, stored in a memory, analyzed, or transmitted.
23. The apparatus of claim 1, wherein the at least one processor is further configured to smooth pixels at the edges between at least a portion of the image of the field of view and the generated pixels outside the field of view.
24. The apparatus of claim 1, wherein the at least one processor implements the generative machine learning model.
25. The apparatus of claim 1, further comprising a communication interface configured to: Send at least a portion of the image and the instruction to change the desired result to a computing device implementing the generative machine learning model; and The modified image is received from the computing device.
26. The apparatus of claim 1, wherein the apparatus further comprises an orientation sensor configured to sense the orientation of the apparatus, and wherein the user interface is configured to interpret rotation of the apparatus as an indication that the desired change includes rotating the field of view.
27. The apparatus of claim 1, wherein one of the following conditions is met: The image is displayed in a vertical format by the user interface, and the altered image includes the generated pixels on at least one side of at least a portion of the image in the field of view; or The image is displayed in landscape format by the user interface, and the altered image includes the generated pixels located above or below at least one portion of the image in the field of view.
28. The apparatus of claim 1, wherein the altered image includes the generated pixel at at least one corner of the altered image.
29. A method for generating image content, the method comprising: Display the image of the field of view at the user interface; The user interface receives user input indicating a desired change to the image, wherein the desired change to the image includes a change to the field of view; Provide at least a portion of the image and an indication of the desired change as input to the generative machine learning model; as well as An altered image is obtained from the generative machine learning model, wherein the altered image includes at least a portion of the image within the field of view and generated pixels outside the field of view.
30. The method of claim 29, wherein receiving the user input comprises interpreting at least one of the following as the user input: Touch gestures; gestures; or Eye position.