Real-time dolly ZOOM effect
The system addresses the challenges of achieving the dolly zoom effect by using subject segmentation and depth sensing in wearable devices to maintain real-time, immersive video creation, overcoming the limitations of conventional techniques.
Patent Information
- Application Number
- PCT/US2023/085757
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-26
AI Technical Summary
Conventional techniques for achieving the dolly zoom effect in videography require complex camera movements and post-production processing, which can be challenging for non-professionals and lack real-time capability.
The described system uses subject segmentation and depth sensing capabilities of wearable devices to provide the dolly zoom effect in real-time by maintaining the initial subject size while moving the camera relative to the subject, using a computer program product and a wearable device with a processor and memory to execute instructions for determining subject image size, removing subject images from the background, and rendering projected images.
This approach allows users to easily achieve the dolly zoom effect in real-time using common devices, enabling immersive and compelling video creation without the need for complex equipment or post-capture processing.
Smart Images

Figure US2023085757_26062025_PF_FP_ABST
Abstract
Description
REAL-TIME DOLLY ZOOM EFFECTTECHNICAL FIELD
[0001] This description relates to videography effects.SUMMARY
[0002] In a general aspect, a computer program product is tangibly embodied on a non-transitory computer-readable storage medium and comprises instructions. When executed by at least one computing device (e.g., by at least one processor of the computing device), the instructions are configured to cause the at least one computing device to determine a size of an initial subject image of a subject in an initial video frame captured by a camera, relative to a background. The instructions, when executed by the at least one computing device, may further cause the at least one computing device to remove subsequent subject images of the subject from the background within subsequent video frames. The instructions, when executed by the at least one computing device, may further cause the at least one computing device to render projected images of the subject, with the size of the initial subject image, within the subsequent video frames and in place of the subsequent subject images, while the camera is moving relative to the subject.
[0003] In another general aspect, a wearable device includes at least one frame for positioning the wearable device on a body of a user, at least one display, at least one processor, and at least one memory storing instructions. When executed, the instructions cause the at least one processor to determine a size of an initial subject image of a subject in an initial video frame captured by a camera, relative to a background. When executed, the instructions cause the at least one processor to remove subsequent subject images of the subject from the background within subsequent video frames. When executed, the instructions cause the at least one processor to render projected images of the subject, with the size of the initial subject image, within the subsequent video frames and in place of the subsequent subject images, while the camera is moving relative to the subject.
[0004] In another general aspect, a method includes determining a size of an initial subject image of a subject in an initial video frame captured by a camera, relative to a background, removing subsequent subject images of the subject from the background within subsequent video frames, and rendering projected images of the subject, with the size of the initial subject image, w ithin the subsequent video frames and in place of the subsequentsubject images, while the camera is moving relative to the subject.
[0005] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] FIG. 1 is a block diagram of a system for achieving the dolly zoom effect using personal devices.
[0007] FIG. 2 is a flowchart illustrating example operations of the system of FIG. 1.
[0008] FIG. 3 illustrates a first example use case scenario.
[0009] FIG. 4 illustrates a second example use case scenario.
[0010] FIG. 5 illustrates a more detailed example implementation.
[0011] FIG. 6 is a flowchart illustrating operations for the example of FIG. 5.
[0012] FIG. 7 is a third person view of a user in an ambient computing environment.
[0013] FIGS. 8A and 8B illustrate front and rear views of an example implementation of a pair of smartglasses.DETAILED DESCRIPTION
[0014] The dolly zoom is a videography technique for imparting dramatic effect in movies or television shows. The term 'dolly’ refers to a camera dolly used to move a camera, so that the ‘dolly zoom’ effect refers to moving (dollying) a camera towards a subject while simultaneously zooming out, or to moving the camera away from the subject while simultaneously zooming in.
[0015] Described systems and techniques enable obtaining the dolly zoom effect in real time, using commonly available wearable and / or personal devices, by instructing a user to move (e.g., walk) towards or away from a subject. The subject is reprojected at a single size (e.g., an initial size is maintained) in subsequent frames, while the background reflects the effects of the user’s movements. Consequently, the user is able to achieve the desired effects easily and in real-time.
[0016] The dolly zoom, also known as "trombone shot" or "contra zoom", was first used in the movie Vertigo by director Alfred Hitchcock, and is therefore also known as the Vertigo effect or the Hitchcock effect. The technique has been used in many movies (e.g., in the movies Jaws and Goodfellas, as described below with respect to FIGS. 3 and 4) to add drama, create tension, convey an emotion, or otherwise facilitate story telling.
[0017] Conventional techniques for achieving the dolly zoom effect involve moving the camera while adjusting a zoom of the camera lens, all while also maintaining a subject in focus. This process can be challenging even for professional filmmakers.
[0018] More recently, post-production processing has been used to achieve the effect. This approach is more attainable and configurable for average users, but has the disadvantage of being unavailable in real time.
[0019] Described techniques use subject segmentation and / or depth sensing capabilities of personal wearable and / or handheld devices to provide the dolly zoom effect in real time. As a result, users may generate desired videos in virtually any context or setting, using commonly available equipment, and without requiring post-capture processing.
[0020] For example, a user with a camera may be instructed to move (e.g., walk) towards or away from a subject in view of the camera. Then, the dolly zoom effect may be provided using a depth map of a field of view (FOV) of the camera, e.g., determined using depth sensors, to determine a first depth of the subject, and then maintaining that depth in subsequent frames. The effect can also be provided in devices that do not have depth sensors using subject segmentation and maintaining an initial subject size in subsequent frames. Put another way, a subject (e.g., a virtual projection of a subject) may be rendered and rerendered at a single size within a series of composite image frames, while a user / camera moves towards / away from the subject.
[0021] Thus, users may achieve the dolly zoom effect in live situations for first- person videography. Resulting videos may be created that provide immersive and compelling effects, for personal or professional use.
[0022] FIG. 1 is a block diagram of a system for achieving the dolly zoom effect using personal devices. In the example of FIG. 1, a head-mounted device (HMD) 102 is illustrated as being worn by a user 104 who is viewing a subject 106 at a depth 108 from the HMD 102. As referenced above, and described in detail, below, the HMD 102 may be configured to provide the dolly zoom effect when capturing video that includes the subject 106.
[0023] For example, the HMD 102 may render a virtual version of the subject 106. while the user 104 moves towards or away from the subject 106. That is, the user 104 may change the physical depth 108 by walking or otherwise moving relative to the subject 106.
[0024] In FIG. 1. a frame 110 represents an image or video frame displayed and / or captured by the HMD 102, which includes the subject 106 as well as any included background 112 that may be present within a given scene or context at a given value of thedepth 108. That is, the subject 106 and the background 112 may be understood to be physically co-present with the user 104, with the subject 106 being positioned at the depth 108 with respect to the user 104. Then, the frame 110 may be understood to represent, for example, a video see-through frame provided by the HMD 102 and positioned by the user 104 to capture the subject 106 and the background 112, as described in more detail below with respect to screen 126 of the HMD 102.
[0025] In other words. FIG. 1 is illustrated to convey that a given background, represented by the background 112, is captured differently as the user 104 reduces or increases the depth 108 (i.e., as the user moves closer to or farther from the subject 106). That is, as the physical depth 108 is reduced, less of the background 112 will be captured but individual captured portions will appear bigger / more prominent (represented by background 112a). As the physical depth is increased, more of the background 112 will be captured while individual captured portions will appear smaller / more distant (represented by background 112b).
[0026] Consequently, as the user 104 changes the depth 108 by moving towards the subject 106, an updated frame 110a will include the reduced background 112a. Conversely, as the user 104 changes the depth 108 by moving away from the subject 106, an updated frame 110b will include the enlarged background 112a. In other words, a captured or imaged background 112 changes dynamically in real-time as the user 104 moves.
[0027] At the same time, using described techniques, the subject 106 may be maintained at a single size within captured images or video frames, even as the depth 108 and the background 112 are changed / changing. For example, the subject 106 in FIG. 1 may represent the physical subject co-present with the user 104 and captured at a first time at the depth 108. Then, if the user moves closer to the subject 106 (reduces the depth 108), a projected subject image 106a represents a rendered or virtual projection of the same subject 106 at the original depth 108 and having a proj ected size that is the same (or approximately the same) size within the frame 110a as in the original frame 110, notwithstanding the fact that the user 104 has moved closer to the subject 106 (reduced the depth 108) when capturing the updated background 112a within the updated frame 110a. Similarly, but conversely, if the user were to move farther from the subject 106 (increase the depth 108), a projected subject image 106b represents a rendered or virtual projection of the same subject 106 at the same depth 108 and the same size within the frame 110b as in the frame 110, notwithstanding the fact that the user 104 has moved farther from the subject 106 (increased the depth 108) when capturing the updated background 112b within the updated frame 110b. As a result, the dollyzoom effect may be achieved and recorded in real-time, with a speed that is related to a speed at which the user 104 moves relative to the subject 106 and the background 112.
[0028] It will be appreciated that the above examples are provided with the frame 110 described as an initial video frame and the frame 110a described as a subsequent frame for purposes of capturing a first type of dolly zoom effect, while the frame 110 is also described as an initial video frame with the frame 110b described as a subsequent frame for purposes of capturing a second t pe of dolly zoom effect. In practice, as described below with respect to FIGS. 2-6, the subsequent frame 110a or the subsequent frame 110b each may represent a plurality of subsequent video frames captured as the user moves in a corresponding direction, for a desired type of dolly zoom effect and for a duration of time chosen by the user 104 that is selected for a desired effect to be achieved.
[0029] The HMD 102 may include, for example, any type of smartglasses or goggles, such as the goggles illustrated in FIG. 5, or the smartglasses shown and described below with respect to FIGS. 7, 8A, and 8B. The HMD may also represent any other type of eyewear, as well as any headset, headband, hat, helmet or any other headwear that may be configured to provide the functionalities described herein.
[0030] The HMD 102 may therefore include any virtual reality (VR), augmented reality (AR), mixed reality (MR), or immersive reality (IR) device, generally referred to herein as an extended reality (XR) device, through which the user 104 may look to view the subject 106. The subject 106 may include any physical (real-world) item or object that the user 104 may wish to view.
[0031] The subject 106 may include, for example, a person or animal. The subject 106 may include two or more items or objects, e.g., that are positioned at (or near) a single depth (e.g.. the depth 108), e.g., as described below with respect to FIG. 4.
[0032] In some implementations, described techniques may be implemented using a handheld or other wearable device, represented in FIG. 1 as a device 102a. For example, the device 102a may represent a smartphone or smartwatch, as described in more detail below, with respect to FIG. 7. In some implementations, the HMD 102 and the device 102a may operate in conjunction with one another to implement the described techniques.
[0033] As shown in the exploded view of FIG. 1, the HMD 102 (and / or the device 102a) may include a processor 114 (which may represent one or more processors), as well as a memory' 116 (which may represent one or more memories (e.g.. non-transitory computer readable storage media)). The processor 114 may represent a Central Processing Unit (CPU) and / or a Graphical Processing Unit (GPU). More detailed examples of the HMD 102 andvarious associated hardware / software resources are provided below, e.g., with respect to FIGS. 7, 8A, and 8B.
[0034] The HMD 102 may include, or have access to, various sensors that may be used to detect, infer, or otherwise determine aspects of the frame 110, e.g., of the subject 106 and / or the background 112. For example, in FIG. 1, the HMD 102 includes a camera 118 and a depth sensor 120. For example, the camera 118 may represent any one or more standard RGB cameras, while the depth sensor 120 may represent any passive or active depth sensor.
[0035] For example, the depth sensor 120 may represent a time-of-flight (ToF) camera or LiDAR sensor. In some implementations, relevant depth data may be determined in whole or in part by the camera 118.
[0036] A depth map generator 122 may be configured to generate a depth map. such as shown in the example of FIG. 5, which captures information characterizing relative depths of detected or rendered objects with respect to a defined perspective or reference point. For example, such a depth map may be used to define or determine the depth 108 to the subject 106, as well as depths of various other elements or aspects of the background 112.
[0037] The depth map generator 122 should be understood to represent and illustrate depth map software stored using the memory 116 and executed using the processor 114, and configured to process depth-related data captured by the camera 118 and / or the depth sensor 120. As described below, the depth map generator 122 may be capable of determining a per- pixel depth of each pixel (and associated object or aspect) in the frame 110. For example, such depth information may be captured and stored as a perspective image containing a depth value instead of a color value in each pixel. A depth map may be generated and stored, e.g., as a 2D array, or as a depth mesh (e.g., a real-time triangulated mesh). Examples of depth map generation and storage are provided in more detail, below, with respect to FIGS. 5 and 6.
[0038] A segmentation manager 124 may be configured to perform object recognition and identification from a captured image frame, such as the frame 110. For example, the segmentation manager 124 may perform color analysis and object recognition using a suitably trained machine learning (ML) model. For example, the segmentation manager 124 may be configured to identify a person as the subject 106. distinct from the background 112. As described in more detail, below, such segmentation may be used to determine a size of the subject 106, e.g., as a number and / or distribution of pixels within the segmented subject and / or relative to an overall frame size of the frame 110.
[0039] A dolly zoom manager 125 may be configured to use data from the camera 118 to provide the dolly zoom effect on a screen 126 of the HMD 102, so that a resultingvideo file may be captured and stored using the memory 116, or otherwise stored or transferred. For example, the screen 126 may represent a display screen such as a glasses lens as in the example of FIGS. 8 A and 8B, or may represent a display of XR goggles. In implementations using the device 102a, the screen 126 may represent a display of the device, e.g., a smartphone display.
[0040] In FIG. 1. the screen 126 illustrates a time lapse of captured image frames 110 for which the described effects are implemented with respect to the subject 106 and the background 112. In the above description, the frame 110 is described as an initial video frame, with either the frame 110a or the frame 110b representing a plurality' of updated frames captured when the user 104 moves in a corresponding direction.
[0041] In the time lapse illustrated in the context of the screen 126. the frame 110a may alternatively be considered to be an initial video frame, so that the frame 110 and then the frame 110b may be considered to represent subsequently captured video frames. Then, in a corresponding implementation viewed from right to left, a first dolly zoom effect is obtained in which the depth 108 progressively increases and a background 112a, 112, 112b progressively grows larger (that is, more background is visible within each successively captured frame, while individual background images / portions thereof appear smaller / more distant) relative to a maintained, singular size of the subject 106a and subsequently projected subjects 106, 106b. A more detailed example of this effect is shown and described with respect to FIG. 3.
[0042] In another implementation, the frame 110b may alternatively be considered to be an initial frame, so that the frame 110 and then the frame 110a may be considered to represent subsequently captured video frames. Then, in a corresponding implementation viewed from left to right, a second dolly zoom effect is obtained in which the depth 108 progressively decreases and a background 112b, 112, 112a progressively grows smaller (that is, less background is visible within each successively captured frame, while individual background images / portions thereof appear larger / closer) relative to a maintained, singular size of the subject 106b and subsequently projected subjects 106, 106a. A more detailed example of this effect is shown and described with respect to FIG. 4.
[0043] With respect to FIG. 1, it will be appreciated that an actual display size of frames 110b, 110, 110a will be dictated by, or rendered in the context of, an available size of the screen 126, and will generally remain constant across the time lapse example of FIG. 1, as shown in FIGS. 3, 4. and 5. In other words, the corresponding backgrounds 112b, 112, 112a will appear to expand or contract relative to the available display / frame size, and relative tothe subject 106, as shown in the more detailed examples of FIGS. 3-5. However, the simplified example of FIG. 1 does not include detailed example background context / details for the background 112, so that the size of the background 112 relative to the subject 106 is conveyed without regard to a frame size or display size of the frames 110b, 110, 110a.
[0044] To obtain the above and related effects, the dolly zoom manager 125 may be configured to provide various functionalities that leverage available information from the camera 118, the depth sensor 120. the depth map generator 122, and / or the segmentation manager 124. The dolly zoom manager 125 may be further configured to provide functionalities related to user-facing interactions that enable the user 104 to obtain and use desired effects in a straightforward, intuitive manner.
[0045] For example, as described in more detail, below, with respect to FIG. 6. the dolly zoom manager 125 may be configured to provide or utilize a suitable user interface (UI) that enables the user 104 to activate and parametrize the desired dolly zoom effect. For example, the user 104 may use such a UI to initiate a dolly zoom effect, identify the subject 106, initiate image capture, and conclude image capture after walking towards / away from the subject 106.
[0046] In FIG. 1, the dolly zoom manager 125 includes a subject identifier 128. The subject identifier 128 may use outputs of the depth map generator 122 and / or the segmentation manager 124 to identify and select the subject 106. For example, the subject identifier 128 may visually highlight the subject 106 on the screen 126 for selection by the user 104. The subject identifier 106 may visually highlight multiple possible subjects and enable selection of one or more of the highlighted subjects by the user 104.
[0047] A depth calculator 130 may be configured to then determine a depth, e.g., the depth 108, of the selected subject 106. Multiple techniques for determining subject depth may be used, alone or in combination, some of which are described in more detail below by way of example. In general, it may be appreciated that the depth map generator 122 represents multiple types of depth map generators that may be used, which produce raw depth data of varying sorts. The depth calculator 130 may thus be understood to represent a common front end component that may be configured to utilize available depth map(s) to provide the types of dolly zoom effects described herein. For example, when the subject 106 includes a human face, the depth calculator 130 may be configured to compute an average depth for each pixel and corresponding facial feature of the face, which may then be used as the depth 108 to maintain a size of the subject face during an ensuing implementation of the dolly zoom effect.
[0048] A rendering engine 132 may be configured to re-proj ect the subj ect 106 at asingle size, corresponding to a single depth, while rendering the background 112 in a way that corresponds to natural reproduction that would occur in conjunction with movements of the user 104 towards / away from the subject 106. For example, an initial captured frame may include a naturally captured representation of the subject 106 and the background 112. Then, as the user 104 moves relative to the subject 106, subsequently captured frames may be modified to construct composite frames in which segmented images of the subject 106 are projected into, or combined with, captured images of the background 112.
[0049] FIG. 2 is a flowchart illustrating example operations of the system of FIG. 1. In the example of FIG. 2, operations 202-206 are illustrated as separate, sequential operations. However, in various example implementations, the operations 202-214 may be implemented in a different order than illustrated, in an overlapping or parallel manner, and / or in a nested, iterative, looped, or branched fashion. Further, various operations or suboperations may be included, omitted, or substituted.
[0050] In FIG. 2, a size of an initial subject image of a subject in an initial video frame captured, relative to a background, by a camera, may be determined (202). For example, the HMD 102 and / or the device 102a may be used by the user 104 to capture an initial subject image of the subject 106 within the frame 110 as an initial video frame, along with the background 112.
[0051] Subsequent subject images of the subject may then be removed from the background within subsequent video frames (204). For example, the subject identifier 128 may identify the subject 106 (i.e., the initial subject image of the subject 106) within an initial frame 110, based on an output of the segmentation manager 124. The rendering engine 132 may be configured to render the subsequent video frames with corresponding actual subject images of the subject 106 removed.
[0052] Then, projected images of the subject may be rendered with a projected size that is based on the size of the initial subject image, within the subsequent video frames and in place of the subsequent subject images, while the camera is moving relative to the subject (206). For example, as rendered on the screen 126 in FIG. 1 and described above, the frame 110 may represent an initialization frame in which the subject 106 and the background 112 are captured normally, and from which a size and depth of the subj ect 106 may be captured in an initial subject image. Then, the user 104 moving away from the subject 106 (increasing the depth 108) may result in the subsequent frame 110b being rendered with the projected subject image 106b having a projected size that is the same size as the original subject image of the subject 106 in the initial frame 110.
[0053] For example, the subsequent frame 110b may be understood to represent a rendered, composite frame that includes a projected subject image 106b of the subject 106 that is maintained at a same size as the subject 106 within the initial frame 110, notwithstanding the change in the depth 108. Similarly, but conversely, if the user 104 were to move toward the subject 106 (decreasing the depth 108), then the frame 110a would represent a subsequent, composite frame with a projected subject image 106a of the subject 106 maintained at a same size as the subject 106 within the initial frame 110. while the background 112a is included in a manner corresponding to the movement of the user 104 towards the subject 106. Thus, the subsequent frame 110a may again be understood to represent a rendered, composite frame, which here includes a projected subject image 106a of the subject 106 that is maintained at a same size as the subject 106 within the initial frame 110, notwithstanding the change in the depth 108.
[0054] Thus, a projected size of the projected subject image 106b may be based on, e.g., may be sized relative to, or in proportion to or as a ratio of, the size of the subject 106. For example, the relative size of the projected size to the original size may be the same or approximately the same, e.g., may be a 1: 1 ratio. In other examples, however, there may be some variation between the proj ected size and the original size, depending on a desired video effect.
[0055] In the examples of FIG. 1, and in many of the following examples, the projected size of the projected subject image 106b is described as being a same / single size as the (initial) image of the subject 106. However, as just referenced, it should be appreciated that projected size(s) of projected subject images may be different than the original subject size. For example, the projected size may be a single projected size throughout remaining frames of a captured video file, which may be defined based on the original subject size. In other examples, there may be some variation of the projected subject size among subsequently captured video frames, depending on specific video effect(s) desired by the user 104.
[0056] FIG. 3 illustrates a first example use case scenario. More specifically, FIG. 3, as noted above, corresponds to an implementation of FIG. 1 in which the time lapse representation illustrated in the context of the screen 126 occurs from right to left. That is, in the example of FIG. 3, the frame 110a is an initial frame and the subject image 106a represents an initial subject image with an initial subject size corresponding to an initial depth, while the frame 110 and then the frame 110b represent subsequent video frames in the context of capturing a first dolly zoom effect.
[0057] FIG. 3 illustrates a reproduction of a first well-known use of the dolly zoom effect, from the movie Jaws. FIG. 3 thus illustrates that a similar effect may be obtained using described techniques.
[0058] In the example of FIG. 3, a subject image 306a is captured in an initial video frame 310a and with an initial background 312a. A second, subsequent video frame 310 thus includes a projected subject image 306, which is maintained at a same or similar size as the original subject image 306a. As shown, the second, subsequent background 312 is larger, in that more of the background elements are visible but individually appear smaller than in the initial video frame 310a.
[0059] A third, subsequent video frame 310b furthers the effect and includes a projected subject image 306b, which is maintained at a same or similar size as the original subject image 306a and the projected subject image 306. As shown, the third, subsequent background 312b is progressively larger, in that, again, more of the background elements are visible but individually appear smaller than in the preceding video frame 310.
[0060] FIG. 4 illustrates a second example use case scenario. More specifically, FIG.4, as noted above, corresponds to an implementation of FIG. 1 in which the time lapse representation illustrated in the context of the screen 126 occurs from left to right. That is, in the example of FIG. 4, the frame 110b is an initial frame and the subject image 106b represents an initial subject image with an initial subject size corresponding to an initial depth, while the frame 110 and then the frame 110a represent subsequent video frames in the context of capturing a second dolly zoom effect.
[0061] FIG. 4 illustrates a reproduction of a second well-known use of the dolly zoom effect, from the movie Goodfellas. FIG. 4 thus illustrates that a similar effect may be obtained using described techniques.
[0062] In the example of FIG. 4, a subject image 406b is captured in an initial video frame 410b and with an initial background 412b. As noted above, the subject 106 may represent 2 or more subjects at a given depth 108, and FIG. 4 provides an example in which the subject 406 includes separate individuals, but at a single camera depth for purposes of the desired effect.
[0063] Then, a second, subsequent video frame 410 includes a projected subject image 406 that includes both individuals, and which is maintained at a same or similar size as the original subject image 406b. As shown, the second, subsequent background 412 is smaller, in that less of the background elements are visible but individually appear larger than in the initial video frame 410b.
[0064] A third, subsequent video frame 410a furthers the effect and includes a projected subject image 406a, which is maintained at a same or similar size as the original subject image 406b and the projected subject image 406. As shown, the third, subsequent background 412a is progressively smaller, in that, again, fewer of the background elements are visible but individually appear larger than in the preceding video frame 410.
[0065] FIG. 5 illustrates a more detailed example implementation. In FIG. 5, HMD in the form of goggles 502 are worn by a user 504. The goggles 502 are used to create a dollyzoom effect with respect to a subject 506 positioned in front of a scene of a background 512 that includes a large globe 507.
[0066] In FIG. 5. the subject 506 and the background 512 are illustrated separately to illustrate that the subject 506 has a subject depth 508a while the background 512 has a background depth 508b. A subject depth map 509a may thus be generated for the subject 506 and the subject depth 508a, while a background depth map 509b may be generated for the background 512 and the background depth 508b.
[0067] As described above, the goggles 502 may then be used to create a dolly zoom effect within a series of video frames 510, including an initial video frame 510a and subsequent video frames 510b. As shown, the initial video frame 510a includes an initial subject image 506a in the context of an initial background 512a. Subsequent video frames 510b include projected subject images 506b that are generated and rendered, using subject segmentation and 3D reprojection, at a single / same size as the subject image 506a. throughout the subsequent video frames 510b.
[0068] Meanwhile, the initial background 512a within the initial video frame 510a is permitted to move in a natural, dynamic manner, as would occur in ty pical situations in which the user 504 might move farther or closer to the subject 506 and the background 512 when using a camera of the goggles 502. In the example of FIG. 5, the initial background 512a is illustrated as getting smaller, so that individual background components, such as the globe 507, appear successively larger within subsequent backgrounds 512b, 512c, 512d of the subsequent video frames 510b, as shown (and as similar to the example of FIG. 4).
[0069] Thus. FIG. 5 illustrates that a first depth (e.g., the subject depth 508a) of the subject 506 in an initial video frame 510a, relative to a camera of the goggles 502, may be determined. Then, a second and subsequent depth(s) of the subject in a second and subsequent video frame(s) 510b may be determined. Projected images 506b of the subject 506 may thus be provided using the first depth and the second depth(s), including generating composite frames 510b of the subsequent video frames that includes the projected images506b of the subject at the first depth 508a.
[0070] FIG. 6 is a flowchart illustrating operations for the example of FIG. 5. In the example of FIG. 6, processing may begin with a user request for a dolly zoom mode (602). For example, the user 104 of FIG. 1 may make an on-screen selection when using any device with, e.g., a video see-through camera and full-screen depth sensing capabilities that provide a depth map with per-pixel depth measurements.
[0071] A subject selection may then be received (604). For example, the HMD 102 of FIG. 1 may highlight one or more potential subjects for inclusion in a dolly zoom effect, depending on a direction of view of the user 104. In some implementations, a subject may be automatically identified and the user 104 may confirm that selection.
[0072] In one or more initial video frames, the selected subject may then be segmented (606), and a depth map may be generated (608). For example, the HMD 102 of FIG. 1 may be used to generate depth maps (such as the depth maps 509a, 509b of FIG. 5). Portrait depth information may be acquired using a Time of Flight (ToF) camera(s), RGB camera(s), and / or trained deep learning models, in combination with depth map(s) obtained using video see-through frames.
[0073] The user may then be instructed to move towards or away from the subject to begin the dolly zoom capture (610). Advantageously, the user may move in either direction at any desired speed to obtain a desired effect.
[0074] An average depth of subject pixels in the initial video frame(s) may then be determined (612). For example, in the example of FIG. 5, an average depth of all pixels in a face of the subject 506 may be calculated. For example, available inputs at the HMD 102 (or the goggles 502 of FIG. 5) may include a camera position (x, y, z) and rotation (quaternion x, y. z, w), a scene depth map (e.g.. stored as a 512x512. 16-bit array), subject segmentation (e.g., stored as a binary mask, 512x512 array), and video see-through RGB images (e.g., as a 512x512x3 RGB image). In such examples, an initial average distance from the face to the camera dO = avg(depthmap (for all pixel in the face)).
[0075] Finally in FIG. 6. the subject may be reprojected at the calculated average depth within subsequent video frames (614). For example, for new frames, the average distance from the face (or other subject) to the camera may be calculated as dl = avg (depthmap (for all pixel in the face in the new / subsequent frames)). Accordingly, a delta of dO - dl may be calculated, and a new depth map may be computed by shifting such delta values as new frames are received. That is, for all depth values in the face (or other subject) in a subsequent frame, a modified depth value of dl + delta may be computed.
[0076] In some implementations, a 3D depth mesh may be constructed, e.g., using techniques referenced above and described in more detail, below. Then, a human texture (RGB values) from the video see-through images may be projected onto the calculated depth mesh (triangle mesh), using new depth values obtained from the recalibrated depth map. This process ensures that the background remains faithful to reality (e.g., dynamically changes as the user 104 / 504 moves), while the subject appears to maintain a relatively constant distance from the device.
[0077] In implementations described above, a background is described as a physically present background or context in which a user is capturing video. In example implementations, however, projection techniques can be used to alter the physically present background into virtually any desired background. For example, in FIG. 1, once the subject 106 is segmented from the background 112, a virtual background 112’ may be generated or simulated by (or accessed by) the dolly zoom manager 125, and used to replace the actual background 112 in the captured dolly zoom video file.
[0078] For example, the virtual background 112’ can be a blurred or otherwise enhanced or distorted version of the actual background 112, and may represent a stationary or active environment. In other examples, the virtual background 112’ can be any scene or context available for insertion, including, e.g., an environmental scene such as mountains, a field, a desert, or may be a city scape, or an interior of a room or building (e.g., in a remote work scenario), any famous or well-known location, or a specific context (such as a roller coaster) selected to convey a specific emotion or otherwise enhance a desired effect in the dolly zoom video being captured.
[0079] As referenced and described above, depth maps and / or depth layers can be used to create the dolly zoom effect for computational photography. Computational photography with depth features improve a user experience by enabling image manipulation at an identified depth, including removing a subject image having a given size / depth and replacing it with a projected subject image having the same size / depth, even when the user has moved to a different depth.
[0080] In these contexts, depth may be understood as a distance from a reference point. Depth can be distance and direction from a reference point. Each pixel in an image can have a depth. Therefore, each pixel in the image can have an associated distance (and optionally a direction) from a reference point. The reference point can be based on a position of a device (e.g.. a mobile device, a tablet, a camera, and / or the like), as referenced above, including a camera. The reference point can be based on a global position (using, e.g., aglobal positioning system (GPS) of the device.
[0081] While the device moves (e.g., changes a perspective) in the real-world space, a depth of view with respect to the device can be tracked in the real-world space using, for example, device sensors and an application programming interface (API) configured to track the depth from the device using the sensors. For example, a subject can be locked to a depth and as the device moves the depth can be tracked (not, for example, maintaining a fixed focal distance). For example, if a subject image is of a subject identified at a depth of six (6) meters and the device is moved to five (5) meters (e.g., a user of the device steps forward), a projected subject image depth may be maintained as if the device were still at six (6) meters, while the remaining background will reflect movement of the device. Contrast this with object tracking where focus is locked on an object and as the object moves (or the device moves) the focus remains with the object so long as the object is within view.
[0082] As referenced above, depth information may include a depth map having a depth value for each pixel in an image. A depth map can be an image associated with a color (e.g., RGB, YUV, and / or the like) image or frame of a video. A depth map can be an image associated with a black and white, greyscale, and the like image or frame of a video. The depth map can store a distance value for each pixel in the image or frame. The depth value can have an 8-bit representation with values between, for example, 0 and 255, where 255 (or 0) represents the closest depth value and 0 (or 255) represents the most distant depth value. In an example implementation, the depth map can be normalized. For example, the depth map can be converted to a range of [0, 1 ]. Normalization can include converting the actual range of depth values. For example, if the depth map includes values between 43 and 203, 43 would be converted to 0, 203 would be converted to 1 and the remaining depth values would be converted to values between 0 and 1.
[0083] A depth map can be generated by a camera including a depth (D) sensor. This type of camera is sometimes called an RGBD camera. A depth map can be generated using an algorithm implemented as a post-processing function of a camera while capturing an image or frame of a video. For example, an image can be generated from a plurality (e.g., 3. 6, 9 or more) images captured by a camera. The algorithm can be implemented in an application programming interface (API) associated with capturing real-world space images (e.g., for augmented reality applications).
[0084] Depth information can be represented as a 3-D point cloud. In other words, a point cloud can be a depth map in three dimensions. A point cloud can be a collection of 3D points (X, Y, Z) that represent the external surface of the scene and can contain colorinformation.
[0085] The depth information can include depth layers each having a number (e.g.. an index number or z-index number) indicating a layer order. The depth information can be a layered depth image (LDI) having multiple ordered depths for each pixel in an image. Color information can be color (e.g., RGB, YUV, and / or the like) for each pixel in an image. A depth image can be an image where each pixel represents a distance from the camera location.
[0086] Using the depth information, modified video frames can be obtained that include the types of projected subject images described herein, using corresponding frame or image editing algorithms. The video frames can be edited during a capture process (e.g., as a function of a camera application executing on the device). Accordingly, the algorithm(s) can be executed on pixels of the image at the selected depth continually (e.g., ever}' frame or every' number of frames) and the projected subject image can be continually updated and displayed. In other words, the described dolly zoom effects can include real-time depthbased image editing.
[0087] Described image editing can be performed in a live image capture of a real- world space. Accordingly, as the device moves, the screen-space is recomputed, the depth map is regenerated, and the projected subject image is re-projected. Video frames are re- edited (e.g., a new composite video frame with the projected subject images and the included background are generated) and rendered on the device. In other words, described techniques may continually repeat (e g., execute in a loop) while the dolly zoom effect is captured.
[0088] Described implementations may use features that can run immediately on a variety of devices by focusing on real time depth map processing, and / or may use techniques requiring a persistent model of the environment generated with 3D surface reconstruction. For example, real-time depth maps provided suitable APIs may be obtained using only a single moving RGB camera to estimate depth. A dedicated depth camera, such as time-of- flight (ToF) cameras can instantly provide depth maps without any initializing camera motion. Additionally, data such as a live camera feed, phone position and orientation, and camera parameters including focal length, intrinsic matrix, extrinsic matrix, and projection matrix for each frame may' be used to establish a mapping betw een the physical world and virtual objects.
[0089] Depth data may be stored in a low-resolution depth buffer (e.g., 160* 120), which is a perspective camera image that contains a depth value instead of color in each pixel. Different categories of data structures may be used.
[0090] For example, a depth array may store depth in a 2D array of a landscape image with 16-bit integers on a CPU. Using phone orientation and maximum sensing range depth may be accessed from any screen point or texture coordinates of the camera image.
[0091] Alternatively, a depth mesh is a real-time triangulated mesh generated for each depth map on both CPU and GPU. In contrast to traditional surface reconstruction with persistent voxels or triangles, depth mesh has little memory and compute overhead and can be generated in real time. Depth texture may be copied to the GPU from the depth array for per-pixel depth use cases in each frame.
[0092] Depth data may be mapped to the camera image and the real-w orld geometry, including adapting to changes in camera orientation, or conversion of points betw een local and global coordinate frames. Depth data may leverage localized depth, surface depth, or dense depth.
[0093] For example, localized depth uses a depth array to operate on a small number of points directly on the CPU, and may be used to compute physical distance, e.g., to the subject 106 / 506. Surface depth leverages the CPU or compute shaders on the GPU to create and update depth meshes in real time, thus enabling collision, physics, texture decal, geometry-aw are shadows, and other effects. Dense depth may be copied to a GPU texture and used for rendering depth-aw are effects with GPU-accelerated bilinear filtering in screen space.
[0094] Every pixel in the color camera image has a depth value mapped to it. which is useful for the real-time projected subject image operations described herein for producing the dolly zoom effect. Screen-space to / from w orld-space conversion may include the following operations.
[0095] For example, for a screen point p = [x, y], a corresponding depth value may be obtained from a depth array Dwxh (e.g., w = 120, h = 160, as in the example above). Then, the screen point may be re-projected to a camera-space vertex vPusing the camera intrinsic matrix K , where vp = D(p) K1[p,l] .
[0096] Given the camera extrinsic matrix C = [R|t], which consists of a 3x3 rotation matrix R and a 3x1 translation vector t, global coordinates gPin the world space may be determined as: gP= C- [vp, 1] . Hence, both virtual objects, such as the projected subject images described herein, and the physical environment (e.g., background 112) may be rendered in the same coordinate system.
[0097] In a reverse process, 3D points may be projected using the camera’s projection matrix P. Then the projected depth values may be normalized and the depth projection may bescaled to the size of the depth map wxh.
[0098] In examples using real-time depth meshes, a mesh refers to, e.g., a set of triangle surfaces that are connected to form a continuous surface, which is the most common representation of a 3D shape. A depth mesh provides the 3D vertex position and the normal vector of surface points to compute world-space texture coordinates. Game and graphics engines are optimized for handling mesh data and provide simple ways to transform, shade, and to detect interactions between shapes.
[0099] In example implementations, variations of screen-space depth meshing techniques may be used that rely on a densely tessellated quad in which each vertex is displaced based on a re-projected depth value. No additional data transfer between CPU and GPU is required during render time, making this method very efficient.
[0100] Although techniques have been described with respect to personal devices, it will be appreciated that described techniques may be implemented in the context of professional filmmaking, as well. Accordingly, a 3D, immersive dolly zoom effect may be obtained in a variety of desired settings and contexts.
[0101] FIG. 7 is a third person view of a user 702 (analogous to the user 104 of FIG. 1) in an ambient environment 7000, with one or more external computing systems shown as additional resources 752 that are accessible to the user 702 via a network 7200. FIG. 7 illustrates numerous different wearable devices that are operable by the user 702 on one or more body parts of the user 702. including a first wearable device 750 in the form of glasses worn on the head of the user, a second wearable device 754 in the form of ear buds worn in one or both ears of the user 702, a third wearable device 756 in the form of a watch worn on the wrist of the user, and a computing device 706 held by the user 702. In FIG. 7, the computing device 706 is illustrated as a handheld computing device but may also be understood to represent any personal computing device, such as a table or personal computer.
[0102] In some examples, the first wearable device 750 is in the form of a pair of smart glasses including, for example, a display, one or more images sensors that can capture images of the ambient environment, audio input / output devices, user input capability, computing / processing capability and the like. Additional examples of the first wearable device 750 are provided below with respect to FIGS. 8A and 8B.
[0103] In some examples, the second wearable device 754 is in the form of an ear worn computing device such as headphones, or earbuds, that can include audio input / output capability, an image sensor that can capture images of the ambient environment 7000, computing / processing capability, user input capability' and the like. In some examples, thethird wearable device 756 is in the form of a smart watch or smart band that includes, for example, a display, an image sensor that can capture images of the ambient environment, audio input / output capability, computing / processing capability, user input capability and the like. In some examples, the handheld computing device 706 can include a display, one or more image sensors that can capture images of the ambient environment, audio input / output capability, computing / processing capability, user input capability, and the like, such as in a smartphone. In some examples, the example wearable devices 750. 754, 756 and the example handheld computing device 706 can communicate with each other and / or with external computing system(s) 752 to exchange information, to receive and transmit input and / or output, and the like. The principles to be described herein may be applied to other types of wearable devices not specifically shown in FIG. 7 or described herein.
[0104] The user 702 may choose to use any one or more of the devices 706, 750, 754, or 756, perhaps in conjunction with the external resources 752, to implement any of the implementations described above with respect to FIGS. 1-6. For example, the user 702 may use an application executing on the device 706 and / or the smartglasses 750 to execute the dolly zoom manager 125 of FIG. 1.
[0105] As referenced above, the device 706 may access the additional resources 752 to facilitate the various dolly zoom techniques described herein, or related techniques. In some examples, the additional resources 752 may be partially or completely available locally on the device 706. In some examples, some of the additional resources 752 may be available locally on the device 706, and some of the additional resources 752 may be available to the device 706 via the network 7200. As shown, the additional resources 752 may include, for example, server computer systems, processors, databases, memory storage, and the like. In some examples, the processor(s) may include training engine(s), transcription engine(s), translation engine(s), rendering engine(s), and other such processors. In some examples, the additional resources may include ML model(s), such as an Al model used by the dolly zoom manager 125 of FIG. 1.
[0106] The device 706 may operate under the control of a control system 760. The device 706 can communicate with one or more external devices, either directly (via wired and / or wireless communication), or via the network 7200. In some examples, the one or more external devices may include various ones of the illustrated wearable computing devices 750, 754, 756. another mobile computing device similar to the device 706, and the like. In some implementations, the device 706 includes a communication module 762 to facilitate external communication. In some implementations, the device 706 includes a sensing system 764including various sensing system components. The sensing system components may include, for example, one or more image sensors 765, one or more position / orientation sensor(s) 764 (including for example, an inertial measurement unit, an accelerometer, a gyroscope, a magnetometer and other such sensors), one or more audio sensors 766 that can detect audio input, one or more image sensors 767 that can detect visual input, one or more touch input sensors 768 that can detect touch inputs, and other such sensors. The device 706 can include more, or fewer, sensing devices and / or combinations of sensing devices. Various ones of the communications modules may be used to control brightness settings among devices described herein, and various sensors may be used individually or together to perform the types of gaze, depth, and / or brightness detection described herein.
[0107] Captured still and / or moving images may be displayed by a display device of an output system 772, and / or transmitted externally via a communication module 762 and the network 7200, and / or stored in a memory 770 of the device 706. The device 706 may include one or more processor(s) 774. The processors 774 may include various modules or engines configured to perform various functions. In some examples, the processor(s) 774 may include, e.g.. training engine(s). transcription engine(s). translation engine(s), rendering engine(s), and other such processors. The processor(s) 774 may be formed in a substrate configured to execute one or more machine executable instructions or pieces of software, firmware, or a combination thereof. The processor(s) 774 can be semiconductor-based including semiconductor material that can perform digital logic. The memory 770 may include any type of storage device or non-transitory computer-readable storage medium that stores information in a format that can be read and / or executed by the processor(s) 774. The memory' 770 may store applications and modules that, when executed by the processor(s) 774, perform certain operations. In some examples, the applications and modules may be stored in an external storage device and loaded into the memory' 770.
[0108] Although not shown separately in FIG. 7, it will be appreciated that the various resources of the computing device 706 may be implemented in whole or in part within one or more of various wearable devices, including the illustrated smartglasses 750. earbuds 754. and smartwatch 756. which may be in communication with one another to provide the various features and functions described herein.
[0109] An example head mounted wearable device 800 in the form of a pair of smart glasses is shown in FIGS. 8A and 8B, for purposes of discussion and illustration. The example head mounted wearable device 800 includes a frame 802 having rim portions 803 surrounding glass portion, or lenses 807, and arm portions 830 coupled to a respective rimportion 803. In some examples, the lenses 807 may be corrective / prescription lenses. In some examples, the lenses 807 may be glass portions that do not necessarily incorporate corrective / prescription parameters. A bridge portion 809 may connect the rim portions 803 of the frame 802. In the example shown in FIGS. 8 A and 8B, the wearable device 800 is in the form of a pair of smart glasses, or augmented reality glasses, simply for purposes of discussion and illustration.
[0110] In some examples, the wearable device 800 includes a display device 804 that can output visual content, for example, at an output coupler providing a visual display area 805, so that the visual content is visible to the user. In the example shown in FIGS. 8A and 8B. the display device 804 is provided in one of the two arm portions 830, simply for purposes of discussion and illustration. Display devices 804 may be provided in each of the two arm portions 830 to provide for binocular output of content. In some examples, the display device 804 may be a see through near eye display. In some examples, the display device 804 may be configured to project light from a display source onto a portion of teleprompter glass functioning as a beamsplitter seated at an angle (e.g., 30-45 degrees). The beamsplitter may allow for reflection and transmission values that allow the light from the display source to be partially reflected while the remaining light is transmitted through. Such an optic design may allow a user to see both physical items in the w orld, for example, through the lenses 807, next to content (for example, digital images, user interface elements, virtual content, and the like) output by the display device 804. In some implementations, waveguide optics may be used to depict content on the display device 804.
[0111] The example w earable device 800, in the form of smart glasses as shown in FIGS. 8A and 8B, includes one or more of an audio output device 806 (such as, for example, one or more speakers), an illumination device 808, a sensing system 810. a control system 812, at least one processor 814, and an outw ard facing image sensor 816 (for example, a camera). In some examples, the sensing system 810 may include various sensing devices and the control system 812 may include various control system devices including, for example, the at least one processor 814 operably coupled to the components of the control system 812. In some examples, the control system 812 may include a communication module providing for communication and exchange of information betw een the wearable device 800 and other external devices. In some examples, the head mounted w earable device 800 includes a gaze tracking device 815 to detect and track eye gaze direction and movement. Data captured by the gaze tracking device 815 may be processed to detect and track gaze direction and movement as a user input. In the example shown in FIGS. 8A and 8B, the gaze trackingdevice 815 is provided in one of two arm portions 830, simply for purposes of discussion and illustration. In the example arrangement shown in FIGS. 8A and 8B, the gaze tracking device 815 is provided in the same arm portion 830 as the display device 804, so that user eye gaze can be tracked not only with respect to objects in the physical environment, but also with respect to the content output for display by the display device 804. In some examples, gaze tracking devices 815 may be provided in each of the two arm portions 830 to provide for gaze tracking of each of the two eyes of the user. In some examples, display devices 804 may be provided in each of the two arm portions 830 to provide for binocular display of visual content.
[0112] The wearable device 800 is illustrated as glasses, such as smartglasses, augmented reality (AR) glasses, or virtual reality (VR) glasses. More generally, the wearable device 800 may represent any head-mounted device (HMD), including, e.g., goggles, helmet, or headband. Even more generally, the wearable device 800 and the computing device 706 may represent any wearable device(s), handheld computing device(s), or combinations thereof.
[0113] Use of the wearable device 800, and similar wearable or handheld devices such as those shown in FIG. 7, enables useful and convenient use case scenarios of implementations of FIGS. 1-6. For example, as shown in FIG. 8B, the display area 805 may be used to display subject 106 and / or background 112 in the example of FIG. 1. More generally, the display area 805 may be used to provide any of the functionality described with respect to FIGS. 1 -6 that may be useful in operating the dolly zoom manager 125.
[0114] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry', integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0115] These computer programs (also known as modules, programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine- readable medium” “computer-readable medium” refers to any computer program product,apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0116] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or LED (light emitting diode)) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well. For example, feedback provided to the user can be any form of sensory’ feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0117] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (“LAN”), a wide area network (“WAN”), and the Internet.
[0118] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0119] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spint and scope of the description and claims.
[0120] In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. In addition, other steps may be provided, or steps may be eliminated, from the described flows, and other components maybe added to, or removed from, the described systems. Accordingly, other implementations are within the scope of the following claims.
[0121] Further to the descriptions above, a user is provided with controls allowing the user to make an election as to both if and when systems, programs, devices, networks, or features described herein may enable collection of user information (e.g., information about a user's social network, social actions, or activities, profession, a user’s preferences, or a user's current location), and if the user is sent content or communications from a server. In addition, certain data may be treated in one or more ways before it is stored or used, so that user information is removed. For example, a user’s identity may be treated so that no user information can be determined for the user, or a user’s geographic location may be generalized where location information is obtained (such as to a city, ZIP code, or state level), so that a particular location of a user cannot be determined. Thus, the user may have control over w hat information is collected about the user, how that information is used, and what information is provided to the user.
[0122] The computer system (e.g.. computing device) may be configured to wirelessly communicate with a netw ork server over a network via a communication link established with the network server using any known wireless communications technologies and protocols including radio frequency (RF), microw ave frequency (MWF), and / or infrared frequency (IRF) wireless communications technologies and protocols adapted for communication over the network.
[0123] In accordance with aspects of the disclosure, implementations of various techniques described herein may be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations of them. Implementations may be implemented as a computer program product (e.g., a computer program tangibly embodied in an information carrier, a machine-readable storage device, a computer-readable medium, a tangible computer-readable medium), for processing by, or to control the operation of, data processing apparatus (e.g., a programmable processor, a computer, or multiple computers). In some implementations, a tangible computer-readable storage medium may be configured to store instructions that when executed cause a processor to perform a process. A computer program, such as the computer program(s) described above, may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may be deployed tobe processed on one computer or on multiple computers at one site or distributed across multiple sites and interconnected by a communication network.
[0124] Specific structural and functional details disclosed herein are merely representative for the purposes of describing example implementations. Example implementations, however, may be embodied in many alternate forms and should not be construed as limited to only the implementations set forth herein.
[0125] The terminology used herein is for the purpose of describing particular implementations only and is not intended to be limiting of the implementations. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises," "comprising," "includes," and / or "including." when used in this specification, specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.
[0126] It will be understood that when an element is referred to as being "coupled," "connected," or "responsive" to, or "on," another element, it can be directly coupled, connected, or responsive to, or on, the other element, or intervening elements may also be present. In contrast, when an element is referred to as being "directly coupled," "directly connected," or "directly responsive" to, or "directly on," another element, there are no intervening elements present. As used herein the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0127] Spatially relative terms, such as "beneath," "below," "lower," "above," "upper," and the like, may be used herein for ease of description to describe one element or feature in relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is turned over, elements described as "below" or "beneath" other elements or features would then be oriented "above" the other elements or features. Thus, the term "below" can encompass both an orientation of above and below. The device may be otherwise oriented (rotated 130 degrees or at other orientations) and the spatially relative descriptors used herein may be interpreted accordingly.
[0128] Example implementations of the concepts are described herein with reference to cross-sectional illustrations that are schematic illustrations of idealized implementations (and intermediate structures) of example implementations. As such, variations from theshapes of the illustrations as a result, for example, of manufacturing techniques and / or tolerances, are to be expected. Thus, example implementations of the described concepts should not be construed as limited to the particular shapes of regions illustrated herein but are to include deviations in shapes that result, for example, from manufacturing. Accordingly, the regions illustrated in the figures are schematic in nature and their shapes are not intended to illustrate the actual shape of a region of a device and are not intended to limit the scope of example implementations.
[0129] It will be understood that although the terms "first," "second," etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. Thus, a "first" element could be termed a "second" element without departing from the teachings of the present implementations.
[0130] Unless otherwise defined, the terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these concepts belong. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and / or the present specification and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0131] While certain features of the described implementations have been illustrated as described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover such modifications and changes as fall within the scope of the implementations. It should be understood that they have been presented by way of example only, not limitation, and various changes in form and details may be made. Any portion of the apparatus and / or methods described herein may be combined in any combination, except mutually exclusive combinations. The implementations described herein can include various combinations and / or sub-combinations of the functions, components, and / or features of the different implementations described.
Claims
WHAT IS CLAIMED IS:
1. A computer program product, the computer program product being tangibly embodied on anon-transitory computer-readable storage medium and comprising instructions that, when executed by at least one computing device, are configured to cause the at least one computing device to: determine a size of an initial subject image of a subject in an initial video frame captured, relative to a background, by a camera; remove subsequent subject images of the subject from the background within subsequent video frames; and render projected images of the subject, with a projected size that is based on the size of the initial subject image, within the subsequent video frames and in place of the subsequent subject images, while the camera is moving relative to the subject.
2. The computer program product of claim 1. wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: determine a first depth of the subj ect in the initial video frame, relative to the camera; determine a second depth of the subject in a second video frame of the subsequent video frames; and determine the projected images of the subject using the first depth and the second depth, including generating a composite frame of the subsequent video frames that includes the projected images of the subject at the first depth.
3. The computer program product of claim 2, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: segment the subject within the initial video frame to obtain a segmented subject image; and determine the first depth as an average depth of pixels of the segmented subject image.
4. The computer program product of any one of the preceding claims, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: render the projected images of the subject within the subsequent video frames with less of the background included in the subsequent video frames as the camera moves closer to the subj ect.
5. The computer program product of any one of the preceding claims, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: render the projected images of the subject within the subsequent video frames with more of the background included in the subsequent video frames as the camera moves away from the subject.
6. The computer program product of any one of the preceding claims, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: store the initial video frame and the subsequent video frames as a video file.
7. The computer program product of any one of the preceding claims, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: segment the subject within the initial video frame to obtain a segmented subject image; and determine the size based on the segmented subject image.
8. The computer program product of any one of the preceding claims, wherein the projected size is the same as the size of the initial subject image.
9. The computer program product of any one of the preceding claims, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: replace the background within the subsequent video frames with a virtual background.
10. The computer program product of any one of the preceding claims, wherein the instructions, when executed by the at least one computing device, are further configured to cause the at least one computing device to: generate a depth map of a field of view of the camera; and determine the size of the initial subject image based on the depth map.
11. A head-mounted device (HMD) comprising: at least one frame for positioning the HMD on a face of a user; at least one camera; at least one display; at least one processor; and at least one memory, the at least one memory storing a set of instructions, which, when executed, cause the at least one processor to: determine a size of an initial subject image of a subject in an initial video frame captured, relative to a background, by the at least one camera; remove subsequent subject images of the subject from the background within subsequent video frames; and render projected images of the subject on the at least one display, with a projected size that is based on the size of the initial subject image, within the subsequent video frames and in place of the subsequent subject images, while the at least one camera is moving relative to the subject.
12. The HMD of claim 11, wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to: determine a first depth of the subject in the initial video frame, relative to the at least one camera; determine a second depth of the subject in a second video frame of the subsequent video frames; and determine the projected images of the subject using the first depth and the second depth, including generating a composite frame of the subsequent video frames that includes the projected images of the subject at the first depth.
13. The HMD of claim 12, wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to:segment the subject within the initial video frame to obtain a segmented subject image; and determine the first depth as an average depth of pixels of the segmented subject image.
14. The HMD of any one of claims 11-13, wherein the projected size is the same as the size of the initial subject image.
15. The HMD of any one of claims 11-14, wherein the set of instructions, when executed by the at least one processor, are further configured to cause the HMD to: generate a depth map of a field of view of the at least one camera; and determine the size of the initial subject image based on the depth map.
16. A method comprising: determining a size of an initial subject image of a subject in an initial video frame captured, relative to a background, by a camera; removing subsequent subject images of the subject from the background within subsequent video frames; and rendering projected images of the subject, with a projected size that is based on the size of the initial subject image, within the subsequent video frames and in place of the subsequent subject images, while the camera is moving relative to the subject.
17. The method of claim 16, further comprising: determining a first depth of the subject in the initial video frame, relative to the camera; determining a second depth of the subject in a second video frame of the subsequent video frames; and determining the projected images of the subject using the first depth and the second depth, including generating a composite frame of the subsequent video frames that includes the projected images of the subject at the first depth.
18. The method of claim 17, further comprising: segmenting the subject within the initial video frame to obtain a segmented subject image; anddetermining the first depth as an average depth of pixels of the segmented subject image.
19. The method of any one of claims claim 16-18, further comprising: rendering the projected images of the subject within the subsequent video frames with less of the background included in the subsequent video frames as the camera moves closer to the subject.
20. The method of any one of claims 16-19, further comprising: rendering the projected images of the subject within the subsequent video frames with more of the background included in the subsequent video frames as the camera moves away from the subject.
Citation Information
Patent Citations
Generating and modifying representations of dynamic objects in an artificial reality environment
US11335077B1
System and method for providing dolly zoom view synthesis
US20210125307A1
Automatic dolly zoom image processing device
US20220358619A1
Automatic dolly zoom image processing device
US20230334619A1