Method and system for displaying composite image data

By combining depth sensors and motion sensors on mobile devices to generate virtual camera transformations and update virtual background content in real time, the problem of non-real-time virtual background updates in existing selfie technologies is solved, and the user experience and social sharing capabilities of selfies are improved.

CN115063523BActive Publication Date: 2025-09-30APPLE INC

Patent Information

Application Number
CN202210700110.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-09-06
Filing Date
2018-09-07
Publication Date
2025-09-30
Estimated Expiration
2038-09-07

AI Technical Summary

Technical Problem

Existing selfie technology makes it difficult to achieve real-time updating of virtual background content on mobile devices, resulting in a poor user experience and an inability to meet the needs of sharing on social networks.

Method used

Image data is captured through the front or rear camera of the mobile device, and a virtual camera transformation is generated by combining the depth sensor and motion sensor to update the virtual background content in real time and synthesize AR selfie videos.

Benefits of technology

It improves the interactive entertainment experience of selfies, allowing users to automatically replace the real-world background with a virtual background on their mobile devices to generate shareable AR selfie images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063523B_ABST
    Figure CN115063523B_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods and systems for displaying composite image data. In one embodiment, a method includes: capturing real-time image data via a first camera of a mobile device, the real-time image data including an image of an object in a physical real-world environment; receiving depth data via a depth sensor of the mobile device, the depth data indicating a distance of the object from the camera in the physical real-world environment; receiving motion data via one or more motion sensors of the mobile device, the motion data indicating at least an orientation of the first camera in the physical real-world environment; generating a virtual camera transform based on the motion data, the camera transform being used to determine an orientation of the virtual camera in the virtual environment; and generating composite image data using the image data, a mask, and virtual background content selected based on the virtual camera orientation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is a divisional application of the invention patent application with international application number PCT / US2018 / 049930, international application date September 7, 2018, date of entry into the Chinese national phase January 20, 2020, Chinese national application number 201880048711.2, and invention name “Augmented Reality Selfie”.

[0003] This patent application claims priority to U.S. Provisional Patent Application No. 62 / 556,297, entitled “Augmented Reality Self-Portraits,” filed on September 8, 2017, and U.S. Patent Application No. 16 / 124,168, entitled “Augmented Reality Self-Portraits,” filed on September 6, 2018. The entire contents of each of the above patent applications are incorporated herein by reference. Technical Field

[0004] The present disclosure generally relates to media editing and augmented reality. Background Art

[0005] Self-photographed digital photos or "selfies" have become a pop culture phenomenon. Selfies are typically taken with a digital camera or smartphone held at arm's length, pointed toward a mirror, or attached to a selfie stick to position the camera farther away from the subject and capture the background scene behind the subject. Selfies are often shared on social networking services (e.g., Augmented reality (AR) is a real-time view of a physical real-world environment whose elements are "augmented" by computer-generated sensory input such as sound, video, or graphics. Summary of the Invention

[0006] The present disclosure relates to creating augmented reality self-videos using machine learning. Disclosed are systems, methods, apparatus, and non-transitory computer-readable storage media for generating AR self-videos or “AR selfies.”

[0007] In one embodiment, a method includes: capturing real-time image data via a first camera of a mobile device, the real-time image data including an image of an object in a physical real-world environment; receiving depth data via a depth sensor of the mobile device, the depth data indicating a distance of the object from the camera in the physical real-world environment; receiving motion data via one or more motion sensors of the mobile device, the motion data indicating at least an orientation of the first camera in the physical real-world environment; generating a virtual camera transform based on the motion data via one or more processors of the mobile device, the camera transform being used to determine the orientation of the virtual camera in the virtual environment; receiving content from the virtual environment via the one or more processors; generating a mask from the image data and the depth data via the one or more processors; generating composite image data using the image data, the mask, and first virtual background content, the first virtual background content being selected from the virtual environment using the camera transform; and causing the composite image data to be displayed on a display of the mobile device via the one or more processors.

[0008] In one embodiment, a method includes: presenting a preview on a display of a mobile device, the preview comprising sequential frames of preview image data captured by a front-facing camera of the mobile device positioned within a close range of an object, the sequential frames of preview image data comprising close-range image data of the object and image data of a background behind the object in a physical real-world environment; receiving first user input for applying a virtual environment effect; capturing depth data via a depth sensor of the mobile device, the depth data indicating a distance of the object from the front-facing camera in the physical real-world environment; capturing orientation data via one or more sensors of the mobile device, the orientation data indicating at least an orientation of the front-facing camera in the physical real-world environment; generating, via one or more processors of the mobile device, a camera transform based on the motion data, the camera transform describing the orientation of a virtual camera in the virtual environment; obtaining, via the one or more processors and using the camera transform, virtual background content from the virtual environment; generating, via the one or more processors, a mask from the sequential frames of image data and the depth data; generating, via the one or more processors, a composite sequential frame of image data, the composite sequential frame comprising the sequential frames of image data, the mask, and the virtual background content; and causing, via the one or more processors, display of the composite sequential frame of image data.

[0009] Other embodiments relate to systems, methods, apparatus, and non-transitory computer-readable media.

[0010] Certain implementations disclosed herein provide one or more of the following advantages. The user experience of creating selfies on a mobile device is improved by allowing a user to capture and record selfie videos using a front-facing or rear-facing camera embedded in a mobile device, and automatically replacing the real-world background captured in a live video preview with user-selected virtual background content that is automatically updated in response to motion data from a motion sensor of the mobile device. Thus, the disclosed implementations provide an interactive and entertaining process for capturing selfie images that can be shared with friends and family via social networks.

[0011] The details of the disclosed implementations are shown in the following drawings and detailed description.Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a conceptual diagram illustrating the basic concept of AR selfie according to an embodiment.

[0013] Figure 2A-2E Mapping of a virtual environment to a mobile device viewport is shown according to an embodiment.

[0014] Figure 3A and Figure 3B A graphical user interface for recording an AR selfie using a front-facing camera is shown according to an embodiment.

[0015] Figure 3C and Figure 3D A graphical user interface is shown with different background scenes selected and showing a recording view and a full-screen playback view according to an embodiment.

[0016] Figure 3E and Figure 3F A graphical user interface for recording and playing back selfies using a rear camera and showing a recording view and a full-screen playback view is shown according to an embodiment.

[0017] Figure 4 is a block diagram of a system illustrating process steps used in creating an AR selfie, according to an embodiment.

[0018] Figure 5 According to an embodiment, a synthesis layer is shown for use in AR selfies.

[0019] Figures 6A-6L A multi-stage process for generating a pre-processed (coarse) mask using depth data is shown according to an embodiment.

[0020] Figures 7A-7C A refinement mask extraction process using video data and a pre-processed (coarse) mask is shown according to an embodiment.

[0021] Figure 8 A post-processing stage for removing artifacts from a refinement mask is shown according to an embodiment.

[0022] Figure 9 is a flow chart of a process for generating an AR selfie, according to an embodiment.

[0023] Figure 10 is a flowchart of a process for generating an AR selfie mask according to an embodiment.

[0024] Figure 11 According to the embodiment shown for realizing reference Figures 1-10 The features and processes of the device architecture are described.

[0025] The same reference symbols used in the various drawings identify similar elements. DETAILED DESCRIPTION

[0026] A "selfie" is a self-portrait image that a user takes in close proximity, often by holding the camera at arm's length or using an extension device such as a "selfie" stick. The selfie subject is often the user's face, or a portion of the user (e.g., the user's upper body), and any background visible behind the user. A front-facing camera is a camera that faces the user when the user is viewing a display screen. Alternatively, a rear-facing camera faces away from the user when the user is viewing a display screen, and captures images of the real-world environment in front of the user and in the opposite direction. Typical mobile devices used to capture selfies are digital cameras, smartphones with one or more embedded digital cameras, or tablets with one or more embedded cameras.

[0027] In one embodiment, the selfie object can be synthesized with virtual background content extracted from the virtual environment data model. The virtual background content may include, but is not limited to, two-dimensional (2D) images, three-dimensional (3D) images, and 360° videos. In the preprocessing stage, a rough mask is generated from the depth data provided by the depth sensor and then refined using video data (e.g., RGB video data). In one embodiment, the depth sensor is an infrared (IR) depth sensor embedded in the mobile device. The mask is synthesized with the video data containing the image of the selfie object (e.g., using alpha synthesis), and the real-world background behind the object is replaced and continuously updated with the virtual background content selected from the virtual environment selected by the user. The virtual background content is selected using a virtual camera transformation, which is generated using motion data from one or more motion sensors (e.g., accelerometers, gyroscopes) of the mobile device. The video data, refined mask, virtual background content, and optionally one or more animation layers are synthesized to form an AR selfie video. The AR selfie video is displayed to the user by the viewport of the mobile device.

[0028] In one embodiment, the mobile device also includes a rear-facing camera that can be used to capture video in front of the user, which can be processed in a similar manner to the video captured by the front-facing camera. A camera flip signal provided by the mobile device's operating system can indicate which camera is capturing video, and this signal can be used to adjust the virtual camera transition to update the virtual background content.

[0029] A mask generation method is disclosed that uses undefined depth data (also referred to herein as "shadow data") to segment a depth image (e.g., a binary depth mask) into foreground and background regions. The mask contains coverage information, including the outline of the object being drawn, making it possible to distinguish between the portion of the binary depth mask where the object is actually drawn and other empty portions of the binary depth mask. In one embodiment, the mask generation process uses a region growing algorithm and / or a 3D facial mesh to identify and fill "holes" (undefined depth data) in the mask caused by sunlight reflecting off sunglasses worn by the subject.

[0030] Although the mask generation process is disclosed herein as part of an AR selfie generation process, the disclosed mask generation process can be used to generate masks from depth data for any image processing application. For example, the disclosed mask generation process can be used to segment images as part of a video / image editing tool.

[0031] In one embodiment, the virtual environment can be any desired environment, such as a famous city (e.g., London, Paris, or New York), and include famous landmarks (e.g., Big Ben, London Bridge, the Eiffel Tower). The virtual environment can also be completely fictional, such as a cartoon environment completed with cartoon characters, flying saucers, and any other desired props. In one embodiment, motion effects (e.g., blur effects, glow effects, cartoon effects) can be applied to one or more of the video data, virtual background content, and mask. Motion effects can also be applied to the final composite video. In one embodiment, one or more animation layers (e.g., an animated particle layer similar to falling snow or sparks) can be composited with the video data, mask, and virtual background content.

[0032] In one embodiment, the selfie GUI includes various controls, such as controls for recording AR selfie videos to a storage device (e.g., flash memory of a mobile device), controls for turning one or more microphones of the mobile device on and off, a camera inversion button for switching between the front camera and the rear camera, and a tray for storing AR selfie video thumbnail images that can be selected to retrieve and play back corresponding videos on the mobile device.

[0033] AR Selfie Concept Overview

[0034] Figure 1 is a conceptual diagram illustrating the concept of an AR selfie, according to an embodiment. A user 100 is shown taking a selfie using the front-facing camera of a mobile device 102. During recording, a viewport on the mobile device 102 displays a live video feed of the user 100 in the foreground with virtual background content 104 extracted from a virtual environment 106. As the user 100 changes the orientation of the mobile device 102 in the real world (e.g., rotates the direction of the camera's field of view), the motion sensors (e.g., accelerometers, gyroscopes) of the mobile device 102 sense the change and generate motion data for updating the virtual background content 104 with new virtual background content extracted from different parts of the virtual environment 106, as shown in FIG. Figure 2A-2E As further described. The portion extracted from the virtual background content 104 depends on how the user 100 is holding the mobile device 102. For example, if the user 100 is holding the mobile device 102 in a "portrait" orientation when taking a selfie, the portion extracted from the virtual background content 104 will have an aspect ratio that will fill the viewport in a portrait or vertical orientation. Similarly, if the user 100 is holding the mobile device 102 in a "landscape" orientation when taking a selfie, the portion extracted from the virtual background content 104 will have an aspect ratio that will fill the viewport in a landscape or horizontal orientation.

[0035] Example mapping of virtual environments

[0036] Figure 2A-2E The mapping of a virtual environment to a mobile device's viewport is shown according to an embodiment. Figure 2A A unit sphere 106 is shown having a viewport 202 projected onto its surface ( Figure 2C ) corner. Figure 2B An equirectangular projection 200 (e.g., a Mercator projection) generated by mapping a projected viewport 202 from a spherical coordinate system to a planar coordinate system is shown. In one embodiment, the horizontal line that divides the equirectangular projection 200 is the equator of the unit sphere 106, and the vertical line that divides the equirectangular projection 200 is the prime meridian of the unit sphere 106. The width of the equirectangular projection 200 is from 0° to 360°, and the height spans 180°.

[0037] Figure 2C The sub-rectangle 203 is shown overlaid on the equirectangular projection 200. The sub-rectangle 203 represents the viewport 202 of the mobile device 102 in plane coordinates. Figure 2E The equirectangular projection 200 is sampled into the viewport 202 using formulas [1] and [2]:

[0038]

[0039] λ=a cos(z c ), longitude. [2]

[0040] Figure 2D A mobile device 102 is shown with a viewport 202 and a front-facing camera 204. The viewing coordinate system (X c ,Y c ,Z c ), where +Z c The coordinates are the viewing direction of the front-facing camera. In computer graphics, a camera analogy is used where a viewer 206 at a view reference point (VRP) observes a virtual environment through a virtual camera 205 and can look around and walk around the virtual environment. This is achieved by defining a view coordinate system (VCS) with the position and orientation of the virtual camera 205, such as Figure 2D and Figure 2E As shown. Figure 2E In FIG, the virtual camera 205 is shown as being fixedly positioned to the origin and having latitude (φ) and longitude (λ) in the virtual world coordinate system. It can be imagined that the virtual camera 205 is looking outward at the unit sphere 106, with the image of the virtual rear camera in the -Z direction, as shown in FIG. Figure 2D As shown. For the front camera 204, the virtual camera 205 (around Figure 2D The Y axis in the image is rotated 180° to generate a front camera view along the +Z direction, which shows the virtual background "over the shoulder" of the observer 206.

[0041] In one embodiment, the pose quaternion generated by the pose processor of the mobile device 102 can be used to determine the field of view direction of the rear camera and the front camera. When the observer 206 rotates the mobile device 102, the motion sensor (e.g., gyroscope) senses the rotation or rotation rate and updates the pose quaternion of the mobile device 102. The updated pose quaternion (e.g., Δ quaternion) can be used to derive a camera transformation for determining the camera field of view direction in the virtual environment for the rear camera, or can be further transformed by 180° for determining the camera field of view direction in the virtual environment for the front camera.

[0042] The mathematical operations used to derive camera transformations are well known in computer graphics and will not be discussed further herein. However, an important feature of the disclosed embodiments is that the real-world orientation of the real-world camera is used to drive the orientation of the virtual camera in the virtual environment. As a result, as the field of view of the real-world camera changes in real time, the field of view of the virtual camera (represented by the camera transformation) also changes in sync with the real-world camera. As will be described below, this technique produces a user's perception of the virtual environment 106 ( Figure 1) and thus capture the illusion that a virtual background behind the user is being captured instead of the real-world background. In one embodiment, when the user first enters the scene, the device orientation (e.g., direction, altitude) can be biased toward a visually impressive portion of the scene (referred to as the "protagonist angle"). For example, as the user looks around the scene, a delta can be applied to the device orientation, where the delta is calculated as the difference between the protagonist angle and the device orientation when the user entered the scene.

[0043] Example GUI for recording an AR selfie

[0044] Figure 3A and Figure 3B A graphical user interface for recording an AR selfie is shown according to an embodiment. Figure 3A , the AR selfie GUI 300 includes a viewport 301 that displays a composite video frame including a selfie object 302a and virtual background content 303a. A "cartoon" effect has been applied to the composite video to create interesting effects and hide artifacts from the alpha synthesis process. Although a single composite video frame is shown, it should be understood that the viewport 301 is displaying a real-time video feed (e.g., 30 frames per second), and if the orientation of the real-world camera's field of view changes, the virtual background 303a will also change seamlessly to display different parts of the virtual environment. This allows the user to "look around" the virtual environment by changing the field of view of the real-world camera.

[0045] In one embodiment, in addition to the orientation, the position of the virtual camera can also be changed in the virtual environment. For example, the position of the virtual camera can be changed by physically moving the mobile device or by using GUI affordances (virtual navigation buttons). In the former, position data (e.g., GNSS data) and / or inertial sensor data (e.g., accelerometer data) can be used to determine the position of the virtual camera in the virtual environment. In one embodiment, the virtual environment can be a 3D video, a 3D 360° video, or a 3D computer-generated image (CGI) that responds to the user's actions.

[0046] GUI 300 also includes several affordances for performing various tasks. Tab bar 304 allows the user to select photo editing options, such as invoking AR selfie recording. Tab bar 305 allows the user to select a camera function (e.g., photo, video, panorama, gallery). Tab bar 304 can be context-sensitive, so that the options in tab bar 304 can change based on the camera function selected in tab bar 305. In the example shown, the "Video" option is selected in tab bar 305, and the AR selfie recording option 311 is selected in tab bar 304.

[0047] To record an AR selfie, the GUI 300 includes a virtual record button 306 for recording the AR selfie to a local storage device (e.g., flash memory). A thumbnail image tray 309 can hold a thumbnail image of the recorded AR selfie, which can be selected to playback the corresponding AR selfie video in the viewport 301. A camera reverse button 307 allows the user to switch between the front camera and the rear camera. A microphone enable button 308 toggles one or more microphones of the mobile device 102 on and off. A done button 310 exits the GUI 300.

[0048] Figure 3B Different special effects are shown applied to the selfie object 302b and different virtual background content 303b. For example, the virtual background content can be a cartoon environment with animated cartoon characters and other objects. It should be understood that any virtual background content can be used in AR selfies. In some specific implementations, animated objects (e.g., animated particles such as snowflakes and sparks) can be inserted between the selfie object and the virtual background content to create a more aesthetically pleasing virtual environment, as shown in FIG. Figure 5 In one embodiment, the selfie object 302b may be given an edge treatment, such as a "glow" or outline around the image, or an "ink" outline. In one embodiment, an animated object may be inserted in front of the selfie objects 302a, 302b. For example, the selfie objects 302a, 302b may be surrounded by a floating text band or other animated object. In one embodiment, the selfie objects 302a, 302b may be layered over an existing real-world photo or video.

[0049] Figure 3C and Figure 3D According to the embodiment, a graphical user interface is shown in which different background scenes are selected and a recording view and a full-screen playback view are shown. Figure 3C , a recording view is shown where user 302c has selected virtual background 303c. Note that during recording, viewport 301 is not full screen to provide space for recording controls. Figure 3D , the full-screen playback view includes a scene selector 313 that can be displayed when user 302d has selected the "Scene" affordance 312. In one embodiment, scene selector 313 is a touch control that user 302d can swipe to select virtual background 303d, which in this example is a Japanese tea garden. Also note that virtual background 303d is now displayed full screen in viewport 311.

[0050] Figure 3E and Figure 3F A graphical user interface for recording and playing back selfies using a rear camera and showing a recording view and a full-screen playback view is shown according to an embodiment. Figure 3E, a recorded view with a virtual background 303e is shown. The virtual background 303e is what the user would see in front of him / her in the virtual environment through the rear camera. The user can select the affordance 307 to switch between the front camera and the rear camera. Figure 3F , the full-screen playback view includes a scene selector 313 that can be replaced when user 302d has selected the "Scene" affordance 312. In one embodiment, scene selector 313 can be swiped by user 302d to select virtual background 303f, which in this example is a Japanese tea garden. Also note that virtual background 303f is now displayed full screen in viewport 314. In one embodiment, when the user first selects the virtual environment, a predefined orientation is presented in the viewport.

[0051] Example system for generating AR selfies

[0052] Figure 4 is a block diagram of a system 400 illustrating the processing steps used in creating an AR selfie, according to an embodiment. The system 400 can be implemented in software and hardware. The front-facing camera 401 generates RGB video and the IR depth sensor 402 generates depth data, which are received by an audio / visual (A / V) processing module 403. The A / V processing module 403 includes software data types and interfaces to efficiently manage queues of video and depth data for distribution to other processes, such as a mask extraction module 409, which performs reference Figures 6A-6L The A / V processing module 403 also provides a foreground video 404 including an image of the self-portrait subject, which can optionally be provided with a motion effect 405a such as Figure 3A The mask extraction module 409 outputs a foreground alpha mask 410 which may optionally be processed by a motion effects module 405b.

[0053] For virtual background processing, one or more of a 2D image source 411, a 3D image source 412, or a 360° video source 413 may be used to generate virtual background content 415. In one embodiment, the 3D image source may be a rendered 3D image scene with 3D characters. Each of these media sources may be processed by a motion source module 412, which selects an appropriate source based on the virtual environment selected by the user. The motion synthesis module 406 generates a composite video from the foreground video 404, the foreground alpha mask 410, and the virtual background content 415, as shown in FIG. Figure 5 A motion effect 407 (eg, a blur effect) may optionally be applied to the composite video output by the motion synthesis module 406 to generate a final AR selfie 408 .

[0054] The accelerometer and gyroscope sensors 416 provide motion data that is processed by the motion processing module 417 to generate camera transformations, as shown in FIG. Figure 2A-2E During recording, real-time motion data from sensor 416 is used to generate an AR selfie and is stored in a local storage device (e.g., in a flash memory). When the AR selfie is played back, the motion data is retrieved from the local storage device. In one embodiment, in addition to the virtual camera orientation, the virtual camera position in the virtual environment can also be provided by the motion processing module 417 based on the sensor data. Using the virtual camera and position information, the user can walk around in a 3D scene with a 3D character.

[0055] Exemplary Synthesis Process

[0056] Figure 5 According to an embodiment, a compositing layer is shown for use in AR selfies. In one embodiment, alpha compositing is used to combine / blend video data containing an image of the selfie subject with virtual background content. An RGB depth mask ("RGB-D mask") includes outline information of the object projected onto a binary depth mask, which is used to combine the foreground image of the object with the virtual background content.

[0057] In the example shown, one or more animation layers 502 (only one layer is shown) are composited on the background content 501. A mask 503 is composited on the one or more animation layers 502, and foreground RGB video data 504 (including objects) is composited on the mask 503, resulting in a final composite AR selfie, which is then displayed through the viewport 301 presented on the display of the mobile device 102. In one embodiment, a motion effect can be applied to the composite video, such as a blur effect to hide any artifacts caused by the compositing process. In one embodiment, the animation layer can be composited in front of or behind the RGB video data 504.

[0058] Example process for generating RGB-D masks

[0059] In one embodiment, the depth sensor is an IR depth sensor. The IR depth sensor includes an IR projector and an IR camera, which may be an RGB camera operating in the IR spectrum. The IR projector projects a pattern of dots using IR light falling on a target in an image scene including an object. The IR camera sends a video feed of the distorted dot pattern to a processor of the depth sensor, and the processor calculates depth data from the displacement of the dots. On close targets, the pattern of dots is dense, and on distant targets, the pattern of dots is spread out. The depth sensor processor constructs a depth image or map that can be read by a processor of the mobile device. If the IR projector is offset relative to the IR camera, some of the depth data may be undefined. Typically, this undefined data is not used. However, in the disclosed mask generation process, the undefined data is used to improve segmentation and contour detection, resulting in a more seamless composite.

[0060] See also Figure 6A and Figure 6B , the mask generation process 600 can be divided into three stages: a pre-processing stage 603, an RGB-D mask extraction stage 604, and a post-processing stage 605. Process 600 takes as input RGB video data 601 comprising an image of a subject and a depth map 602 comprising depth data provided by an IR depth sensor. It should be noted that the depth map 602 includes shadow areas where the depth data is undefined. Note that the shadows along the left outline of the subject's face are thicker (more undefined data) than along the right outline of the subject's face. This is due to the offset between the IR projector and the IR camera. Each of stages 603-605 will be described below in turn.

[0061] See also Figure 6C , shows the steps of the pre-processing stage 603, which include histogram generation 606, histogram threshold segmentation 607, outer contour detection 608, inner contour detection 609, and coarse depth mask generation 610, iterative region growing 612, and 3D face mesh modeling 613. Each of these pre-processing steps will now be described in turn.

[0062] The histogram generation 606 places the depth data into bins. The histogram threshold segmentation step 607 is used to segment the foreground depth data from the background depth data by finding the “peaks and valleys” in the histogram. Figure 6D As shown, a histogram 614 is generated from the absolute distance data, where the vertical axis indicates the number of depth data values ​​(hereinafter referred to as "depth pixels") in each bin, and the horizontal axis indicates the distance values ​​provided by the depth sensor, which are absolute distances in this example. Note that in this example, the distance values ​​are in bins with indices that are multiples of 10.

[0063] from Figure 6D6. It can be seen that foreground pixels are clustered together in adjacent bins centered approximately at 550 mm, and background pixels are clustered together in adjacent bins centered approximately at 830 mm. Note that if an object is inserted between the object and the background or in front of the object, there may be additional clusters of distance data. A distance threshold (shown as line 615) can be established that can be used to segment pixels into foreground pixels and background pixels based on distance to create a binary depth mask. For example, each pixel with a distance less than 700 mm is designated as foreground and assigned a binary value of 255 for white pixels in the binary depth mask (e.g., assuming an 8-bit mask), and each pixel with a distance greater than 700 mm is designated as background and assigned a binary value of 0 for black pixels in the binary depth mask.

[0064] See also Figure 6E , a threshold 615 (e.g., at about 700 mm) is applied to the histogram 614 to generate two binary depth masks 616a, 616b for finding the inner and outer contours of the object, respectively. In one embodiment, the threshold 615 can be selected to be the average distance between the outermost bin of the foreground bin (the bin containing the pixels with the longest distance) and the innermost bin of the background bin (the bin containing the pixels with the shortest distance).

[0065] Although the pixel segmentation described above uses a simple histogram threshold segmentation method, other segmentation techniques may also be used, including but not limited to balanced histogram threshold segmentation, k-means clustering, and Otsu's method.

[0066] See again Figure 6E Steps 608 and 609 extract the inner and outer contours of the object from the binary depth masks 616a and 616b, respectively. A contour detection algorithm is applied to the depth masks 616a and 616b. An example contour detection algorithm is described in Suzuki, S. and Abe, K., Topological Structural Analysis of Digitized Binary Images by Border Following. CVGIP 30 1, pp. 32-46 (1985).

[0067] Depth mask 616a is generated using only defined depth data, and depth mask 616b is generated using both defined depth data and undefined depth data (shadow data). If depth masks 616a, 616b were to be combined into a single depth mask, the resulting combined depth mask would be similar to Figure 7CThe three-dimensional image 704 is shown, where the gray area between the inner and outer contours (called the "blended" area) includes undefined depth data that may include important contour details that should be included in the foreground. After the inner and outer contours are extracted, they can be smoothed using, for example, a Gaussian blur kernel. After the contours are smoothed, they are combined 618 into a rough depth mask 619, as shown in FIG. Figure 6F-Figure 6I As stated.

[0068] Figure 6F The use of a distance transform to create a rough depth mask 619 is shown. Outer contours 621 and inner contours 622 define a mixed region of undefined pixels (without defined depth data) between the contours. In some cases, some of the undefined pixels may contain important contour information that should be assigned to the foreground (assigned white pixels). To generate the rough depth mask 619, the object is vertically divided into left and right hemispheres, and a distance transform is performed on the undefined pixels in the mixed region.

[0069] In one embodiment, the vertical distance between the pixels of the inner contour 622 and the outer contour 621 is calculated, such as Figure 6F and Figure 6G Then, the probability density function of the calculated distance is calculated for the left hemisphere and the right hemisphere respectively, as shown in Figure 6H and Figure 6I As shown. The left and right hemispheres have different probability density functions because, as previously described, due to the offset between the IR projector and the IR camera, the shadows on the left side of the subject's face are coarser than the shadows on the right side of the subject's face. In one embodiment, a Gaussian distribution model is applied to the distances to determine the mean μ and standard deviation σ for each of the left and right hemispheres. The standard deviation σ or a multiple of the standard deviation (e.g., 2σ or 3σ) can be used as a threshold to compare with the distance in each hemisphere. Pixels in the undefined area (gray area) in the left hemisphere are compared with the threshold for the left hemisphere. Pixels with a distance less than or equal to the threshold are included in the foreground and are assigned a white pixel value. Pixels with a distance greater than the threshold are included in the background and are assigned a black pixel value. The same process is performed for the right hemisphere. The result of the above distance conversion is a coarse depth mask 619, which ends the pre-processing stage 603.

[0070] Example region growing / face mesh process

[0071] In some cases, the coarse mask 619 will have islands of undefined pixels in the foreground. For example, when taking a selfie outdoors in sunlight, the performance of the IR depth sensor degrades. Specifically, if the selfie subject is wearing sunglasses, the resulting depth map will have two black holes where their eyes are located due to sunlight reflecting off the sunglasses. These holes are visible in the coarse depth mask 619 and are filled with white pixels using an iterative region growing segmentation algorithm. In one embodiment, a histogram of the foreground RGB video data 601 can be used to determine an appropriate threshold for the region membership criterion.

[0072] See also Figure 6J-6L , a 3D facial mesh model 625 can be generated from the RGB video data 623. The facial mesh model 625 can be used to identify the locations of facial landmarks on the subject's face, such as sunglasses 624. The facial mesh model 625 can be overlaid on a coarse depth mask 626 to identify the location of the sunglasses 624. Any islands 628 of undefined pixels identified by the facial mesh model 625 in the foreground region 627 are filled with white pixels so that these pixels are included in the foreground region 627.

[0073] Figure 7A and Figure 7B According to an embodiment, a process for RGB-D mask extraction using a combination of RGB video data and a pre-processed depth mask 619 is shown. Figure 7A , the tripartite map module 701 generates a tripartite map 704 from the rough depth mask 619. In one embodiment, the tripartite map module 704 uses the same segmentation process used to generate the rough depth mask 619, or some other known segmentation technique (e.g., k-means clustering). The tripartite map 704 has three regions: a foreground region, a background region, and a mixed region. The tripartite map 704 is input into a Gaussian mixture model (GMM) 702 along with the RGB video data 601. The GMM 702 classifies the foreground region and the background region by a probability density function approximated by a Gaussian mixture (see Figure 7B ) is modeled as shown in formula [3]:

[0074]

[0075] The probability density function is used by the graph cut module 703 to perform segmentation using an iterative graph cut algorithm. An example graph cut algorithm is described in DMGreig, BT Porteous, and AH Seheult (1989), Exact maximum aposteriori estimation for binary images, Journal of the Royal Statistical Society Series B, 51, pp. 271-279. The refined depth mask 705 output by the graph cut module 703 is fed back into the triplanar module 701, and the process continues for N iterations or until convergence.

[0076] Figure 7C The results of the first two stages of the mask generation process 600 are shown. The depth map 602 is preprocessed into binary depth masks 616a and 616b, where depth mask 616a is generated using only defined depth data, while depth mask 616b is generated using both defined and undefined depth data. The binary depth masks 616a and 616b are then combined into a coarse depth mask 619 using a distance transform. The coarse depth mask 619 is input to the RGB-D mask extraction process 604, which uses an iterative graph cut algorithm and GMM to model the foreground and background regions of the tripartite map 704. The result of the RGB-D mask extraction process 604 is a refined mask 705.

[0077] Figure 8 According to an embodiment, a post-processing stage 605 is shown for removing artifacts added by the refinement process. In the post-processing stage 605, the distance conversion module 803 uses the reference Figure 6F-Figure 6I The same technique described above calculates the distance between the contours in the rough depth mask 619 and the refined mask 705. The distance check module 804 then compares the distance to a threshold. Any undefined pixels that are farther from the inner contour than the threshold are considered artifacts and assigned to the background area. In the example shown, the depth mask 805 includes artifacts 806 before post-processing. The final result of the post-processing stage 606 is the final AR selfie mask 808 used to synthesize the AR selfie, as shown in Figure 8. Figure 5 Note that the artifact 806 has been removed from the AR selfie mask 808 due to the post-processing described above.

[0078] Exemplary Process

[0079] Figure 9 is a flow chart of a process 900 for generating an AR selfie according to an embodiment. The process 900 may be performed using, for example, a reference Figure 11 The device architecture is implemented.

[0080] Process 900 may begin by receiving image data (e.g., video data) and depth data from an image capture device (e.g., a camera) and a depth sensor, respectively (901). For example, the image data may be RGB video data provided by a red, green, and blue (RGB) camera that includes an image of an object. The depth sensor may be an IR depth sensor that provides a depth map that can be used to generate an RGB-D mask, as described in reference to FIG. Figure 10 As stated.

[0081] Process 900 continues by receiving motion data from one or more motion sensors (902). For example, the motion data may be acceleration data and orientation data (e.g., angular rate data) provided by an accelerometer and a gyroscope, respectively. The motion data may be provided in the form of a coordinate transformation (e.g., a body-fixed quaternion). The coordinate transformation describes the orientation of the camera's field of view in a real-world reference coordinate system, which may be transformed into a virtual-world reference coordinate system using a camera transformation.

[0082] Process 900 continues by receiving virtual background content from a storage device (903). For example, the virtual background content can be a 2D image, a 3D image, or a 360° video. The virtual background content can be selected by the user through a GUI. The virtual background content can be extracted or sampled from any desired virtual environment, such as a famous city or a cartoon environment with animated cartoon characters and objects.

[0083] Process 900 continues by generating a virtual camera transform from the motion data ( 904 ).

[0084] The process 900 continues by generating a mask from the image data and the depth data (905). For example, the RGB-D mask may be as described in reference Figure 6I-6L The generated RGB-D mask includes contour information of the object and is used to synthesize the RGB video with the virtual background content.

[0085] Process 900 may continue with compositing image data, RGB-D mask, and virtual background content (905), as described in reference Figure 5 During this step, the camera transform is used to extract or sample appropriate virtual background content for compositing with the image data and the RGB-D mask (906). In one embodiment, one or more animation layers are also composited to provide, for example, animated particles (e.g., snowflakes, sparkles, fireflies). In one embodiment, the camera transform is adjusted to account for camera flips caused by the user flipping between the front camera and the rear camera, or vice versa, as described in reference to FIG. Figure 3A As stated.

[0086] Process 900 may continue with rendering for displaying the composite media (e.g., composite video) in the viewport of the mobile device (907). During the recording operation, the composite media is presented as a real-time video feed. As the user changes the field of view of the real-world camera, the virtual camera transitions synchronize with the real-world camera to update the virtual background content in real time. The recorded AR selfie video can be played back from the storage device through the viewport and can also be shared with others, for example, on a social network.

[0087] Figure 10 is a flow chart of a process 1000 for generating an AR selfie mask according to an embodiment. The process 1000 may be performed, for example, using reference Figure 11 The device architecture is implemented.

[0088] Process 1000 may begin by generating a histogram of depth data ( 1001 ) and applying one or more thresholds to the histogram to segment the depth data into foreground and background regions ( 1002 ).

[0089] Process 1000 continues by generating an outer contour and an inner contour of the object as a binary depth mask (1003). For example, the inner contour can be generated in a first binary depth mask using a contour detection algorithm and only defined depth data, and the outer contour can be generated in a second binary depth mask using a contour detection algorithm and depth data including both defined depth data and undefined depth data.

[0090] Process 1000 continues with optionally smoothing the inner and outer contours (1004). For example, the inner and outer contours may be smoothed using a Gaussian blur kernel.

[0091] Process 1000 continues by combining the outer contour and the inner contour to generate a coarse mask (1005).For example, a distance transform using a Gaussian distribution can be used to combine the first binary depth mask and the second binary depth mask into a combined coarse mask.

[0092] Process 1000 may continue to generate a refined mask (e.g., an RGB-D mask) using the coarse depth mask, the image data, and the depth data (1006). For example, an iterative graph cut algorithm may be used on the triplicate map generated by the coarse mask and the GMM to generate the RGB-D mask.

[0093] Process 1000 may continue by removing undefined areas and artifacts from the refined mask (1007). For example, islands of undefined pixels in the foreground region of the RGB-D mask caused by sunlight reflected off sunglasses may be identified and filled with white foreground pixels using an iterative region growing algorithm and / or a 3D facial mesh model, as shown in FIG. Figure 6J-6L As stated.

[0094] Exemplary device architecture

[0095] Figure 11 According to the embodiment shown for realizing reference Figures 1-10 The device architecture of the features and processes described herein. Architecture 1100 may include a memory interface 1102, one or more data processors, video processors, coprocessors, image processors, and / or other processors 1104, and a peripheral device interface 1106. Memory interface 1102, one or more processors 1104, and / or peripheral device interface 1106 may be separate components or may be integrated into one or more integrated circuits. The various components in architecture 1100 may be coupled via one or more communication buses or signal lines.

[0096] Sensors, devices, and subsystems can be coupled to the peripheral device interface 1106 to facilitate multiple functions. For example, one or more motion sensors 1110, light sensors 1112, and proximity sensors 1114 can be coupled to the peripheral device interface 1106 to facilitate motion sensing (e.g., acceleration, rotation rate), lighting, and proximity functions of the mobile device. A location processor 1115 can be connected to the peripheral device interface 1106 to provide geolocation and process sensor measurements. In some specific implementations, the location processor 1115 can be a GNSS receiver, such as a Global Positioning System (GPS) receiver chip. An electronic magnetometer 1116 (e.g., an integrated circuit chip) can also be connected to the peripheral device interface 1106 to provide data that can be used to determine magnetic north. The electronic magnetometer 1116 can provide data to an electronic compass application. The one or more motion sensors 1110 can include one or more accelerometers and / or gyroscopes configured to determine changes in the speed and direction of movement of the mobile device. A barometer 1117 can be configured to measure the atmospheric pressure surrounding the mobile device.

[0097] The camera subsystem 1120 and one or more cameras 1122 (e.g., front-facing and rear-facing cameras) for capturing digital photos and recording video clips include a video and image for generating AR selfies, as shown in FIG. Figures 1-10 As stated.

[0098] Communication functionality may be facilitated by one or more wireless communication subsystems 1124, which may include radio frequency (RF) receivers and transmitters (or transceivers) and / or optical (e.g., infrared) receivers and transmitters. The specific design and implementation of the communication subsystems 1124 may depend on the communication network or networks in which the mobile device is intended to operate. For example, the architecture 1100 may include a device designed to communicate over a GSM network, a GPRS network, an EDGE network, a Wi-Fi network, or a similar network. TM or Wi-Max TM Network and Bluetooth TMThe wireless communication subsystem 1124 may include a host protocol that allows the mobile device to be configured as a base station for other wireless devices.

[0099] The audio subsystem 1126 may be coupled to a speaker 1128 and a microphone 1130 to facilitate voice-enabled functions such as voice recognition, voice replication, digital recording, and telephony functions. The audio subsystem 1126 may be configured to receive voice commands from a user.

[0100] The I / O subsystem 1140 may include a touch surface controller 1142 and / or one or more other input controllers 1144. The touch surface controller 1142 may be coupled to a touch surface 1146 or pad. The touch surface 1146 and touch surface controller 1142 may detect contact and its movement or interruption, for example, using any of a variety of touch-sensitive technologies, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the touch surface 1146. The touch surface 1146 may include, for example, a touch screen. The I / O subsystem 1140 may include a haptic engine or device for providing tactile feedback (e.g., vibration) in response to commands from the processor.

[0101] One or more other input controllers 1144 may be coupled to other input / control devices 1148, such as one or more buttons, rocker switches, a thumb wheel, an infrared port, a USB port, and / or a pointer device such as a stylus. One or more buttons (not shown) may include an up / down button for volume control of the speaker 1128 and / or the microphone 1130. The touch surface 1146 or other controllers 1144 (e.g., buttons) may include or be coupled to fingerprint identification circuitry for use with fingerprint authentication applications to authenticate a user based on one or more of the user's fingerprints.

[0102] In one embodiment, pressing a button for a first duration may unlock touch surface 1146; and pressing the button for a second duration, which is longer than the first duration, may turn the mobile device on or off. A user may customize the functionality of one or more buttons. For example, touch surface 1146 may also be used to implement virtual or soft buttons and / or a virtual touch keyboard.

[0103] In some implementations, the computing device can present recorded audio and / or video files, such as MP3, AAC, and MPEG files. In some implementations, the mobile device can include the functionality of an MP3 player. Other input / output and control devices can also be used.

[0104] The memory interface 1102 can be coupled to a memory 1150. The memory 1150 may include high-speed random access memory and / or non-volatile memory, such as one or more magnetic disk storage devices, one or more optical storage devices, and / or flash memory (e.g., NAND, NOR). The memory 1150 may store an operating system 1152, such as iOS, Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or an embedded operating system such as VxWorks. The operating system 1152 may include instructions for handling basic system services and for performing hardware-related tasks. In some implementations, the operating system 1152 may include a kernel (e.g., a UNIX kernel).

[0105] The memory 1150 may also store communication instructions 1154 that facilitate communication with one or more additional devices, one or more computers, and / or one or more servers, such as, for example, instructions of a software stack for implementing wired or wireless communication with other devices. The memory 1150 may include: graphical user interface instructions 1156 that facilitate graphical user interface processing; sensor processing instructions 1158 that facilitate sensor-related processing and functions; telephony instructions 1160 that facilitate telephony-related processes and functions; electronic messaging instructions 1162 that facilitate electronic messaging-related processes and functions; web browsing instructions 1164 that facilitate web browsing-related processes and functions; media processing instructions 1166 that facilitate media-related processes and functions; GNSS / location instructions 1168 that facilitate general GNSS and location-related processes and instructions; and camera instructions 1170 that facilitate camera-related processes and functions for both the front-facing camera and the rear-facing camera.

[0106] The memory 1150 also includes a memory for executing reference Figures 1-10 The memory 1150 may also store other software instructions (not shown), such as security instructions, network video instructions for facilitating processes and functions related to network video, and / or online shopping instructions for facilitating processes and functions related to online shopping. In some implementations, the media processing instructions 1166 are divided into audio processing instructions and video processing instructions to facilitate processes and functions related to audio processing and processes and functions related to video processing, respectively.

[0107] Each of the instructions and applications identified above may correspond to an instruction set for performing one or more of the functions described above. These instructions need not be implemented as separate software programs, processes, or modules. Memory 1150 may include additional instructions or fewer instructions. Furthermore, various functions of the mobile device may be implemented in hardware and / or software, including in one or more signal processing and / or application specific integrated circuits.

[0108] The described features can advantageously be implemented in one or more computer programs executable on a programmable system, the programmable system comprising at least one input device, at least one output device, and at least one programmable processor coupled to receive data and instructions from a data storage system and to send data and instructions to the data storage system. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform an activity or produce a result. A computer program can be written in any programming language, including compiled and interpreted languages ​​(e.g., SWIFT, Objective-C, C#, Java), and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, browser-based web application, or other unit suitable for use in a computing environment.

[0109] For example, suitable processors for executing a program of instructions include both general-purpose and special-purpose microprocessors, as well as one or the sole processor of multiple processors or cores in any type of computer. Generally, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer will also include, or be operatively coupled to communicate with, one or more mass storage devices for storing data files; such devices include: magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include: all forms of non-volatile memory, including, for example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, an ASIC (Application Specific Integrated Circuit).

[0110] To provide for interaction with a user, these features may be implemented on a computer having a display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or retina display device for displaying information to the user. The computer may have a touch surface input device (e.g., a touch screen) or a keyboard and pointing device such as a mouse or trackball through which a user can provide input to the computer. The computer may have a voice input device for receiving voice commands from the user.

[0111] These features can be implemented in a computer system that includes back-end components such as a data server, or that includes middleware components such as an application server or an internet server, or that includes front-end components such as a client computer with a graphical user interface or an internet browser, or any combination thereof. The components of the system can be connected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include, for example, a LAN, a WAN, and the computers and networks that form the internet.

[0112] A computing system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The relationship between the client and the server arises by virtue of computer programs running on the respective computers and having a client-server relationship with each other. In some implementations, the server transmits data (e.g., an HTML page) to a client device (e.g., to display data to a user interacting with the client device and to receive user input from the user interacting with the client device). Data generated at the client device (e.g., results of user interactions) may be received at the server from the client device.

[0113] A system of one or more computers may be configured to perform a particular action by having software, firmware, hardware, or a combination thereof installed on the system that, when in operation, causes the system to perform the particular action. One or more computer programs may be configured to perform the particular action by including instructions that, when executed by a data processing device, cause the device to perform the particular action.

[0114] An application programming interface (API) may be used to implement one or more features or steps of the disclosed embodiments. An API may define one or more parameters passed between a calling application and other software code (e.g., an operating system, a library program, a function) that provides a service, provides data, or performs an operation or calculation. An API may be implemented as one or more calls in program code that send or receive one or more parameters through a parameter list or other structure based on a calling convention defined in an API specification document. A parameter may be a constant, a key, a data structure, an object, an object class, a variable, a data type, a pointer, an array, a list, or another call. API calls and parameters may be implemented in any programming language. A programming language may define the vocabulary and calling conventions that a programmer will use to access functionality that supports the API. In some implementations, an API call may report to an application the capabilities of a device running the application, such as input capabilities, output capabilities, processing capabilities, power capabilities, communication capabilities, and the like.

[0115] As mentioned above, some aspects of the subject matter of this specification include the collection and use of data from various sources to improve the services that mobile devices can provide to users. This disclosure contemplates that, in some cases, this collected data may identify a specific location or address based on device usage. Such personal information data may include location-based data, addresses, subscriber account identifiers, or other identifying information.

[0116] The present disclosure also envisions that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information data will comply with established privacy policies and / or privacy practices. Specifically, such entities should implement and adhere to privacy policies and practices that are recognized as meeting or exceeding industry or government requirements for maintaining the privacy and security of personal information data. For example, personal information from users should be collected for the entity's legitimate and reasonable purposes and not shared or sold outside of these legitimate purposes. In addition, such collection should only be carried out with the user's informed consent. In addition, such entities should take any necessary steps to safeguard and protect access to such personal information data and ensure that others who can access personal information data comply with their privacy policies and procedures. In addition, such entities may subject themselves to third-party assessments to demonstrate their compliance with widely accepted privacy policies and practices.

[0117] With respect to ad delivery services, the present disclosure also contemplates implementations where users can selectively block the use or access of personal information data. Specifically, the present disclosure contemplates providing hardware and / or software components to prevent or block access to such personal information data. For example, with respect to ad delivery services, the techniques of the present disclosure may be configured to allow users to opt-in or opt-out of the collection of personal information data during service registration.

[0118] Thus, while the present disclosure broadly covers the use of personal information data to implement one or more of the various disclosed embodiments, the present disclosure also contemplates that various embodiments may be implemented without access to such personal information data. That is, the various embodiments of the disclosed technology will not fail to function due to the absence of all or a portion of such personal information data. For example, content may be selected and delivered to a user by inferring preferences based on non-personal information data or an absolute minimum amount of personal information, such as content requested by a device associated with the user, other non-personal information available to a content delivery service, or publicly available information.

[0119] Although this specification includes many specific implementation details, these specific implementation details should not be understood as limiting the scope of any invention or the content that may be claimed, but should be understood as descriptions of the features of the specific embodiments of specific inventions. Certain features described in this specification in the context of different embodiments may also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in the form of any suitable sub-combination. In addition, although certain features may be described above as working in certain combinations and even initially claimed in this way, one or more features of the claimed combination may be removed from the combination in some cases, and the claimed combination may relate to a sub-combination or a variation of the sub-combination.

[0120] Similarly, although operations are shown in a particular order in the accompanying drawings, this should not be understood as requiring that such operations be performed in a sequential order or in the particular order shown, or that all of the operations shown be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the division of each system component in the above-described embodiment should not be understood as requiring such division in all embodiments, and it should be understood that the program components and system can generally be integrated together in a single software product or packaged in multiple software products.

Claims

1. A method comprising: capturing image data of an object in a real-world environment using a camera of a mobile device; capturing depth data using one or more sensors of the mobile device, the depth data indicating a distance of the object from the camera in the real-world environment, the depth data comprising first depth data and second depth data; Generating a mask from the depth data, comprising: projecting the depth data into a first binary depth mask having a foreground area and a background area, the first binary depth mask including the first depth data; projecting the depth data into a second binary depth mask having a foreground area and a background area, the second binary depth mask including the second depth data; obtaining inner contour data corresponding to the inner contour of the foreground object from the first binary depth mask; obtaining outer contour data corresponding to the outer contour of the foreground object from the second binary depth mask, the inner contour data and the outer contour data defining a mixed region of depth data; and generating a mask by combining the inner contour data and the outer contour data, and including the depth data in the mixed area in one of a foreground area or a background area of ​​the mask; generating composite image data using at least the image data and the mask; and The composite image data is caused to be displayed on a display of the mobile device. 2 . The method of claim 1 , wherein the second depth data comprises undefined depth data provided by the one or more sensors, and wherein the first depth data does not include the undefined depth data.

3. The method of claim 1 , wherein projecting the depth data onto the first binary depth mask or the second binary depth mask comprises: generating a histogram of the depth data; as well as A threshold is applied to the histogram to segment the depth data into the foreground region or the background region corresponding to at least one of the first binary depth mask or the second binary depth mask.

4. The method of claim 1 , wherein generating the mask by combining the inner contour data and the outer contour data comprises: calculating a first set of distances between the inner contour and the outer contour; (i) calculating a first probability density function for the first set of distances for a first subset, and (ii) calculating a second probability density function for the first set of distances for a second subset; comparing a first set of hybrid depth data located in the hybrid region to one or more characteristics of the first probability density function; comparing a second set of hybrid depth data in the hybrid region to one or more characteristics of the second probability density function; identifying, based on a result of the comparison, depth data in the mixed area belonging to the foreground area; as well as The identified depth data is added to the foreground area.

5. The method of claim 1 , wherein the mask is a coarse mask, the method further comprising: A refined mask is generated by applying an iterative segmentation method to the coarse mask.

6. The method according to claim 5, further comprising: identifying one or more holes in the foreground region including the second depth data using an iterative region growing method and a threshold determined from the image data; and The second depth data is assigned to the foreground area.

7. The method according to claim 5, further comprising: identifying one or more holes in the foreground region comprising undefined depth data using a facial mesh model generated from the image data; and The second depth data is assigned to the foreground area.

8. The method of claim 7, wherein identifying the one or more holes in the foreground area further comprises: using the facial mesh to match the aperture to sunglasses worn by the subject in the foreground area; determining an area in the foreground area that overlaps with the facial mesh; as well as The holes are filled based on the determined overlap.

9. The method according to claim 5, further comprising: combining the coarse mask and the refined mask into a combined mask; identifying one or more artifacts in the combined mask; as well as A distance transform is used to remove the one or more artifacts from the combined mask.

10. A system comprising: one or more processors; A memory, coupled to the one or more processors and storing instructions, which, when executed, cause the one or more processors to perform operations including: capturing image data of an object in a real-world environment using a camera of a mobile device; capturing depth data using one or more sensors of the mobile device, the depth data indicating a distance of the object from the camera in the real-world environment, the depth data comprising first depth data and second depth data; Generating a mask from the depth data, comprising: projecting the depth data into a first binary depth mask having a foreground area and a background area, the first binary depth mask including the first depth data; projecting the depth data into a second binary depth mask having a foreground area and a background area, the second binary depth mask including the second depth data; obtaining inner contour data corresponding to the inner contour of the foreground object from the first binary depth mask; obtaining outer contour data corresponding to the outer contour of the foreground object from the second binary depth mask, the inner contour data and the outer contour data defining a mixed region of depth data; and generating a mask by combining the inner contour data and the outer contour data, and including the depth data in the mixed area in one of a foreground area or a background area of ​​the mask; generating composite image data using at least the image data and the mask; and The composite image data is caused to be displayed on a display of the mobile device. 11 . The system of claim 10 , wherein the second depth data comprises undefined depth data provided by the one or more sensors, and wherein the first depth data does not include the undefined depth data.

12. The system of claim 10, wherein projecting the depth data onto the first binary depth mask or the second binary depth mask comprises: generating a histogram of the depth data; as well as A threshold is applied to the histogram to segment the depth data into the foreground region or the background region corresponding to at least one of the first binary depth mask or the second binary depth mask.

13. The system of claim 10, wherein generating the mask by combining the inner contour data and the outer contour data comprises: calculating a first set of distances between the inner contour and the outer contour; (i) calculating a first probability density function for the first set of distances for a first subset, and (ii) calculating a second probability density function for the first set of distances for a second subset; comparing a first set of hybrid depth data located in the hybrid region to one or more characteristics of the first probability density function; comparing a second set of hybrid depth data in the hybrid region to one or more characteristics of the second probability density function; identifying, based on a result of the comparison, depth data in the mixed area belonging to the foreground area; as well as The identified depth data is added to the foreground area.

14. The system of claim 10, wherein the mask is a coarse mask, the operations further comprising: A refined mask is generated by applying an iterative segmentation method to the coarse mask.

15. The system of claim 14, the operations further comprising: identifying one or more holes in the foreground region comprising undefined depth data using a facial mesh model generated from the image data; and The second depth data is assigned to the foreground area.

16. The system of claim 15, wherein identifying the one or more holes in the foreground area further comprises: using the facial mesh to match the aperture to sunglasses worn by the subject in the foreground area; determining an area in the foreground area that overlaps with the facial mesh; as well as The holes are filled based on the determined overlap.

17. The system of claim 15, wherein the operations further comprise: combining the coarse mask and the refined mask into a combined mask; identifying one or more artifacts in the combined mask; as well as A distance transform is used to remove the one or more artifacts from the combined mask.

18. One or more non-transitory storage media storing instructions that, when executed, cause one or more processors to perform operations comprising: capturing image data of an object in a real-world environment using a camera of a mobile device; capturing depth data using one or more sensors of the mobile device, the depth data indicating a distance of the object from the camera in the real-world environment, the depth data comprising first depth data and second depth data; Generating a mask from the depth data, comprising: projecting the depth data into a first binary depth mask having a foreground area and a background area, the first binary depth mask including the first depth data; projecting the depth data into a second binary depth mask having a foreground area and a background area, the second binary depth mask including the second depth data; obtaining inner contour data corresponding to the inner contour of the foreground object from the first binary depth mask; obtaining outer contour data corresponding to the outer contour of the foreground object from the second binary depth mask, the inner contour data and the outer contour data defining a mixed region of depth data; and generating a mask by combining the inner contour data and the outer contour data, and including the depth data in the mixed area in one of a foreground area or a background area of ​​the mask; generating composite image data using at least the image data and the mask; and The composite image data is caused to be displayed on a display of the mobile device. 19 . The non-transitory storage medium of claim 18 , wherein the second depth data includes undefined depth data provided by the one or more sensors, and wherein the first depth data does not include the undefined depth data.

20. The non-transitory storage medium of claim 18, wherein generating the mask by combining the inner contour data and the outer contour data comprises: calculating a first set of distances between the inner contour and the outer contour; (i) calculating a first probability density function for the first set of distances for a first subset, and (ii) calculating a second probability density function for the first set of distances for a second subset; comparing a first set of hybrid depth data located in the hybrid region to one or more characteristics of the first probability density function; comparing a second set of hybrid depth data in the hybrid region to one or more characteristics of the second probability density function; identifying, based on a result of the comparison, depth data in the mixed area belonging to the foreground area; as well as The identified depth data is added to the foreground area.

21. The non-transitory storage medium of claim 18, wherein the mask is a coarse mask, the operations further comprising: generating a refined mask by applying an iterative segmentation method to the coarse mask; combining the coarse mask and the refined mask into a combined mask; identifying one or more artifacts in the combined mask; as well as A distance transform is used to remove the one or more artifacts from the combined mask.

Citation Information

Patent Citations

  • Modulation of background substitution based on camera attitude and motion

    US20090315915A1

  • Portable electronic devices with integrated image / video compositing

    US20160057363A1

Cited By

  • Method and system for displaying composite image data

    CN120782938A