Apparatus, method, and program
The apparatus and method address the issue of obscured faces in VR/MR by reconstructing hidden facial regions using pre-capture images and machine learning, ensuring full 3D face visibility and improved immersion.
Patent Information
- Application Number
- JP2025519870
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-14
- Filing Date
- 2023-10-20
- Publication Date
- 2025-11-27
AI Technical Summary
Existing virtual reality and mixed reality systems obscure the upper part of a user's face when using headsets, preventing full 3D face visibility.
An apparatus and method for receiving images before and after headset obstruction, determining the headset's orientation and position, and performing region swaps to reconstruct the obscured facial regions using alignment and machine learning techniques.
Reconstructs the obscured facial regions in real-time, providing a complete 3D face representation in virtual reality environments, enhancing user interaction and immersion.
Smart Images

Figure 2025538281000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 380,452, filed October 21, 2022, and U.S. Provisional Patent Application No. 63 / 383,583, filed November 14, 2022, both of which are incorporated herein in their entireties. background Technical Field FIELD OF THE DISCLOSURE This disclosure relates generally to video image processing in virtual reality environments. [Background technology]
[0002] 2. Description of Related Art Recent advances in mixed reality have made it practical for people to participate in virtual meetings and gatherings using headsets or head-mounted displays (HMDs) and see each other's 3D faces in real time. Such gatherings are becoming increasingly necessary as some situations, such as pandemics and other disease outbreaks, prevent people from meeting in person.
[0003] To see each other's 3D faces using virtual reality and / or mixed reality, a headset is required. However, when the headset is placed on a user's face, no one can actually see the entire 3D face of the other person because the upper part of the face is obscured by the headset. Therefore, finding a way to recover the upper part of the face that is obscured from the 3D face upon removal of the headset is important to the overall performance of virtual reality and / or mixed reality. Summary of the Invention
[0004] summary An embodiment of the present disclosure provides an apparatus and method for receiving a first image of a user during a pre-capture process, receiving a second image of the user that is partially obscured by a wearable device, determining the orientation and position of the wearable device, identifying the position of the wearable device in the received second image, performing a region swap on the second image by replacing the obscured portion of the user with a corresponding region taken from the first image, and generating a third image composed of the second image and the first image as output to a display of the wearable device.
[0005] In other embodiments, which may include other embodiments described herein, the received first image includes one or more items of facial information, and the region swap is performed by using the determined orientation and position of the wearable device in the received second image to determine a correspondence between the orientation and position of the wearable device and one or more of the one or more items of facial information from the first image.
[0006] In other embodiments, which may include other embodiments described herein, the one or more items of facial information include one or more yaw, pitch, and / or roll orientations, one or more facial expressions, or blinks.
[0007] In other embodiments, which may include other embodiments described herein, the apparatus and method further create a third image by combining one or more two-dimensional images of areas in the first image that correspond to hidden portions of the second image into the third image as output to a display of the wearable device.
[0008] In other embodiments, which may include other embodiments described herein, the apparatus and method further include providing the received first images in real time to a first image processing channel that uses alignment and position information of the wearable device to select each of the first images as a candidate replacement image based on facial information associated with the first images, providing the received second images in real time to a second image processing channel that extracts regions of the second images that correspond to the occluded regions, and performing region swaps using portions of the candidate replacement images that correspond to the regions extracted from the second images based on orientation, facial expression, or blink information extracted from the second images.
[0009] In other embodiments, which may include other embodiments described herein, the apparatus and method further extract regions of the second image by providing the second image to a trained machine learning model that is trained based on multiple images of general users wearing wearable devices and that classifies regions in the second image that represent wearable devices.
[0010] In other embodiments, which may include other embodiments described herein, the candidate replacement images include eye or nose regions, and execution of the stored instructions further causes the one or more processors to perform the region swap by inpainting the extracted region from the second image with the eye or nose region of the first image.
[0011] In other embodiments, which may include other embodiments described herein, the generated third image includes eye and nose regions inpainted with the first image having an orientation, expression, or blink substantially similar to the orientation, expression, or blink of the second image.
[0012] According to other embodiments, which may include other embodiments described herein, there is provided an apparatus and method for acquiring an image of a human face, detecting landmarks in the acquired image of the human face, acquiring landmarks in a reference image of the human face, aligning the acquired landmarks in the image of the human face with the landmarks in the reference image of the human face, generating features based on the aligned landmarks, classifying the acquired image of the human face using a trained machine learning model, and identifying the presence or absence of respective facial motion units in the acquired image of the human face based on the generated features.
[0013] In other embodiments, which may include other embodiments described herein, the reference image of the human face is either an image generated from the acquired image of the human face or an image different from the acquired image of the human face.
[0014] Other embodiments may include other embodiments described herein, and further include: after acquiring an image of a human face, determining whether a reference image of the human face in the acquired image is stored in memory; if it is determined that a reference image corresponding to the acquired image is not stored, acquiring a standard image of a human face to be used as a reference image for alignment; and if it is determined that a reference image corresponding to the acquired image is stored, using the stored reference image for alignment.
[0015] In other embodiments, which may include other embodiments described herein, the apparatus and method further use the stored reference image for alignment if it is determined that the stored reference image was generated based on a threshold number of image frames of a human face in the acquired image.
[0016] In other embodiments, which may include other embodiments described herein, the apparatus and method further include, after acquiring the image of the human face, determining whether a reference image of the human face in the acquired image is stored in the memory, and if it is determined that a reference image is not stored, using a standard image of the human face for alignment for a predetermined number of acquired image frames including the image of the human face for alignment, storing the acquired human face images from consecutive frames, and averaging the stored images to generate a user-specific reference image.
[0017] In other embodiments, which may include other embodiments described in this specification, the device and method further provide a classified image in which the specific facial motion unit is present to the image processing application in response to the image processing application determining that the live captured image contains the specific facial motion unit, and the image processing application replaces a portion of the live captured image with a corresponding portion of the provided classified image.
[0018] Other embodiments, which may include other embodiments described herein, provide an apparatus and method for obtaining information representing the positions of the upper and lower eyelids from a series of facial images, obtaining information representing the positions of the upper and lower parts of the face from the series of facial images to determine the length of the face, determining the occurrence of a user's blink in the series of images based on the relative positions of the upper eyelid and the lower eyelid in relation to the length of the face, extracting a first frame including a blink and a second frame not including a blink from the series of images, and replacing a facial area in the second series of images with the first frame or the second frame based on a predetermined replacement rule.
[0019] In other embodiments, which may include other embodiments described herein, the apparatus and method further remove baseline data from the series of images that represent differences in relative position greater than a predetermined distance; compare the baseline-removed series of images to a first threshold that indicates the likelihood of a blink occurring; compare the baseline-removed series of images to a second threshold that is less than the first threshold and represents the duration of a blink; and identify segments in the series of images that exceed both the first and second thresholds as blink segments.
[0020] In other embodiments, which may include other embodiments described herein, the device and method further determine a blink occurrence principle based on one or more features associated with the determined blink occurrence of the user, and replace the face region in the second series of images with the first frame or the second frame based on the determined blink occurrence principle.
[0021] In other embodiments, which may include other embodiments described herein, the blinking principle is determined using a statistical analysis of the occurrence of blinks within at least one series of images of the user.
[0022] In other embodiments, which may include other embodiments described herein, the one or more characteristics include at least one or both of a time interval between the determined blink occurrences and a blink duration during the determined blink occurrences.
[0023] Other embodiments, which may include other embodiments described herein, include identifying eye regions for the second series of images based on one or more facial landmarks, generating eye meshes for the identified eye regions, and replacing the eye mesh in the second series of images with the first frame if it is determined that a blink is likely to occur, and replacing the eye mesh in the second series of images with the second frame if it is determined that a blink is not likely to occur.
[0024] In other embodiments, which may include other embodiments described herein, the apparatus and method further store the extracted first and second frames in a memory device, and retrieve the extracted first and second frames from the memory device to replace the region in the second series of images.
[0025] According to other embodiments, which may include other embodiments described herein, there is provided an apparatus and method for receiving position and orientation information from a wearable device worn by a user and having a first time signature, capturing an image of the user wearing the wearable device using an imaging device having a second time signature, determining an offset between the first and second time signatures using the position and orientation information of the wearable device and the orientation and position information extracted from the captured image, and using the determined offset as a reference time to synchronize timing between the imaging device and the wearable device.
[0026] In other embodiments, which may include other embodiments described herein, the method and apparatus further generate, for each captured image of a user wearing the wearable device, a bounding box surrounding the wearable device in the captured image, obtain coordinates of the generated bounding box in the captured image, and determine an offset using the position and orientation information received from the wearable device and the obtained coordinates of the bounding box in the captured image.
[0027] In other embodiments, which may include other embodiments described in this specification, the method and apparatus further determine the offset by performing a cross-correlation process using the received orientation information of the wearable device and information about the position of the wearable device in the captured image.
[0028] In other embodiments, which may include other embodiments described herein, the orientation information is first orientation information of the wearable device, and the position of the wearable device in the captured image is based on specific coordinates associated with a bounding box surrounding the wearable device in the captured image.
[0029] In other embodiments, which may include other embodiments described in this specification, the first orientation information of the wearable device is a pitch value of the wearable device, and the specific coordinate associated with the bounding box in the captured image is a Y coordinate value.
[0030] In other embodiments, which may include other embodiments described in this specification, the first orientation information of the wearable device is a yaw value of the wearable device, and the specific coordinate associated with the bounding box in the captured image is an X coordinate value.
[0031] In other embodiments, which may include other embodiments described herein, the method and apparatus further generate a cross-correlation coefficient between a signal representing position and orientation information received from the wearable at a first timestamp and a captured image of a user wearing the wearable at a second timestamp over a predetermined number of captured image frames, and use the generated cross-correlation coefficient as an offset value to synchronize the timestamps.
[0032] In other embodiments, which may include other embodiments described herein, the method and apparatus further shift the frames of the captured image forward by a predetermined number of frames or shift backward by a predetermined number of frames based on the determined offset.
[0033] In other embodiments, which may include other embodiments described in this specification, the method and apparatus further determine the offset by performing a cross-correlation process using multiple items of position and orientation information of the received wearable device and multiple items of position information of the wearable device in the captured image.
[0034] In other embodiments, which may include other embodiments described herein, the plurality of items of orientation information include at least two or more of pitch, yaw, roll, X, Y, and Z values received from the wearable device, and the plurality of items of position of the wearable device in the captured image are based on at least two coordinate values associated with a bounding box surrounding the wearable device in the captured image.
[0035] In other embodiments, which may include other embodiments described herein, the apparatus and method further include generating a first data set having a single dimension, the first data set including a plurality of received position and orientation data items, generating a second data set having a single dimension, the second data set including at least two position items of the wearable device in the captured image, determining a relative entropy between the first and second data sets, and obtaining an offset value based on the determined maximum relative entropy value.
[0036] In other embodiments, which may include other embodiments described herein, the apparatus and method further receive a series of images of a user wearing a wearable device, determine a position and orientation of the wearable device based on position and orientation information obtained from one or more sensors of the wearable device and the position and orientation of the wearable device determined from the received series of images, and estimate a pose of the user in the received series of images based on the determined position, location, and orientation.
[0037] In other embodiments, which may include other embodiments described herein, the apparatus and method further determine the position and orientation of the wearable device based on information located on the wearable device.
[0038] In other embodiments, which may include other embodiments described herein, the apparatus and method further determine the location and orientation of the wearable device by estimating one or more landmarks of the user that are obscured by the wearable device, and obtain the location and orientation information based on the one or more estimated landmarks.
[0039] In other embodiments, which may include other embodiments described herein, the apparatus and method further generate a bounding box surrounding the wearable device in each image of the sequence of images, inpaint the generated bounding box in each image, estimate one or more facial landmarks occluded by the wearable device, and obtain coordinate and orientation information representing the one or more landmarks inpainted in the images.
[0040] In other embodiments, which may include other embodiments described herein, the apparatus and method further estimate the user's pose by inpainting one or more features of the user that are obscured by the wearable device, feeding the inpainted image to a trained machine learning model trained to predict facial features, predicting landmarks on the user in areas not covered by the wearable device, and verifying the estimated pose by feeding the inpainted image to a trained machine learning model trained to predict facial landmarks.
[0041] In other embodiments, which may include other embodiments described herein, the apparatus and method further include generating input images by inpainting one or more features of the user that are obscured by the wearable device into each image of the sequence of images, feeding the generated input images to a trained machine learning model trained to generate facial landmarks, obtaining estimated facial landmarks using output from the trained machine learning model, obtaining projected facial landmarks on the user's face in areas not obscured by the wearable device from each image of the sequence of images, and estimating the user's head pose based on the estimated facial landmarks and the projected facial landmarks.
[0042] In other embodiments, which may include other embodiments described herein, the apparatus and method further generate a user interface displayable on the wearable device indicating a target position for the wearable device in the virtual reality environment, instruct the user to move the wearable device so that the wearable device is at the target position by displaying one or more graphical elements within the generated user interface, and estimate the user's posture using the coordinates of the target position and position and orientation information obtained from one or more sensors on the wearable device.
[0043] In other embodiments, which may include other embodiments described herein, the target location is determined based on a predetermined first target area displayed within the user interface and a second target area corresponding to a bounding box surrounding the wearable device.
[0044] Other embodiments may include other embodiments described herein, and further use the current position of the wearable device, as determined by the received position and orientation of the wearable device, to generate one or more image elements to instruct the user to move to a substantial center point of both the first target area and the second target area.
[0045] In other embodiments, which may include other embodiments described herein, the target location is determined based on the orientation of the wearable device as determined by the orientation received from the wearable device.
[0046] In other embodiments, which may include other embodiments described herein, the apparatus and method further include taking a series of images of a user wearing the wearable device, estimating positions of one or more facial landmarks in the face obscured by the wearable device, and extracting actual facial landmarks in areas of the face that are not obscured by the wearable device; using a predetermined three-dimensional model of the face of the human wearing the wearable device, the predetermined three-dimensional model including known positions of facial landmarks in areas obscured by the wearable device and known positions of facial landmarks in areas of the face that are not obscured by the wearable device; and obtaining head pose information by aligning the estimated positions of the one or more facial landmarks in the areas of the face that are not obscured by the wearable device with the known positions of the facial landmarks in the areas that are not obscured by the wearable device from the three-dimensional model.
[0047] According to other embodiments, which may include other embodiments described herein, there is provided an apparatus and method for determining one or more regions to recolor in a target image using corresponding one or more regions in a source image, performing a color conversion on the source image from a first color space to a second color space, performing a color conversion on the target image from the first color space to the second color space, and if it is determined that there is a correspondence between one or more features in one or more regions of the source image and one or more features in one or more regions of the target image, performing a recolor process on one or more regions in the target image using the color-converted source image in the second color space, and performing a color conversion on the target image from the second color space to the first color space.
[0048] In other embodiments, which may include other embodiments described herein, the apparatus and method further include, when it is determined that one or more features in one or more regions of the source image do not correspond to one or more features in one or more regions of the target image, using a reference image color converted from a first color space to a second color space, performing a first recolor operation using one or more regions of the source image and the reference image in the second color space, and performing a second recolor operation using one or more regions of the target image and one or more regions of the recolored reference image.
[0049] In other embodiments, which may include other embodiments described herein, the apparatus and method further identify one or more shared regions in the source image and the target image, obtain a reference image correlating the source image and the target image, perform a color transformation on the source image based on at least one shared region in the source image that is common with the reference image, and replace a region in the target image with a color-transformed source image that has been transformed based on the commonality of features between the source image and the reference image.
[0050] In other embodiments, which may include other embodiments described herein, the source image is an image previously captured by an imaging device and is used to replace at least a portion of the target image.
[0051] In other embodiments, which may include other embodiments described herein, the target image is a live captured image of a user wearing a wearable device that obscures at least a portion of the user's face.
[0052] In other embodiments, which may include other embodiments described herein, the source image includes segment information identifying multiple segments, each having a respective color value, and further identifies one or more segments in a live-captured target image, and performs recolor processing of the target image using segment information from the source image that corresponds to the one or more segments identified in the live-captured target image.
[0053] In other embodiments, which may include other embodiments described herein, the apparatus and method further include acquiring a source image, acquiring a target image, acquiring a reference image, converting the source image from a first color space to a second color space, converting the target image from the first color space to the second color space, converting the reference image from the first color space to the second color space, performing a color transform on a region of the target image in the second color space based at least in part on the reference image in the second color space, and performing a color transform on a region of the source image in the second color space based at least in part on the reference image in the second color space.
[0054] According to other embodiments, which may include other embodiments described herein, there is provided an apparatus and method for acquiring an image of a user wearing a head-mounted display device that partially obscures areas of the user's face, inferring facial landmarks in the partially obscured areas of the acquired image based on a facial model and the head-mounted display device worn on the face, and generating an image that includes the inferred facial landmarks and actual facial landmarks acquired from areas of the face that are not obscured by the head-mounted display device.
[0055] In other embodiments, which may include other embodiments described herein, the apparatus and method further estimate facial landmarks in partially occluded areas based on the orientation and position of the head-mounted display.
[0056] In other embodiments, which may include other embodiments described herein, the model includes a 3D point cloud of the face and a 3D point cloud of the head-mounted display projected onto the image plane to obtain a 2D point cloud representing a face wearing the head-mounted display.
[0057] In other embodiments, which may include other embodiments described herein, the apparatus and method further utilize an image segmentation process to generate a first bounding box surrounding the head-mounted display in the acquired image, generate a second bounding box for the head-mounted display based on the model by projecting a cloud of three-dimensional model points from the model onto the two-dimensional image plane, and align the first and second bounding boxes to generate a target image that includes estimated facial landmarks in areas obscured by the head-mounted display device.
[0058] In other embodiments, which may include other embodiments described herein, the apparatus and method further include acquiring a 3D model of the face, acquiring a 3D model of the head-mounted display, acquiring a 3D model of the face wearing the head-mounted display, acquiring an orientation of the face in the image, generating a first bounding box for the face in the image, projecting the 3D model of the face and the model of the head-mounted display onto the image plane, generating a second bounding box for the projected model of the head-mounted display, generating a transformation based on the first bounding box and the second bounding box, and inferring landmarks in the image of the face based on the transformation.
[0059] According to other embodiments, which may include other embodiments described herein, methods and apparatus are provided for obtaining image frames that are encoded and communicated to an external device over a network, encoding the image frames for network communication by including transparency information in the color space information of the image frames, and transmitting the encoded image frames to an external device that renders the image frames using the color space information and the transparency information.
[0060] According to other embodiments, which may include other embodiments described herein, the encoding includes determining, for each pixel in the image frame, whether the pixel is a foreground pixel or a background pixel, and for pixels determined to be foreground pixels, setting transparency information equal to or greater than a predetermined threshold that indicates to an external device that the determined foreground pixel should be displayed using color space information associated with the pixel, and for pixels determined to be background pixels, setting transparency information equal to or less than a predetermined threshold that indicates to an external device that the determined background pixel should be displayed as transparent.
[0061] According to other embodiments, which may include other embodiments described in this specification, for pixels determined to be foreground pixels, the foreground pixels are reduced to generate transparency information so that the transparency information and color space information do not overlap.
[0062] According to other embodiments, which may include other embodiments described herein, the encoding further comprises splitting the image frame into a first tensor containing color space information and a second tensor containing transparency information, and upon splitting, adding timestamp information to each of the first tensors, and separately transmitting the first tensor and the second tensor to an external device that synchronizes the transparency information and color space information for display using the timestamps associated with the first tensor and the second tensor.
[0063] In another embodiment, an apparatus is provided that includes one or more memories that store instructions and one or more processors, the stored instructions, when executed, causing the one or more processors to perform operations in accordance with any of the embodiments described herein.
[0064] In another embodiment, a server is provided that includes one or more memories that store instructions and one or more processors, where the stored instructions, when executed, cause the one or more processors to perform operations in accordance with any of the embodiments described herein.
[0065] In another embodiment, a computer-readable storage medium is provided having stored thereon instructions that, when executed by one or more processors, are configured to cause an apparatus or device to perform one or more of the methods of any of the embodiments described herein.
[0066] These and other objects, features, and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure, taken in conjunction with the accompanying drawings and claims. [Brief explanation of the drawings]
[0067] [Figure 1] FIG. 1 illustrates a virtual reality imaging and display system according to the present disclosure. [Figure 2] FIG. 2 illustrates an embodiment of the present disclosure. [Figure 3] FIG. 3 illustrates a virtual reality environment rendered to a user according to the present disclosure. [Figure 4] FIG. 4 illustrates a block diagram of an exemplary system according to the present disclosure. [Figure 5] FIG. 5 illustrates the 3D perception of a 2D human image in a 3D virtual environment according to the present disclosure. [Figure 6A] , [Figure 6B] 6A and 6B illustrate an HMD removal process according to the present disclosure. [Figure 7] FIG. 7 is a flowchart illustrating an exemplary HMD removal process according to this disclosure. [Figure 8] FIG. 8 illustrates an exemplary orientation during the pre-capture process according to the present disclosure. [Figure 9] FIG. 9 illustrates an example IMU alignment with respect to time shift and orientation according to the present disclosure. [Figure 10] FIG. 10 illustrates an exemplary 3D geometry of a human head, HMD, and camera system according to the present disclosure. [Figure 11] FIG. 11 is a flow diagram of a real-time HMD removal process performed in accordance with the present disclosure. [Figure 12]FIG. 12 is a flow diagram of a real-time HMD removal process performed in accordance with the present disclosure. [Figure 13] FIG. 13 is a flow diagram of a process for determining a user's position within a series of captured image frames. [Figure 14] FIG. 14 shows boundaries defined to determine the user's position within the image frame. [Figure 15] FIG. 15 is a graph showing the state vector associated with the user's position within an image frame. [Figure 16] FIG. 16 is a diagram showing a user wearing an HMD device and ready to take a picture. [Figure 17] FIG. 17 is a diagram showing a user wearing an HMD in a position where the user should be guided to face an imaging device in the present disclosure. [Figure 18] FIG. 18 is a diagram illustrating an overview of a standard face classifier used in this disclosure. [Figure 19] FIG. 19 illustrates an exemplary algorithm used in this disclosure to determine a standard face. [Figure 20] FIG. 20 illustrates an exemplary balancing algorithm in the present disclosure. [Figure 21] FIG. 21 illustrates an exemplary balancing algorithm according to the present disclosure. [Figure 22] FIG. 22 is a diagram illustrating a blink detection process according to the present disclosure. [Figure 23] FIG. 23 is a graphical illustration of a normalized determination of blinks in a series of image frames according to the present disclosure. [Figure 24] FIG. 24 is a flow diagram of the blink detection process according to the present disclosure. [Figure 25] FIG. 25 is a diagram illustrating a three-stage blink detection process according to the present disclosure. [Figure 26] FIG. 26 is a histogram illustrating a new set of artificially generated user blinks according to the present disclosure. [Figure 27]FIG. 27 is a flow diagram detailing the generation of a video with desired blinking in accordance with the present disclosure. [Figure 28] FIG. 28 is a diagram illustrating the results of the processing performed in FIG. 27 in the present disclosure. [Figure 29] FIG. 29 is a graph illustrating the correlation between a signal from an HMD device and information from a video frame displayed within the HMD device according to the present disclosure. [Figure 30] FIG. 30 is a graph illustrating the maximum offset between signals according to the present disclosure. [Figure 31] FIG. 31 is a flow diagram showing in detail the first alignment process performed in the present disclosure. [Figure 32] FIG. 32 is a histogram obtained from the second alignment process performed in the present disclosure. [Figure 33] FIG. 33 is a flow diagram showing in detail the second alignment process performed in the present disclosure. [Figure 34] FIG. 34 is a schematic diagram of the alignment process for aligning the HMD with the image captured by the imaging device. [Figure 35] FIG. 35 is a diagram illustrating a positioning process for aligning the HMD according to the present disclosure with an image captured by an imaging device. [Figure 36] FIG. 36 illustrates an exemplary approach for estimating an offset constant according to the present disclosure. [Figure 37] FIG. 37 is a diagram showing images processed by the alignment process of FIGS. 34 to 36 according to the present disclosure. [Figure 38] FIG. 38 is a diagram illustrating an exemplary GUI used in the registration process according to the present disclosure. [Figure 39] FIG. 39 is a diagram illustrating an exemplary GUI used in the registration process according to the present disclosure. [Figure 40] FIG. 40 is a schematic diagram of an initial setup of a 3D model of a head wearing an HMD device according to the present disclosure. [Figure 41]FIG. 41 is a schematic diagram of a geometric readjustment process according to the present disclosure. [Figure 42] FIG. 42 is a diagram illustrating one embodiment of the recolor process performed in the present disclosure. [Figure 43] FIG. 43 is a flow diagram detailing an exemplary recolor process performed in the present disclosure. [Figure 44] FIG. 44 illustrates another embodiment of the recolor process performed in the present disclosure. [Figure 45] FIG. 45 illustrates another embodiment of the recolor process performed in the present disclosure. [Figure 46A] , [Figure 46B] 46A and 46B illustrate another embodiment of the recolor process performed in the present disclosure. [Figure 47] FIG. 47 illustrates another embodiment of the recolor process performed in the present disclosure. [Figure 48] FIG. 48 is a flow diagram detailing an exemplary recolor process performed in the present disclosure. [Figure 49] FIG. 49 illustrates another embodiment of the recolor process performed in the present disclosure. [Figure 50] FIG. 50 illustrates another embodiment of the recolor process performed in the present disclosure. [Figure 51] FIG. 51 illustrates another embodiment of the recolor process performed in the present disclosure. [Figure 52] FIG. 52 illustrates an embodiment of the landmark inference process performed in the present disclosure. [Figure 53] FIG. 53 illustrates an embodiment of the landmark inference process performed in the present disclosure. [Figure 54] FIG. 54 illustrates an embodiment of the landmark inference process performed in the present disclosure. [Figure 55] FIG. 55 is a diagram illustrating the types of transformations used in landmark inference processing performed in accordance with the present disclosure. [Figure 56A] , [Figure 56B]56A and 56B are flow diagrams detailing an exemplary landmark inference process performed in the present disclosure. [Figure 57] , [Figure 58] 57 and 58 illustrate an exemplary CNN model architecture according to the present disclosure. [Figure 59] FIG. 59 is a flow diagram illustrating image pre-processing for CNN training in accordance with the principles of the invention. [Figure 60] FIG. 60 is a flow diagram illustrating image post-processing in accordance with the principles of the invention. [Figure 61A] , [Figure 61B] 61A and 61B are diagrams illustrating an exemplary method for creating the perception of 3D content. [Figure 62] FIG. 62 illustrates the difference in 2D projections from different camera models between a person's shot and the 3D virtual environment. [Figure 63] FIG. 63 is a diagram illustrating a correction process for a camera system according to the present disclosure. [Figure 64] FIG. 64 is a diagram illustrating image size adjustment and shift correction processing according to the present disclosure. [Figure 65A] , [Figure 65B] Figures 65A and 65B show two cases where the height and width of a person changes as the person moves. [Figure 66] FIG. 66 is a diagram illustrating a pinhole camera used to resize and shift the height and width of a user in accordance with the principles of the present invention. [Figure 67] FIG. 67 is a diagram illustrating an example image frame before encoding according to this disclosure. [Figure 68] FIG. 68 illustrates an exemplary image frame after encoding according to this disclosure. [Figure 69] FIG. 69 illustrates an example encoding scheme according to this disclosure. [Figure 70] FIG. 70 illustrates an exemplary encoding scheme according to this disclosure.
[0068] Throughout the figures, the same reference numerals and characters, unless otherwise stated, are used to denote like features, elements, components, or portions of the illustrated embodiments. Moreover, while the present disclosure will be described in detail with reference to the figures, it is done so in connection with the exemplary embodiments. It is intended that changes and modifications can be made to the exemplary embodiments described without departing from the true scope and spirit of the present disclosure, as defined by the appended claims. DETAILED DESCRIPTION OF THE INVENTION
[0069] Description of the embodiment Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Note that the following embodiments are merely examples for implementing the present disclosure and can be appropriately modified or changed depending on the individual configuration of the device to which the present disclosure is applied and various conditions. Therefore, the present disclosure is not limited to the following embodiments, and the following drawings and embodiments allow the described embodiments to be applied / implemented in situations other than those exemplified below. Furthermore, when multiple embodiments are described, the respective embodiments can be combined with each other unless otherwise specified. This includes the possibility of substituting various steps and functions between embodiments as deemed appropriate by those skilled in the art.
[0070] Section 1: Environment Overview The present disclosure below describes systems and methods for providing virtual reality-based immersive calling.
[0071] FIG. 1 illustrates a virtual reality imaging and display system 100. The virtual reality imaging system includes an imaging device 110. The imaging device may be, for example, a camera with a sensor and optics designed to capture 2D RGB images or video. In one embodiment, imaging device 110 is a smartphone equipped with front- and rear-facing cameras and capable of displaying captured images on a display screen. Some embodiments use specialized optics, such as a binocular view or light-field camera, to capture multiple images from different perspectives. Some embodiments include one or more such cameras. In some embodiments, the imaging device may include a range sensor that effectively captures RGBD (red, green, blue, depth) images, either directly or through software / firmware fusion of multiple sensors, such as an RGB sensor and a range sensor (e.g., a lidar system or a point-cloud-based depth sensor). The imaging devices can be connected via a network 160 to local or remote (e.g., cloud-based) systems 150 and 140 (hereinafter referred to as server 140), respectively. Imaging device 110 is configured to communicate with server 140 via network connection 160 such that imaging device transmits a sequence of images (eg, a video stream) to server 140 for further processing.
[0072] Also shown in FIG. 1 is a user 120 of the system. In this embodiment, user 120 is wearing a virtual reality (VR) device 130 configured to transmit stereo video to user 120's left and right eyes. As an example, the VR device may be a headset worn by the user. The terms VR device and head-mounted display (HMD) are used interchangeably herein. Other examples include a stereoscopic display panel or any display device that enables implementation of embodiments described in this disclosure. The VR device is configured to receive data input from server 140 via second network 170. In some embodiments, network 170 may be the same physical network as network 160, although the data transmitted from imaging device 110 to server 140 may differ from the data transmitted between server 140 and VR device 130. Some embodiments of the system do not include VR device 130, as described below. The system may also include a microphone 180 and a speaker / headphone device 190. In some embodiments, the microphone and speaker device are part of VR device 130.
[0073] 2 illustrates an embodiment of a system 200 in which two users 220 and 270 reside in two user environments 205 and 255, respectively. In this embodiment, users 220 and 270 include imaging devices 210 and 260 and VR devices 230 and 280, respectively, connected to a server 250 via networks 240 and 270, respectively. In some cases, only one user may have imaging device 210 or 260, while the other user may have only a VR device. In this case, from a video capture perspective, one user environment may be considered a transmitter and the other user environment may be considered a receiver. However, in embodiments in which the roles of transmitter and receiver are clearly distinguished, audio content may be transmitted and received between only the transmitter and receiver, both, or vice versa.
[0074] Figure 3 shows a virtual reality environment 300 rendered for a user. This environment includes a computer graphic model 320 of a virtual world accompanied by a computer graphic projection of a photographed user 310. For example, user 220 of Figure 2 can view the virtual world 320 and an image 310 of second user 270 of Figure 2 through their respective VR devices 230. In this example, imaging device 260 captures images of user 270, processes them on server 250, and renders them in virtual reality environment 300.
[0075] In the example of Figure 3, user image 310 of user 270 of Figure 2 shows the user without VR device 280. This disclosure describes several algorithms that, when executed, cause the representation of user 270 to appear as if it were captured naturally without VR device 280 and without wearing the VR device. In some embodiments, the user is shown wearing VR device 280. In other embodiments, user 270 does not use wearable VR device 280. Furthermore, in some embodiments, the captured image of user 270 is captured with the wearable VR device, but the user image is processed to remove the wearable VR device and replace it with a likeness of the user's face.
[0076] Additionally, adding the user's image 310 to the virtual reality environment 300 along with the VR content 320 may include a lighting adjustment step to adjust the lighting of the photographed user 310 and the drawn user 310 to better match the VR content 320.
[0077] In this disclosure, a first user 220 of Figure 2 is shown the VR image 300 of Figure 3 via their respective VR device 230. Thus, the first user 220 sees a user 270 and virtual environment content 320. Similarly, in some embodiments, a second user 270 of Figure 2 sees the same VR environment 320 from a different perspective, such as the perspective of a virtual character image 310.
[0078] To achieve the immersive calling described above, it is important to depict each user in the VR environment as if they were not wearing a headset experiencing the VR content. The following describes a real-time process for acquiring an image of each user in the real world wearing a virtual reality device 130 (hereinafter also referred to as a head-mounted display (HMD) device).
[0079] Section 2: Hardware 4 illustrates an exemplary embodiment of a system for a virtual reality immersive communication system. The system includes two user environment systems 400 and 410, which are specially configured computing devices, and two corresponding virtual reality devices 404 and 414, and two imaging devices 405 and 415. In this embodiment, the two user environment systems 400 and 410 communicate over one or more networks 420. The networks 420 may include wired networks, wireless networks, LANs, WANs, MANs, and PANs. In some embodiments, the devices also communicate over other wired or wireless channels.
[0080] The two user environment systems 400 and 410 each include one or more processors 401 and 411, one or more I / O components 402 and 412, and one or more storage units 403 and 413. The hardware components of the two user environment systems 400 and 410 communicate via one or more buses or other electrical connections. Examples of buses include a Universal Serial Bus (USB), an IEEE 1394 bus, a PCI bus, an Accelerated Graphics Port (AGP) bus, a Serial AT Attachment (SATA) bus, and a Small Computer System Interface (SCSI) bus.
[0081] The one or more processors 401 and 411 include one or more central processing units (CPUs), which may include one or more microprocessors (e.g., single-core microprocessors, multi-core microprocessors), one or more graphics processing units (GPUs), one or more tensor processing units (TPUs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more digital signal processors (DSPs), or other electronic circuitry (e.g., other integrated circuits). The I / O components 402 and 412 include communication components (e.g., graphics cards, network interface controllers) that communicate with the respective virtual reality devices 404 and 414, the respective imaging devices 405 and 415, the network 420, and other input or output devices (not shown), which may include a keyboard, mouse, printing device, touchscreen, light pen, optical storage device, scanner, microphone, drives, and game controllers (joysticks, gamepads, etc.).
[0082] The storage units 403 and 413 include one or more computer-readable storage media. In this specification, a computer-readable storage medium includes, for example, a magnetic disk (e.g., a floppy disk, a hard disk), an optical disk (e.g., a CD, a DVD, a Blu-ray), a magneto-optical disk, a magnetic tape, a semiconductor memory (e.g., a non-volatile memory card, a flash memory, a solid-state drive, an SRAM, a DRAM, an EPROM, an EEPROM), and other manufactured products. The storage units 403 and 413 may include both ROM and RAM and can store computer-readable data or computer-executable instructions.
[0083] The two user environment systems 400 and 410 also include communication modules 403A and 413A, imaging modules 403B and 413B, drawing modules 403C and 413C, positioning modules 403D and 413D, and user image drawing modules 403E and 413E, respectively. The modules include logic, computer-readable data, or computer-executable instructions. In the embodiment shown in FIG. 4 , the modules are implemented in software (e.g., assembly language, C, C++, C#, Java, BASIC, Perl, Visual Basic, Python, Swift). However, in some embodiments, the modules are implemented in hardware (e.g., customized circuitry) or a combination of software and hardware. If the modules are implemented at least partially in software, the software can be stored in memory units 403 and 413. Also, in some embodiments, the two user environment systems 400 and 410 may include additional or fewer modules, or the modules may be combined into fewer modules or divided into more modules. One environment system may be similar to or different from another in terms of its modular containment or composition.
[0084] Each of the imaging modules 403B and 413B includes operations programmed to perform imaging, as shown at 110 in FIG. 1 and 210 and 260 in FIG. 2 . Each of the rendering modules 403C and 413C includes operations programmed to perform functions associated with rendering captured images of one or more users participating in the VR environment. Each of the positioning modules 403D and 413D includes operations programmed to perform processes including identifying and determining the position of each user in the VR environment. Each of the user image rendering modules 403E and 413E includes operations programmed to perform user image rendering, as shown in the following figures. The pre-learning module 403F includes operations programmed to estimate the nature and type of images captured prior to participating in the VR environment for use in head-mounted display removal processing. In some embodiments, some modules are stored and executed on an intermediary system, such as a cloud server. In other embodiments, each of the imaging devices 405 and 415 includes one or more modules stored in its respective memory that, when executed, perform specific operations described below.
[0085] Section 3: HMD Removal Process Overview As mentioned above, given the advances in augmented virtual reality, it is becoming more common for users to participate in immersive communication sessions in VR environments, where each user wears a headset or head-mounted display (HMD) to participate in the virtual reality while remaining in their own location. However, HMD devices hinder the achievement of a better user experience because, unless the HMD is removed, you cannot see the entire face of others in VR, and others cannot see your entire face.
[0086] Thus, the present disclosure advantageously provides a system and method for removing an HMD device from a 2D facial image of a user wearing an HMD and participating in a VR environment. Removing the HMD from the 2D image of the user's face, but not from the 3D object, is advantageous because a human can perceive a 3D effect from the 2D image of the human by inserting the 2D image of the human into a 3D environment.
[0087] Specifically, in a 3D virtual environment, if a person's image is created in 3D or with depth information, the 3D effect of the person can be perceived. However, the 3D effect of the person's image can also be perceived without depth information. An example is shown in Figure 5. Here, a photographed 2D image of a person is placed in a 3D virtual environment. Despite the absence of 3D depth information, the resulting 2D image is perceived as a 3D image because human perception automatically complements the depth information. This is similar to the phenomenon of blind spot "complementation" in human vision.
[0088] In augmented reality and / or virtual reality, a user wears an HMD device. Once in a virtual reality environment or application, the user may be depicted as an avatar or replica of themselves in animated form, but this does not represent an actual image of themselves captured in real time. The present disclosure remedies this deficiency by providing a real-time live view of the user in physical space while the user is experiencing the virtual environment. To capture an image of the user for others to see within the VR environment, an imaging device is a camera placed in front of the user that captures the user's image. However, due to the HMD device worn by the user, others can only see the lower part of the user's face, as the upper part is hidden by the HMD device.
[0089] To ensure full visibility of the user's face captured by the imaging device, an HMD removal process is performed, replacing the HMD region with the upper portion of the face in the image. An example illustrating the effect of HMD removal is shown in FIG. 6. The HMD region 602 in the HMD image shown in FIG. 6A is replaced with one or more pre-captured or artificially generated images of the user to form the full-face image 604 shown in FIG. 6B. In generating the full-face image 604, facial features that are normally hidden when wearing the HMD 602 are captured and used to generate the full-face image 604 in a virtual reality environment, with the eye region visible. HMD removal is an important element in any augmented or virtual reality environment because it improves visual perception. The HMD region 602 is replaced with a streamlined image of the user captured previously.
[0090] The pre-captured images used as replacement images during the HMD removal process are images acquired using an imaging device, such as a cell phone camera or other camera, that allow a user to position themselves within the imaging area, move their face in specific ways, and create different facial expressions via instructions displayed on a display device. These pre-captured images may be still or video data and stored in a storage device. The pre-captured images may be cataloged and labeled by a pre-capture application and stored in a database in association with user-specific authentication information (e.g., user ID). One or more of these pre-captured images may then be retrieved and used as replacement images to replace the top portion of the facial image, including the HMD. This process is described further below.
[0091] The flow diagram of FIG. 7 illustrates an HMD removal process performed in accordance with the present disclosure. The HMD removal process employs a first pre-capture stage 700 and a second stage 710 for real-time removal of the HMD device from images captured by an image capture device. According to the present disclosure, the first stage 700 is a pre-capture stage in which an image of the user is captured prior to entering a VR environment, without the HMD positioned over the user's face and partially obscuring it. The second stage 710 is a live, real-time HMD removal process performed when the user wishes to participate in a virtual reality environment, such as an immersive call. The second stage 710 identifies the upper portion of the face in a real-time captured image that includes the HMD, retrieves a corresponding image from one or more images acquired during the pre-capture stage 700, and uses the user's pre-capture image to generate a replacement image that includes the lower portion of the face captured in real time. This process is performed frame-by-frame on video image data in real time. However, this process can also be performed on still images.
[0092] In a first stage 700 of image pre-capture, a series of images are captured with the user's face unobstructed. A first pre-capture process is performed, which includes a multi-stage pre-capture process 701. In this process, images of the user are captured (1) in multiple orientations, (2) making multiple facial expressions, and (3) performing multiple eye movements and blinks. Figure 8 shows an example, showing three sample images pre-captured from the left, middle, and right viewpoints.
[0093] This is merely an example; in reality, the facial orientations, positions, and expressions captured and used for the HMD removal process in the second processing stage are diverse. For example, if a user is instructed by a pre-capture application to move their head left, right, up, down, etc., the pre-captured user image may cover 20 different orientations. In another embodiment, the pre-captured image may represent orientations ranging from tilted left to tilted right. It should be understood that the multiple user orientations captured during the pre-capture process cover all three directions (X, Y, and Z) in world space with reasonable resolution. Ideally, higher orientation resolution improves human perception in terms of performance, but also increases the difficulty of the pre-capture process for the user. A specific example would be to divide the yaw angle range from -60 to 60 degrees into nine intervals, the pitch angle range from -40 to 40 degrees into nine intervals, and leave the roll angle unchanged, since roll images can be created programmatically using rotation. The pre-captured images are labeled based on orientation information obtained from one or more sensors in the HMD device. In other embodiments, a pre-capture application running on the image capture device may generate a label for each captured image that provides orientation information identifying the particular orientation from which the image was captured.
[0094] The multi-stage pre-capture process 701 further includes capturing facial images of the user, typically expressing various expressions, such as happiness, sadness, surprise, anger, and fear, via a pre-capture application running on the image capture device. These are provided for illustrative purposes only; any facial expression corresponding to the type of expression a user may make while communicating with other users may be captured during the pre-capture process. This allows the facial expression of the upper portion of the face to be aligned with the facial expression of the lower portion of the face during the HMD removal process. This aspect of the pre-capture stage 701 is described in more detail in Section 4 of this disclosure below.
[0095] The pre-capture stage 701 further includes capturing images of the user speaking spontaneously or reading a predetermined set of text and extracting information associated with the user's eye and gaze movements while the user is speaking, a process described below in Section 5 of this disclosure. This eye movement information, used in combination with orientation and facial expression information, advantageously provides the real-time HMD removal stage 710 with the ability to generate a natural eye and gaze perception when such eye and gaze information needs to be artificially simulated during the replacement process, as described below.
[0096] Once the capture portion of the pre-capture phase is complete, the pre-capture image data is automatically processed, and one or more trained machine learning models are generated by processing the pre-capture data to label the data with information necessary for real-time HMD removal and replacement in the HMD removal phase 710. The information necessary to label each image includes face orientation, facial motion units detected from the face, and blink metrics from the face, so that the HMD removal phase can provide the live-captured images to the trained machine learning models to retrieve from storage the pre-captured images that are best suited for replacement for a particular user.
[0097] The second stage 710 of the HMD removal algorithm includes three sub-stages: an IMU (Inertial Measurement Unit) data alignment stage 711, a geometry configuration stage 712 (e.g., 3D geometry of the camera, HMD, and human head), and a real-time HMD removal stage 713. The first two sub-stages are calibration processes that configure and / or align various components of the system so that the third sub-stage, real-time HMD removal 713, can occur. The components that are configured and aligned in the first two sub-stages include the HMD worn by the user, an image capture device that captures images of the user wearing the HMD, and a processing device such as a cloud server that receives captured images of the user wearing the HMD, including pre-captured image data, which all need to be synchronized to enable real-time identification, extraction, and replacement of portions of the captured images that include the HMD.
[0098] In an IMU data alignment stage 711, a first calibration process is performed to align the IMU data generated from the HMD device with the real-time images generated by the image capture device during the capture process. This includes timestamp alignment and orientation alignment. Timestamp alignment ensures that the timestamps of the IMU data match the timestamps of the captured images. Without this time alignment, specific IMU sensor data and one particular image cannot be correlated because they are acquired from two different devices with two different clock systems. This process is described below in Section 6 of this disclosure.
[0099] Additionally, in the IMU data alignment stage 711, a second calibration process is performed to align the orientation readings from the IMU sensor to match the user orientation estimated from live images of the user wearing the HMD. Because each component defines its 3D coordinate system differently, without this calibration, the orientation of one data set cannot be used with the other data set. An example is shown in Figure 9 and described in Section 7 below. Figure 9 shows three images with different yaw orientations. If the center image, with its optical axis passing through the center of the face, is defined as the origin, the orientation estimates from the three images may be +10, 0, and -10. However, the IMU sensor may output yaw orientations of 30, 20, and 10 for these three images. Therefore, the second calibration process aligns the orientations to ensure that the readings can be transferred to each other.
[0100] According to the second sub-stage 712 of the HMD removal process shown in FIG. 10, the alignment of the 3D geometry of the human head, HMD device and camera is performed together with the orientation and time alignment in the first stage.
[0101] Despite the assumption that all users wear the same HMD device, each user's head shape and size are different. Therefore, a single numerical value cannot represent the various values. Furthermore, the degree to which users wear the HMD device may also vary, and this difference must be considered and compensated for during the geometric matching process. This geometric configuration is crucial for the HMD removal process because it is desirable to logically infer a 2D human face based on the position and orientation of the HMD device relative to the human head. Without accurate 3D configuration data, it is difficult to achieve high accuracy in inferring the position of the 2D human face. This is because this inference relies solely on inferred information from the HMD device.
[0102] A real-time HMD removal process step 713 is performed after the calibration steps are completed. In one embodiment, the calibration processes of steps 711 and 712 are performed at the beginning of the session before entering the virtual reality environment. In another embodiment, the calibration process is performed on each frame as it is acquired before removing the HMD in real time for captured images and providing the images to the remote user in the virtual reality environment. The third step 713 of Figure 7 is described in more detail with reference to the flow diagram of Figure 11.
[0103] The HMD removal module receives an input image of a user wearing an HMD in a captured image at 1100. The image is passed to two parallel channels: a first pre-captured image channel and a second live image channel. The pre-captured image channel represents processing related to one or more pre-captured images. First, at 1101, the xyz position and RPY orientation are automatically generated from the HMD device's IMU sensor and calibrated based on the IMU alignment described below in Section 6. The calibrated orientation is then used at 1103 to find a pre-captured image from all pre-captured image data acquired from the user that matches the orientation in the live-captured image. At 1105, this image is cropped, resized, and recolored at 1107 in preparation for the image replacement process at 1110. The recoloring process is described in more detail in Section 8.
[0104] The live image capture channel acquires a series of live images of a user wearing an HMD through imaging processing performed by an imaging device. These live captured images are processed to identify an area in the live captured images that includes the HMD, extract that area from the live captured images, and replace that area with a portion of the pre-captured image based on the output of the pre-capture channel. At 1102, the HMD area is segmented and its bounding box is extracted based on a trained machine learning model trained using images of different users wearing the HMD. This process is described in more detail in Section 9. At 1104, candidate eye / nose areas are determined from the segmented area based on the IMU orientation and the user-based 3D geometric configuration of the head, HMD, and camera described above in FIG. 10, and are projected and inpainted onto the HMD area. This process is described in more detail in Section 10. At 1106, initial landmarks from the face region are estimated, secondary eye and nose regions are regenerated, and updated landmarks for the face region are estimated at 1108, as described below. A triangulation of these landmarks is also generated, enabling 3D geometric processing. Based on these updated landmarks and triangulation from the live channel and the images returned from the pre-captured channel, a replacement process 1110 is performed, which includes one or more of three image region exchanges: head image exchange, facial expression exchange, and eye region exchange. After image processing to exchange at least one or all of these three regions, an updated image is generated at 1112. In this image, the worn and live-captured HMD is removed by replacing it with the aligned pre-captured image. This live output is described below in Section 12, and a live output image with the HMD removed can be obtained.
[0105] In another embodiment, as shown in FIG. 12, instead of the three separate replacement steps shown in FIG. 11, a single replacement step is performed. Therefore, processing steps 1100-1108 remain the same and need not be further described herein. In this embodiment of FIG. 12, 1210 represents one step for finding a single pre-captured image with all the correct information regarding the head, facial expression, blinks, and gaze. In another embodiment, improved performance with respect to human perception can be achieved by searching for an image of the head in the correct orientation, then replacing its face region with another image containing the correct facial motion units, and finally replacing the eye region of the face with a third image containing the correct blinks. The output of the above approach is a combination of three pre-captured images. To further improve image quality with respect to the appearance of unnatural boundaries between the head and face and between the face and eyes, this embodiment uses the determined parameters to search for and find a single image with all the correct information regarding orientation, facial motion units, and blinks, and uses it as a replacement image to reduce unnatural transitions between regions. Thereafter, the same process as in 1112 is executed.
[0106] The real-time HMD removal algorithm described above advantageously replaces the HMD region in the facial image with a reasonable orientation, facial expression, blink, and gaze. Performing this process in real time advantageously allows other users in the VR environment to see that the user is not wearing the HMD device required to enter and enjoy the VR environment. This creates a more realistic connection between users in VR, resulting in an immersive and connected communication experience. This advantageously utilizes an image pre-capture process to collect one or more facial images of the user in which the HMD does not obscure any part of the face. Orientation is adjusted, and such orientation information associated with the HMD is passed from the HMD to the image. Configuration is performed to construct and align a 3D model of the user's head, HMD, and a camera capturing an image of the user wearing the HMD, resulting in the head, face, and eye regions being swapped to replace the HMD region in the live-captured facial image.
[0107] Some of the benefits of the HMD removal process described in the above and subsequent subsections of this specification depend on the successful transmission of both the data used in the HMD removal process as well as information that assists the user wearing the HMD in being in the desired position so that capture by the imaging device can be performed properly and the resulting data can be transmitted at a size that allows for the real-time processing described in this disclosure.
[0108] With regard to the ability to ensure a user is in a desired position to capture an image of themselves wearing an HMD, the present disclosure advantageously uses an alpha channel to determine whether a person has moved and the amount of movement, or whether they are entering or leaving the camera frame. The alpha channel is an information parameter transmitted with each image frame that characterizes one or more aspects of the associated image frame. According to the present disclosure, the alpha channel defines regions within the associated image frame as either background or foreground image regions. For example, the information in the alpha channel includes a binary representation, where a "1" indicates that the pixel should be understood as foreground and a "0" indicates that the pixel should be understood as background. Because the HMD removal process is performed on a user wearing an HMD device, it is advantageous for the HMD removal algorithm to identify only the desired region in which the HMD removal process is to be performed—in this case, the foreground region containing the user within the 2D image captured by the imaging device.
[0109] According to the present disclosure, in some embodiments, this may be derived from a background detection / segmentation network. In some embodiments, this alpha channel may be further post-processed, for example using connected component analysis (CCA), to filter out small false positive regions and select only the largest connected foreground region as a mask for the detected person. In other embodiments, post-processing, for example in a video context, may include extracting a bounding box from the alpha channel that outlines the foreground object detected in one frame. In the next frame, this bounding box (which may be expanded by a fixed or adaptive factor in some embodiments) is used as a starting point, and CCA is performed only within the bounding box, automatically marking the rest of the frame as background, improving the speed and robustness of the algorithm. When multiple foreground objects of similar size are present in a frame, analyzing the entire frame can cause the CCA to jump between objects. However, by restricting the CCA to only track within the (expanded) bounding box from the previous frame, the jumping of the CCA is reduced and the first detected object is instead tracked. If the bounding box becomes smaller than a certain threshold, or in some embodiments after a certain number of frames, the bounding box may be reinitialized to cover the entire frame, so that CCA does not get "stuck" on tracking one foreground component when there is a possibility that another larger component exists outside the bounding box.
[0110] FIG. 13 illustrates an exemplary algorithm for the above algorithm. At 1302, a captured image frame is input and a background detection process is performed to identify an area of interest (e.g., a foreground object) and a background area, which is the area of the frame that contains everything but the foreground object. At 1304, a query is performed on the received image frame to determine whether a bounding box exists within the frame that encloses the foreground object. If the result of 1304 indicates that a bounding box does not exist, the algorithm automatically sets a bounding box with the dimensions of the input image frame at 1306. If the query at 1304 is positive, indicating that a bounding box exists within the input frame, a query is performed at 1305 to determine whether a threshold has been met. In one embodiment, the threshold indicates the size of the bounding box. In another embodiment, the threshold is whether the algorithm has received a predetermined number of input frames. In yet another embodiment, a threshold can be set to determine the size of the bounding box and whether a certain number of frames have been received and processed by the algorithm. If the query at 1305 returns a negative result, processing continues at 1306, where a bounding box approximately the same size as the input frame is generated. If the query at 1305 returns a positive result, or if the bounding box is set to the size of the frame, processing continues at 1308. At 1308, processing such as CCA is performed only within a predefined extended region associated with the set bounding box. To speed up processing, the area outside the bounding box is ignored. Once the foreground has been reliably extracted, processing continues at 1310 and 1312, which are described in the following sections.
[0111] At this stage, the alpha channel either has a continuously varying opacity (e.g., 0 to 255 if the image is stored as an unsigned 8-bit integer array, or 0 to 1 if the image is stored as a floating-point array), or it is a binary mask, with 1 being foreground and 0 being background, to identify foreground and background regions. If the alpha channel is not a binary mask, it may be converted to one by choosing a threshold that separates foreground from background (e.g., all pixels with an alpha value greater than 128 are foreground).
[0112] At 1312, the algorithm calculates a state vector for the given frame as a set of four state values, each ranging from 0 to 3. These state values represent the degree to which the user is on the left side of the frame, the right side of the frame, too close to the camera, and too far from the camera. In some embodiments, these state values are restricted to discrete values, such as integers, and can be interpreted as discrete states. For example, if the state values indicating the degree to which the user is on the left side of the frame are categorized into the integers 0, 1, 2, and 3, these may be identified as 0 = "not at all left," 1 = "slightly left," 2 = "moderately left," 3 = "extremely left," or similar categories. Based on these state values, various actions may be performed, such as displaying a message on the graphical user interface (GUI) of the HMD device (or other device visible to the user). In another embodiment, the state vector is addressed, such as by changing the functionality of an application.
[0113] We will now describe the processing involved in calculating the state vector. For each state value, we define a region of interest (ROI) within the frame and calculate the average percentage of foreground pixels within the ROI. This ROI is fixed for each value of the state vector. In some embodiments, this value is defined by a fixed number of rows / columns / pixels, while in other embodiments, it is defined by a fixed percentage of rows / columns / pixels in a particular portion of the frame.
[0114] For example, a region of interest (ROI) for determining how far to the left of the frame a user is can be defined as the leftmost 10 columns of the frame. Similarly, a ROI for determining how far to the right of the frame a user is can be defined as the rightmost 10 columns of the frame. To determine how far forward a user is, an ROI can be defined as the top 10 rows and the bottom 10 rows, or a combination thereof, e.g., a logical OR of two separate ROIs containing a binary 1 if either the corresponding pixel in the top 10 rows or the corresponding pixel in the bottom 10 rows is a foreground pixel. A ROI for determining how far back a user is can be the entire frame. This is shown in FIG. 14, where each labeled ROI is assigned a state value, which is used as feedback to determine the user's position in the frame and guide the user to the appropriate position within the frame for successful imaging and subsequent HMD removal processing of the resulting 2D image.
[0115] In an exemplary embodiment, as described in 1310, the state value s may be calculated from the average proportion M of foreground pixels within the ROI according to the following formula: TIFF2025538281000002.tif955Here, P is the total number of pixels in the ROI, and p is the binary value of each pixel (1 for foreground, 0 for background). This average value is mapped to the interval [0,3] based on the start and end thresholds a and b for each state value. The function f(x), which is 0 when x is less than a and 1 when x is greater than b, and changes linearly from 0 to 1 between x = a and x = b, is defined as follows: TIFF2025538281000003.tif1886
[0116] Next, for each of the left, right, and nearby state values, calculate:
[0117] s(M; a,b) = 3 · f(M; a, b) …(3) For the far value, it is desirable that s(0) = 3 and s(1) = 0, which is the opposite of the other state values (i.e., if the number of foreground pixels is very small, the user is too far away), so in this case s(M) = 3 f(1-M) is defined and plotted in Figure 15.
[0118] In another embodiment, one or more sensors of the HMD device can be used to obtain specific state vector information. In this embodiment, when the tracked person is wearing a head-mounted device (HMD) or other device that includes an inertial measurement unit (IMU) with position data, the algorithm performs an out-of-frame detection using the IMU after an initial alignment step (described below) that converts IMU coordinates to camera axis coordinates and then detects the initial distance between the user and the camera, as described below. This information can be used to accurately track the user's position in three dimensions and check when the user leaves the frustum of the camera viewport.
[0119] In addition to using information contained in the alpha channel associated with the captured image to ensure proper positioning within the image frame, additional controls are performed to ensure that the user wearing the HMD is properly facing the imaging device that captures the image on which the HMD removal process is performed.
[0120] To perform a capture of a user wearing a headset so that the user can enter and interact with the virtual reality environment, the user, with the headset covering their eyes, must stand substantially within the field of view and facing the camera to initialize subsequent procedures such as IMU alignment (described below). The user's pose orientation is estimated from this initial capture and used to guide the user's pose adjustments during interaction with the virtual reality environment. The most important angle for this purpose is the yaw angle of the body pose relative to the camera, as shown in FIG. 16. In FIG. 16, a user 1602 wearing an HMD device 1604 is shown standing on a surface, such as a floor 1606, within the field of view of an imaging device 1608. The angle denoted λ is the user's angle relative to the imaging device. The value of λ is used by the algorithm to correctly guide the user in the movements required to achieve the desired position.
[0121] A human pose skeleton estimator is implemented to estimate the user's joints, including, for example, shoulders, hips, elbows, etc. The estimated joint positions are expressed in x, y, z coordinates in a predefined coordinate system. Here, since the region of interest is in the pose orientation relative to the camera, the joint coordinates need to be transformed into world-centric coordinates.
[0122] Next, the orientation of the upper body is estimated. To do this, the human upper body is approximated as a planar surface, which is used to estimate the orientation of the surface normal relative to the camera. All joint coordinates are projected onto a normal vector, and the sum of the residuals is minimized. The normal vector that minimizes the sum of the residuals is the normal vector used to calculate the angle. In an exemplary operation, a sufficiently accurate angle estimate is obtained by using a predetermined number of joints, such as the shoulders and hips. In this case, the normal vector is calculated by taking the cross product of one of the two vectors connecting the three joints (the two shoulders and the hips). This angle is used to guide the user to approximately face the camera. An example of this is shown in Figure 17. Figure 17 shows the user facing away from the camera, with a yaw angle of 204.8 degrees (180 degrees means facing directly away from the camera). In this case, the algorithm controls the HMD's display to guide the user to turn accordingly.
[0123] Returning to the HMD processing depicted in Figure 7, the following sections referenced therein will now be described in more detail. Each of these sections describes the specific algorithmic processing steps performed to obtain the desired output used in the real-time HMD removal process. We will begin by describing aspects of the pre-capture process 700.
[0124] Section 4: Video-based Facial Action Unit (FAU) Detection Detecting facial action units (FAUs) is a challenging task in computer vision. Most existing detection models attempt to determine the FAUs represented in a single image. In some situations, such as video recording, multiple images of the same subject are often present but are underutilized. One dataset used to train FAU detection models is BP4D, which contains many FAU labels. However, within the dataset, some FAUs are more prevalent than others, while many FAUs are present in only a small fraction of images. This can lead to data imbalance issues, distorting training results and reducing overall performance. This is particularly problematic in the current state of HMD removal algorithms, because to replace images of a user wearing an HMD, it is important to verify that the images selected for replacement contain the correct FAUs represented in the pre-captured images by comparing them with the live-captured images.
[0125] Some embodiments of video-based FAU detection include at least four components. The first component is a tool that infers a standard set of facial landmarks from a particular frame (image) of video. Examples include Mediapipe's Face Mesh, Google's MLKit, or dlib's facial landmark detector. Because a larger number of facial landmarks, i.e., higher resolution, typically improves performance, Mediapipe's 468-landmark set may provide satisfactory results. The second component is a framework that records the facial landmarks detected across multiple frames and enables frame-to-frame comparison. The third component is a "standard" face, which includes facial landmarks for a standardized, frontal, neutral face to which all detected landmarks can be compared and analyzed. This face can be individual-specific (e.g., the average of normalized landmarks detected across multiple images of the same person) or generic (e.g., a "typical" 2D or 3D model of a human face). The fourth component is a classifier or a collection of classifiers that takes as input features derived from landmark coordinates and outputs a binary classification for a single FAU or multiple FAUs with a shared architecture. Some embodiments of the fourth component include a support vector machine (SVM), some include an artificial neural network (ANN), and some use other applicable types of classifiers or classification algorithms.
[0126] As mentioned above, one feature of these embodiments is the use of a "standard" face to augment the detected facial landmarks. An exemplary alignment process is shown in Figure 18. In this process, an input image 1802 of a human face is provided, and human facial landmarks are determined 1804 in a frontal orientation and / or one or more side orientations. These determined landmarks are aligned 1806 with the standard face and provided to train a classifier 1808, which outputs a binary classification vector 1810.
[0127] Because every person's face has a unique shape and features, using only raw landmark coordinates may not fully account for natural population variations. Because using the displacement of landmarks from a reference point (e.g., the location of the landmark when the face is in a "rest" position) may yield better results than using the absolute detected location of the landmark, some embodiments of the classifier are trained using the coordinate difference between (i) the landmark on a standard face and (ii) a version of the detected location of the landmark that has been normalized and aligned to the standard face.
[0128] For example, some embodiments of standard facial landmarks are based on a set of two-dimensional or three-dimensional coordinates {c i} N i=1 where N is the number of landmarks that can be detected for each face. The landmarks detected from the input image are stored in a set {l i} N i=1 To align the detected landmarks to standard landmarks, both can first be centered around zero and normalized to unit standard deviation. For example, μ l and μ c are the averages of all detected standard landmarks, respectively, and σ l and σ c If x is the corresponding standard deviation, then the landmarks normalized around zero can be written (and calculated) as: TIFF2025538281000004.tif1232where c^ i is a standard landmark normalized around zero, and l^ iare the detected landmarks normalized around zero, and the division is performed component-wise. Then, in some embodiments, a linear transformation matrix M is generated that maps the detected landmarks to standard landmarks. For example, some embodiments of the linear transformation matrix M can be represented by (and calculated according to) the following equation: TIFF2025538281000005.tif1086
[0129] Equation 4 represents the least squares regression, which selects the transformation that minimizes the mean square of the norm of the difference between the normalized coordinates of the detected landmarks after the transformation and the normalized coordinates of the standard landmarks. Then, the transformed features of each landmark, i.e., the features {Ml i -c i} N i=1 may be input to the classifier (e.g., one by one). In other words, the features of all landmarks are combined as input to the classifier, thereby representing the entire face at once, rather than landmark by landmark (1808).
[0130] Similarly, some embodiments may use other features besides raw landmark coordinates, such as the cosine similarity of all edges in the face mesh, the interior angles of each triangle in the mesh, or a combination of the above, and embodiments using these features may use feature differences between the input face and a standard face.
[0131] In embodiments for identifying FAUs in a single image of a person's face (face image) where there is no and will be no access to other face images or pre-defined 3D models of the same person, the standard face can be a generic 3D model of a human face, such as Mediapipe's Face Mesh standard face model. In other embodiments, a collection of pre-recorded face images can be accessed. In these embodiments, a registration procedure can be used to align all face images in a collection to a generic 3D model of a face, and then the average of all face images can be used as the standard face. This has the added advantage of being person-specific and generally leads to better results than using a generic model as the standard face.
[0132] FIG. 19 illustrates an exemplary algorithm used to determine a standard face according to the present disclosure. At 1902, a first frame of video data is received, but no facial landmarks are detected on the received image frame, and a decision is made to use a first standard face, which is a generic standard face, as described above. At 1904, FAU landmark detection is performed on the received input image, and the detected landmarks are aligned with the generic standard face. From there, at 1906, the standard face is subtracted, and the result is provided to a trained machine learning classifier (e.g., a neural network) to classify the FAU based on the aligned landmarks. After classification, at 1908, it is determined whether the FAU has been classified for a threshold number of frames. If the determination at 1908 is positive, at 1910, the algorithm proceeds to analyze the next image frame and then returns to 1904. If the determination at 1908 is negative, the standard face used at 1902 is updated at 1912 using the average of the alignments detected so far, until the threshold number of alignments occur. After the standard face update, processing continues to 1910 for the next frame.
[0133] Some embodiments detect FAU in video content. While FAU does not initially have access to multiple facial images of a subject, it can derive and add multiple individual-specific facial images to a collection for each new frame captured (see 1906-1910). These embodiments can start with a generic model of a standard face (1902) and calculate an individual-specific average face after a predetermined threshold of facial image frames is reached. Alternatively, a continuous adaptive process can be employed, in which the initial image is a generic model standard face, and each new facial image frame processed is averaged with the existing standard face to generate a new standard face for processing subsequent facial image frames. In some of these embodiments, a threshold for the maximum number of facial image frames averaged to generate the standard face can be set, as shown in FIG. 19. The original generic model may or may not be retained in the cumulative average standard face. The threshold here is the maximum number of faces to be averaged. If the threshold is not reached, the current face is added to the average before proceeding to the next frame (1912). After the threshold is reached, the average update is stopped and the current face is used for the remaining frames.
[0134] Balancing the training dataset Many existing data sampling and augmentation methods exist that can help balance positive and negative labels in single-label datasets. As described herein, a label indicates whether each FAU appears in an image. For each image in the dataset, the label is a list of binary yes / no responses for all FAUs. More generally, the labels of a dataset are the target variable, i.e., what we want the model to predict. However, these methods may not be effective when the dataset contains multiple labels for each sample, as is the case with the BP4D dataset, which is often used to train FAU classifiers. Therefore, embodiments of at least one of the following two methods may be used to balance multi-label datasets and improve training performance. Both methods use undersampling to construct a training set, but differ in how each sample is determined based on past selections.
[0135] Most unbalanced class implementation FIG. 20 illustrates an exemplary balancing algorithm. This first method prioritizes the most imbalanced class (the class whose proportion of positive samples is furthest from 50%). Note that classes with a prevalence rate greater than 50% are over-represented, and classes with a prevalence rate less than 50% are under-represented. Here, "class" refers to each FAU individually. For example, if 90% of images are labeled as representing FAU12 and only 10% are labeled as not representing FAU12, the FAU12 class is said to be imbalanced. A perfectly balanced class is one in which 50% of the samples (i.e., images) are labeled as positive and 50% are labeled as negative. To begin data selection for the training set, an embodiment of the first method first selects, at 2002, a single random sample with a positive label for the least prevalent class. The first method then selects the next training sample until a threshold (e.g., a user-defined threshold) is reached regarding the number or proportion of samples, or until no samples meeting the criteria can be selected.
[0136] In 2004 and 2006, the most imbalanced class in the dataset is determined by calculating the occurrence rate of positive samples for each class in the already selected teacher set and selecting the most over- or under-represented class. If the class determined in part (1) is over-represented, in 2008, a negatively labeled sample of this class is selected from the dataset. If the class is under-represented, in 2007, a positively labeled sample of this class is selected from the dataset. If no such sample exists in 2010, steps 2004–2008 are repeated for the second most imbalanced class. If no such sample exists again, the process is repeated for the third, fourth, and so on, imbalanced classes, until a matching sample is found. If no such match is found for any class, the balancing process is terminated (2012).
[0137] Hamming Distance Implementation Example The first method, shown in Figure 20, considers only the most imbalanced class in each selection. However, basing selection solely on this may further exacerbate the imbalance rate of selected samples, since the selected samples may contain labels from other classes. For example, if the most imbalanced class is under-represented, samples with a positive label for that class are selected. If the positive label for that class is correlated with the labels of other classes (e.g., in the case of FAUs, this may mean that certain motion units are often co-represented), selecting samples with a positive label for the most imbalanced class may unintentionally increase the imbalance rate of the correlated class.
[0138] To solve this problem, the second method embodiment shown in FIG. 21 incorporates a selection criterion based on Hamming distance from information theory. Assuming that the current positive occurrence rate of each class in the selected teacher set is a binary bit string, with 0 corresponding to an under-represented class (occurrence rate less than 50%) and 1 corresponding to an over-represented class (occurrence rate greater than 50%), these embodiments calculate the Hamming distance of each sample in the remaining unselected dataset (0 = negative label, 1 = positive label) to a similar binary representation. (Although face images are similarly used in this disclosure, the balancing method is valid for any type of dataset with multi-class labels.) Then, the sample with the largest Hamming distance to the current teacher dataset is selected. This maximizes the guarantee that the imbalance rate of any class does not erroneously increase with each new selection. This process includes selecting an initial sample with a positive label for the least frequently occurring class from the dataset at 2102 and determining a binary representation of the imbalance rate in the current teacher set at 2104. At 2106, the Hamming distance to all other unselected samples is calculated and the sample with the largest Hamming distance is selected at 2108. At 2110, it is determined whether a threshold number of samples have been calculated and selected, and if so, the algorithm ends at 2112. If the determination at 2110 is negative, the process returns to 2104 and repeats until the threshold is reached.
[0139] FAU detection using a standard face can use differential information to determine facial expressions rather than using raw coordinates. This has the advantage of better accounting for differences in facial structure among people. Furthermore, the balancing method provides a more appropriately acquired dataset, further improving training performance. This allows for the use of relative / differential features for detecting FAUs, video-based FAU detection, and balancing of a multi-label dataset with undersampling to advantageously provide the model with features used to determine which pre-captured images can be selected for a given HMD removal in a given frame, thereby further improving the training set of the model used in the HMD removal process.
[0140] Continuing with the description of the processing performed during the pre-capture stage 700, the functions of determining and simulating blinks and gaze will now be described.
[0141] Section 5: User-based blink simulation and gaze generation As described above, in augmented virtual reality, a user wears an HMD device to experience virtual reality. When a user wishes to be seen by others in virtual reality, a camera in front of the user captures an image of the user. However, because the user is wearing an HMD device, the upper part of the user's face is obscured by the HMD device, preventing others from seeing the user's entire face. To enable the user's entire face to be seen, HMD removal, as described herein, is often performed on the HMD region in the user's HMD face image. Typically, the HMD region can be replaced with a pre-photographed image or an artificially generated image. This replacement completely obscures the eye region, making it difficult to accurately determine the eye region. The algorithm described herein improves this drawback by generating a series of images with natural blinking and gaze.
[0142] Blink and gaze data are collected from the user during the pre-capture stage of the HMD removal algorithm in Section 3. This has the advantage that it allows the system to artificially generate natural blink and gaze images. The collected eye information includes, but is not limited to, blink frequency, which indicates how frequently the user blinks; blink interval, which indicates the time between two adjacent (consecutive) blinks; and eye size, which indicates how widely the eyes open and close. This eye information is collected using a blink indicator module. The design and operation of the blink indicator is shown in Figure 22.
[0143] For each photographed face, landmarks representing the upper and lower eyelids of each eye are determined and located, and the distance between the upper and lower eyelids relative to the entire face length is measured. The first column in both (A) and (B) shows two images, one with the eyes open and one with the eyes closed. Facial landmarks are estimated using a facial feature library such as Mediapipe. The landmarks obtained for the entire face are shown in the second column in each of (A) and (B). Once all facial landmarks are determined and collected, landmarks related to blink characteristics are identified. For example, four landmarks representing the upper and lower eyelids, two from the left and two from the right, and two landmarks representing the face length are identified. Specifically, a predetermined number of landmarks are selected from each eye to represent the eyes, and eye information is determined and acquired. As shown here, four landmarks are selected from both the left and right eyes: two from the left eye, including LU (upper left) and LL (lower left), and two from the right eye, including RU (upper right) and RL (lower right). Furthermore, a predetermined number of facial landmarks are obtained, representing the upper and lower parts of the entire face. These are designated as FU (upper face), and the others are designated as FL (lower face). The predetermined number of eye and facial landmarks are described as examples only, and any number of landmarks may be used in this process depending on the computation time and capabilities of the processing device executing the algorithm. The blink index is defined using the distance between the upper eyelid landmark and the lower eyelid landmark and the length of the face, or the distance from the upper part of the face to the lower part of the face.
[0144] The blink index can be expressed as follows:
[0145] left eyeblink inicator = (LU · LL) / (FU · FL)*100 …(5) right eyeblink inicator = (RU · RL) / (FU · FL)*100 …(6)
[0146] The blink index is adjusted to identify a range of eye positions. The distance is multiplied by 100 to adjust the index to a value between 0 and 100. This adjusted determination is graphically illustrated in FIG. 23, which depicts the blink index performed on a series of 512 video frames. FIG. 23 shows the blink index from the left eye for a video spanning a predetermined number of frames, here over 512 frames. As can be seen, the blink index value for this user varies from 6.0 to 3.0 based on the percentage of the entire face. Note that 6.0 here means that the distance between the upper and lower eyelids is 6 percent of the entire face length. Observations show that when the user's eyes are open, the distance between the upper and lower eyelids is approximately 5%, and when the eyes are closed, the distance between the upper and lower eyelids is approximately 3.5%. As shown, the blink index successfully identifies seven blinks out of the 512 frames, as indicated by the numbers next to the plot line in FIG. 23.
[0147] 24 is a flow diagram illustrating a blink detection algorithm used in generating the plots described above and as part of an HMD removal algorithm according to the present disclosure. At 2402, a facial image is acquired and provided to a landmark detection model at 2404 to detect all facial landmarks present in the acquired image. Once all facial landmarks present in the image have been determined, a blink metric is calculated at 2406. The resulting calculated blink metric is provided as input to the blink detection process described with reference to FIG. 25.
[0148] The blink detection process includes three stages. The first stage, 2408, performs a baseline removal process for the blink metric. This process removes large or sudden changes between eye open and eye closed. The baseline is first estimated based on a moving average of the blink metric over a predetermined window. A specific example is 60 frames or 2 seconds at a frame rate of 30 frames per second. Next, the baseline is subtracted from the raw data, resulting in the baseline-removed blink metric shown in Figure 25A. The baseline-removed metric is then applied to two thresholds at 2410. The first threshold, shown as a solid line, enables reliable and effective blink detection. The second threshold, shown as a dashed line, enables reliable and effective blink duration detection. The blink detection threshold is set high to be less susceptible to residual baseline noise, while the duration determination threshold is set low to provide a slightly wider range of blink coverage for eye region replacement. The result of the thresholding process is shown in Figure 25B. Based on the results shown in (B), a segment detection algorithm is applied at 2412 to identify all detected blinks. The segment detection algorithm works by first identifying negative transitions from zero to negative and positive transitions from negative to zero, respectively, and then combining two corresponding transitions, positive to one and negative to one, into a single segment. As shown in (C), at 2414, the algorithm correctly identified all seven blinks in the input video as output. Once blinks are identified in the image sequence, the algorithm obtains and collects statistics of features that describe the blinks. The first feature is the time interval between two adjacent blinks, and the second feature is the duration of the blink. These two features determine when the next blink should be generated and how long the generated blink will last.
[0149] FIG. 26 illustrates histograms for artificially generating new sequences of a user's blinks according to the present disclosure. Histograms (A) and (B) show first and second features based on 71 blinks identified by the blink index. It is important to note that different users may have different histogram values depending on their natural blink patterns, which are user-dependent. In operation, the algorithm simulates the user's blinks by ensuring that the simulated blinks match the statistical information previously collected from the user, typically during the pre-capture phase described above in Section 3. The time interval and duration of each simulated blink are determined by sampling them from the distribution of actually collected data. Histograms (C) and (D) show the distribution of the time intervals and durations of 200 sampled blinks. While the above description includes the first and second features described above, this is merely exemplary, and other features may be used in place of or in combination with the first and second features. Other features that can be used to simulate blinks may include, but are not limited to, blink strength, correlation between two consecutive blinks, or other higher order statistical information related to blinks.
[0150] After determining the time interval and duration of each blink, a series of images containing expected blinks from the user is artificially generated, as shown in FIG. 27. Given an input video 2702, all blinks are detected, and a histogram of the blink-related features is estimated at 2704. Based on the blink features, an image pre-selection process is performed to pre-select and save certain images. At 2705, some of the pre-selected images contain blinks, while others do not. These pre-selected images are used for gaze generation, as described below. A population of these blink features is generated by sampling from a histogram of actual experimental data at 2706. An existing series of frames is identified. Alternatively, several segments of frames are artificially collected. These sequences are combined at 2708 as a baseline sequence of images. This baseline is then replaced at 2710 with pre-recorded images with or without the desired blinks from 2705. After the replacement, a video of facial images containing the desired blinks, i.e., with or without blinks, is generated and output for display at 2712.
[0151] With reference to FIG. 28, the exchange and generation of output images will be described. To exchange the eye regions of two images, a mesh of the eye regions as shown in FIG. 28 is obtained. Here, two images, image 2801 with the eyes open and image 2802 with the eyes closed, shown in column A of FIG. 28, are used. Given two input images, landmarks representing the eye regions and a mesh (e.g., triangles) associated with these landmarks are obtained by an estimation process using a trained machine learning model. The results of the estimation process are shown in images 2803 and 2804 in column (B) of FIG. 28. Because these triangles estimated from the two images are related to each other, pixels from each triangle can be exchanged between the two images based on their corresponding positions. The results are shown in column (C) of FIG. 28. This result has the advantage that even if the eyes are closed in the raw image, the eyes can be made open in the output image, as shown in 2806, and even if the eyes are open in the raw image, the eyes can be made closed in the output image, as shown in 2805.
[0152] The blink detection algorithm detects the presence or absence of blinks from the input image, and features associated with blinks and facial images are collected and used to artificially generate a video of a user with a natural perception of blinks by simulating a blink sequence that is naturally perceived by the human mind, and by swapping the eye areas between the two images, the person in the image appears to be blinking naturally to the user viewing the image, even if the user was not blinking at the time the image was taken.
[0153] Sections 4 and 5 described the respective algorithms used to identify and process pre-capture images during the pre-capture stage 700 of Figure 7. Once sufficiently processed, these images can be used as source images for use in the replacement process described below. As a first part of the removal pipeline, a configuration process must be performed to properly correlate the images being captured by the image capture device with the images being displayed to the user in the HMD. This begins in Section 6, which describes a time-shift alignment process between the images in the HMD and one or more sensors in the HMD.
[0154] Section 6: Time shift alignment between HMD images and HMD IMU data Returning to the second stage of the HMD removal algorithm in Section 3, which performs real-time HMD removal, the relevant first step is the alignment process between the HMD image and the inertial measurement unit (IMU) data acquired from the HMD. Given video frames of a user wearing a virtual reality (VR) headset and headset position and orientation data provided by the inertial measurement unit (IMU) sensor, it is desirable to accurately map (or replace) the image onto the user. This requires computations that rely on accurately detecting the headset in the video frames and its orientation based on the IMU data. These two sets of data must be properly aligned; otherwise, the image may be mapped to the correct location but with a different orientation, or vice versa. The problem with this is that the sets of data are acquired from different devices, and there is no way to ensure that the corresponding data in each set is properly aligned. This also applies when all data is time-stamped, but there is no guarantee that the clocks of the two devices are synchronized. The alignment algorithm estimates the offset using two different techniques: cross-correlation estimation and common information estimation.
[0155] Both methods require a certain number of video data frames to calculate the offset. In one embodiment, at least several hundred frames are used to accurately estimate the offset. The calculation can be performed over 200 to 1,000 frames, but is typically performed over an intermediate period of approximately 600 frames to maximize accuracy and time-consuming sampling times. A high frame count means that offset correction is not performed while the sample frames are being collected. Thus, for 600 frames at 30 frames per second, no correction is applied for at least 20 seconds until the estimation is complete. While a larger sample size helps ensure sufficient data for analysis, sample frames can be used to test their effectiveness, as explained below.
[0156] Each video frame is accompanied by IMU data, including the X, Y, and Z position of the headset, as well as its pitch, yaw, and roll. This IMU data forms the first dataset. Within the video frame, a bounding box is detected around the headset, and the coordinates of this bounding box are used as the second dataset (this means that only frames in which the headset can be detected are valid for both approaches). The IMU data is aligned with the video frames by timestamp; however, there is no guarantee that the video and IMU times will match. The solution below uses the motion described by both datasets to estimate an offset that synchronizes the two datasets.
[0157] The first alignment process used here is based on cross-correlation. Cross-correlation compares specific signals from each data set, specifically signals that should have a higher correlation with each other. This is most often done by comparing the IMU pitch reading with the Y coordinate of the headset's bounding box in the video frame. If the user tilts their head up or down while looking into the headset, the headset will also move up or down relative to the video frame. For the same reason, comparing the IMU yaw with the X coordinate of the bounding box works equally well. Depending on the user's movement, using the IMU's X or Y position may also be a viable option to refine the estimate.
[0158] This is shown in Figure 29, which compares a sample of pitch values from the headset with the Y coordinate of the headset's bounding box in a video frame. Both signals are normalized as part of the process to estimate the offset and make the visual comparison clearer. In Figure 29, the normalized data from the headset pitch values is plotted as a solid line, and the normalized landmark detection Y coordinate is plotted as a dotted line. In Figure 29, the Y axis is the normalized value for each data set, and the X axis is the frame number. From the output, we can determine the frame number at which we offset the coordinate (dotted line) data to most accurately match the IMU data. The criterion for determining the score for each offset is the cross-correlation coefficient of the two signals.
[0159] To determine the optimal offset, we calculate the cross-correlation coefficient between the two signals for a range of offsets. Here, we want to identify smaller offsets between -5 and 10 frames (shifting the data 5 frames backward or 10 frames forward). For example, the data shown in Figure 29 visualizes approximately 150 frames of video and IMU data. To calculate the cross-correlation coefficient, two equal-length arrays are required. To allow enough room to shift the data 5 frames backward or 10 frames forward, frames 5 through 140 of the IMU data (a total of 135 frames) are used. In this way, if we offset the video data 5 frames backward, IMU frames 5 through 140 can be compared with video frames 0 through 135. Similarly, if we shift the video data 10 frames forward, IMU frames 5 through 140 can be compared with video frames 15 through 150. For each value in the -5 through 10 range, the video data is offset by that value, as shown by the dotted lines in Figure 24, and the cross-correlation coefficient is calculated using NumPy, an external library that provides fast methods for array arithmetic. Rather than simply taking the offset with the highest cross-correlation coefficient, we take the coefficient of the peak offset and the coefficients of its two adjacent offsets (on either side) and perform a quadratic interpolation as shown in Figure 30. Note that Figure 30 shows the offset on the x-axis and the cross-correlation coefficient on the y-axis, with the maximum coefficient interpolated.
[0160] As shown in Figure 30, the offset is maximized (x-axis) at 1.0. Looking at neighboring offsets, the coefficient at 0 frames is significantly higher than at 2 frames. To find a more accurate offset, we take the two neighboring offsets on either side and their coefficients, and perform a quadratic interpolation. The peak of this interpolation, just below 1.0 frame (shown as the peak of the parabola in Figure 30), is taken as the final offset for pitch and Y coordinate. This process is repeated across other data sets, including (a) yaw to X coordinate, (b) IMU X to X coordinate, and (c) IMU Y to Y coordinate. It is determined whether the calculated offset should be considered an outlier (because the offset is outside a specified range or the cross-correlation coefficient score is below a certain threshold), and if it is outside the offset range, it is filtered out. In this case, the calculation is performed over the range of frames -5 to 10 of the segment of video data. The remaining calculated values are averaged, and this average is returned as the final result. This approach can be summarized as shown in Figure 31.
[0161] The second alignment process is based on mutual information and differs in how the acquired data is interpreted. With mutual information, each dataset is analyzed as a whole, rather than as individual parts. Principal component analysis (PCA) is then used on both datasets to determine the dimension with the greatest variation. For example, rather than comparing the video coordinate Y to the IMU pitch, coordinate X to the IMU yaw, etc., as in the cross-correlation alignment process, mutual information alignment compares coordinate X and Y together with all the IMU data (pitch, yaw, roll, X, Y, Z). PCA accomplishes this by taking all six dimensions of the IMU data and reducing them to the single dimension with the greatest variation. The same is done for the two dimensions of the video coordinate data (X and Y). We calculate PCA using functions provided by scikit-learn, a library of data analysis methods for Python. This allows the algorithm to compare the two datasets as a single dimension, rather than all of their original dimensions.
[0162] Similar to the cross-correlation process, the two signals are compared at different frame offsets. However, instead of using the cross-correlation coefficient, a new criterion is calculated to determine the optimal offset. For each offset, two one-dimensional data sets are input into a 2D histogram, with each axis representing one of the data sets. The number of bins in each histogram is determined based on the size of the given sample data. Currently, with 600 frame samples, we use a 10x10 histogram. This is shown in Figure 32, where the histogram shows the calculated entropy at 0 offset.
[0163] Once the histogram is filled, the next step is to calculate the relative entropy of the histogram. This histogram is the baseline for this approach. From here, the procedure is similar to the first approach. We take the offset with the highest entropy value and the entropy around it, and re-interpolate the peak offset using the same quadratic interpolation method. Because we look at each data set as a whole and find the dimension with the most variation, there is no need to average across different dimensions; this interpolation is the final result. This process is shown in the flow diagram in Figure 33.
[0164] It is important to note that each of the two alignment methods has its own method for determining the validity of a sample frame. These methods cannot accurately estimate the offset if there is no motion in the recorded frames. Without motion, the signal being compared will not change. While it is important to have a large enough sample size of frames (200-1,000), it is also possible to check for motion in each case. Once the frames and data are collected, the data is evaluated before any calculations are performed to estimate the amount of motion depicted in the sample. In the case of cross-correlation, the signal is examined directly, so the amount of motion can be determined by directly examining the signal value and fluctuation. By comparing the absolute value of the data in each frame with the average position, it is possible to determine whether the amount of fluctuation meets a certain threshold. In the case of mutual information, when the dataset is reduced to one dimension using PCA, the variance ratio described above is used to determine the proportion of variance in one dimension. This again determines whether the motion exceeds a certain threshold, i.e., whether the frame is usable.
[0165] This allows for up to 10 frames of misalignment between the resulting video data and the headset's IMU data to be corrected. While a misalignment of a few frames may seem small, it can result in a noticeable delay between the raw video and the frame-locked, corrected image in the final output. Because IMU data is required to measure the user's orientation relative to the headset for headset removal, and there is no way to guarantee that the video and headset clocks are synchronized, this solution accounts for the discrepancy between the two data streams. This alignment algorithm allows a comprehensive program to run continuously while collecting sample sets. It then uses samples from these two input data sets to determine the offset between the two data streams and how one signal corresponds to the other. Once this is complete, the program can continue running using the estimated offset, resulting in more accurate calculations.
[0166] Continuing with the description of other aspects of the alignment process that is performed, Section 7 describes the alignment process between the IMU device and the imager.
[0167] Section 7: Alignment between the IMU device and the imaging device Returning to the second stage of the HMD removal algorithm from Section 3, when real-time HMD removal is performed, the second step involved is the alignment process between the IMU device and the image capture device that captures live images of the user wearing the HMD.
[0168] Head pose estimation is a computer vision task that predicts the head orientation (yaw, pitch, roll) of a camera viewpoint. This prediction is necessary for successful face swapping, gaze estimation, face recognition applications, and augmented reality (AR) applications. A first exemplary head pose estimation is a two-stage approach in which a trained machine learning model first predicts facial landmarks from an input image, and then estimates head pose by aligning the predicted facial landmarks with those of a standard face. In another embodiment, a trained machine learning model is trained to directly estimate head pose from the input image. A challenge associated with both of these approaches is unique to the situation when a user is wearing a VR device that occludes the face. In these approaches, the model architecture includes a backbone subnetwork—a convolutional neural network (CNN), such as MobileNet, ResNet, or VGG—to extract feature maps from the input image. However, these models require the input image to contain a clear face that is not significantly obscured by a head-mounted device (HMD). Therefore, if these models are fed images in which the user's face is somehow occluded, they will not produce the desired prediction results if key facial features are obscured by an opaque head-mounted device (HMD), such as the Oculus Quest-2 headset. The HMD orientation process below solves these problems and provides an algorithm that can estimate the position even when part of the face is occluded.
[0169] During the VR experience, the user is fixed to the head and is positioned at [X imu , Y imu , Z imu ,] and orientation measurement [Yaw imu , Pitch imu , Roll imu The user wears an HMD equipped with an IMU (Inertial Measurement Unit) that provides the camera's viewpoint [Yaw , ]. During operation, the image capture device captures real-time images of the user wearing the HMD and provides them as input. cam , Pitch cam , Roll cam ,], which is used to estimate head pose data. imu , Pitch imu , Rolli imu ,] to [Yaw cam , Pitch cam , Roll cam ,] can be established. Then, the linear transformation between the two parameter systems is as follows:
[0170] Yaw cam = Yaw imu - Yaw0 Pitch cam = Pitch imu - Pitch0 Roll cam = Roll imu - Roll0
[0171] Here, the offset constants [Yaw0, Pitch0, Roll0] are the static posture of the head facing the camera [Yaw cam , Pitch cam , Roll cam ,] = [0, 0, 0]. Head pose estimation using IMU data can be obtained from the rest pose in the alignment step.
[0172] In one embodiment, the alignment process is performed using indicia attached to the HMD. In one example, a QR code is attached approximately in the center of the front panel of the HMD, as shown in FIG. 34. During the alignment process, a QR code processor such as pyzbar decodes the QR code for each frame and obtains a current set of coordinates for the four corners of the QR code, represented as [[x1, y1], [x2, y2], [x3, y3], [x4, y4]]. This is obtained while the user makes small head pose adjustments in front of the camera. When the four corners form an upright square, indicating a resting pose, the current IMU readings are considered offset constants [Yaw0, Pitch0, Roll0], and alignment between the IMU and the camera is performed.
[0173] One method for determining whether a QR code is an upright square includes determining one or more of the following: (a) the left and right sides are vertically oriented, (b) all sides are equal in length, and (c) the four interior angles are 90 degrees. An exemplary alignment process is shown in Figure 34. As shown in Figure 34, the four circles indicate the current coordinates of the four corners decoded by the QR code module.
[0174] In another embodiment, alignment can be performed using inpainting. As used herein, inpainting refers to the process of adding features that exist in real space but are not visible in virtual space because they are obscured by an object. For example, upper face inpainting refers to the process of adding upper facial features (e.g., eyes, nose, forehead) onto the HMD in the captured image. In this embodiment, a QR code is not required. The alignment process consists of the following four elements and will be explained using Figure 35. Figure 35 is a schematic diagram of the inpainting process. An input image 3501 is provided. In A, a bounding box surrounding the HMD is determined in 3502. In 3503, a bounding box that surrounds and includes candidate facial features is obtained. In B, alignment between bboxB and bboxA is performed.
[0175] To provide further detail to the depiction in Figure 35, the following algorithm is used in 3504 to generate the final output: In 3501, an input image of a user wearing an HMD, captured by an imaging device, is provided. In 3502, a bounding box (bboxA) of the head mounted device (HMD) in the image is generated. This is done by segmenting the image, where the bounding box (bboxA) is defined by its top left corner [x min , y min ] and the bottom right corner [x max , y max [x, y, z, ]. In 3503, a predefined bounding box (bboxB) is generated that prompts the user to adjust its position relative to the camera [x, y, z, ] for alignment. In 3504, bboxB and bboxA are aligned by performing an affine transformation (scaling and translating bboxB to fit bboxA). Next, the affine transformation can be applied to a predefined image of the eyes and nose for inpainting the upper part of the face. From this image, a Perspective-n-Point (PnP) method is used to estimate the head pose using landmarks in the lower part of the face and inpaint the hidden upper part of the face. Once the bounding box alignment and inpainting are complete, a trained machine learning model is used to estimate facial landmarks, providing an image to confirm the estimated landmarks in the visible lower part of the face, and the head pose is estimated.
[0176] By collecting one data point close to the resting attitude, the offset constants [Yaw0, Pitch0, Roll0,] can be estimated using the following formula:
[0177] Yaw0= Yaw imu - Yaw cam Pitch0= Pitch imu - Pitch cam Roll0= Roll imu - Roll cam
[0178] If multiple data points (e.g., k = 1, 2, 3, 4, 5) are collected around the resting pose, the offset constants [Yaw0, Pitch0, Roll0, ] are estimated via the following linear regression:
[0179] loss(Yaw0) = sum of |Yaw k imu - Yaw k cam - Yaw0| 2 loss(Pitch0) = sum of |Pitch k imu - Pitch k cam - Pitch0| 2 loss(Roll0) = sum of |Roll k imu - Roll k cam - Roll0| 2
[0180] [Yaw cam , Yaw imu ], [Pitch cam , Pitch imu ], [Roll cam , Roll imu If a one-dimensional scan of data points around the resting posture is collected for each of [Yaw 0, Pitch 0, Roll 0], the offset constants [Yaw 0, Pitch 0, Roll 0] are estimated by line fitting and zero crossing, respectively, as shown in Figure 36. Exemplary approaches 3601, 3602, and 3603 for estimating the offset constants 3604 are shown in Figure 36. To demonstrate option 3 (2603), the offset constants [Yaw 0, Pitch 0, Roll 0] are estimated by line fitting and zero crossing, respectively, as shown in Figure 36. cam , Yaw imu Collect a one-dimensional scan of data points around the resting position of []. A linear transformation between the two parameter systems means estimating the offset constant Yaw0 and estimating the intercept by line fitting and zero crossings.
[0181] Current AI models are designed and trained to estimate landmarks from input images in which the face is clearly visible. Therefore, if key facial features are obscured by an opaque head-mounted device (HMD) in the image, the model may fail to predict landmarks. To obtain facial landmarks, we inpaint the eyes and nose on the worn headset, and then call an AI model such as mediapipe face mesh to predict facial landmarks. Since only the lower part of the face is visible to verify the results, we extract landmarks from the lower part of the face to estimate the head pose.
[0182] The left image in Figure 37 shows a typical input image and the output landmarks of the landmark AI model, while the right image shows an invalid input image where the landmark AI model cannot recognize key facial features. The middle image shows that by inpainting the upper part of the face, it is possible to estimate landmarks for the lower part of the face, resulting in a corrected image that allows the AI model to estimate landmarks even if the entire face is not visible in the input image.
[0183] In another embodiment, as part of the process of aligning the HMD's IMU axes to a set of camera-referenced axes, a process is performed to guide the user to look directly at the camera. At this point, the IMU orientation and position are recorded as a reference, and given the distance from the user to the camera as described below, a complete 3D environment can be constructed, and the reference orientation and position can be used to map the IMU orientation and position to the camera axes. Thus, during the alignment phase, a graphical user interface (GUI) is presented to the user to guide the user to look directly at the camera.
[0184] In one embodiment, the generated GUI may include, for example, a rectangle anchored to the frame at the target position, which will allow satisfactory results from the orientation detection algorithm described in the IMU alignment section of this specification. In another embodiment, it may include a rectangle derived from the bounding box of the user's HMD within the frame. In other embodiments, any graphic (logo, image, shape, etc.) or text (e.g., description) may be placed within or around any of the above rectangles to aid the user in the alignment process by improving visualization of the desired position and orientation. The GUI may also include a means for determining the user's movements to make to align the rectangle with the target rectangle, as shown in FIG. 38.
[0185] As shown in Figure 38, two rectangles are generated based on the HMD's position in the face image and displayed in the HMD's GUI. As shown in Figure 38, the destination rectangle is the inner rectangle, and the moving rectangle derived from the HMD is the larger outer rectangle. The image elements represented by arrows indicate the instructed direction of movement for the user and are continually updated to further prompt the user's head movement direction based on the positions of the inner and outer rectangles. An additional embodiment is shown in Figure 39, where instead of the image elements being arrows, there are image elements that match both the outer and inner rectangles, and the user is instructed to match the image elements to each other to determine positions to obtain alignment information.
[0186] Given the rectangles in Figures 38 and 39, alignment is achieved by determining the centroid, or average (x,y) coordinate of each rectangle's corners. The difference between the x and y coordinates is normalized in some embodiments to [0,1] or some other useful scale and used to determine in which direction the user's rectangle needs to be moved to align with the target rectangle. For example, if the y coordinate of the center of the user box is greater than the y coordinate of the center of the target box, the user can be instructed to move their head downward in the frame.
[0187] In addition to having the correct (x,y) coordinates, the user must be the correct distance from the camera to get reasonable results in orientation detection. Using two rectangles, their intersection and union (IOU) calculation can be interpreted as a distance requirement. If the user is too far from the camera, the outer rectangle will be much smaller than the target, and the IOU will be small as well. On the other hand, if the user is too close to the camera, rectangle #2 will be much larger than the target, again resulting in a small IOU. Only when the user is the desired distance from the camera will the IOU be large and orientation detection be reliable.
[0188] Because it is not enough for the user to be in the correct position in the frame; they must also be oriented correctly, some embodiments also consider the orientation detected by the face IMU alignment unit and artificially shift the position of the user's bounding box accordingly. For example, if the user is exactly in the target position (and therefore the bounding box is perfectly aligned with the target box) but is looking left instead of toward the camera, the displayed bounding box will move to the left to indicate a misalignment. Similarly, the bounding box and any graphics and / or text it contains may rotate clockwise or counterclockwise to indicate that the user is not aligned correctly. Once the user successfully moves their head to a position where the two boxes line up, the reference angle becomes reliable and can be recorded for use in other applications, such as HMD removal.
[0189] In another embodiment, the alignment process described in this section involves using a 3D model reshaping process. As described in this embodiment, to estimate facial landmarks from input images / frames in which the face is largely occluded by a head-mounted device (HMD), 3D models of the face and headset are created offline as a default configuration. First, a standard face (standard media-pipe face: 468 facial landmarks) and HMD (e.g., Oculus Quest-2) model are manually aligned and assembled, using a default head model (e.g., SMPL head) as a reference head. In online processing, the standard face can be automatically replaced with a user-specific face using a non-rigid alignment approach (the least-squares method is used to best match one facial landmark with the other). A schematic of the initial model setup is shown in Figure 40, where 3D models of the face and headset are generated offline prior to the processing performed here. The default head model can serve as a visual reference head for manually aligning the face and headset models. This may be partially generated using a face mesh generation library such as Media-Pipe, a landmark inference machine learning model that can predict (media-pipe face: 468 facial landmarks) from face images.
[0190] In online processing, it is also desirable to adjust the HMD position according to a frontal reference image, since different users may have slightly different HMD positions and cover the upper part of their faces in slightly different ways. Here, the frontal image used to complete the IMU alignment step to adjust the HMD position relative to the face model can be used. The adjustment algorithm is shown in Figure 41 and involves generating a bounding box for the head-mounted device (HMD) in the reference image (bboxB). This is a basic computer vision task that can be achieved by image segmentation, and the bounding box is defined by the upper left corner [X min , Y min ] and the bottom right corner [X max , Y max]. The generation of facial landmarks in the reference image can be achieved by inpainting the eyes / nose into the HMD region defined by bboxB and invoking a landmark estimation machine learning model such as the media-pipe face mesh solution. The 3D point cloud [X i , Y i , Z i ] onto the image plane. i , Y i ]. The bounding box of a 2D headset point (bboxA) is calculated as [X min = min(X i ), X max = max(Xi)] and [Y min = min(Y i ), Y max = max(Y i )]. Then, (2) and (4) in Figure 41 are aligned using landmarks on the lower part of the face, which can be achieved by affine transformation (scaling and translation). By comparing the two centers of bboxA and bboxB, the adjustment of the HMD model in the X or Y direction can be derived. As shown in Figure 41, A represents the projection onto the image plane by projective transformation, B represents image segmentation, eye / nose inpainting, and landmark estimation, and (1) and (3) in Figure 41 represent the 3D model and reference image, respectively.
[0191] In addition to aligning the HMD device with the camera capturing the user wearing the HMD, the captured image is used to generate an updated image that includes a replacement portion obtained from the pre-captured image. However, because the pre-captured image was acquired before the current capture of the user wearing the HMD device, a color correction process must be performed in real time between the candidate image for replacing the current captured image including the HMD device and the current captured image including the HMD device. This process is described in Section 8.
[0192] Section 8: Recoloring pre-photographed images Being able to correctly relight or transform the color of an image of an object or environment is useful in augmented virtual reality. This processing is part of the processing performed on pre-captured images described above in Section 3. Images of a virtual environment or images of different users are often taken in different locations, at different times, and therefore under different lighting conditions. Because humans easily perceive lighting differences, combining images without matching the lighting between them can cause the scene to appear unnatural. Furthermore, in augmented virtual reality, users often wear HMD devices, which necessitates HMD removal, i.e., replacement of the HMD region, in order for others to see the user's entire face. When replacing the HMD region with a pre-captured or artificially generated image, the lighting conditions of the pre-captured image must match the lighting conditions of the HMD image.
[0193] In one embodiment, the recoloring process includes a region-based color transformation, which is described below with reference to FIG. 42. Some systems, devices, and methods perform (e.g., implement algorithms that perform) relighting or transforming the colors of a source image (e.g., transforming the colors of a target image to a source image) based on specific regions of the target image. In FIG. 40, the goal is to transform or relight the colors of the face in the source image, shown in region A, based on the face in the target image, shown in region B. Different aspects may be used to transform or relight the colors of the face. In some embodiments, the CIELAB color space is used, while in other embodiments, a covariance matrix of the RGB channels is used. In some relighting embodiments, the goal is to decorrelate the information between the red, green, and blue channels in the RGB color space using the CIELAB color space or covariance matrix. Furthermore, the RGB color space is device-dependent, meaning that different devices have different color reproduction capabilities. As such, the RGB color space may not be ideal for color transformation between images with different illuminations.
[0194] An embodiment using the CIELAB color space is shown in Figure 43, which illustrates an exemplary embodiment of a color conversion method using a first color space (CIELAB color space). After obtaining a source image 4301 and a destination image 4302, these two images are converted from the RGB color space to the CIELAB color space (4303, 4304). Next, regions requiring color conversion are estimated from both the source image 4305 and the destination image 4306, and the mean and standard deviation (STD) of the Lab components of the CIELAB color space in these regions are also calculated for each of the source image 4307 and the destination image 4308. After obtaining the mean and STD for both the source image and the destination image, in some embodiments, the LAB components of the source image are adjusted in 4310 so that they can be described by equation (7) (e.g., according to equation (7)). TIFF2025538281000006.tif11132, where x is one of the three components L, a, or b in CIELAB space. The adjustments determined to be made are then converted to RGB color space at 4312 and used to recolor the source image (e.g., the image used for substitution) at 4314.
[0195] An embodiment using covariance-based decorrelation may follow a similar operational flow to Figure 43. Although the covariance matrix-based embodiment may appear to be more data-driven than the CIELAB-based embodiment, decorrelation of the RGB color space is often simpler in CIELAB because it is designed entirely based on human vision. Therefore, the CIELAB-based embodiment may be faster and more flexible in multi-gamut-based color transformation.
[0196] In another embodiment, the recoloring can use a reference image-based color transform. In other embodiments, multiple recoloring processes can be employed. The reference image recoloring uses a reference image based on a multi-region color transform. This is described with reference to Figure 44, which illustrates a case where region-based object relighting cannot be directly applied.
[0197] When applying region-based color transformation, both the source and target images must share the same features or contain the same feature information. This is because, if the images share the same features, the difference in the mean and STD between regions of the source and target images can be assumed to be due to differences in the illumination of the images, rather than differences in features. Figure 44 shows two cases where region-based color transformation cannot be applied directly.
[0198] In the first case, the color of region A1 is transformed using the color of region B1. Region A1 and region B1 contain different features because the people in the image are wearing different clothes. Therefore, the difference in the mean and standard deviation between region A1 and region B1 cannot be assumed to be due to lighting alone. If region A1 is directly renormalized based on the mean and standard deviation of region B1, it may be perceived as a very unnatural color from a user's perspective.
[0199] In the second case, the color transformation of region A2 is performed using region B2. Region A2 is obtained from the eye area of a human face, but region B2 represents a head-mounted display (HMD). Recoloring region A2 using the average and standard deviation of region B2 is meaningless because they are obtained from a completely different object. Note that the second case is typically when the HMD device is replaced with a pre-photographed image, and the color of the eye area in the upper part of the face in the pre-photographed image is changed to match the color of the lower part of the face.
[0200] Using the same color conversion technique as shown in Figure 43, some embodiments use a reference image to convert colors between regions that do not share features. An example is shown in Figure 45. Note that the reference image can be, for example, the source image or a predefined image for a particular effect. In Figure 45, the source image, reference image, and target image all share the same features in regions SS, RS, and TS. Because these regions share the same features, embodiments of the above method can be used to convert colors between images even if some features are not shared by all images.
[0201] Figures 46A and 46B show exemplary embodiments of color conversion methods. Figure 46A shows a first scheme, and Figure 46B shows a second scheme, in which a reference image R is used for color conversion between a source image S and a destination image T. Figure 46A uses a reference image (region R2) as the destination image for color conversion of region S2. In some embodiments, both the source image S and the destination image T are converted with the color of reference image R. Because both the source image S and the destination image T undergo color conversion from the same reference image R, the converted color of region S2 should match the color of regions T and T1. Therefore, region T2 can be replaced with the color-converted region S2.
[0202] FIG. 46B relates source image S and target image T using reference image R. First, the entire reference image R is color transformed based on the shared regions RS and TS, and then a new transformation is derived based on the shared regions S2 and R2 of source image S and reference image R. After color transforming region S2, the recolored region S2 should match the color of regions TS and T1. Then, region T2 can be replaced with the color-transformed region S2.
[0203] In another embodiment, the recolor process includes a multi-region color transformation process, which will be described below with reference to FIG. 47. The region intensity average and STD are two numbers that provide a quick summary (e.g., statistical information) of the overall lighting conditions or color of the entire region. These two numbers may be insufficient to represent the change in lighting across the region (dynamic lighting changes within the region). However, because objects are three-dimensional and lighting may vary across different regions, lighting may change across the region (lighting may change dynamically). Therefore, some embodiments do not use only one or two numbers to represent lighting across the region (i.e., some embodiments use more than one or two numbers to represent lighting across the region).
[0204] To better represent different lighting across a region, some embodiments use a multi-region-based color transformation. As shown in Figure 47, a mesh can be used to divide an object (a face in this example) or a region in an image into multiple subregions based on feature landmarks. Each mesh in the source image has a corresponding mesh in the target image, and each subregion in the source image has a corresponding subregion in the target image. For example, a triangular subregion S in the source image corresponds to a triangular subregion T in the target image. Because these subregions share the same features and are therefore related to each other (e.g., correspond to each other), the flow diagram shown in Figure 48 can be used to perform color transformation for each subregion.
[0205] FIG. 48 illustrates an exemplary embodiment of a multi-region-based color conversion method. Similar to FIG. 43, this flow begins by converting a source image 4801 and a target image 4802 from RGB color space to CIELAB color space (4803, 4804). The algorithm generates multiple subregions based on triangles or meshes generated from feature landmarks in the source image 4805 and the target image 4806, respectively. In one exemplary embodiment, the subregions were generated based on 468 landmarks in Mediapipe. Based on the 468 landmarks, this embodiment generated 898 meshes, dividing the entire face into 898 subregions. If a subregion is too small, it can be merged with other subregions. After calculating (e.g., estimating) the mean and STD of the merged subregions for the source image 4807 and the target image 4808, the region-based mean and STD of the subregions are interpolated to become pixel-based mean and STD (4809, 4810). These images are used to renormalize the entire source image based on the interpolated mean and STD of each pixel of the target image at 4812. Finally, at 4814, the source image is converted back to the RGB color space and the recolored source image is saved or output.
[0206] In this manner, device, system, and method embodiments transform image colors according to human perception, and transform image colors to account for large changes in lighting conditions or color distribution. Additionally, device, system, and method embodiments enable color transformation of regions that do not share common features by using a reference image, and enable color transformation using multiple regions to account for changes in color distribution across the image.
[0207] According to another embodiment, the recoloring process involves performing a color transformation between images based on the detected corresponding regions of the images. Given video frames of a user wearing a virtual reality (VR) headset, it is effective to render an image of the user without the headset by stitching the image over the frames where the face is obscured by the headset. This is the real-time HMD removal process discussed herein. To maintain the representation of the video frames, only the areas of the face obscured by the headset are replaced. As mentioned above, the stitched images are obtained from a set of sample pre-capture images taken before the user enters the virtual reality environment. Assuming the stitched images correctly replace the missing areas of the face, the color of the pre-capture images must also be changed to match the video frames to seamlessly stitch the two together. The colors of the video frames and pre-capture images showing the same face are correctly matched, but other areas of the two images, such as the background, are ignored because they do not necessarily need to match. This solution requires detecting corresponding face regions in the two images and matching the colors of the images based solely on those regions.
[0208] The first step in the color matching algorithm is to find corresponding candidate regions in the current video frame and the pre-captured image for color matching. Because the VR headset obscures most of the upper face region (eyes, nose, etc.), this process is performed using the lower face region in both the current video image and the pre-captured image. To consider only the lower face region during color matching, a mask is created for each image that includes only the image regions used for comparison. In exemplary operation, this mask is generated to cover the areas of the face not obscured by the HMD device. To do this, a face mesh is created (as described above) based on the currently captured video frame and one or more pre-captured images from the pre-capture stage 700 of FIG. 3. A set of landmarks believed to encompass the lower face is determined, and a mask of the lower face is created by creating a path (i.e., a point-connecting path) between the selected landmarks and including the area within this path as a mask. This process is illustrated using the captured video frame 4901 shown in FIG. 49. A mask is created based on estimated landmarks shown on the user's face in video frame 4902, creating mask 4904, shown as the shaded area in third image 4903. Because headset 4905 obscures much of the face, the library alone cannot estimate these landmarks. However, this set of landmarks, already determined, is used to generate mask 4904. In FIG. 49, the original video frame 4901 is shown alongside frame 4902 with the estimated landmarks painted on it. Finally, in 4903, we see how the mask created by connecting the landmarks in the lower part of the face appears. In operation, the landmarks selected to create the mask are chosen to surround the headset. As can be seen in FIG. 49, mask 4904 is generated such that the left side of the mask is highest near the ear, slopes downward toward the mouth, and rises again on the right side. Because the user may be facing at various angles, the headset 4905 may obscure parts of the mask (such as the right side of the face in FIG. 49).As will be discussed later in this article, it is important to ensure that the headset 4905 is not included in the color-matching mask 4904, as this may result in poor color matching. It is possible to create a smaller mask with fewer landmarks that is less likely to include the headset 4905, but depending on the angle, any areas of the lower part of the face may be obscured by the headset 4905. Therefore, it is desirable that the mask 4904 be large enough to be visible from most angles. In contrast, the algorithm described here preserves the mask 4904 but removes the portions obscured by the headset 4905. This is possible because the algorithm has already detected where the headset 4905 is located in the image being replaced and color-matched.
[0209] Figure 50 illustrates the process of removing the headset region and generating a final mask for use in the region recolor process. Figure 50 shows the landmark-based mask 4904 obtained in the process described above with reference to Figure 49, the previously obtained headset mask 5002, and the desired final mask 5004. The final mask 5004 is generated by removing from the original mask 4904 the portion of the mask 4904 that overlaps with the headset mask 5002.
[0210] Using a pre-capture image representing the user without an HMD makes this process much easier, as there is no headset obscuring the user's face. Thus, Figure 51 shows how a mask 5104 appears in a candidate pre-capture image (the mask is shown as the shaded area). Once the final video frame mask 5004 and pre-capture mask 5104 are generated, color matching is performed between the two images based on these regions. To paste the pre-capture image onto the video frame, the color of the entire pre-capture image is modified to match the current video frame. More specifically, the area considered for color matching is used for color matching, but the color of the entire image is modified because the top portion of the face in the pre-capture image of the user without an HMD is pasted onto the current video frame. To match the images, the color of the video frame is converted to the color of the pre-capture image used for substitution. Both images are analyzed using the LAB color space, which is designed to approximate human vision (as opposed to other color spaces such as RGB or CMYK). Next, for each image, the mean and standard deviation of all three color channels in the masked area are obtained. From here, a process is performed to replace the value with the result value for each color channel according to Equation 8 (the video frame is the source image, and the pre-captured image is the target image). TIFF2025538281000007.tif9144Finally, values outside the standard range of color values (0 to 255) are clipped to ensure facial colors match between the two images.
[0211] This solution has the advantage of being able to perform a more accurate color conversion between the two images, preventing mismatches in the desired areas of the video frame, especially when the face is partially obscured by a headset. This solution recognizes the area of the face displayed in each frame and matches the pre-recorded image to it.
[0212] Section 9: Landmark estimation from occluded face images Facial landmark detection is a computer vision task in which a computer (e.g., a machine learning model implemented by the computer) identifies or predicts landmarks (e.g., points) for eyes, eyebrows, nose, lips, and other facial structures in one or more input images. The results of facial landmark detection can be passed to other functions that perform other computer vision tasks, including face replacement, head pose estimation, gaze direction identification, and augmented reality applications. However, some landmark machine learning models are designed and trained to infer landmarks from images in which the entire face is clearly visible. These models may fail to identify landmarks or return distorted landmarks if key facial features are obscured, for example, by an opaque head-mounted device (HMD) in the image. Therefore, to obtain facial landmarks (e.g., lines connecting landmarks or a face) as output indicative of key facial features (e.g., eyes, eyebrows, nose, lips, and jawline) from input images, some landmark machine learning models (e.g., deep learning models) require input images in which the entire face is visible, for example, not significantly obscured by an opaque head-mounted device (HMD).
[0213] An overview of facial landmark detection is shown in Figure 41. In Figure 41, landmarks output by a trained landmark machine learning model are displayed on an input image 5202. Image 5204 is an input image in which key facial features are hidden by an HMD and therefore invisible to the landmark machine learning model. If facial landmarks can be inferred from occluded faces when video communication between two or more users takes place via an HMD, a wider range of applications will be possible. To infer facial landmarks from input images in which the face is largely obscured by an HMD, some devices, systems, and methods can first generate 3D models (e.g., point clouds) of the face and headset offline and then load these models as runtime data usable for fast online landmark inference.
[0214] 3D model (3D point cloud) A 3D point cloud is a collection of data points that resemble a three-dimensional, real-world object. Each point is defined by its location and (possibly) color. These points can be plotted to create an accurate 3D model of the object. While LiDAR is a common scanning technology for creating point clouds, not all point clouds are created using LiDAR. For example, output landmarks (e.g., point clouds, keypoints) can be generated by a landmark machine learning model from one or more input images. An example of a face mesh solution is shown in Figure 53. In Figure 53, (A) represents a landmark machine learning model that can identify (e.g., infer) 468 three-dimensional (3D) facial landmarks. Next, for landmark inference (e.g., online landmark inference, real-time (or near-real-time) landmark inference), the 3D models of the face and HMD are projected onto the image plane, and a bounding box (bboxB) defined by the projected 2D point cloud of the HMD is obtained. The HMD bounding box (bboxA) in the image is also detected by image segmentation. By aligning bboxB with bboxA, for example by an affine transformation, the facial points can be transformed accordingly to point to landmarks in the image. Thus, correspondences between the projected 2D points and the 2D points detected in the image can be determined. This generates a geometric reference for transforming the facial points that can be aligned with facial features in the image. In addition to the HMD bounding box, other markers such as the corners of a QR code or the HMD camera can also be used to establish correspondences.
[0215] An exemplary algorithm for estimating landmarks will be described with reference to FIG. 54. That is, the online landmark estimation process includes the following four stages: In the first stage, obtain the bounding box of the head-mounted device (HMD) in the image (bboxA). This is a basic computer vision task and can be achieved by image segmentation. Also, for example, the bounding box is obtained by dividing the upper left corner [X min , Y min] and the bottom right corner "X max , Y max ]. This is shown in Figure 54 at 5401 and 5403. In the second stage, the orientation of the hidden face or head is obtained based on data output from the HMD's inertial measurement unit (IMU). This data may be provided as part of the image (e.g., frame) annotation (e.g., as metadata). The HMD's IMU may provide position and orientation measurements [X, Y, Z, yaw, pitch, roll]. In the third stage, the 3D point cloud [X i , Y i , Z i ] onto the image plane, i , Y i This can be achieved by obtaining the bounding box defined by the 2D points (bboxB) of the HMDs 5403 and 5404 through a projection transformation, which is X min = min(X i ), X max = max(X i ), Y min = min(Y i ), Y max = max(Y i ) The projection transformation of the virtual camera can be modeled as a perspective or orthographic transformation. The points described here represent a face wearing an HMD, with one point cloud labeled as the face and the other labeled as the HMD. Perspective projection or perspective transformation is a linear projection in which a 3D object is projected onto the image plane. One effect is that distant objects appear smaller than closer objects. Because the camera lens and the human eye function similarly, perspective projection appears more realistic to the viewer. When a 3D object is placed away from the image plane, perspective projection can be approximated as a weak perspective projection, which is essentially an orthographic projection with a scaling factor. A schematic of perspective and orthographic transformations is shown in Figure 55 below, where a perspective projection 5502 is shown, f is the focal length (the axial distance from the camera center to the image plane), and 5504 is the orthographic projection.
[0216] In the fourth step, bboxB and bboxA are aligned, as shown in 5405. This can be achieved by an affine transformation, which involves scaling and translating bboxB to fit bboxA. This affine transformation can then be applied to the facial points to estimate landmarks in the facial image. Figure 54 shows landmark estimation before the correction process. The process performed between 5403 and 5404 represents projection onto the image plane using a projective transformation, the process performed between 5401 and 5402 represents determining the bounding box through image segmentation, and the process performed using 5404 and 5402 represents aligning bboxB and bboxA using an affine transformation to generate 5405.
[0217] Landmark correction by inpainting If the accuracy of the facial landmarks obtained above is not sufficient (e.g., the lower face landmarks do not align well with the lower face features in the image performed after inpainting with a trained model for detecting lower face landmarks, which can serve as ground truth for calculating the difference), a correction process is performed to update and improve the landmarks. Conventionally, the corners of the bounding box are used as reference points. In this embodiment, the detected lower face landmarks are used as reference points. Instead of the HMD, the lower face is the actual direct target (ground truth) for alignment. An offline procedure can collect a set of pre-photographed images of the user in which the user's face is clearly shown, and the pre-photographed images can be annotated with landmarks and orientation. First, the upper face is inpainted with the upper face from the pre-photographed images. Then, a machine learning model (e.g., mediapipe's face mesh solution) is invoked to detect facial landmarks that closely match the lower face features. In summary, the upper face is inpainted with the upper face from a pre-photographed image, and a landmark machine learning model is invoked to detect facial landmarks, which ensure that the detected lower face landmarks closely match the lower face features in the image. The landmarks are then updated and improved by realigning using the detected lower face landmarks as new geometric references. In addition to using pre-photographed images, a drawn image (e.g., an image obtained by drawing a 3D face model) may also be used for the upper face inpainting. Here, reference points are required to derive an affine transformation. Conventionally, the corners of the bounding box are used as reference points. In this embodiment, the detected lower face landmarks are used as reference points. A new affine transformation can then be applied to the facial landmarks (either from a previous face model or pre-obtained landmarks) to improve alignment with the face.
[0218] An example of the landmark estimation process is shown below in Figures 56A and 56B. The flow diagram repeats the steps described above and shows the timing at which they are performed.
[0219] Thus, some embodiments enable landmark inference even when the target (e.g., a face) is largely occluded in the image, thereby enabling a wider range of applications. Additionally, some embodiments perform time-consuming requirements and operations offline, such as building a 3D model so that it can be loaded as runtime data for online landmark inference, thereby enabling real-time frame processing.
[0220] Thus, in some embodiments, landmarks are inferred from images where the target (e.g., a face) is largely obscured, for example, by a device (e.g., an HMD), or by projecting a pre-built 3D model onto the image plane (thus avoiding the need to invoke a landmark machine learning model for every frame), or for real-time frame processing.
[0221] Section 10: HMD landmark detection processing using CNN Detecting key landmarks in images of people and other objects is a relatively common task in computer vision, yet most machine learning models are trained using images of people whose faces are not occluded. As a result, pre-trained machine learning models tend to perform poorly on images of people wearing head-mounted devices (HMDs) that obscure the upper part of their faces. Furthermore, existing machine learning models that can be trained to detect general landmarks may not perform satisfactorily when trained on datasets that include labeled HMD landmarks. If machine learning models that track landmarks on the HMD itself are available, these machine learning models can be combined with or used to infer results from other landmark detection machine learning models. Some HMDs have a distinctly colored body (e.g., a white body) and multiple cameras (e.g., four cameras) on the front. The cameras serve as excellent landmarks for detecting, tracking, and locating the HMD's orientation within an image sequence, such as an image or video.
[0222] To detect these camera landmarks, some embodiments of a machine learning model (e.g., a convolutional neural network (CNN)) accept an RGB image as input (the following description uses an example image size of 224x224, but other embodiments have other sizes). This machine learning architecture may bear some similarities to U-Net, a CNN that attempts to perform image-to-image operations such as segmentation through a "classical" convolutional backbone extended with short convolutional branches that preserve spatial information at various scales from the input image.
[0223] For example, for an HMD embodiment with four cameras, some embodiments of the CNN have a similar architecture with two convolutional branches that preserve full-resolution spatial information from the input image. These two branches are concatenated with a main feature extraction backbone before a final convolution and sigmoid activation is applied to the output, resulting in a four-channel heat map the same size as the input image with values in [0, 1]. Each of the four channels contains a heat map corresponding to the position of one of the four HMD cameras. The number of channels may also be equal to the number of cameras in the HMD (e.g., 2, 3, 4, 5, or 6). Thus, if the HMD has two cameras, the number of channels may be two.
[0224] Figures 57 and 58 illustrate an exemplary CNN architecture. Figure 37 shows an overview of the entire CNN, while Figure 38 shows a detailed view of the two main components of the backbone. The magnitude of the output of each layer or unit is shown in parentheses. Arrows indicate the path of the input image through the network, with branching and recombination occurring as necessary. The embodiment shown in Figures 57 and 58 includes a total of 2,146,340 parameters, of which 2,142,404 are trainable. Other embodiments may include more or fewer parameters, depending on the "length" of the backbone, the number of convolutions in each secondary branch, the number of convolutions within each unit of the backbone (e.g., U-Net unit), or similar modifications.
[0225] Training set preparation and image preprocessing To train the CNN, some embodiments use a set of images of various people wearing HMDs. This is shown in FIG. 59, which shows how an input image 5901 is preprocessed, e.g., first using a segmentation machine learning model (e.g., a neural network) to remove the background 5902, and then using a different segmentation machine learning model (e.g., a neural network) to segment the HMD region and identify (e.g., identify, find) a bounding box 5903. A square region containing the bounding box is selected, and the image within the bounding box is scaled 5904 to the 224x224 size required for the input of the exemplary CNN (scaling may be different in other CNN embodiments depending on the respective input size). To generate a target output for each image, the location (e.g., (x,y) coordinates) of each of the four HMD cameras is manually labeled in each image. The labels distinguish between the top-left, top-right, bottom-left, and bottom-right cameras. If a camera is not visible in the image, an estimated camera location may be used. In this way, the CNN learns to include spatial relationships between visible landmarks, allowing it to accurately predict the locations of hidden landmarks.
[0226] The determined positions (e.g., (x,y) coordinates) are then converted into a four-channel image compatible with the CNN output. In this example, we created an all-zero array of size (224,224,4) and then generated a 3x3 square centered on the (x,y) coordinates of the top-left camera in the first channel of the output. This was repeated for the top-right (channel 2), bottom-left (channel 3), and bottom-right (channel 4) cameras. Note that the four-channel image should generally have the same number of channels as the HMD cameras / features being detected. Each channel corresponds to a camera / feature of the HMD and is fed with data as described above. This four-channel image serves as the target output for the input image. The network accepts a preprocessed 224x224 RGB image as input and generates a corresponding 224x224 four-channel image generated from the manually identified camera (x,y) coordinates as output.
[0227] Learning Loop The CNN can be trained using preprocessed images (e.g., input RGB images cropped to fit the HMD bounding box, as shown in Figure 59 above) with random operations such as translation, scaling, and rotation on both the input image and the target heat map. Training can use a modified binary cross-entropy loss function, applied pixel-by-pixel and using constant multipliers (α, β) to emphasize correcting false negatives over false positives. This is because a heavily biased output toward zero makes it more likely to simply output zero everywhere. Training can also include a weighting term, γ, that penalizes positive outputs in multiple channels for the same pixel, encouraging the CNN to distinguish between landmarks for each camera, rather than training to output 1 in all channels for each landmark. Some embodiments of the loss function can be described in Equation 9. TIFF2025538281000008.tif10161where y p,c ^ is the predicted value of channel c for pixel p, and y p,c is the target value. The 4 and 224 in this function are specific to the 4-camera network in this example, which accepts images of size 224x224, where the terms are explained above. The alpha term penalizes false negatives, the beta term penalizes false positives, and the gamma term penalizes positive outputs in multiple channels.
[0228] The weights of these three terms were changed in stages, first by equalizing the weights of false positives and false negatives (α = β), then by increasing the weight of the false negative term (α > β), and then by increasing the weight of the multi-channel penalty γ before the final stage. In this training and CNN embodiment, each stage took approximately 50 epochs.
[0229] Post-processing of heatmap output After training and evaluation, the output of the CNN is still four separate heatmaps rather than a set of four (x,y) coordinates for each individual landmark. To obtain the coordinates, some embodiments perform the following post-processing operations shown in the output heatmap 6001, as shown in Figure 60. These include thresholding and binarizing the output 6002, clustering the resulting non-zero pixel locations in each channel 6003, calculating the condition number of the 2x2 covariance matrix for each cluster 6004 and the total number of pixels in each cluster, selecting the best cluster based on shape (1 / condition number) and size (number of pixels) in the ratio 70-30 6004, and using the center of mass (COM) of the pixels in the best cluster to determine the (x,y) coordinate 6005.
[0230] The machine learning model (CNN) can reliably detect all cameras on the HMD, even if they are occluded. Furthermore, this machine learning model can easily be trained to detect different landmarks using other datasets. By gradually changing the loss function parameters, the network can learn incrementally, refining its knowledge at each stage rather than trying to learn everything at once. Adding clustering and covariance stages in post-processing enables more robust false positive detection and removal.
[0231] Thus, some embodiments include a machine learning model (e.g., CNN) for landmark detection, gradually varying the loss function parameters, and generating a post-processed heatmap output including clustering and covariance matrices.
[0232] Section 11: Adjusting the scale and shift when embedding 2D images into 3D virtual reality Returning to the real-time HMD removal process, after all extractions and replacements have been performed, a live output image containing the replaced parts is provided and displayed to another user in the VR environment. A challenge associated with generating the live output image is related to the process of inserting a 2D image into a 3D environment. More specifically, an algorithm is needed to correct for the shift and scale issues that arise when inserting a 2D image into a 3D image.
[0233] The problem solved by this algorithm can be understood by looking at Figures 61A and 61B, which show different methods for creating the perception of 3D content. Figure 61A shows that the 3D effect of a person can be perceived when the person is created in 3D or with depth information in a 3D virtual environment. However, as shown in Figure 61B, the 3D effect of a person can be perceived even without depth information. In Figure 61B, a 2D image of a person is placed in a 3D virtual environment. Although there is no 3D depth information, the resulting 2D image is perceived as a 3D person by automatically filling in the depth information. This is similar to the "filling in" phenomenon of blind spots in human vision.
[0234] Figure 42 illustrates the difference in 2D projections from different camera models between a portrait and a 3D virtual environment, which is the source of the problem solved by the algorithm of this disclosure. Here, a pinhole camera model is used to illustrate the 2D projection of a 3D object. The pinhole camera model uses triangles to simplify the mathematical relationship between the coordinates of a point in the 3D physical world and its projection onto the image plane.
[0235] Placing a 2D person image in a 3D virtual environment may allow humans to perceive it in 3D, but one or more adjustments are required to insert it naturally into the 3D virtual environment. One reason for the adjustment process is that the camera model used to capture the real person differs from the camera model used in the 3D virtual environment. A 2D image can be thought of as a projection of a 3D object within a physical camera model. Here, the camera model refers to factors that affect the projection of a 3D object, such as focal length, angle of view, image size, and resolution. The same 3D object will result in different 2D projection images using different camera models. Conversely, even if the same 2D image is used, humans may perceive the 3D object differently if the assumed camera model is different.
[0236] The challenge is illustrated in Figure 62, which shows two different camera models, each with its own specifications and used in two specific environments. The first camera system, shown at 6210 in row (A), is from a 2D imaging environment in which a real person moves in a real 3D world, and is identified as camera C as shown in Figure 62. The second camera system 6220 is from a 3D virtual environment displayed on the display of an HMD, and a user is assumed to move through this 3D virtual world, as shown in Figure 62 by camera V. When these models are different, it is difficult to compensate for these differences, as will be shown below.
[0237] When a person of height h moves a distance d, it can be modeled as a person represented by a solid line MN moving from z1 to z2 on camera C (6210), or a person represented by a solid line MM-NN moving from zz1 to zz2 on camera V (6220). Because both models represent the same person's movement in 3D space, the distance between z1 and z2 and the distance between zz1 and zz2 are the same, and the length of MN and the length of MM-NN are also the same. If the focal length of camera C (6210) is fc and the focal length of camera V (6220) is fv, the 2D projections of the same person, i.e., the 2D images drawn or photographed by these two models, are represented by a movement from y1 to y2 on camera C and a movement from yy1 to yy2 on camera V. Because fc is different from fv, the change in size in the 2D photographed environment differs from the change in size in the 3D virtual environment. In other words, the ratio of y1 to y2 differs from the ratio of yy1 to yy2. Furthermore, the difference is not only related to the change in size, but also to the position or deviation from the optical axis. Therefore, a 2D image taken in the camera C environment 6210 cannot be directly placed in the camera V environment 6220. This is because the human mind will not correctly perceive the 3D effect from the 2D image if the image is not corrected.
[0238] The process performed to correct for this scale and shift is described with reference to FIG. 63. This process adjusts the 2D image captured by the user's camera to achieve a natural perception of a reasonable scale and shift in the virtual environment's camera system. FIG. 63 illustrates a correction algorithm for correcting the inaccurate / inconsistent perception between a 3D person and a 3D virtual environment that occurs when a 2D person image is simply placed on a different camera system. This correction algorithm advantageously restores the expected change in the 2D image in the virtual camera system based on the recorded change in the 2D image in the camera used to capture the user's live view in real space. This is particularly useful for the real-time HMD displacement algorithm described herein, which must ensure that the scale matches the live-captured image, especially when performing HMD displacement of pre-captured images.
[0239] Referring to Figure 63, there are two camera systems, one for the real space 6210 and one for the virtual space 6220. As a person MN moves from z1 to z2, a change in the 2D image from y1 to y2 is recorded on the imaging plane at the imaging focal plane fc. To enable the user to perceive the same movement of the person from zz1 to zz2 on the virtual camera, the delivered or rendered image must change from yy1 to yy2. For any position yx between y1 and y2 resulting from movement to zx, the adjusted image on the imaging plane of the virtual camera shifts from yy1 to yyx, allowing for a consistent perception of movement from zz1 to zzx in the virtual 3D reality.
[0240] The correction algorithm, as shown in (1) of Figure 63, involves eliminating changes in scale and shift in the 2D image due to the person's movement, manually aligning the 2D image information y1 in pixel units with the 2D image plane information yy1 in meters or other physical units in the 3D virtual world, and adjusting the scale and shift from yy1 to yyx using a pinhole virtual camera model V. The second step only needs to be performed once throughout the entire process, because the purpose of this step is simply to associate the same person between the image coordinates and the physical world; once the starting positions z1 and zz1 for the two camera systems are selected, this relationship does not change. The transfer function derived in this step is implemented as a global parameter for the input scale and shift adjustment. In one embodiment, the scale and shift adjustment in the 3D virtual space is performed automatically if the movement distance x in the physical, real 3D space is known.
[0241] In one example, a virtual imaging plane MM-NN is placed in 3D virtual space. This is a physical plane in the virtual world, with an initial position zz1. This ensures that the physical units of this virtual plane match the physical units of the image information. For example, if a person is 2 meters tall in the real 3D world and z1 is 100 pixels in the captured 2D image, an imaging plane is created that shows a 2D person with a height of 2 meters located at zz1. Then, as the person moves from z1 to zx, the image plane also moves from zz1 to zzx. This is shown in Figure 64.
[0242] Based on the above explanation, the only remaining step is to remove the scale and shift changes and reset yx to y1 when the user moves to position zx. This is shown in Figures 65A and 65B, which illustrate why directly changing the scale of a person's image is not ideal. Scale and shift adjustments should not be applied directly to a person's height or width. An example is shown in Figures 65A and 65B, which illustrate two cases where a person's height and width change when they perform some action, such as raising their hand as shown in Figure 65A or bending over as shown in Figure 65B. Thus, directly using a person's height and width is not suitable for resizing an image.
[0243] Figure 66 shows a pinhole camera used to rescale and shift the height and width dimensions of a user, illustrating the realignment of the image to compensate for the movement of the person being photographed in real time. In Figure 66, a floor is included in the camera model to further adjust the image dimension and shift. The person's height is denoted by h, and the distance between the floor and the optical axis is g. Note that g can be negative depending on how the image coordinates are defined. Considering a pinhole camera model with focal length f, P1, Px, and P2 are the 2D projections of a person of height h at positions z1, zx, and z2, and the endpoints of P1, Px, and p2 are y1 and b1, yx and bx, and y2 and b2. From this, the following equations can be derived with reference to Figure 66:
[0244] To obtain the relationship between the person's position in the real world and its 2D projection on the focal plane, we use the relationship between triangle M1-z1-O and triangle y1-fO according to Equation 10.
[0245] y1 / f=(g+h) / z1 …(10)
[0246] Similarly, according to Equation 11, we use triangle Mx-zx-O and triangle yx-fO to recalculate the relationship between the 3D person and its 2D projection when the person moves from position z1 to position zx.
[0247] yx / f=(g+h) / zx=(g+h) / (z1+d) …(11) By dividing equation (11) by equation (10), the ratio of y1 to yx can be obtained.
[0248] y1 / yx=(z1+d) / z1=s …(12) Here, this ratio is the scale change at one end of the 2D projection of the person. Similarly, the scale change at the other end of the 2D projection can be obtained from equations (13), (14), and (1514).
[0249] b1 / f=g / z1 …(13) bx / f=g / zx=g / (z1+d) …(14) b1 / bx=(z1+d) / z1=s …(15)
[0250] The lengths of the two-dimensional projections of the person at two positions z1 and zx on the focal plane are (y1-b1) and (yx-bx). Therefore, the magnitude change of the person's length on the focal plane can be determined by equation (16).
[0251] (y1-b1) / (yx-bx) =(s*yx-s*bx) / (yx-bx)=s …(16) And the scale adjustment is s = (z1 + d) / z … (17) is.
[0252] Since the change in size of the person shown in equation (17) is the same as the change in size of the image shown in equations (12) and (15), the size of the image changes, and therefore the size of the person must also change accordingly. Next, to change the scale of the person, if we know the starting position z1 and the distance the person has moved in the real world, we can simply multiply the image by a scale factor of (z1+d) / z1 to determine how the image will look when the person is not moving.
[0253] Returning to Figure 66, the value of d is often obtained from the IMU sensor of the head-mounted display device. For example, the distance traveled in the x, y, and z directions in the 3D real world is obtained using the IMU sensor of the head-mounted display device. Furthermore, before running this algorithm, z1 is estimated by transforming Equation (12) as follows:
[0254] z1=d*bx / (b1-bx) …(18)
[0255] Here, z1 is the position of the person selected in the physical world to align the person's 2D projection in the focal plane with the 3D virtual world, d is the distance the person moved from z1 to zx, and finally, b1 and bx are the 2D projections of the foot plane estimate at z1 and zx. The scale factor s is obtained from the z1 and d readings from the IMU. Similarly, we estimate the image shift based on the same pinhole camera described in Figure 66. As the person moves from z1 to zx, the projected person also changes from P1 to Px. Clearly, the change from P1 to Px involves a shift in position as well as magnitude. The shift is estimated by subtracting bx from b1.
[0256] Mathematically, the shift is obtained by subtracting equation (14) from equation (13) on both the left and right extremes, as shown in equation (19). shift=b1-bx=f*g / z1-f*g / (z1+d) =f*g*(1 / z1-1 / (z1+d)) =f*g*d / (z1*(z1+d)) …(19)
[0257] After obtaining the values of f, g, and z1, the shift value is estimated based on the movement of d. The value of z1 is estimated based on equation (18), but there is no need to estimate f and g separately; only the value of f*g derived from equation (13) needs to be estimated. Rearranging equation (13) gives f*g=b1*z1 …(20) where z1 is the position of the person selected in the physical world to align the 2D projection at the focal plane to the 3D virtual world, and b1 is the 2D projection of the foot plane estimate at the position of z1.
[0258] Substituting equation (20) into equation (19) gives equation (21).
[0259] shift=b1*z1*d / (z1*(z1+d)) …(21)
[0260] Because b1 and z1 can be estimated, the shift can be calculated based solely on the movement of d. By calculating the shift value to apply to resizing live captured images from the first camera system to the second camera system, a real-time live HMD removal process can be performed, allowing users to see each other in an HMD-free VR environment even though the images used to generate those images were taken while wearing an HMD and were taken with a different camera system than the camera system of the VR device on which they are viewing the images.
[0261] Section 12: Network-based alpha channel transmission to enable HMD removal processing There are a variety of application programmer interfaces (APIs) that can be used to transfer media data between applications over an Internet connection. However, some APIs are limited to transmitting only the red, green, and blue (RGB) color channels over the network. To enable full or partial transparency in video, an alpha channel is required.
[0262] Consider a rasterized video frame with four dimensions (r, g, b, α) that is h units high and w units wide. These dimensions represent color values (0-255) in the r, g, and b dimensions, respectively, and transparency (hereafter referred to as alpha) in the α dimension. In this scenario, there are two devices on the network: a source device that generates the video frames and performs the majority of the visual processing, and a receiving device that receives the video frames and performs relatively little visual processing. Because the API in this scenario is limited to RGB, a method of transmitting transparency data over the network is desirable to enable the above processing.
[0263] Chroma Key The first method for transmitting transparency data over a network is to use a chromakey transfer method. One way to transmit transparency data within an RGB image is to encode it within the RGB color space itself. The receiving device on the network renders pixels with that color code as transparent. In the source image frame, all pixel values that are to be transparent are designated as having an alpha value below a certain threshold, and pixels with alpha values above the threshold are designated as foreground. These foreground pixels are then scaled down by a percentage S∈[0,1]. Scaling the foreground pixels ensures that the transparent and foreground color spaces do not overlap. Here, foreground pixel values are clamped between [0,S×255]. The RGB values of transparent pixels are set to 255 for the r, g, and b pixel values. The alpha channel is then discarded, and the RGB frame is transmitted over the network. On the receiving device, the RGB values of the pixels are examined individually. If they fall within the range [0,S×255] of the foreground pixels, they are rescaled by dividing them by the scale factor S and rendered as is. If the RGB pixel values are all within the range [(1-S) x 255,255], they will be drawn as transparent. The reason for this range rather than 255 is that lossy compression (RGB → YCrCb → RGB) can cause pixel values to change slightly when transmitting frames over a network.
[0264] In this example, the scaling for foreground pixels is chosen to be S=0.95. Therefore, the RGB range for foreground pixel values is [0,242], and the RGB values of the background pixels are all set to 255. Figure 67 shows the initial image frame before alpha color encoding, and Figure 68 shows the frame after alpha color encoding. On the receiving device, all pixels that are not 255 are divided by 0.95 to restore their original color, and pixels with an RGB value of 255 are drawn as transparent.
[0265] Data Channel Another way to achieve transparency over a media network without using an alpha channel is to use a separate data channel. Some real-time communication (RTC) protocols have a data channel in addition to the media channel used for video and audio data. In the original frame, the RGBA frame can be split into two tensors: an RGB tensor and an alpha tensor. The RGB tensor contains the original RGB data and can be transmitted over the network without any problems. The alpha tensor is then transmitted over the data channel. However, using a data channel presents some challenges, which are resolved by the methods described below.
[0266] When two data streams are sent over the network from the source device, synchronizing the correct RGB tensor and its constituent alpha tensor becomes a challenge. Network latency can cause variations in the arrival timing of each tensor. The underlying network API can also buffer data packets together for optimal signaling. This buffering scheme can differ between video and data channels, further disrupting the timing of the RGB and alpha tensors. Another constraint when using non-media channels is bandwidth. These data channels often have limited bandwidth compared to media channels, so it is important to utilize that bandwidth effectively so that the alpha tensor can be transmitted without issue. Consider a 1080p frame. The alpha tensor segmented from this frame is over 259 kilobytes in size. With a target frame rate of 30 frames per second, the bandwidth requirement balloons to over 7 megabytes per second. These values can increase with larger frames and faster target frame rates. This is particularly problematic during the time-sensitive HMD removal process mentioned above.
[0267] synchronization One way to synchronize an RGB frame with its component alpha frame is to add a timestamp to the alpha layer. On the source device, when an RGBA frame is created, a timestamp value, in this case a 64-bit unsigned integer, is associated with the frame. When the frame is split into an RGB tensor and an alpha tensor, a timestamp is added to both tensors. Because the RGB tensor is transmitted over the video channel, the timestamp is already attached to the frame. For the alpha tensor, the data is first serialized to a binary string, and the timestamp value is also serialized to a binary and added to the alpha tensor's binary string. This is shown in Figure 69, which illustrates the serialization. On the receiving device, when the alpha tensor is received, it is stored in a hashmap-type data structure using the timestamp as a key. When an RGB frame is received, the timestamp can be used to match it with its component alpha tensor, allowing the alpha tensor to be synchronized with the correct RGB frame.
[0268] Due to bandwidth constraints, the alpha tensor must be properly compressed. Two aspects of the alpha tensor dictate the compression method described here. Alpha tensors are generally sparse, meaning that the data is mostly zero and the data is binary (true or false). Because it is binary data, the bytes of the alpha tensor are packed into a single bit. This compresses the alpha tensor by a factor of eight. The sparsity of the data is particularly advantageous. On the receiving device, once the RGB tensor is aligned with its constituent alpha tensors, their data values are scanned pixel by pixel. However, a slight optimization can be performed by reading the first byte of the alpha tensor as zero, since we know that this first byte actually represents the transparency of the 8 bytes of the RGB frame. Therefore, the algorithm proceeds 8 steps without performing any operations. Because the alpha tensor is mostly zero, this significantly speeds up the process.
[0269] An additional method is to additionally store the index at which the alpha value changes. Instead of storing the values of each individual pixel, we store the index at which the value changes. These index values are stored as 16-bit unsigned integers. The value 25535 is reserved for a new row. In this encoding method, we start with the first row of the RGB frame. The first value of the alpha tensor represents the first foreground pixel, and the second value represents the start of the background pixels. The value 25535 moves to the next row. In this compression method, the size of the alpha tensor changes, so this value must be transmitted along with the tensor. Example ending and starting index values for the alpha tensor are shown in Figure 70.
[0270] It should be noted that both alpha channel transmission methods are not mutually exclusive and can be used simultaneously if network bandwidth and computational constraints allow. In one embodiment, the data channel method is primarily used to transmit transparency data, while in another embodiment, the transparency data is encoded within the color space of the RGB frame, with the chromakey method used as a fallback if the alpha tensor frame is missing.
[0271] The above algorithms represent one or more embodiments of a head mounted display removal process. These embodiments may be performed individually or in combination. The described methods are understood to represent steps stored in a memory that, when executed by a processing device, configures the processing device to perform the described steps.
[0272] Some embodiments of the method include receiving a first image of a user during a pre-capture process, receiving a second image of the user that is partially obscured by the wearable device, determining the orientation and position of the wearable device, identifying the position of the wearable device in the received second image, obtaining a three-dimensional model of the user and the wearable device, performing a region swap on the second image by replacing the obscured portion of the user with a corresponding region obtained from the first image, and generating a third image composed of the second image and the first image as output to a display of the wearable device.
[0273] Some embodiments of the method include acquiring images of the object, detecting landmarks in the images, acquiring landmarks of a reference object, and aligning the landmarks in the images with the landmarks of the reference object. Some embodiments of the method further include generating features based on the aligned landmarks, and some embodiments further include inputting the features into a machine learning model. Some embodiments further include generating landmarks for the reference object based on the set of images of the reference object.
[0274] Some embodiments of the method include obtaining information representing the positions of the upper and lower eyelids from a series of facial images, obtaining information representing the positions of the upper and lower parts of the face from the series of facial images to determine the length of the face, determining the occurrence of blinks by the user in the series of images based on the relative positions of the upper eyelid and the lower part of the eye to the length of the face, extracting a first frame including a blink and a second frame not including a blink from the series of images, and replacing the facial area in the second series of images with the first frame or the second frame based on a predetermined replacement rule.
[0275] Some embodiments of the method include receiving position and orientation information from a wearable device worn by a user having a first time signature, capturing an image of the user wearing the wearable device using an imaging device having a second time signature, determining an offset between the first and second time signatures using the position and orientation information of the wearable device and the orientation and position information extracted from the captured image, and using the determined offset as a reference time to synchronize timing between the imaging device and the wearable device.
[0276] Some embodiments of the method receive a series of images of a user wearing a wearable device, determine the position and orientation of the wearable device based on position and orientation information obtained from one or more sensors of the wearable device and the position and orientation of the wearable device determined from the received series of images, and estimate the user's pose in the received series of images based on the determined position, location, and orientation.
[0277] Some embodiments of the method include acquiring a source image, acquiring a target image, acquiring a reference image, converting the source image from a first color space to a second color space, converting the target image from the first color space to the second color space, converting the reference image from the first color space to the second color space, performing a color transform on a region of the target image in the second color space based at least in part on the reference image in the second color space, and performing a color transform on a region of the source image in the second color space based at least in part on the reference image in the second color space.
[0278] Some embodiments of the method include obtaining a set of training images showing a head-mounted display including one or more cameras, inputting the training images into a machine learning model that outputs respective locations of the one or more cameras, and modifying the machine learning model based on the respective positions output by the machine learning model and the labeled locations of the one or more cameras. In some embodiments of the method, the modification is performed based on a loss function. Also, in some embodiments, the machine learning model has at least one convolutional backbone and at least one convolutional branch that preserves full-resolution spatial information from any input image. Also, some embodiments include preprocessing the training images, where the preprocessing includes removing background, identifying an image region that includes the head-mounted device, or generating a bounding box around the head-mounted device.
[0279] Some embodiments of the method include acquiring a face model, acquiring a head-mounted display model, acquiring an image of a face wearing the head-mounted display, acquiring a face orientation in the image, generating a first bounding box for the face in the image, projecting the face model and the head-mounted display model onto an image plane, generating a second bounding box for the projected model of the head-mounted display, generating a transformation based on the first bounding box and the second bounding box, and inferring landmarks in the image of the face based on the transformation. Some embodiments also include inpainting the head-mounted display in the image with a previously captured image of the face, detecting landmarks in the inpainted face, and refining the estimated landmarks based on the detected landmarks in the inpainted face.
[0280] At least some of the above-described devices, systems, and methods may be realized, at least in part, by providing one or more computer-readable media containing computer-executable instructions for implementing the above-described operations to one or more computing devices configured to read and execute the computer-executable instructions. The system or device performs the operations of the above-described embodiments when executing the computer-executable instructions. Additionally, an operating system on one or more systems or devices may implement at least some of the operations of the above-described embodiments.
[0281] Furthermore, some embodiments use one or more functional units to perform the devices, systems, and methods described above. The functional units may be implemented solely in hardware (e.g., customized circuitry) or in a combination of software and hardware (e.g., a microprocessor running software).
[0282] Furthermore, some embodiments of the devices, systems, and methods combine features of two or more of the embodiments described herein. Also, as used herein, the conjunction "or" generally refers to an inclusive "or," but may refer to an exclusive "or" if expressly indicated, or if the context indicates that the "or" must be an exclusive "or."
[0283] While the present invention has been described with reference to exemplary embodiments, it is to be understood that the invention is not limited to the disclosed exemplary embodiments.
Claims
1. one or more memories for storing instructions; One or more processors that, upon execution of the instructions, receiving a first image of a user during a pre-photography process; receiving a second image of the user partially obscured by the wearable device; determining an orientation and position of the wearable device and identifying a position of the wearable device in the received second image; performing a region swap on the second image by replacing the hidden portion of the user with a corresponding region obtained from the first image; generating a third image composed of the second image and the first image as output to a display of the wearable device; An apparatus characterized in that
2. the received first image includes one or more items of facial information; Region swapping is performed by using the determined orientation and position of the wearable device in the received second image to determine a correspondence between the orientation and position of the wearable device and one or more of the one or more items of facial information from the first image.
2. The device of claim 1 .
3. The one or more items of facial information include one or more of yaw, pitch, and / or roll orientation, one or more facial expressions, or blinks.
3. The device according to claim 2.
4. Executing the stored instructions further causes the one or more processors to create the third image by combining one or more two-dimensional images of the area in the first image corresponding to the hidden portion of the second image with the third image for output to a display of the wearable device.
2. The device of claim 1 .
5. Execution of the stored instructions causes the one or more processors to further: providing the received first images in real time to a first image processing channel that uses alignment and position information of the wearable device to select each of the first images as a candidate replacement image based on facial information associated with the first images; providing the received second image in real time to a second image processing channel that extracts a region of the second image that corresponds to the hidden region; performing the region swap using candidate portions of the replacement image that correspond to the extracted regions from the second image based on orientation, facial expression, or blink information extracted from the second image; 2. The device of claim 1 .
6. Executing the stored instructions further causes the one or more processors to extract the region of the second image by providing the second image to a trained machine learning model that is trained based on a plurality of images of general users wearing the wearable device and that classifies the region in the second image that represents the wearable device.
6. The device according to claim 5.
7. The candidate replacement image includes an eye or nose region, and executing the stored instructions causes the one or more processors to further perform a region swap by inpainting the extracted region from the second image with the eye or nose region of the first image.
6. The device according to claim 5.
8. The generated third image includes the inpainted eye and nose regions from the first image having an orientation, expression, or blink substantially similar to an orientation, expression, or blink of the second image.
8. The device according to claim 7.
9. receiving a first image of a user during a pre-capture process; receiving the second image of the user partially obscured by the wearable device; determining an orientation and position of the wearable device and identifying a position of the wearable device in the received second image; performing a region swap on the second image by replacing the hidden portion of the user with a corresponding region obtained from the first image; generating a third image composed of the second image and the first image as output to a display of the wearable device; A method characterized by:
10. the received first image includes one or more items of facial information; Region swapping is performed by using the determined orientation and position of the wearable device in the received second image to determine a correspondence between the orientation and position of the wearable device and one or more of the one or more items of facial information from the first image.
10. The method of claim 9.
11. The one or more items of facial information include one or more of yaw, pitch, and / or roll orientation, one or more facial expressions, or blinks.
11. The method of claim 10.
12. and creating the third image by combining one or more two-dimensional images of the area in the first image corresponding to the hidden portion of the second image with the third image as output to a display of the wearable device.
10. The method of claim 9.
13. Furthermore, providing the received first images in real time to a first image processing channel that uses alignment and position information of the wearable device to select each of the first images as a candidate replacement image based on facial information associated with the first images; providing the received second image in real time to a second image processing channel that extracts a region of the second image that corresponds to the hidden region; performing the region swap using candidate portions of the replacement image that correspond to the extracted regions from the second image based on orientation, facial expression, or blink information extracted from the second image; 10. The method of claim 9.
14. Further, extracting the region of the second image by providing the second image to a trained machine learning model that has been trained based on a plurality of images of a general user wearing the wearable device and that classifies the region in the second image that represents the wearable device.
14. The method of claim 13.
15. The candidate replacement image includes an eye or nose region, and executing the stored instructions causes the one or more processors to further perform a region swap by inpainting the extracted region from the second image with the eye or nose region of the first image.
14. The method of claim 13.
16. The generated third image includes the inpainted eye and nose regions from the first image having an orientation, expression, or blink substantially similar to an orientation, expression, or blink of the second image.
16. The method of claim 15.
17. A computer readable storage medium having stored thereon instructions which, when executed, cause an apparatus to carry out the method of any one of claims 9 to 16.
18. one or more memories for storing instructions; One or more processors that, upon execution of the instructions, Obtaining an image of a human face, detecting landmarks in the captured image of the human; Obtain landmarks on a reference image of a human face; aligning the obtained landmarks in the image of the human face with landmarks in the reference image of the human face; generating features based on the aligned landmarks; Using a trained machine learning model, classify the acquired human face images and identify the presence or absence of each facial motion unit in the acquired human face images based on the generated features. An apparatus characterized in that
19. The reference image of a human face is either an image generated from the acquired image of a human face or an image different from the acquired image of a human face.
20. The device of claim 18.
20. Execution of the stored instructions causes the one or more processors to further: After capturing the image of a human face, determining whether a reference image of the human face in the captured image is stored in a memory; If it is determined that a reference image corresponding to the acquired image is not stored, a standard image of a human face is acquired to be used as the reference image for alignment; If it is determined that a reference image corresponding to the acquired image is stored, the stored reference image is used for alignment.
20. The device of claim 18.
21. Execution of the stored instructions causes the one or more processors to further: using the stored reference image for alignment when it is determined that the stored reference image was generated based on a threshold number of image frames of the human face in the acquired image; 21. The apparatus of claim 20.
22. Execution of the stored instructions causes the one or more processors to further: After capturing the image of a human face, determining whether a reference image of the human face in the captured image is stored in a memory; If it is determined that the reference image is not stored, using a standard image of a human face for alignment for a predetermined number of acquired image frames including the image of the human face for alignment; Storing images of the human face taken from successive frames and averaging the stored images to generate a user-specific reference image.
20. The device of claim 18.
23. Execution of the stored instructions causes the one or more processors to further: In response to the image processing application determining that the live captured image contains a specific facial motion unit, a classified image in which the specific facial motion unit exists is provided to the image processing application, and the image processing application replaces a portion of the live captured image with a corresponding portion of the provided classified image.
20. The device of claim 18.
24. Obtaining an image of a human face, detecting landmarks in the captured image of the human; Obtain landmarks on a reference image of a human face; aligning the obtained landmarks in the image of the human face with landmarks in the reference image of the human face; generating features based on the aligned landmarks; Using a trained machine learning model, classify the acquired human face images and identify the presence or absence of each facial motion unit in the acquired human face images based on the generated features. A method characterized by:
25. The reference image of a human face is either an image generated from the acquired image of a human face or an image different from the acquired image of a human face.
25. The method of claim 24.
26. Furthermore, After capturing the image of a human face, determining whether a reference image of the human face in the captured image is stored in a memory; If it is determined that a reference image corresponding to the acquired image is not stored, a standard image of a human face is acquired to be used as the reference image for alignment; If it is determined that a reference image corresponding to the acquired image is stored, the stored reference image is used for alignment.
25. The method of claim 24.
27. using the stored reference image for alignment when it is determined that the stored reference image was generated based on a threshold number of image frames of the human face in the acquired image; 27. The method of claim 26.
28. Furthermore, After capturing the image of a human face, determining whether a reference image of the human face in the captured image is stored in a memory; if it is determined that the reference image is not stored, using a standard image of a human face for alignment for a predetermined number of acquired image frames that include the image of the human face for alignment; Storing images of the human face taken from successive frames and averaging the stored images to generate a user-specific reference image.
25. The method of claim 24.
29. Furthermore, In response to the image processing application determining that the live captured image contains a specific facial motion unit, a classified image in which the specific facial motion unit exists is provided to the image processing application, and the image processing application replaces a portion of the live captured image with a corresponding portion of the provided classified image.
25. The method of claim 24.
30. 30. A computer readable storage medium having stored thereon instructions which, when executed, cause an apparatus to carry out the method of any one of claims 24 to 29.
31. one or more memories for storing instructions; One or more processors that, upon execution of the instructions, Obtain information representing the positions of the upper and lower eyelids from a series of facial images, obtaining information representing the positions of upper and lower parts of a face from the series of facial images to determine the length of the face; determining the occurrence of a blink of the user in the series of images based on the relative positions of the upper eyelid and the lower eye with respect to the length of the face; extracting a first frame including a blink and a second frame not including a blink from the series of images; replacing the facial region in the second series of images with the first frame or the second frame based on a predetermined replacement rule; An apparatus characterized in that
32. Execution of the stored instructions causes the one or more processors to further: removing baseline data from the series of images that represent differences in relative position greater than a predetermined distance; comparing the baseline-removed series of images to a first threshold indicative of the likelihood of a blink occurring; comparing the baseline-removed series of images to a second threshold that is less than the first threshold and represents the duration of a blink; Identifying segments in the series of images that exceed both the first and second thresholds as blink segments. The device according to claim 31, wherein the occurrence of a blink is determined by:
33. Execution of the stored instructions causes the one or more processors to further: determining a blink occurrence principle based on one or more characteristics associated with the determined blink occurrence of the user; Replacing the face region in the second series of images with the first frame or the second frame based on the determined blink principle.
32. The apparatus of claim 31 .
34. The blinking principle is determined using a statistical analysis of the occurrence of blinks in at least one series of images of the user.
34. The apparatus of claim 33.
35. The one or more characteristics include at least one or both of a time interval between determined blink occurrences and a blink duration during the determined blink occurrences.
34. The apparatus of claim 33.
36. Execution of the stored instructions causes the one or more processors to further: For said second series of images: Identifying an eye region based on one or more facial landmarks; generating an eye mesh for the identified eye region; replacing the eye mesh in the second series of images with the first frame if it is determined that a blink is likely to occur; If it is determined that a blink is unlikely to occur, the eye mesh in the second series of images is replaced with the second frame.
32. The apparatus of claim 31 .
37. Execution of the stored instructions causes the one or more processors to further: storing the extracted first and second frames in a memory device; Retrieving the extracted first and second frames from the memory device and replacing the region in the second series of images.
32. The apparatus of claim 31 .
38. Obtain information representing the positions of the upper and lower eyelids from a series of facial images, obtaining information representing the positions of upper and lower parts of a face from the series of facial images to determine the length of the face; determining the occurrence of a blink of the user in the series of images based on the relative positions of the upper eyelid and the lower eye with respect to the length of the face; extracting a first frame including a blink and a second frame not including a blink from the series of images; replacing the facial region in the second series of images with the first frame or the second frame based on a predetermined replacement rule; A method characterized by:
39. removing baseline data from the series of images that represent differences in relative position greater than a predetermined distance; comparing the baseline-removed series of images to a first threshold indicative of the likelihood of a blink occurring; comparing the baseline-removed series of images to a second threshold that is less than the first threshold and represents the duration of a blink; Identifying segments in the series of images that exceed both the first and second thresholds as blink segments. The method according to claim 38, wherein the occurrence of a blink is determined by:
40. Furthermore, determining a blink occurrence principle based on one or more characteristics associated with the determined blink occurrence of the user; Replacing the face region in the second series of images with the first frame or the second frame based on the determined blink principle.
39. The method of claim 38.
41. The blinking principle is determined using a statistical analysis of the occurrence of blinks in at least one series of images of the user.
41. The method of claim 40.
42. The one or more characteristics include at least one or both of a time interval between determined blink occurrences and a blink duration during the determined blink occurrences.
41. The method of claim 40.
43. Furthermore, For said second series of images: Identifying an eye region based on one or more facial landmarks; generating an eye mesh for the identified eye region; replacing the eye mesh in the second series of images with the first frame if it is determined that a blink is likely to occur; If it is determined that a blink is unlikely to occur, the eye mesh in the second series of images is replaced with the second frame.
39. The method of claim 38.
44. Furthermore, storing the extracted first and second frames in a memory device; Retrieving the extracted first and second frames from the memory device and replacing the region in the second series of images.
39. The method of claim 38.
45. 45. A computer readable storage medium having stored thereon instructions which, when executed, cause an apparatus to carry out the method of any one of claims 38 to 44.
46. one or more memories for storing instructions; One or more processors that, upon execution of the instructions, receiving location and orientation information from a wearable device worn by a user having a first time signature; capturing an image of the user wearing the wearable device using an imaging device having a second time signature; determining an offset between the first and second temporal signatures using the position and orientation information of the wearable device and the orientation and position information extracted from the captured image; Using the determined offset as a reference time, synchronize the timing between the imaging device and the wearable device. An apparatus characterized in that
47. Execution of the stored instructions causes the one or more processors to further: For each captured image of the user wearing the wearable device, generate a bounding box that surrounds the wearable device in the captured image; obtaining coordinates of the generated bounding box within the captured image; determining the offset using the received position and orientation information from the wearable device and the obtained coordinates of the bounding box in the captured image; 47. The apparatus of claim 46.
48. Execution of the stored instructions causes the one or more processors to further: determining the offset by performing a cross-correlation process using the received information about the orientation of the wearable device and the information about the position of the wearable device in the captured image; 47. The apparatus of claim 46.
49. The orientation information is first orientation information of the wearable device, and the position of the wearable device in the captured image is based on specific coordinates associated with a bounding box surrounding the wearable device in the captured image.
49. The apparatus of claim 48.
50. The first orientation information of the wearable device is a pitch value of the wearable device, and the specific coordinate associated with the bounding box in the captured image is a Y coordinate value.
50. The apparatus of claim 49.
51. The first orientation information of the wearable device is a yaw value of the wearable device, and the specific coordinate associated with the bounding box in the captured image is an X coordinate value.
50. The apparatus of claim 49.
52. Execution of the stored instructions causes the one or more processors to further: generating a cross-correlation coefficient between the signal representing the received position and orientation information from the wearable at the first timestamp and the captured image of the user wearing the wearable at the second timestamp over a predetermined number of captured image frames; The generated cross-correlation coefficient is used as the offset value to synchronize the timestamps.
47. The apparatus of claim 46.
53. Execution of the stored instructions causes the one or more processors to further: Shifting the captured image frames forward by a predetermined number of frames or backward by a predetermined number of frames based on the determined offset.
47. The apparatus of claim 46.
54. Execution of the stored instructions causes the one or more processors to further: The offset is determined by performing a cross-correlation process using the received multiple items of position and orientation information of the wearable device and the multiple items of position information of the wearable device in the captured image.
47. The apparatus of claim 46.
55. The plurality of items of orientation information include at least two or more of pitch, yaw, roll, X, Y, and Z values received from the wearable device, and the plurality of items of position of the wearable device in the captured image are based on at least two coordinate values associated with a bounding box surrounding the wearable device in the captured image, and by executing the stored instructions, the one or more processors further: generating a first data set including the plurality of received position and orientation data items and having a single dimension; generating a second data set having a single dimension, the second data set including entries for the at least two positions of the wearable device in the captured images; determining a relative entropy between the first and second data sets; Obtaining the offset value based on the determined maximum relative entropy value.
55. The apparatus of claim 54.
56. receiving location and orientation information from a wearable device worn by a user having a first time signature; capturing an image of the user wearing the wearable device using an imaging device having a second time signature; determining an offset between the first and second temporal signatures using position and orientation information of the wearable device and the orientation and position information extracted from the captured image; Using the determined offset as a reference time, synchronize the timing between the imaging device and the wearable device. A method characterized by:
57. Furthermore, For each captured image of the user wearing the wearable device, generate a bounding box that surrounds the wearable device in the captured image; obtaining coordinates of the generated bounding box within the captured image; determining the offset using the received position and orientation information from the wearable device and the obtained coordinates of the bounding box in the captured image; 57. The method of claim 56.
58. Execution of the stored instructions causes the one or more processors to further: determining the offset by performing a cross-correlation process using the received information about the orientation of the wearable device and the information about the position of the wearable device in the captured image; 57. The method of claim 56.
59. The orientation information is first orientation information of the wearable device, and the position of the wearable device in the captured image is based on specific coordinates associated with a bounding box surrounding the wearable device in the captured image.
57. The method of claim 56.
60. The first orientation information of the wearable device is a pitch value of the wearable device, and the specific coordinate associated with the bounding box in the captured image is a Y coordinate value.
60. The method of claim 59.
61. The first orientation information of the wearable device is a yaw value of the wearable device, and the specific coordinate associated with the bounding box in the captured image is an X coordinate value.
60. The method of claim 59.
62. Furthermore, generating a cross-correlation coefficient between the signal representing the received position and orientation information from the wearable at the first timestamp and the captured image of the user wearing the wearable at the second timestamp over a predetermined number of captured image frames; The generated cross-correlation coefficient is used as the offset value to synchronize the timestamps.
57. The method of claim 56.
63. Furthermore, Shifting the captured image frames forward by a predetermined number of frames or backward by a predetermined number of frames based on the determined offset.
57. The method of claim 56.
64. Furthermore, The offset is determined by performing a cross-correlation process using the received multiple items of position and orientation information of the wearable device and the multiple items of position information of the wearable device in the captured image.
57. The method of claim 56.
65. The plurality of items of orientation information include at least two or more of a pitch value, a yaw value, a roll value, an X value, a Y value, and a Z value received from the wearable device, and the plurality of items of position of the wearable device in the captured image are based on at least two coordinate values associated with a bounding box surrounding the wearable device in the captured image.
65. The method of claim 64.
66. Furthermore, generating a first data set including the plurality of received position and orientation data items and having a single dimension; generating a second data set having a single dimension, the second data set including entries for the at least two positions of the wearable device in the captured images; determining a relative entropy between the first and second data sets; Obtaining the offset value based on the determined maximum relative entropy value.
66. The method of claim 65.
67. 67. A computer readable storage medium having stored thereon instructions which, when executed, cause an apparatus to carry out the method of any one of claims 56 to 66.
68. one or more memories for storing instructions; One or more processors that, upon execution of the instructions, receiving a series of images of a user wearing the wearable device; determining a position and orientation of the wearable device based on position and orientation information obtained from one or more sensors of the wearable device and the position and orientation of the wearable device determined from the received series of images; Estimating a pose of the user in the received sequence of images based on the determined position, location, and orientation. An apparatus characterized in that
69. Execution of the stored instructions causes the one or more processors to further: Determining the position and orientation of the wearable device based on information located on the wearable device.
69. The apparatus of claim 68.
70. Execution of the stored instructions causes the one or more processors to further: determining the location and orientation of the wearable device by estimating one or more landmarks of the user that are obscured by the wearable device; Obtaining location and orientation information based on the one or more estimated landmarks.
69. The apparatus of claim 68.
71. Execution of the stored instructions causes the one or more processors to further: generating a bounding box surrounding the wearable device in each image of the sequence of images; Inpainting the generated bounding box into each image, Estimating the one or more landmarks of a face that is occluded by the wearable device; Obtain coordinate and orientation information representing the one or more landmarks inpainted on the image.
71. The apparatus of claim 70.
72. Execution of the stored instructions causes the one or more processors to further: inpainting one or more features of the user that are obscured by the wearable device; feeding the inpainted image to a trained machine learning model trained to predict facial features; estimating the pose of the user by predicting landmarks on the user in areas not covered by the wearable device; Validate the estimated pose by feeding the inpainted image to a trained machine learning model trained to predict facial landmarks.
69. The apparatus of claim 68.
73. Execution of the stored instructions causes the one or more processors to further: generating an input image by inpainting one or more features of the user that are obscured by the wearable device onto each image in the sequence of images; feeding the generated input image to a trained machine learning model trained to generate facial landmarks; obtaining estimated facial landmarks using output from the trained machine learning model; From each image in the sequence of images, obtain projected facial landmarks on the user's face in an area not obscured by the wearable device. Estimating a user's head pose based on the estimated facial landmarks and the projected facial landmarks.
69. The apparatus of claim 68.
74. Execution of the stored instructions causes the one or more processors to further: generating a user interface displayable on the wearable device that indicates a target position of the wearable device in a virtual reality environment; instructing the user to move the wearable device to the target position by displaying one or more graphical elements within the generated user interface; Estimating the posture of the user using the coordinates of the target location and the position and orientation information obtained from one or more sensors of the wearable device.
69. The apparatus of claim 68.
75. The target location is determined based on a predetermined first target area displayed within the user interface and a second target area corresponding to a bounding box surrounding the wearable device.
75. The apparatus of claim 74.
76. Execution of the stored instructions causes the one or more processors to further: generating one or more image elements for instructing the user to move to a substantial center point of both the first target area and the second target area using a current position of the wearable device determined by the received position and orientation of the wearable device; 76. The apparatus of claim 75.
77. The target position is determined based on an orientation of the wearable device determined by the received orientation from the wearable device.
75. The apparatus of claim 74.
78. Execution of the stored instructions causes the one or more processors to further: capturing the series of images of the user wearing the wearable device; estimating the positions of one or more facial landmarks in the face obscured by the wearable device to extract actual facial landmarks in areas of the face not obscured by the wearable device; using a predetermined three-dimensional model of the face of the human wearing the wearable device, the predetermined three-dimensional model including known positions of facial landmarks in the area obscured by the wearable device and known positions of the facial landmarks in the area of the face not obscured by the wearable device; and obtaining head pose information by aligning the estimated positions of one or more facial landmarks in areas of the face not occluded by the wearable device with the known positions of facial landmarks in the areas not occluded by the wearable device from the three-dimensional model.
69. The apparatus of claim 68, wherein the position of the wearable device is determined by:
79. receiving a series of images of a user wearing the wearable device; determining a position and orientation of the wearable device based on position and orientation information obtained from one or more sensors of the wearable device and the position and orientation of the wearable device determined from the received series of images; Estimating a pose of the user in the received sequence of images based on the determined position, location, and orientation. A method characterized by:
80. Furthermore, Determining the position and orientation of the wearable device based on information located on the wearable device.
80. The method of claim 79.
81. Furthermore, determining the location and orientation of the wearable device by estimating one or more landmarks of the user that are obscured by the wearable device; Obtaining location and orientation information based on the one or more estimated landmarks.
80. The method of claim 79.
82. Furthermore, generating a bounding box surrounding the wearable device in each image of the sequence of images; Inpainting the generated bounding box into each image, Estimating the one or more landmarks of a face that is occluded by the wearable device; Obtain coordinate and orientation information representing the one or more landmarks inpainted on the image.
80. The method of claim 79.
83. Furthermore, inpainting one or more features of the user that are obscured by the wearable device; feeding the inpainted image to a trained machine learning model trained to predict facial features; estimating the pose of the user by predicting landmarks on the user in areas not covered by the wearable device; Validate the estimated pose by feeding the inpainted image to a trained machine learning model trained to predict facial landmarks.
80. The method of claim 79.
84. Furthermore, generating an input image by inpainting one or more features of the user that are obscured by the wearable device onto each image of the sequence of images; feeding the generated input image to a trained machine learning model trained to generate facial landmarks; obtaining estimated facial landmarks using output from the trained machine learning model; From each image in the sequence of images, obtain projected facial landmarks on the user's face in an area not obscured by the wearable device. Estimating a user's head pose based on the estimated facial landmarks and the projected facial landmarks.
80. The method of claim 79.
85. The target position is determined based on a predetermined first target area displayed within the user interface and a second target area corresponding to a bounding box surrounding the wearable device; and generating a user interface displayable on the wearable device that indicates a target position of the wearable device in a virtual reality environment; instructing the user to move the wearable device to the target position by displaying one or more graphical elements within the generated user interface; Estimating the posture of the user using the coordinates of the target location and the position and orientation information obtained from one or more sensors of the wearable device.
80. The method of claim 79.
86. Furthermore, generating one or more image elements for instructing the user to move to a substantial center point of both the first target area and the second target area using a current position of the wearable device determined by the received position and orientation of the wearable device; 86. The method of claim 85.
87. The target position is determined based on an orientation of the wearable device determined by the orientation received from the wearable device.
86. The method of claim 85.
88. Furthermore, capturing the series of images of the user wearing the wearable device; estimating the positions of one or more facial landmarks in the face obscured by the wearable device to extract actual facial landmarks in areas of the face not obscured by the wearable device; using a predetermined three-dimensional model of the face of the human wearing the wearable device, the predetermined three-dimensional model including known positions of facial landmarks in the area obscured by the wearable device and known positions of the facial landmarks in the area of the face not obscured by the wearable device; and obtaining head pose information by aligning the estimated positions of one or more facial landmarks in areas of the face not occluded by the wearable device with the known positions of facial landmarks in the areas not occluded by the wearable device from the three-dimensional model.
80. The method of claim 79.
89. 89. A computer readable storage medium having stored thereon instructions which, when executed, cause an apparatus to perform the method of any one of claims 79 to 88.
90. one or more memories for storing instructions; One or more processors that, upon execution of the instructions, determining one or more regions in the target image to recolor using the corresponding one or more regions in the source image; performing a color transformation on the source image from a first color space to a second color space; performing a color transformation from the first color space to the second color space on a target image; recoloring one or more regions of the target image using the color-transformed source image in the second color space when it is determined that there is a correspondence between one or more features of the one or more regions of the source image and one or more features of the one or more regions of the target image; performing a color transformation from the second color space to the first color space on the target image; An apparatus characterized in that
91. Execution of the stored instructions causes the one or more processors to further: using a reference image color-converted from the first color space to the second color space when it is determined that the one or more features in the one or more regions of the source image do not correspond to the one or more features in the one or more regions of the target image; performing a first recolor operation using the one or more regions of the source image and the reference image in the second color space; performing a second recoloring process using the one or more regions of the target image and the one or more regions of the recolored reference image; 91. The apparatus of claim 90.
92. Execution of the stored instructions causes the one or more processors to further: identifying the one or more shared regions of the source image and the target; obtaining a reference image that correlates with the source image and the target image; performing a color transformation on the source image based on at least one shared region in the source image that is in common with the reference image; replacing the region in the target image with the color-transformed source image, the color-transformed source image being transformed based on commonalities of features between the source image and the reference image; 91. The apparatus of claim 90.
93. The source image is an image previously captured by an imaging device and is used to replace at least a portion of the target image.
91. The apparatus of claim 90.
94. The target image is a live captured image of the user wearing a wearable device that obscures at least a portion of the user's face.
94. The apparatus of claim 93.
95. The source image includes segment information identifying a plurality of segments, each having a respective color value, and executing the stored instructions causes the one or more processors to further: Identifying one or more segments in a live-captured target image; recoloring the target image using segment information from the source image that corresponds to one or more segments identified in the live-captured target image; 91. The apparatus of claim 90.
96. Execution of the stored instructions causes the one or more processors to further: obtaining the source image; acquiring the target image; Obtain a reference image, converting the source image from a first color space to a second color space; converting the target image from the first color space to the second color space; converting the reference image from the first color space to the second color space; performing a color transformation on a region of the target image in the second color space based at least in part on the reference image in the second color space; performing a color transformation on a region of the source image in the second color space based at least in part on the reference image in the second color space; 91. The apparatus of claim 90.
97. determining one or more regions in the target image to recolor using the corresponding one or more regions in the source image; performing a color transformation on the source image from a first color space to a second color space; performing a color transformation from the first color space to the second color space on a target image; recoloring the one or more regions of the target image using the color-transformed source image in the second color space when it is determined that there is a correspondence between one or more features in the one or more regions of the source image and one or more features in the one or more regions of the target image; performing a color transformation from the second color space to the first color space on the target image; A method characterized by:
98. Furthermore, using a reference image color-converted from the first color space to the second color space when it is determined that the one or more features in the one or more regions of the source image do not correspond to the one or more features in the one or more regions of the target image; performing a first recolor operation using the one or more regions of the source image and the reference image in the second color space; performing a second recoloring process using the one or more regions of the target image and the one or more regions of the recolored reference image; 98. The method of claim 97.
99. Furthermore, identifying the one or more shared regions of the source image and the target; obtaining a reference image that correlates with the source image and the target image; performing a color transformation on the source image based on at least one shared region in the source image that is in common with the reference image; replacing the region in the target image with the color-transformed source image, the color-transformed source image being transformed based on commonalities of features between the source image and the reference image; 98. The method of claim 97.
100. The source image is an image previously captured by an imaging device and is used to replace at least a portion of the target image.
98. The method of claim 97.
101. The target image is a live captured image of the user wearing a wearable device that obscures at least a portion of the user's face.
101. The method of claim 100.
102. the source image includes segment information identifying a plurality of segments each having a respective color value; and Identifying one or more segments in a live-captured target image; recoloring the target image using segment information from the source image that corresponds to one or more segments identified in the live-captured target image; 98. The method of claim 97.
103. Furthermore, obtaining the source image; acquiring the target image; Obtain a reference image, converting the source image from a first color space to a second color space; converting the target image from the first color space to the second color space; converting the reference image from the first color space to the second color space; performing a color transformation on a region of the target image in the second color space based at least in part on the reference image in the second color space; performing a color transformation on a region of the source image in the second color space based at least in part on the reference image in the second color space; 98. The method of claim 97.
104. A computer readable storage medium having stored thereon instructions which, when executed, cause an apparatus to perform the method of any one of claims 97 to 103.
105. one or more memories for storing instructions; One or more processors that, upon execution of the instructions, acquiring an image of a user wearing a head-mounted display device that partially obscures a facial region of the user; estimating facial landmarks in the partially occluded region of the captured image based on a face model and a head-mounted display device attached to the face; generating an image including the inferred facial landmarks and actual facial landmarks obtained from areas of the face that are not occluded by the head-mounted display device; An apparatus characterized in that
106. Execution of the stored instructions causes the one or more processors to further: Inferring the facial landmarks in the partially occluded region based on the orientation and position of the head-mounted display.
106. The apparatus of claim 105.
107. The model includes a 3D point cloud of a face and a 3D point cloud of a head-mounted display projected onto an image plane to obtain a 2D point cloud representing a face wearing the head-mounted display.
106. The apparatus of claim 105.
108. Execution of the stored instructions causes the one or more processors to further: generating a first bounding box in the acquired image that surrounds the head mounted display using an image segmentation process; generating a second bounding box for the head mounted display based on the model by projecting three-dimensional model points from the model onto a two-dimensional image plane; Aligning the first and second bounding boxes to generate a target image that includes the estimated facial landmarks in the area obscured by the head-mounted display device.
108. The apparatus of claim 107.
109. Execution of the stored instructions causes the one or more processors to further: Get a 3D model of your face, Obtain a 3D model of the head-mounted display, Obtain a 3D model of the face wearing a head-mounted display, obtaining a face orientation in the image; generating a first bounding box for the face in the image; projecting the 3D model of the face and the model of the head mounted display onto an image plane; generating a second bounding box of the projected model of the head mounted display; generating a transformation based on the first bounding box and the second boundary; Inferring landmarks in the facial image based on the transformation.
106. The apparatus of claim 105.
110. acquiring an image of a user wearing a head-mounted display device that partially obscures a facial region of the user; estimating facial landmarks in the partially occluded region of the captured image based on a face model and a head-mounted display device attached to the face; generating an image including the inferred facial landmarks and actual facial landmarks obtained from areas of the face that are not occluded by the head-mounted display device; A method characterized by:
111. Furthermore, Inferring the facial landmarks in the partially occluded region based on the orientation and position of the head-mounted display.
111. The method of claim 110.
112. The model includes a 3D point cloud of a face and a 3D point cloud of a head-mounted display projected onto an image plane to obtain a 2D point cloud representing a face wearing the head-mounted display.
111. The method of claim 110.
113. Furthermore, generating a first bounding box in the acquired image that surrounds the head mounted display using an image segmentation process; generating a second bounding box for the head mounted display based on the model by projecting three-dimensional model points from the model onto a two-dimensional image plane; Aligning the first and second bounding boxes to generate a target image that includes the estimated facial landmarks in the area obscured by the head-mounted display device.
113. The method of claim 112.
114. Furthermore, Get a 3D model of your face, Obtain a 3D model of the head-mounted display, Obtain a 3D model of the face wearing a head-mounted display, obtaining a face orientation in the image; generating a first bounding box for the face in the image; projecting the 3D model of the face and the model of the head mounted display onto an image plane; generating a second bounding box of the projected model of the head mounted display; generating a transformation based on the first bounding box and the second boundary; Inferring landmarks in the facial image based on the transformation.
111. The method of claim 110.
115. A computer readable storage medium having stored thereon instructions that, when executed, cause an apparatus to perform the method of any one of claims 110 to 114.
116. Obtaining image frames that are encoded and communicated to an external device over a network; encoding the image frame for network communication by including transparency information in color space information of the image frame; transmitting the encoded image frame to the external device, which renders the image frame using the color space information and transparency information; A method characterized by:
117. The encoding is For each pixel in the image frame, determining whether the pixel is a foreground pixel or a background pixel; For pixels determined to be foreground pixels, setting the transparency information to a value equal to or greater than a predetermined threshold value that indicates to the external device that the determined foreground pixel should be displayed using the color space information associated with the pixel; For pixels determined to be background pixels, setting the transparency information below a predetermined threshold that indicates to the external device to display the determined background pixels as transparent.
117. The method of claim 116.
118. Furthermore, For the pixels determined to be foreground pixels, the transparency information is generated by reducing the foreground pixels so that the transparency information and color space information do not overlap.
118. The method of claim 117.
119. The encoding further comprises: Dividing the image frame into a first tensor including color space information and a second tensor including transparency information, and adding timestamp information to each of the first tensors at the time of division; Separately transmitting the first and second tensors to the external device, which uses the first and second tensors and the associated timestamps to synchronize transparency information and color space information for display.
117. The method of claim 116.
120. 120. An apparatus comprising one or more processors and one or more memories storing instructions that, when executed, cause the one or more processors to perform a method according to any one of claims 116 to 119.
121. 120. A computer readable storage medium having stored thereon instructions that, when executed, cause an apparatus to perform the method of any one of claims 116 to 119.