System and method for automatically estimating height of a user

The apparatus uses machine learning to segment and remove headsets, estimating user height for accurate full-body display in VR, addressing obstructed face visibility and privacy concerns.

WO2025147248A1PCT designated stage expired Publication Date: 2025-07-10CANON KK +3
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/010272
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-04
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

In virtual reality environments, headsets obstruct the upper part of a user's face, preventing full 3D face visibility, necessitating a method to recover the blocked region for enhanced performance.

Method used

An apparatus using machine learning models to segment images, identify and remove headsets, and estimate user height based on bounding box ratios, adjusting display height accordingly.

Benefits of technology

Enables accurate and automatic user height estimation without user input, ensuring full-body display at correct proportions in VR environments, enhancing immersion and privacy by avoiding manual height entry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024010272_10072025_PF_FP_ABST
    Figure US2024010272_10072025_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus is provided and includes one or more memories storing instructions; and one or more processors that, upon execution of the instructions, are configured to determine, from an image captured by an image capture device, a first region in the captured image; determine, from within the first region, a second region; calculate an actual height of the first region based on a size ratio of the second region relative to the first region; and use the calculated height to cause display of an object contained within the first region at the calculated height.
Need to check novelty before this filing date? Find Prior Art

Description

TITLESystem and Method for Automatically Estimating Height of a UserBACKGROUNDTechnical Field

[0001] The present disclosure relates generally to video image processing in a virtual reality environment.Description of Related Art

[0002] Given the progress that has been recently made in mixed reality, it is becoming practical to use a headset or Head Mounted Display (HMD) to join a virtual conference or a get-together meeting and be able to see each other with 3D faces in real-time. The need for these gatherings has been made more important because, in some scenarios such as a pandemic or other disease outbreaks, people cannot meet together in person.

[0003] Headsets are needed so we are able to see the 3D faces of each other using virtual and / or mixed reality. However, with the headset positioned on the face of a user, no one can really see the entire 3D face of others because the upper part of the face will be blocked by the headset. Therefore, to find a way to remove the headset and recover the blocked upper face region from the 3D faces is critical to the overall performance in virtual and / or mixed reality.SUMMARY

[0004] According the present disclosure, an apparatus is provided and includes one or more memories storing instructions; and one or more processors that, upon execution of the instructions, are configured to determine, from an image captured by an image capture device, a first region in the captured image; determine, from within the first region, a second region; calculate an actual height of the first region based on a size ratio of the second region relative to the first region; and use the calculated height to cause display of an object contained within the first region at the calculated height.

[0005] According to another embodiment, the apparatus is configured to receive the captured image from the image capture device, wherein the captured image includes a subject wearing a head mount display device.

[0006] In another embodiment, the apparatus is configured to determine the first region by providing the captured image to a trained machine learning model trained to segment an image toidentify, from within the image, a human body; generating a bounding box surrounding the identified human body; and identifying, a height in pixels, representing the generated bounding box representing the first region, wherein the identified height of the first region is used in calculating the actual height of the first region.

[0007] In a further embodiment, the apparatus is configured to determine the second region by providing the captured image to a trained machine learning model trained to recognize a head mount display device being worn by a user; generate a bounding box surrounding the identified head mount display device; and identify, a height in pixels, representing the generated bounding box representing the second region, wherein the identified height of the second region is used in calculating the actual height of the first region.

[0008] In another embodiment, the captured image includes a subject wearing a head mount display device, and the apparatus is configured to determine a position of a head of a user within the captured image based on orientation information derived from one or more sensors on the head mount display device; and use a height of the head mount display device as a height of the second region in response to determining that the position indicates that the head of a user is in the neutral position.

[0009] In further embodiments, the captured image includes a subject wearing a head mount display device, and the apparatus is configured to determine a position of a head of a user within the captured image based on orientation information derived from one or more sensors on the head mount display device; and correct a height of the head mount display device in response to determining that the position indicates that the head of a user is rotated from a neutral position; and use the corrected height as the height of the second region.

[0010] In other embodiments, the captured image includes a subject wearing a head mount display device, and the apparatus is configured to correct a height of the second region based on a position of camera that captured the image to ensure that the determined second region represents only a head mount display and excludes regions of a face of the user not occluded by the head mount display device.

[0011] In yet other embodiments, the apparatus is configured to determine a height in pixels of the first region; determine a height in pixels of the second region; calculate an actual height of a subject using the height in pixels of the first region and an actual height of an object in the second region relative to the determined height in pixels of the second region.

[0012] Other embodiments include the method executed by the apparatus described in any of the embodiments of the present disclosure. Further embodiments include non-transitory computer readable storage medium storing instructions that, when executed by one or more processors of an apparatus, configures the apparatus to perform a method according to the present disclosure. Other embodiments include a system comprising a head mount display device configured to be worn by a user; an image capture device configured to capture real time images of the user wearing the head mount display device; and an apparatus according to any of the embodiments of the present disclosure.

[0013] These and other objects, features, and advantages of the present disclosure will become apparent upon reading the following detailed description of exemplary embodiments of the present disclosure, when taken in conjunction with the appended drawings, and provided claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Fig. 1 illustrates a virtual reality capture and display system according the present disclosure.

[0015] Fig. 2 shows an embodiment of the present disclosure.

[0016] Fig. 3 shows a virtual reality environment as rendered to a user according to the present disclosure.

[0017] Fig. 4 illustrates a block diagram of an exemplary system according to the present disclosure.

[0018] Fig. 5 is a flow diagram detailing height estimation processing according to the present disclosure.

[0019] Fig. 6 illustrates an exemplary image captured and processed according to the height estimation processing according to the present disclosure.

[0020] Fig. 7 illustrates an exemplary image captured and processed according to the height estimation processing according to the present disclosure.

[0021] Figs. 8A - 8C illustrate HMD devices for use in the height estimation processing according to the present disclosure.

[0022] Fig. 9 illustrates an exemplary image captured and processed according to the height estimation processing according to the present disclosure.

[0023] Figs. 10A & 10B illustrates exemplary image capture processing and positions of an image capture device according to the present disclosure.

[0024] Throughout the figures, the same reference numerals and characters, unless otherwise stated, arc used to denote like features, elements, components or portions of the illustrated embodiments. Moreover, while the subject disclosure will now be described in detail with reference to the figures, it is done so in connection with the illustrative exemplary embodiments. It is intended that changes and modifications can be made to the described exemplary embodiments without departing from the true scope and spirit of the subject disclosure as defined by the appended claims.DESCRIPTION OF THE EMBODIMENTS

[0025] Exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It is to be noted that the following exemplary embodiment is merely one example for implementing the present disclosure and can be appropriately modified or changed depending on individual constructions and various conditions of apparatuses to which the present disclosure is applied. Thus, the present disclosure is in no way limited to the following exemplary embodiment and, according to the Figures and embodiments described below, embodiments described can be applied / performed in situations other than the situations described below as examples. Further, where more than one embodiment is described, each embodiment can be combined with one another unless explicitly stated otherwise. This includes the ability to substitute various steps and functionality between embodiments as one skilled in the art would see fit.

[0026] Environment Overview

[0027] FIG. 1 shows a virtual reality capture and display system 100. The virtual reality capture system comprises a capture device 110. The capture device may be a camera with sensor and optics designed to capture 2D RGB images or video, for example. In one embodiment, the image capture device 110 is a smartphone that has front and rear facing cameras and which can display images captured thereby on a display screen thereof. Some embodiments use specialized optics that capture multiple images from disparate view-points such as a binocular view or a light-field camera. Some embodiments include one or more such cameras. In some embodiments the capture device may include a range sensor that effectively captures RGBD (Red, Green, Blue, Depth) images either directly or via the software / firmware fusion of multiple sensors such as an RGB sensor and a range sensor (e.g., a lidar system, or a point-cloud based depth sensor). The capture device may be connected via a network 160 to a local or remote (e.g., cloud based) system 150and 140 respectively, hereafter referred to as the server 1 0. The capture device 110 is configured to communicate via the network connect 160 to the server 140 such that the capture device transmits a sequence of images (e.g., a video stream) to the server 140 for further processing.

[0028] Also, in FIG 1, a user 120 of the system is shown. In the example embodiment the user 120 is wearing a Virtual Reality (VR) device 130 configured to transmit stereo video to the left and right eye of the user 120. As an example, the VR device may be a headset worn by the user. As used herein, the VR device and head mounted display (HMD) device may be used interchangeably. Other examples can include a stereoscopic display panel or any display device that would enable practice of the embodiments described in the present disclosure. The VR device is configured to receive incoming data from the server 140 via a second network 170. In some embodiments the network 170 may be the same physical network as network 160 although the data transmitted from the capture device 110 to the server 140 may be different than the data transmitted between the server 140 and the VR device 130. Some embodiments of the system do not include a VR device 130 as will be explained later. The system may also include a microphone 180 and a speaker / headphone device 190. In some embodiments the microphone and speaker device are part of the VR device 130.

[0029] FIG 2 shows an embodiment of the system 200 with two users 220 and 270 in two respective user environments 205 and 255. In this example embodiment, each user 220 and 270 are equipped with a respective capture devices 210 and 260, respective VR devices 230 and 280, and are connected via respective networks 240 and 270 to a server 250. In some instances, only one user has a capture device 210 or 260, and the opposite user may only have a VR device. In this case, one user environment may be considered as a transmitter and the other user environment may be considered the receiver in terms of video capture. However, in embodiments with distinct transmitter and receiver roles, audio content may be transmitted and received by only the transmitter and receiver or by both, or even in reversed roles.

[0030] FIG 3 shows a virtual reality environment 300 as rendered to a user. The environment includes a computer graphic model 320 of the virtual world with a computer graphic projection of a captured user 310. For example, the user 220 of FIG 2, may see via the respective VR device 230, the virtual world 320 and a rendition 310 of the second user 270 of FIG 2. In this example, the capture device 260 would capture images of user 270, process them on the server 250 and render them into the virtual reality environment 300.

[0031] In the example of FIG 3, the user rendition 310 of user 270 of FIG 2, shows the user without the respective VR device 280. The present disclosure sets forth a plurality of algorithms that, when executed, cause the display of user 270 to appear without the VR device 280 and as if they were captured naturally without wearing the VR device. Some embodiments show the user with the VR device 280. In other embodiments the user 270 does not use a wearable VR device 280. Furthermore, in some embodiments the captured images of user 270 capture a wearable VR device, but the processing of the user images remove the wearable VR device and replace it with the likeness of the users face.

[0032] Additionally, the addition of the user rendition 310 into the virtual reality environment 300 along with VR content 320 may include a lighting adjustment step to adjust the lighting of the captured and rendered user 310 to better match the VR content 320.

[0033] In the present disclosure, the first user 220 of FIG 2, is shown via the respective VR device 230, the VR rendition 300 of FIG 3. Thus, the first user 220, sees user 270 and the virtual environment content 320. Likewise, in some embodiments, the second user 270 of FIG 2, will see in the same VR environment 320, but from a different view-point, e.g. the view-point of the virtual character rendition of 310 for example.

[0034] In order to achieve the immersive calling as described above, it is important to render each user within the VR environment as if they were not wearing the headset in which they are experiencing the VR content. The following describes the real-time processing performed that obtains images of a respective user in the real world while wearing a virtual reality device 130 also referred to hereinafter as the head mount display (HMD) device.

[0035] Hardware

[0036] FIG. 4 illustrates an example embodiment of a system for virtual reality immersive calling system. The system includes two user environment systems 400 and 410, which are specially- configured computing devices; two respective virtual reality devices 404 and 414, and two respective image capture devices 405 and 415. In this embodiment, the two user environment systems 400 and 410 communicate via one or more networks 420, which may include a wired network, a wireless network, a LAN, a WAN, a MAN, and a PAN. Also, in some embodiments the devices communicate via other wired or wireless channels.

[0037] The two user environment systems 400 and 410 include one or more respective processors 401 and 411, one or more respective I / O components 402 and 412, and respective storage 403 and413. Also, the hardware components of the two user environment systems 400 and 410 communicate via one or more buses or other electrical connections. Examples of buses include a universal serial bus (USB), an IEEE 1394 bus, a PCI bus, an Accelerated Graphics Port (AGP) bus, a Serial AT Attachment (SATA) bus, and a Small Computer System Interface (SCSI) bus.

[0038] The one or more processors 401 and 411 include one or more central processing units (CPUs), which may include one or more microprocessors (e.g., a single core microprocessor, a multi-core microprocessor); one or more graphics processing units (GPUs); one or more tensor processing units (TPUs); one or more application-specific integrated circuits (ASICs); one or more field-programmable-gate arrays (FPGAs); one or more digital signal processors (DSPs); or other electronic circuitry (e.g., other integrated circuits). The I / O components 402 and 412 include communication components (e.g., a graphics card, a network-interface controller) that communicate with the respective virtual reality devices 404 and 414, the respective capture devices 405 and 415, the network 420, and other input or output devices (not illustrated), which may include a keyboard, a mouse, a printing device, a touch screen, a light pen, an optical- storage device, a scanner, a microphone, a drive, and a game controller (e.g., a joystick, a gamepad).

[0039] The storages 403 and 413 include one or more computer-readable storage media. As used herein, a computer-readable storage medium includes an article of manufacture, for example a magnetic disk (e.g., a floppy disk, a hard disk), an optical disc (e.g., a CD, a DVD, a Blu-ray), a magneto-optical disk, magnetic tape, and semiconductor memory (e.g., a non-volatile memory card, flash memory, a solid-state drive, SRAM, DRAM, EPROM, EEPROM). The storages 403 and 413, which may include both ROM and RAM, can store computer-readable data or computerexecutable instructions.

[0040] The two user environment systems 400 and 410 also include respective communication modules 403 A and 413 A, respective capture modules 403B and 413B, respective rendering module 403C and 413C, respective positioning module 403D and 413D, and respective user rendition modules 403E and 413E. A module includes logic, computer-readable data, or computer-executable instructions. In the embodiment shown in FIG. 4, the modules are implemented in software (e.g., Assembly, C, C++, C#, Java, BASIC, Perl, Visual Basic, Python, Swift). However, in some embodiments, the modules are implemented in hardware (e.g., customized circuitry) or, alternatively, a combination of software and hardware. When the modules are implemented, at least in part, in software, then the software can be stored in the storage403 and 413. Also, in some embodiments, the two user environment systems 400 and 410 includes additional or fewer modules, the modules arc combined into fewer modules, or the modules arc divided into more modules. One environment system may be similar to the other or may be different in terms of the inclusion or organization of the modules.

[0041] The respective capture modules 403B and 413B include operations programed to carry out image capture as shown in 110 of FIG 1, 210 and 260 of FIG 2. The respective rendering module 403C and 413C contain operations programed to carry out the functionality associated with rendering images that are captured to one or more users participating in the VR environment. The respective positioning module 403D and 413D contain operations programmed to carry out the process including identifying and determining position of each respective user in the VR environment. The respective user rendition modules 403E and 413E contains operations programmed to carry out user rendering as illustrated in the following figures described hereinbelow. The prior-training module 403F contains operations programmed to estimate the nature and type of images that were captured prior to participating in the VR environment that are used for the head mount display removal processing. In some embodiments the some modules are stored and executed on an intermediate system such as a cloud server. In other embodiments, the capture devices 405 and 415, respectively, include one or more modules stored in memory thereof that, when executed perform certain of the operations described hereinbelow.

[0042] Height Estimation Processing For Correcting Display in a VR Environment

[0043] As noted above, in view of the progress made in augmented virtual reality, it is becoming more common to enter into an immersive communication session in a VR environment where each user is in their own location wearing a headset or Head Mounted Display (HMD) to join together in virtual reality. In order for the users participating in these immersive communication sessions, it is beneficial if each of the participating users appear at the correct height scale within the VR environment. As such, the estimation of user height is critical in many camera-based virtual reality applications. For instance, when placing a 2D image of a person onto an image plane within a 3D virtual reality environment, it is important to scale or resize the 2D image so that the person’s proportions match our perception of them in the real world. According to one embodiment, this includes making use of three dimensional point clouds or meshes of a human figure which provide a set of 3D landmarks of a user in space and making use of this 3D data to ensure proper height calculation. By knowing the person’s height allows us to adjust these 3D point clouds to ensureour perception of them is consistent with their actual scale or dimensions in the real world. However, it is not always possible to obtain the known height of a user because, to obtain user height, may require requesting that the user input their height directly. However, this is not an ideal solution because there is no guarantee that the height provided by the user is accurate and incorrect (or erroneous) input thereof negatively impacts the real-world perception experience. Further, in some instances, there may be privacy concerns by directly inputting a height of a user as it represent personally identifying information. Finally, requiring users to input their height could introduce an inconvenient step in the process whereby so many steps are required before actually engaging in the VR communication session. The automatic height determination algorithm according to the present disclosure remedies these issues because it will estimate user height automatically without needing to request information from the user and do with a high degree of precision.

[0044] In the exemplary immersive VR communication environments such as those described with respect to Figs. 1 - 4, each of the users are wearing a Head Mounted Display (HMD) device (e.g. Oculus 2), in order to experience the virtual reality or augmented virtual reality generated by a virtual reality communication application executing on a server which is in communication with the HMD. In exemplary embodiment, the VR application receives images captured from an information processing device such as a mobile phone which has a camera and is positioned in front of the user to capture images of the user, in real-time, wearing the HMD. These images are provided to the VR application which processes them to remove the HMD from the image and replace at least a portion the image where the HMD is present with one or more precaptured images of the user that were captured without the HMD in position. As discussed above, for a good VR experience it is crucial to be able to display the full size image of the user’s entire body at the correct height perspective in the VR environment. A flow diagram as described in Fig. 5 describes an algorithm for automatically determining a height of a user to assign to and cause display of the use in a VR environment. The algorithmic steps described in Fig. 5 are a set of instructions that are stored in memory device and executed by one or more processors. In one embodiment, the automatic height adjustment algorithm described herein is embodied as a stand-alone application. In other embodiments, the algorithm described herein is embodied as a module executing within the VR communication application whereby the module receives input images and processes them to provide an output representative of a user’s height to be displayed in the VR environment.However, for purposes of example and ease of understanding the processing being performed, the present disclosure will refer to this processing as being performed by the automatic height determination algorithm according to the present disclosure.

[0045] At step 502, images of the user captured by the information processing device (e.g. mobile phone) are received by the height determination algorithm. An image capturing application executing on the information processing device is configured to capture the images of the user wearing the HMD and provide those captured images, in real-time, the VR communication application so that they can be further processed. In step 504, each captured image is provided to a first deep learning-based neural network we trained to perform HMD detection. The trained model has been trained using images different users of different sizes that were wearing the HMD with the HMD identified as an HMD such that, upon images being input into the trained model, the output thereof is a bounding box surrounding the HMD being worn by the user. In step 506, a full body bounding box that is determined based on the captured image being input into and processed by a second trained machine learning model trained to perform background segmentation to identify a human body within a captured image is determined and positioned around the entire body of the user. In one embodiment, the first and second models are separate trained machine learning models. In other embodiments, the first and second trained machine learning models are embodied as a single model trained to output bounding boxes representing an HMD and the user. The results of the operations of 502 - 506 are shown in Fig. 6 whereby an image of a user 600 wearing an HMD 601 is captured. After being processed the image frame containing the user 601 wearing the HMD 601 is processed to have an HMD bounding box 602 (e.g. a first bounding box) that substantially surrounds the HMD 601 is generated. Further the image fame is processed such that u user bounding box 604 (e.g. second bounding box) that substantially surrounds the entire user 600 is generated. As a result, and as shown in Fig. 6, steps 500-506 estimate the bounding boxes of both the HMD region and the entire human body for a 2D input HMD image.

[0046] In step 508, the height of the user 600 in the captured image is automatically determined using information from the generated bounding boxes 602 and 604 in Fig. 7 and 8. Fig. 7 illustrates only a portion of the bounding box in order to further explain how the respective bounding boxes are used to determine the user’ s height but is should be understood that the entire bounding box asshown in Fig. 6 is present as part of the processing. The height determination processing will be described with respect to Figs. 7.

[0047] The bounding box information 602 and 604 in the 2D image, from both HMD and the user, enables the height determination algorithm to estimate the ratio of the height of the bounding box and the user in the real world using pixel-based heights determined from the image shown in Fig. 7. Because the actual height of the HMD bounding box in the real world is known, the algorithm calculates the user’s actual height in the real world using this ratio. Based the determination of bounding boxes 602 and 604, heights, in pixels, of each of the bounding boxes 602,604 are determined. The determined height, in pixels of the HMD bounding box 602 is represented in Fig. 7 as HMD^y and the determined height, in pixels of the user bounding box 604 is represented in Fig. 7 as User^y. Further, an actual height, in the real-world, of the HMD 601 is obtained and is represented in Fig. 7 as HMD_h. From these determined data values, the height determination algorithm determines the actual height of the user represented in Fig. 7 as User_h. For example, a HMD with a height of 10 centimeters, we can calculate the user’s height (User_h) using the following equations:User_h / User_y = HMD_h / HMD_y (1)User_h — User_y * HMD_h / HMD_y (2)

[0048] In one embodiment, actual height of the HMD information is known in advance and stored in a repository. In another embodiment, the repository includes a plurality of information items representing different HMD devices and, in response to identifying the HMD device, retrieves the actual height information from the set of candidate height information values. In another embodiment, it is determined based on object recognition processing what type of HMD 601 is being worn by the user 600 and height information associated with the determined HMD device is retrieved from the repository. In another embodiment, HMD dimension information is provided as metadata with the image captured by the information processing device and the actual height information of the particular HMD 601 is obtained therefrom. In a further embodiment, an HMD device type identifier is provided at some point during the processing (e.g. during initial setup with the VR communication application or as metadata with the captured images) and that information is used to obtain, from the repository, actual height information HMD h.

[0049] The algorithm also corrects for the position of the HMD 601 on the user’s face. More specifically, in a case where a user’s head is in a neutral position (e.g. looking directly at the imagecapture device), the calculation as described above may be used to obtained the value of User_h. However, often times this positioning is not realistic and / or, during the interaction using the VR communication application, the user 601 is continually moving their head in various positions and orientations. To remedy this, the algorithm uses orientation information generated by one or more sensors on the HMD device 602. The one or more sensors may include an Inertial Measurement Unit (IMU) which tracks the orientation of the HMD in space over time. This orientation information is represented using Pitch, Yaw and Roll as shown in Fig.8A. If a user wears the HMD upright and the head of the user is in the neutral position such that the face is looking through the imaging plane of the information processing apparatus capturing the images, a HMD_h, the actual height of HMD in real space is set equal to the size of HMDj as illustrated in Fig. 8B. However, if the user rotates their head while wearing the HMD 601, HMDJi cannot be used as HMD . Instead, a rotational HMD value, HMD_h_rot, is set and used as HMD^y. To obtain HMD_h_rot, three dimensional point cloud data representing a 3D points of the head of the user wearing the HMD device. The 3D point cloud data is stored in a repository and accessible by the algorithm. The default 3D point cloud information is obtained prior to any VR communication and, for example, is generated by the manufacturer of HMD device. . During this process, images of the user wearing the HMD are captured and processed to generate an updated3D point cloud of the user’s head and face having the HMD positioned thereon. From the obtained 3D point cloud data, the actual y-values correlated with the HMD device 601 are determined and assigned as y_max and y_min where y_max represents a top boundary 802 of the HMD device 601 shown in Fig. 8B and y_min is a bottom boundary 804 of the HMD device shown in Fig. 8B. Using the obtained 3D point cloud information, first rotate the original 3D point cloud in the HMD CAD model, and then determine the maximum and minimum y values in rotated 3D point cloud. As used herein, the rotation being described is a vertical rotation whereby the maximum rotation is when the user’ s head is rotated looking upwards and the minimum rotation is when the user’s head is rotated looking downwards. The maximum rotation is denoted as y_rot_max and the minimum rotation is denoted as y_rol_min. By using the original maximum and minimum y coordinates in the 3D point cloud y_max and y_min, a rotation correction value is HMD_h_rot is calculated based on the Equation 3:HMD_h_rot = HMD_h * (^y_rot_max — y_rot_min) / y_max — y_min) (3)By calculating a rotation correction value, the size of the HMD can he compensated for and a correct height value can be used in Equations 1 and 2 to obtain the proper ratio to calculate the actual height of the user.

[0050] In another embodiment, a rotation correction value can be calculated by locating two predefined landmarks in the 3D point cloud of the HMD CAD model. From these points a matrix with translation and scale factor is estimated to align the bounding box of HMD point cloud with the HMD bounding box in 2D image. Next, this matrix to project these two predefined landmarks onto the 2D images. The actual length between two landmarks in the real world and their corresponding pixel distance are used to determine the actual height in the HMD bounding box. The actual height of HMD bounding box and its pixel length should have a proportional relationship identical to that between the actual length and pixel distance of these two landmarks. This relationship forms an equation similar to Equation 1. By determining three of these components, the fourth can be computed. For rotation correction, we rotate the selected two landmarks in opposite directions for Pitch, Yaw and Roll, and then re-project these rotated landmarks back onto the 2D image. These re-projected landmarks will help us calculate the actual height of the HMD bounding box with rotated angles.

[0051] According to another embodiment, in addition to the height of the HMD bounding box (602 in Fig. 6) and the known actual height of the HMD 601 in real space, a width of the bounding box 902 in Fig. 9 along with a known actual width of the HMD 601 in real space can be used to further determine the actual height of the user User_h. Referring back to Equation 1 wherein the “HMD_h / HMD_y” is used to determine the real-world representative dimension for each pixel. However, to estimate its value, there is no need to be limited by the y coordinates. Using the x- coordinates and the associated information from the x-axis, or the width of a HMD device 601, Equations 1 and 2 are rewritten as Equations 4 and 5 as follows:Userji / Useryy = HMD_w / HMD_x (4)Userji = User_y * HMD_h / HMD_y (5)

[0052] In exemplary operation, the user height estimation from both the height and width of a HMD device is combined based on a weighted average using a weighted coefficient. Ideally, calculations in both the y and x directions should yield identical results for the user’s actual height. However, in practice, there may be differences between these two methods due to numerical discrepancies or other estimation errors. Consequently, averaging these two estimates provides amore accurate determination of the user height. The weighted coefficient enables the application of different weighting factors to these estimates. Combining HMD height and width improves the estimation accuracy of the user height in real space. A weighted average of these two estimates is used as the final human height estimate for this single image.

[0053] Turning back to Fig. 5, after the height estimation processing as described above in 508 is performed, camera position correction processing 510 is performed whereby the algorithm advantageously considers a position of the camera relative to user that is being captured by the camera in order to ensure that the height estimation performed in 508 is correct and should be used. Camera position correction processing 510 will be described with respect to Figs. 10A and 10B.

[0054] The bounding box of the HMD (602 in Fig. 6) is used as a reference to determine the user’s height. However, the meaning of the HMD bounding box may vary depending on the camera position as shown in Figs. 10A and 10B. For instance, in Fig. 10A, if a user places the camera at the same height as the HMD device 601, the bounding box of HMD 602 in the 2D image may represent only the front face of the HMD device. Conversely, if a user positions the camera at a lower height, as shown in Fig. 10B, the bounding box of the HMD 602 may include both the front face (having the portion of the face occluded by the HMD) and the bottom face portion as well as the HMD device. Therefore, HMD_h is adjusted based on the camera position. This adjustment could be implemented using some experimental data collected in advance by placing the camera at different predefined height positions. A factor fcam_pos isintroduced into the equation, as presented in Equation 6, which accounts for the impact of the camera’s position relative to the user height. This factor fcam_pos should equal 1. However, due to the HMD not maintaining the same distance from the camera compared to the distance from the user body to the camera, fcam_posmay fluctuate around 1 based on the camera’s positioning. In practice, we adjust the camera to various positions and calculate the corresponding fcam_pos values using Equation 6, by determining all the measurement for Userji, User_y, HMDJi, and HMD_y.User_h / User_y = fcam-pOs * HMD_h / HMD_y (6)

[0055] After correcting for camera position relative to the user in 512, a value of User_h is associated with the particular image frame and stored in memory. In one embodiment, it may be stored in a buffer configured to store a predetermined number of values of User_h for a predetermined number of successive image frames. From the stored plurality of User_h values,outliers are identified based on their mean and standard deviation, for example, in 514. Once the outliers arc removed, further statistical processing is performed in 516 on the remaining User_h values over a predetermined number of image frames to obtained the final height value representing a true value of the human figure in the image being captured by the information processing device. Typically, users need to manually input their heights to scale their 2D images or 3D point cloud appropriately, ensuring they align proportionally with the surrounding objects in virtual reality. With our approach, manual entry of height by the user is unnecessary. We can accurately estimate their height with high confidence. This eliminates the possibility of random or incorrect height input and ultimately enhancing the user experience in various virtual reality- related applications.

[0056] The processing 512-514 further aids in providing an accurate estimate of the user’s height using the HMD image captured by the camera because, in many situations, the images of the user captured may not be ideal thereby impacting the above height determination processing. For example, the user might raise their hands or bend their body causing the resulting bounding box of the human figure to be shifted and not reflect their actual height. Therefore, estimating the human height from multiple images of the user, and then use the mean or median value of these estimates along with applying statistical processing techniques, such as Ransac, to better deal with the outliers in refining our estimate of human height improves the final output for the user’s height calculation.

[0057] In one embodiment, this height estimation and determination algorithm is executed over a first series of image frames captured during VR communication and the final value in 516 is determined and used for a duration of the VR communication session. In other embodiments, the determination is stored in association with a user profile information so that, during subsequent VR communication sessions, the VR communication application can obtain user height information from the user profile thereby reducing the processing needing to be performed to execute the VR communication session between users. In other embodiments, as a quality check, this algorithm can be performed at predetermined intervals during a VR communication session to ensure that a current height of the user and its associated display within the VR environment is still accurate thereby improving the quality of the VR communication session between users.

[0058] At least some of the above-described devices, systems, and methods can be implemented, at least in part, by providing one or more computer-readable media that contain computer-executable instructions for realizing the above-described operations to one or more computing devices that arc configured to read and execute the computer-executable instructions. The systems or devices perform the operations of the above-described embodiments when executing the computer-executable instructions. Also, an operating system on the one or more systems or devices may implement at least some of the operations of the above-described embodiments.

[0059] Furthermore, some embodiments use one or more functional units to implement the abovedescribed devices, systems, and methods. The functional units may be implemented in only hardware (e.g., customized circuitry) or in a combination of software and hardware (e.g., a microprocessor that executes software).

[0060] Additionally, some embodiments of the devices, systems, and methods combine features from two or more of the embodiments that are described herein. Also, as used herein, the conjunction “or” generally refers to an inclusive “or,” though “or” may refer to an exclusive “or” if expressly indicated or if the context indicates that the “or” must be an exclusive “or.”

[0061] While the present disclosure has been described with reference to exemplary embodiments, it is to be understood that the invention is not limited to the disclosed exemplary embodiments.

Claims

CLAIMSWc Claim,1. An apparatus comprising: one or more memories storing instructions; and one or more processors that, upon execution of the instructions, are configured to: determine, from an image captured by an image capture device, a first region in the captured image; determine, from within the first region, a second region; calculate an actual height of the first region based on a size ratio of the second region relative to the first region; and use the calculated height to cause display of an object contained within the first region at the calculated height.

2. The apparatus according to claim 1, wherein execution of the stored instructions further configures the apparatus to: receive the captured image from the image capture device, wherein the captured image includes a subject wearing a head mount display device.

3. The apparatus according to claim 1, wherein execution of the stored instructions further configures the apparatus to determine the first region by providing the captured image to a trained machine learning model trained to segment an image to identify, from within the image, a human body; generating a bounding box surrounding the identified human body; and identifying, a height in pixels, representing the generated bounding box representing the first region, wherein the identified height of the first region is used in calculating the actual height of the first region.

4. The apparatus according to claim 1, wherein execution of the stored instructions further configures the apparatus to determine the second region byproviding the captured image to a trained machine learning model trained to recognize a head mount display device being worn by a user; generate a bounding box surrounding the identified head mount display device; and identify, a height in pixels, representing the generated bounding box representing the second region, wherein the identified height of the second region is used in calculating the actual height of the first region.

5. The apparatus according to claim 1, wherein the captured image includes a subject wearing a head mount display device, and the execution of the stored instructions further configures the apparatus to determine a position of a head of a user within the captured image based on orientation information derived from one or more sensors on the head mount display device; and use a height of the head mount display device as a height of the second region in response to determining that the position indicates that the head of a user is in the neutral position.

6. The apparatus according to claim, wherein the captured image includes a subject wearing a head mount display device, and the execution of the stored instructions further configures the apparatus to determine a position of a head of a user within the captured image based on orientation information derived from one or more sensors on the head mount display device; and correct a height of the head mount display device in response to determining that the position indicates that the head of a user is rotated from a neutral position; and use the corrected height as the height of the second region.

7. The apparatus according to claim 1, wherein the captured image includes a subject wearing a head mount display device, and the execution of the stored instructions further configures the apparatus to correct a height of the second region based on a position of camera that captured the image to ensure that the determined second region represents only a head mount display and excludes regions of a face of the user not occluded by the head mount display device.

8. The apparatus according to claim 1, wherein execution of the stored instructions further configures the apparatus to determine a height in pixels of the first region; determine a height in pixels of the second region; calculate an actual height of a subject using the height in pixels of the first region and an actual height of an object in the second region relative to the determined height in pixels of the second region.

9. A method of determining a height of an object in an image captured by an image capture device, the method comprising: determining, from the image captured by an image capture device, a first region in the captured image; determining, from within the first region, a second region; calculating an actual height of the first region based on a size ratio of the second region relative to the first region; and using the calculated height to cause display of an object contained within the first region at the calculated height.

10. The method according to claim 9, further comprising receiving the captured image from the image capture device, wherein the captured image includes a subject wearing a head mount display device.

11. The method according to claim 9, further comprising determining the first region by providing the captured image to a trained machine learning model trained to segment an image to identify, from within the image, a human body; generating a bounding box surrounding the identified human body; and identifying, a height in pixels, representing the generated bounding box representing the first region, wherein the identified height of the first region is used in calculating the actual height of the first region.

12. The method according to claim 9, further comprising determining the second region by providing the captured image to a trained machine learning model trained to recognize a head mount display device being worn by a user; generating a bounding box surrounding the identified head mount display device; and identifying, a height in pixels, representing the generated bounding box representing the second region, wherein the identified height of the second region is used in calculating the actual height of the first region.

13. The method according to claim 9, wherein the captured image includes a subject wealing a head mount display device, and further comprising determining a position of a head of a user within the captured image based on orientation information derived from one or more sensors on the head mount display device; and using a height of the head mount display device as a height of the second region in response to determining that the position indicates that the head of a user is in the neutral position.

14. The method according to claim 9, wherein the captured image includes a subject wealing a head mount display device, and further comprising determining a position of a head of a user within the captured image based on orientation information derived from one or more sensors on the head mount display device; and correcting a height of the head mount display device in response to determining that the position indicates that the head of a user is rotated from a neutral position; and using the corrected height as the height of the second region.

15. The method according to claim 9, wherein the captured image includes a subject wearing a head mount display device, and further comprising correcting a height of the second region based on a position of camera that captured the image to ensure that the determined second region represents only a head mount display and excludes regions of a face of the user not occluded by the head mount display device.

16. The method according to claim 9, further comprising determining a height in pixels of the first region; determining a height in pixels of the second region; calculating an actual height of a subject using the height in pixels of the first region and an actual height of an object in the second region relative to the determined height in pixels of the second region.

17. A non-transitory computer readable storage medium storing instructions that, when executed by one or more processors of an apparatus, configures the apparatus to perform a method according to any of claims 9 - 16.

18. A system comprising: a head mount display device configured to be worn by a user; an image capture device configured to capture real time images of the user wearing the head mount display device; and an apparatus according to any of claims 1 - 8.

Citation Information

Patent Citations

  • Image correction method and device

    US20190370943A1

  • XR device and method for controlling the same

    US20190378280A1

  • Object size estimation using camera map and / or radar information

    US20210209785A1

  • Augmentation of unmanned-vehicle line-of-sight

    US20210240986A1

  • Occlusion detection and object coordinate correction for estimating the position of an object

    US20230206643A1