Image-based reconstruction of 3D landmarks for use in generating personalized head transfer functions.
The method uses smartphone cameras to reconstruct anthropometric features through image-based 3D reconstruction, addressing complexity and hardware limitations for personalized HRTF generation, enabling high-quality binaural audio on devices with limited capabilities.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-07
- Publication Date
- 2026-03-17
AI Technical Summary
Existing methods for obtaining user-specific anthropometric features for personalized Head-Related Transfer Function (pHRTF) generation are complex and require dedicated hardware, making them difficult to implement on devices with limited sensing and processing capabilities like smartphones.
A method for 3D reconstructing anthropometric features from images using a smartphone camera, involving image capture, landmark detection, head pose estimation, and distance measurement to generate personalized HRTFs, utilizing techniques like deep neural networks and multi-angulation to achieve accurate 3D coordinates.
Enables efficient and accurate reconstruction of anthropometric features for personalized HRTF generation on devices with limited capabilities, facilitating high-quality binaural audio rendering without requiring specialized hardware.
Smart Images

Figure 2026509208000001_ABST
Abstract
Description
Technical Field
[0001] Cross - References to Related Applications This application claims the benefit of priority from PCT Application No. PCT / CN2023 / 080861, filed on March 10, 2023, PCT Application No. PCT / CN2023 / 082495, filed on March 20, 2023, U.S. Provisional Application No. 63 / 499,759, filed on May 3, 2023, and U.S. Provisional Application No. 63 / 507,498, filed on June 12, 2023, each of which is incorporated herein by reference in its entirety.
[0002] Technical Field The present invention relates to image - based 3D reconstruction of anthropometric features, particularly features of the head, face, and ears. Such anthropometric features can be used for the generation of a personalized head - related transfer function (pHRTF).
Background Art
[0003] The Head - Related Transfer Function (HRTF) is a set of functions that describe how a human ear receives sound from sound sources in various directions of arrival. The functions typically describe a linear filtering process that reflects the acoustic effects of the ear, head, and torso on incoming sound waves. Rendering audio using HRTF is called "binaural rendering", and the resulting audio is called "binaural audio".
[0004] HRTFs can be defined in many ways, such as as time-domain impulse responses or as frequency-domain responses. HRTFs are typically grouped into pairs to provide a response for each ear. HRTF filter pairs can be used to provide listeners with an experience that mimics the sound (in each ear) that would occur if an audio signal were presented from a particular direction of arrival. Different HRTF filter pairs produce different illusions of source direction, timbre, and externalization.
[0005] A user's anatomical structure, particularly features of the head and ears, influences how they perceive sound, including timbre, localization, and externalization. Therefore, tailoring the HRTF to a specific individual is generally desirable. Such an HRTF is referred to herein as a personalized HRTF (pHRTF).
[0006] Personalized HRTFs can be obtained through experimental measurement procedures or modeled using personalized information about the user. Dimensions and orientation of ear and head features are examples of anthropometric information that can be used to generate personalized pRHTFs. Patent Document 1 discloses an example of such a technique. [Patent Document 1] U.S. Patent Application Publication No. 2021 / 0211825 [Overview of the project] [Problems that the invention aims to solve]
[0007] From the above, accurately and reliably obtaining user-specific anthropometric features (e.g., landmarks and distances between landmarks) is crucial for determining and maintaining the quality of the pHRTF. Existing methods for obtaining anthropometric features for pHRTF generation are complex and often require dedicated hardware, making their implementation difficult on devices with limited sensing and / or processing capabilities, such as smartphones.
[0008] The objective of the present invention is to facilitate the reconstruction of anthropometric features from user images. Such anthropometric features can then be used for pHRTF generation. [Means for solving the problem]
[0009] A first aspect of the present invention relates to a method for 3D reconstructing anthropometric features from images of a subject, comprising the steps of: obtaining a set of images of a subject from a feed of images captured by an image capture device from multiple viewpoints; detecting the 2D positions of multiple anthropometric landmarks for each image in the set; and estimating the head pose based on those 2D positions. The method further comprises the steps of: generating relative 3D coordinates corresponding to at least some of the anthropometric features captured in the set of images based on the detected 2D positions and associated head poses for at least some of the images in the set of images; and generating scaled 3D coordinates from the relative 3D coordinates based on characteristics associated with the image capture device and the estimated distance between the subject and the image capture device.
[0010] This approach provides an efficient method for obtaining scaled 3D anthropometric features using a series of images acquired by simple means, such as using a mobile camera device like a smartphone. Anthropometric features are obtained by detecting ear landmarks and / or facial landmarks.
[0011] The set of images may be selected based on a desired uniform angular sampling frequency, meaning that the images in the set were obtained from a set of observation angles relative to the subject (e.g., the subject's face), and that these observation angles are uniformly distributed across the relevant observation angle range. As an example, the relevant yaw angle range may be 20 to 55 degrees on each side.
[0012] As a result, if the subject is moving relatively fast, the time between consecutive images in the set should be relatively short, and if the subject is moving relatively slowly, the time between consecutive images in the set should be relatively long.
[0013] One way to achieve a uniform angular sampling frequency is to monitor the subject's angular velocity (for example, relative to the image acquisition device) and adjust the sampling from the image feed. The motion metric representing the angular velocity may be based on data from sensors worn by the subject (e.g., IMU, accelerometer, gyroscope), but alternatively, it may be provided by software within the image acquisition device.
[0014] Another method for achieving a uniform angular sampling frequency is simply to use the head pose acquired for each image. In this case, the head pose may be acquired for a larger set of images, and then a smaller set of images is selected for further processing.
[0015] In some embodiments, a subset of anthropometric landmarks is detected only for a subset of images where the head posture is within a given head posture range (i.e., for observation angles relative to the head within a certain range). As a specific example, ear landmarks may be detected only for a subset of images where the head posture is within a given azimuth angle (yaw angle). For azimuth angles (yaw angles) substantially directly in front of the subject, the ear may not be clearly visible, and landmark detection may be difficult or unreliable.
[0016] In some embodiments, head pose estimation involves estimating rotational transformations (yaw, pitch, roll) that convert the positions of 2D anthropometric landmarks into a general 3D face mesh at the center of the image. Matching may be performed in a least-squares sense.
[0017] Alternatively, head pose estimation can be achieved by making a request (call) to software within an image acquisition device and obtaining a response that includes head pose data. Such a response may also include 2D positions for at least a subset of anthropometric landmarks.
[0018] The estimated distance between the subject and the image acquisition device can be detected using any suitable distance detection technique, including structured light imaging, time-of-flight imaging, or any other hardware-based distance / depth sensor. Alternatively, it is typically possible to obtain a satisfactory distance estimate using statistical information about the subject's anthropometric features, such as iris size or interpupillary distance. The process for obtaining the estimated distance can be performed on a set of images or on different sets of images acquired in a separate process. For example, it may be beneficial to acquire a subset of images of a subject looking straight ahead and estimate the distance based on this subset of images.
[0019] Scaled 3D coordinates can be used to obtain personalized HRTFs. In one embodiment, a default set of head-related transfer functions (HRTFs) is modified based on the scaled 3D coordinates. In another embodiment, the set of scaled 3D coordinates is applied to a generative model to generate data corresponding to one or more pHRTFs fitted to the subject.
[0020] A further aspect of the present invention relates to a method for determining the head radius r of a subject, comprising the steps of: acquiring at least one image of the subject acquired using an image-capturing device; estimating the distance from the image-capturing device to the subject's face plane; identifying the 2D coordinates of the subject's tragus in the at least one image; projecting the 2D coordinates onto the image plane of the image-capturing device; and estimating the head radius based on the distance between the projections of the tragus in the image plane, and the estimated ratio between the head radius r and the distance between the face plane to which the estimated distance relates and the ear plane on which the tragus is located.
[0021] This method has certain advantages in relation to the first aspect and may also be advantageous in other situations.
Brief Description of the Drawings
[0022] The present invention will be described in more detail with reference to the accompanying drawings showing the current preferred embodiments of the invention.
[0023] [Figure 1] A - C show a video capture process for obtaining a set of images using a handheld device according to an embodiment of the present invention.
[0024] [Figure 2] It is a block diagram of a system according to an embodiment of the present invention.
[0025] [Figure 3] It is a flowchart of a method according to an embodiment of the present invention. <000This disclosure describes techniques that enable reliable and accurate extraction of 3D position and dimensions of anthropometric landmarks and features (e.g., those related to ears and heads) from image data (e.g., two-dimensional or 2D images) captured by a wide variety of camera systems (e.g., mobile phones, tablets, gaming devices, laptops, etc., which include cameras) using diverse capture modalities. While these techniques are described with respect to anthropometric landmarks and features most relevant to pHRTF generation, it is understood that such techniques are generally useful for efficiently obtaining 3D dimensional and / or positional data corresponding to objects captured in 2D images.
[0031] The techniques described herein include strategies for inferring the 3D coordinates of anthropometric landmarks (e.g., head, ear, or torso landmarks) using a set of images showing the subject's head and / or torso (e.g., a sequence of images or frames, which may or may not be obtained from video data). The 3D coordinates are then used as input information to generate a personalized PHRTF, which can be used for audio playback customized for a particular user.
[0032] The following explanations of terms are illustrative and provided solely to aid in understanding the techniques described, and should not be interpreted as limiting. • Camera orientation – This can refer to the camera's position in the world coordinate system, expressed as (yaw, pitch, roll) for a frame / image (for example, from video). • Head posture – This can refer to the head orientation of the user's head in relation to the camera, expressed as yaw, pitch, and roll. • Image plane – This can refer to the plane on which an image is formed.
[0033] Overview / Summary In some embodiments, the 3D reconstruction process is achieved by a technique that includes the following steps: 1. Detect the 2D coordinates (x,y) of the upper body anthropometric landmarks in each image / frame within the set of images. a. In some embodiments, upper body anthropometric landmark detection includes, for example, head landmark detection, torso landmark detection, face landmark detection, ear landmark detection, etc. b. In some embodiments, an additional step is performed to detect the presence and location of a face or ear prior to 2D anthropometric landmark detection. In some embodiments, ear landmark detection is limited to images / frames in which an ear is detected (e.g., each image / frame in a set of images / frames contains a detected ear). In some embodiments, face landmark detection is limited to images / frames in which a face is detected (e.g., each image / frame in a set of images / frames contains a detected face). c. In some embodiments, the detection of ear landmark coordinates is limited to images / frames corresponding to head pose values within a fixed range (e.g., an azimuthal range of [-10°, -45°] or an azimuthal angle range of [10°, 45°]). d. In some embodiments, the set of video images / frames to which landmark detection is applied is selected to uniformly sample the azimuthal space. In other words, the angular separation between images in the set is approximately equal, thereby ensuring that the images in the set are evenly distributed around faces. This can be beneficial as it avoids the processing of very similar images (which contain little additional information). e. In some embodiments, ear landmark detection includes determining data corresponding to ear landmarks that describe the anatomical features of the auricle, helix, antihelix, tragus, fossa, and concha. 2. By utilizing the head orientation of the user's head relative to the camera in each image / frame, the camera orientation for each image / frame in the video image / frame set is estimated (for example, a 3D reconstruction assuming the camera is moving and the head remains stationary, or the head rotates and the camera remains stationary). a. In some embodiments, head pose information is obtained via API calls to external software. b. In some embodiments, head pose information is estimated by calculating an instantaneous face mesh of the user's face and estimating rotational translations (yaw, pitch, roll) that rotate and translate (transform) the instantaneous user face mesh points to a general 3D face mesh (yaw==pitch==roll==0) at the center of the camera frame. The instantaneous face mesh is the face mesh calculated for each frame in a feed, for example, a video feed. c. In some embodiments, instantaneous facial mesh information is used to detect the position of the face or ear and improve the accuracy of 2D landmark detection. 3. 2D anthropometric landmarks are projected onto a 3D world from multiple images / frames of a set of video images / frames, and multi-angulation is performed to obtain 3D facial landmark coordinates (x,y,z). In some embodiments, 3D points are identified as points that minimize the error when intersecting projection lines. In some embodiments, different images are selected according to the determination that the projection error exceeds a threshold, and steps 1-3 are repeated. In some embodiments, additional video data is captured according to the determination that the projection error exceeds a threshold, and steps 1-3 are repeated for a set of images from the additional video data. 4. For each image / frame in the set of images / frames, estimate the distance of the face from the camera (e.g., camera distance), and use this absolute distance to scale the distance. a. In some embodiments, the camera distance is estimated using a hardware depth sensor (e.g., a time-of-flight sensor). In some embodiments, the camera distance is estimated without a hardware depth sensor (e.g., using depth from motion algorithms, multiple non-depth cameras, infrared cameras / dot projectors, etc.). b. In some embodiments, the camera distance is estimated by measuring the user's iris diameter, assuming that iris size is constant across the human population, and scaling the world relative distance using the iris dimension as an absolute scaling factor. c. In some embodiments, the camera distance is estimated by combining different strategies for distance scaling (e.g., iris diameter, pupillary distance, tragus length, reference object, etc.). In some embodiments, the results from multiple distance estimation strategies are combined using Bayesian inference. 5. Optionally, estimate the head radius for each image / frame in a set of images / frames based on the projection of 2D face mesh information onto the image plane and the estimated distance of the head from the camera (e.g., camera distance). a. In some embodiments, the head radius is estimated based on the distance between the ear plane (e.g., a coronal plane relative to the head that includes two tragus) and the face plane (e.g., a plane parallel to the ear plane that minimizes the distance between the face plane and the face mesh landmark). 6. Extract features (of the ears, face, and head) from 3D anthropometric landmarks (e.g., distance, angle, area), and use the extracted features to generate one or more PHRTFs tailored to a specific user.
[0034] The first part of the process involves acquiring a set of images of subject (user) 1 that include one or more anatomical features, such as ears, eyes, etc., in order to enable the determination of anthropometric features. The images may, in principle, be acquired as separate images, but otherwise they may be selected from a video feed. The image acquisition device may be a mobile acquisition device 1 such as a phone or tablet, or a fixed device such as a gaming device or workstation.
[0035] Figures 1a-c illustrate the video capture process for acquiring a feed of images of subject (user) 1 using a handheld image capture device 2. This process is similar when a fixed device is used. For the purposes of this disclosure, the acquired images should include at least the head 3 of subject 1.
[0036] First, in Figure 1a, video acquisition is initialized with Subject 1 holding Camera 2 in front of their face and looking straight at the camera. Next, in Figure 1b, while video acquisition is in progress, the subject turns their head to the right (or left) and then returns to the starting position. Finally, in Figure 1c, the subject turns their head to the left (or right) and then returns to the starting position.
[0037] By obtaining a video feed of images in this manner, the feed will include images spanning a wide azimuth range (e.g., 180 degrees, 110 degrees, 90 degrees, etc.). Furthermore, by starting with a front-facing orientation, the process can begin by identifying anthropometric features in a front view where many relevant features are visible. Additionally, by stopping at a starting position between left and right rotation, it is possible to ensure a "reset" of any offsets caused by rotational motion.
[0038] Figure 2 shows a block diagram of system 10 according to several embodiments. System 10 is configured to process a set of input image frames 11 of a person's upper body to extract 3D head, face, and ear landmark coordinates.
[0039] As shown in Figure 2, the set of images 11 is provided to three parallel modules: a 2D anthropometric landmark detection module 12, a head pose detection module 13, and a face distance estimation module 14. Although modules 12, 13, and 14 are shown in parallel, it should be noted that the processing of these modules may be at least partially sequential. For example, as discussed below, head pose may be obtained by request to external software (e.g., available in an image acquisition device), and such a request may also provide a face mesh useful for landmark detection in module 12.
[0040] The results from modules 12 and 13 are provided to a multi-angulation module 15, which outputs relative 3D coordinates of anthropometric features. These coordinates are then provided to a scaling module 16, which generates scaled 3D coordinates using the estimated distance D from module 14 and parameters P (such as focal length) related to the image acquisition device.
[0041] Referring to the flowchart in Figure 3, we will now briefly outline the operation of the system (for example, 10).
[0042] First, in stage S11 (image acquisition), a set of images of the subject is acquired from a feed of images captured from multiple viewpoints by an image acquisition device. The process then proceeds to stage S12, which, together with stages S13 and S14, forms a loop that is performed for each image in the set. Specifically, in stage S13 (detection of 2D position), the 2D positions of multiple anthropometric landmarks are detected. The landmarks may be, for example, face landmarks or ear landmarks. In stage S14 (estimation of head posture), the head posture is estimated based on the 2D positions.
[0043] Further selection of image subsets may be made based on the estimated head pose for all images in the set. Such selection can ensure a subset of images with a uniform angular sampling frequency, i.e., images acquired at observation angles relative to the subject's head that are evenly distributed in angular space. Also, anthropometric landmarks detected in a particular image may depend on the estimated head pose. For example, it may be impossible (or unreliable) to detect ear landmarks for a head pose that is substantially facing straight ahead (facing the camera).
[0044] In step S15 (generating relative 3D coordinates), relative 3D coordinates are generated for a set of anthropometric landmarks based on the detected 2D positions of corresponding landmarks and associated head poses for the set of images. Note that the set of anthropometric landmarks for which the 3D coordinates are determined may include all or only some of the 2D positions detected in each image. For example, as mentioned above, some landmarks may not be very relevant depending on the estimated head pose.
[0045] In step S16 (generating scaled 3D coordinates), scaled 3D coordinates are generated from relative 3D coordinates based on characteristics associated with the image acquisition device and the estimated distance between the subject and the image acquisition device. Relevant characteristics may include focal length, pixel density, and field of view. The estimated distance can be obtained by appropriate distance / depth detection hardware or by relating some of the detected anthropometric landmarks to known statistical data about certain anatomical features, such as pupil diameter or interpupillary distance.
[0046] 2D landmark detection In module 12, 2D anthropometric landmark detection is performed on image frames (e.g., individual images or images constituting a frame). Landmark detection can be achieved using different strategies. In some embodiments, 2D landmark detection is performed using a deep neural network (DNN). In some embodiments, the DNN includes a MobileNet V1 or V2 backbone with a fully connected network on top and a regression activation function. The DNN is trained using the mean Euclidean error between estimated (x,y) coordinates and annotated coordinates as the cost function. A large dataset of annotated images characterized by human head, torso, and ear landmarks is used for training. Strategies such as image augmentation, regularization, dropout, and pooling can be used to improve performance on test data. In some embodiments, conventional computer vision strategies for landmark detection may be used in conjunction with or instead of the exemplary strategies described above.
[0047] In certain applications, an optional step is performed. For example, according to some embodiments, a specific feature detector (e.g., an ear or face detector) is used to identify the presence of relevant features (e.g., ears or faces) within an image / frame. In some embodiments, the region associated with a specific relevant feature (e.g., ear or face) may be estimated using instantaneous faceMesh and head pose information obtained from an external source (e.g., via an API call or otherwise from external software), for example, from software available in the image acquisition device. This step is typically performed before ear and face landmark detection and may be used to reduce the complexity of subsequent processes and / or improve accuracy.
[0048] Figure 4 shows exemplary results of 2D ear landmark detection according to several embodiments. In Figure 4, multiple ear landmarks 18 are shown in the image of the head 3. Similar anthropometric landmark annotations can be applied to the head, face, or other upper body parts such as the torso.
[0049] Camera pose estimation In module 13, the head pose of each image in the set is determined. The head pose can be thought of as a coordinate transformation between the camera coordinate system and the head coordinate system. For practical purposes, it may be sufficient to consider rotational transformations in two or three degrees of freedom, such as yaw, pitch, and roll. In some embodiments, camera pose information (camera yaw, pitch, and roll in a global reference system) is obtained (e.g., determined, estimated, accessed, received, etc.) by measuring (e.g., determined, estimated, accessed, received, etc.) the head pose information (yaw, pitch, and roll position of the head relative to the image frame).
[0050] In some embodiments, camera pose information is obtained, for example, by calling a head pose API available within the image acquisition device, the head pose API returns information about the head pose, and this information is used in the algorithms described herein.
[0051] In some embodiments, head pose information is measured directly. To achieve this, known strategies and techniques for facial landmark detection are utilized, such as open-source libraries or other DNNs trained for this purpose.
[0052] By detecting facial landmarks, the matrix
number
number
[0053] Also, a facial mesh of a typical human face
number
number
[0054] This is done for images within the video, obtaining yaw, pitch, and roll for each image / frame. In some embodiments, camera pose estimation is performed on a subset of images / frames in the video. For example, according to some embodiments, in frames where the estimated camera pose is determined to be missing, unreliable, or incorrect, the camera pose values are corrected or re-estimated using machine learning or signal processing algorithms such as regression or Kalman filtering.
[0055] Multi-angulation of landmark projection Module 15 determines relative 3D coordinates for a set of anthropometric features. Specifically, using head pose (rotational transformation) expressed as yaw, pitch, and roll, as well as the 2D (x,y) coordinates of landmarks in the image plane, anthropometric landmarks can be projected into a global reference frame, and 3D points in space where the projections approximately intersect can be identified, as shown in Figure 5.
[0056] Figure 5 shows the 2D (x,y) coordinates P of a landmark in an image, using head pose information for a set of images acquired from different observation angles (k-1, k, k+1). j,k-1 ,P j,k ,P j,k+1 This shows that the points are projected onto a 3D world reference frame. As mentioned above, the approximate intersection of the projections is the 3D (x,y,z) coordinates P of the landmark in the world reference frame. j To identify the subject. Note that Figure 5 shows several cameras to represent images of the subject taken from different observation angles.
[0057] Several algorithmic strategies can be used to approximate the intersections of projection lines. In some embodiments, the strategy includes identifying a point  ̄p (sometimes denoted as p with a bar) that minimizes the mean squared distance between  ̄p and the projection line in a least squares sense. In some embodiments, the strategy includes identifying a point  ̄p that minimizes the mean Euclidean error between the (x,y) coordinates of the original 2D landmark and the (x,y) coordinates obtained by projecting  ̄p back onto the image frame, averaged across the image viewpoint.
[0058] Multiple strategies can be used to remove noisy projections or outliers. For example, projections where the distance from  ̄p lies in the tail of the distribution describing the distance between the projection line and  ̄p (e.g., below the 5th percenter or above the 95th percenter) can be removed, and  ̄p can be recursively recalculated.
[0059] Other strategies can be used to improve multi-angulation accuracy. For example, head angular velocity can be estimated, and available image viewpoints can be undersampled if the angular velocity is below a certain threshold, or upsampled if the angular velocity is above a certain threshold. By adjusting the sampling frequency (i.e., the selection of images from the feed to the set), a consistent angular spread between images in the set can be achieved. As mentioned above, a consistent angular spread (uniform angular sampling frequency).
[0060] Furthermore, angular velocity can be used to estimate the accuracy of the multi-angulation process. Higher values of angular velocity are associated with faster head rotation, which may be associated with blurred images (e.g., images unsuitable for one or more processing steps). Slower values of angular velocity, on the other hand, are associated with less blur and higher image resolution. Thus, in some embodiments, angular velocity is used to efficiently select a subset of images / frames for 2D / 3D landmark recognition that is more likely to provide subtle data for feature extraction and ultimately pHRTF generation. Such selection may be made when acquiring a set of images in step S11, or before processing in step S15 in Figure 3 to select a subset of the set of images.
[0061] Multi-angulation error can be used as a substitute for multi-angulation accuracy. In some embodiments, different metrics of multi-angulation error are used for this purpose. For example, multi-angulation error can be measured as the average distance of  ̄p from the projection line, or as the average Euclidean distance between the original (x,y) landmark in the image plane and the corresponding (x,y) coordinates obtained by back-projecting  ̄p into the image plane, averaged over all image viewpoints.
[0062] In some embodiments, multi-angulation accuracy is used to navigate subsequent steps in the pHRTF generation process. For example, if it is determined that the multi-angulation accuracy for one ear is low (e.g., below a threshold), additional data may be acquired (e.g., the user may be prompted to recapture an image corresponding to that side of the user's head), and the overall landmark detection and extraction process is repeated with the newly acquired data. In some embodiments, if it is determined that the multi-angulation accuracy is low (e.g., below a threshold or remains low) after a second capture (e.g., subsequent recapture), data corresponding to the contralateral ear (i.e., the other ear) is used to generate the corresponding pHRTF (i.e., by assuming symmetry in the size and shape of the user's ear).
[0063] World standard scaling The multi-angulation performed in module 15 obtains relative 3D coordinates. To convert relative world coordinates to absolute world coordinates (e.g., mm or inches), the reference frame is scaled to absolute coordinates in module 16.
[0064] In some embodiments, depth information (distance D) of the image / frame is obtained (e.g., via an API call) and used with known camera and sensor-specific parameters P (e.g., focal length, AOV, FOV, etc.) to determine absolute scaling. In some embodiments, module 14 obtains depth information that describes the distance of each object in the image from the camera image plane, calculated pixel by pixel.
[0065] In some embodiments, module 14 determines the iris size in the image and uses that measurement to estimate the distance D of the head from the camera image plane. In some embodiments, the iris size is determined based on one or more images / frames corresponding to a subject facing directly towards the camera (for example, based on the detected head posture angle). Using the iris makes it possible to measure / estimate the distance of the face from the camera with an error of less than 10% without requiring any special hardware. This relies on the characteristic that the horizontal iris diameter of the human eye is approximately constant at 11.7 ± 0.5 mm across a wide population.
[0066] Once the distance D from the camera to the face is determined, the relative reference frame is scaled using intrinsic camera parameters. For example, by knowing the camera's angle of view (aov), the camera's field of view (fov) can be estimated using the following formula: fov = tan(aov / 2). The camera's focal ratio is obtained as foc == fov / 2.
[0067] The scaling factor for converting any pixel distance to mm is given by px2mm = depth / (max(imsize)·foc), where depth is the distance of a point from the camera plane (obtained through a depth map or alternative strategy) and max(imsize) represents the maximum dimensions of the image.
[0068] Estimation of head radius Once the distance of the face from the camera is known, it can be used to estimate the head radius, which is the distance between the edges of the ears, also known as the tragus. The challenge is that the tragus lies within the ear plane, and the ear plane is slightly further from the camera than the face plane (to which the distance has been estimated).
[0069] Figure 6 shows techniques for estimating head radius according to several embodiments.
[0070] The head bitragion is the distance from the two tragus. The tragus is a small, pointed protrusion of the outer ear, located in front of the concha, overlapping the ear canal and projecting backward. The head bitragion is assumed to be twice the length of the head radius r. The distance between the face plane and the ear plane (for example, the respective planes described in the overview section above) correlates with the head radius by a predetermined correlation coefficient α. In some embodiments, the correlation coefficient α is approximately 0.6. Thus, the distance between the camera and the ear plane is D + αr. By projecting the 2D position of the tragus onto the image plane, the distance x between the tragus (when it appears on the face plane) is determined. Using the focal length f (i.e., one of the camera's intrinsic parameters P), the head radius is estimated by these parameters along with several simple geometric arguments.
[0071] Specifically, considering two congruent triangles with bases x and 2r, and heights f and D+αr respectively, we obtain the following equation.
number
[0072] PHRTF generation Once absolute 3D coordinates are obtained (e.g., estimated), inductive features (distance, angle, and area) describing the anthropometric features that affect the pHRTF are calculated. In some embodiments, these inductive features are used to generate the pHRTF using a generative model (e.g., a model trained to generate pHRTF data from data describing anthropometric features). Another approach is to use the technique described in PCT / US2024 / 016014, which is incorporated herein.
[0073] Implementation details Figure 7 shows a schematic block diagram of an exemplary electronic device or architecture 200 (e.g., Apparatus 200) suitable for implementing an exemplary embodiment of the present disclosure. Architecture 200 may form part of a mobile device such as a smartphone 2, or it may be a standalone device. Architecture 200 may, but is not limited to, the system described in relation to Figure 2.
[0074] As illustrated, the architecture 200 includes a central processing unit (CPU) 201 capable of executing various processes according to, for example, a program stored in read-only memory (ROM) 202, or a program loaded from a memory unit 208 into random access memory (RAM) 203. The CPU 201 may be an electronic processor 201, which may include one or more processor cores, and in some examples, the processor 201 may consist of multiple processors. The RAM 203 also stores data necessary for the CPU 201 to execute various processes, as needed. The CPU 201, ROM 202, and RAM 203 are interconnected via a bus 204. An input / output (I / O) interface 205 is also connected to the bus 204.
[0075] The following components are connected to the I / O interface 205: an input unit 206 which may include a keyboard, mouse, etc.; an output unit 207 which may include a display such as a liquid crystal display (LCD) and one or more speakers; a storage unit 208 which may include a hard disk or another suitable storage device; and a communication unit 209 which may include a network interface card such as a network card (e.g., wired or wireless).
[0076] The input unit 206 also includes a camera that enables the capture of an image feed.
[0077] In some implementations, the output unit 207 includes a system with varying numbers of speakers. The output unit 207 can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other appropriate formats).
[0078] In some embodiments, the communication unit 209 is configured to communicate with other devices (for example, via a network). The drive 210 is also connected to the I / O interface 205, if necessary. A removable medium 211, such as a magnetic disk, optical disk, magneto-optical disk, flash drive, or other suitable removable medium, is mounted on the drive 210, thereby allowing computer programs read therefrom to be installed in the storage unit 208, if necessary. Those skilled in the art will understand that although the apparatus 200 is described as including the above-described components, in actual applications it is possible to add, remove, and / or replace some of these components, and all such modifications or changes will fall within the scope of this disclosure.
[0079] According to exemplary embodiments of the present disclosure, the processes described above may be implemented as a computer software program or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product which includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing the method. In such embodiments, the computer program may be downloaded and mounted from a network via a communication unit 209, as shown in Figure 7, and / or installed from a removable medium 211.
[0080] In general, various exemplary embodiments of this disclosure may be implemented in hardware or dedicated circuitry (e.g., control circuits), software, logic, or any combination thereof. For example, the various steps in Figure 3 may be performed by a control circuit (e.g., CPU 201 in combination with the other components in Figure 7), and thus the control circuit may perform the actions described in this disclosure.
[0081] Some aspects may be implemented in hardware, while others may be implemented in firmware or software, which may be executed by other computing devices, including controllers, processors, and / or control circuits. Various aspects of the exemplary embodiments of this disclosure are illustrated and described using block diagrams, flowcharts, or any other pictorial representation, but it will be understood that any blocks, apparatus, systems, techniques, or methods described herein may, in non-limiting examples, be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or any combination thereof.
[0082] Furthermore, the various blocks shown in the flowchart can be viewed as method steps and / or as operations resulting from the operation of computer program code and / or as a plurality of coupled logic circuit elements constructed to perform related functions. For example, embodiments of the present disclosure include a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code configured to perform the methods described above.
[0083] Computer program code for performing the methods of this disclosure may be written in any combination of one or more programming languages. These computer program codes may be provided to one or more processors of a general-purpose computer, a dedicated computer, or other programmable data processing device having control circuits, so that when the program code is executed by one or more processors of the computer or other programmable data processing device, it implements the functions / operations specified in the flowcharts and / or block diagrams. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer, partially on a remote computer, or entirely on a remote computer or server, or distributed across one or more remote computers and / or servers.
[0084] One or more processors may operate as standalone devices or be connected to other processors, for example, to a network. Such a network may be built on a variety of different network protocols and may be the internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.
[0085] Software may be distributed on computer-readable media, which may include computer storage media (or non-temporary media) and communication media (or temporary media). As is well known to those skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, various forms of physical (non-temporary) storage media, such as ROM, PROM, EPROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other media that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media (temporary) typically include any information delivery medium, which embodies computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transport mechanisms.
[0086] The implementations of the technology disclosed in the figures are merely illustrative examples, and the present invention is not limited in that way. For example, illustrated partitions such as the block in Figure 2 are merely illustrative logical partitions for the sake of clarity, and such partitions may be divided into additional partitions, combined into fewer partitions, supplemented by additional partitions, or reduced by deleting partitions, without departing from the spirit of the present invention. In the illustrative flowchart of Figure 3, the partitions of operational stages, which may also be called functions, stages, operations, processes, or steps, may be combined into fewer stages or divided into additional stages, where stages may be rearranged or deleted, in whole or in part, without departing from the spirit of the present disclosure.
[0087] Unless otherwise specified, as is evident from the following description, any description using terms such as “process,” “calculate,” “calculate,” “determine,” and “analyze” throughout this disclosure is understood to refer to the functions, actions, stages, and / or processes of computer hardware or computing systems or similar electronic computing devices that manipulate and / or transform data that is represented as a physical quantity, such as an electronic quantity, into other data that is similarly represented as a physical quantity.
[0088] In the above description of exemplary embodiments of the present invention, it should be understood that various features of the invention are grouped together in a single embodiment, figure, or description for the purpose of improving the flow of disclosure and aiding in the understanding of one or more of the various aspects of the invention. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are explicitly described in each claim. Rather, as reflected in the following claims, the aspects of the invention are fewer than all the features of a single, aforementioned disclosed embodiment. Thus, the claims following the detailed description are explicitly incorporated into this detailed description, with each claim standing independently as a distinct embodiment of the invention. Furthermore, some embodiments described herein include some features included in other embodiments, but not others, and combinations of features of different embodiments are intended to be within the scope of the invention and form different embodiments, as will be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments may be used in any combination.
[0089] Furthermore, some embodiments are described herein as methods or combinations of elements of methods that can be implemented by a processor of a computer system or by other means of performing the function. Thus, a processor having instructions for performing such methods or elements of methods forms means for performing the methods or elements of methods. Note that if a method includes several elements, for example, several steps, the ordering of such elements is not implied unless specifically stated. Furthermore, the elements of apparatus embodiments described herein are examples of means for performing the function performed by those elements for the purpose of performing embodiments of the present invention. Numerous specific details are described in the description provided herein. However, it is understood that embodiments of the present invention can be carried out without these specific details. On the other hand, well-known methods, structures, and arts are not shown in detail so as not to obscure the understanding of this paper.
[0090] Those skilled in the art will understand that the present invention is by no means limited to the preferred embodiments described above. Rather, many modifications and variations are possible within the scope of the appended claims.
[0091] Various aspects and implementations of this disclosure can also be understood from the following bulleted exemplary embodiments (EEEs) which are not part of the claims. [EEE1] In electronic devices: The steps include: acquiring multiple images (e.g., multiple images including the subject and the subject's anatomical features (e.g., ears), a sequence of images, a video, etc.) capturing the subject (e.g., a person, an object) from multiple viewpoints; For at least two of the aforementioned images: Perform landmark detection to find the location (e.g., 2D coordinates) corresponding to multiple landmarks; Head posture data is estimated from the position corresponding to the landmark; The step of generating camera posture data from the aforementioned head posture data; The steps include generating relative 3D coordinate data corresponding to the positions (e.g., 3D relative coordinates) of a set of anatomical features captured in the plurality of images from camera pose data corresponding to at least two of the plurality of images; From the aforementioned relative 3D coordinate data, Characteristics related to the camera used to capture the aforementioned multiple images (e.g., AOV, FOV, focal length, resolution, etc.), and Depth data associated with the aforementioned multiple images Using at least one of the following, the step of generating scaled 3D coordinate data A method that includes performing [something]. [EEE2] The method according to EEE1, further comprising the step of generating data corresponding to a head transfer function fitted to the subject, at least in part, based on the scaled 3D coordinate data. [EEE3] A method according to EEE1 or 2, The process further includes determining motion metrics (e.g., angular or linear, velocity or acceleration) associated with the aforementioned plurality of images; In accordance with the determination that the motion metric satisfies a first criterion (for example, above a first threshold, below a first threshold), the at least two images include a first amount of images. In accordance with the determination that the motion metric satisfies a second criterion (for example, below a second threshold, above a second threshold), the at least two images include a second amount of images that is greater than a first amount of images. In some embodiments, the first and second criteria are thresholds corresponding to a common value (e.g., angular values related to head or camera orientation). In some embodiments, the first and second criteria are thresholds corresponding to different values. In some embodiments, the motion metric is based on data from sensors of an electronic device (e.g., IMU, accelerometer, gyroscope, etc.). In some embodiments, the motion metric is based on image data from multiple images. [EEE4] A method according to any one of EEE1 to 3, The process further includes determining the frame rate associated with the plurality of images (for example, the fps of the video associated with the plurality of images), In accordance with the determination that the frame rate satisfies a third criterion (for example, above a third threshold, below a third threshold), the at least two images include a first quantity of images. In accordance with the determination that the frame rate satisfies a fourth criterion (for example, below a fourth threshold, above a fourth threshold), the at least two images include a second amount of images that is greater than a first amount of images. In some embodiments, the third and fourth criteria are thresholds corresponding to the same value. In some embodiments, the first and second criteria are thresholds corresponding to different values. In some embodiments, the frame rate is obtained from metadata associated with the plurality of images. [EEE5] The method according to any one of EEE1 to 4, wherein the at least two images are selected at least in part on the determination that the estimated head posture angle of each of the plurality of images falls within a predetermined head posture range (e.g., azimuth range). [EEE6] The method according to any one of EEE1 to 5, wherein the at least two images are selected at least in part on the basis of a desired uniform angular sampling frequency (for example, N = the number of images for sampling a given azimuth range at a given frequency). [EEE7] The method according to any one of EEE1 to 6, wherein detecting multiple landmarks includes performing at least one of ear detection and face detection. [EEE8] The method according to EEE7, wherein ear detection includes determining data corresponding to ear landmarks that describe the anatomical features of the auricle, helix, antihelix, tragus, fossa, and concha. [EEE9] The method according to any one of EEE1 to 8, wherein estimating head posture data includes obtaining head posture data via a request (for example, an API call; sending a request for head posture data for each image and receiving head posture data for each image in response to said request). [EEE10] Estimating head posture data: The steps include: calculating the instantaneous facial mesh of the subject's face; The step involves estimating rotational translation (yaw, pitch, roll) to rotate and translate the aforementioned instantaneous facial mesh points of the subject onto a general 3D facial mesh at the center of the image / frame. The method described in any one of the EEE1 to 9, including the method described in any one of the EEE1 to 9. [EEE11] Landmark detection is: The steps include: calculating the instantaneous facial mesh of the subject's face; The steps include using the aforementioned instantaneous face mesh to detect the position of the face or ear in each of at least two images, and The method described in any one of the EEE1 to 9, including the method described in any one of the EEE1 to 9. [EEE12] Generating relative 3D coordinate data is: A step of projecting a corresponding detected landmark from each of the at least two images into the 3D world using camera pose data or head pose data associated with each of the at least two images; In order to obtain the relative 3D face landmark coordinates for each detected landmark, the process involves performing multi-angulation. The method described in any one of the EEE1 to 11, including the method described in any one of the EEE1 to 11. [EEE13] The method according to EEE12, wherein each 3D facial landmark coordinate is identified as a point that minimizes the error in intersecting the projection line. [EEE14] To generate scaled 3D coordinate data from the aforementioned relative 3D coordinate data: Of the aforementioned plurality of images, at least two images, and A reference subset of the plurality of images (for example, a set of one or more images from the plurality of images that does not include the at least two images) that is different from the subset formed by the at least two images of the plurality of images. The method according to any one of EEE1 to 13, comprising estimating the distance from the subject to the camera for at least one of the following: In some embodiments, the reference subset corresponds to one or more images from the plurality of images that capture a front view of the subject (e.g., images associated with a head posture within one or more central threshold angle ranges (e.g., yaw, pitch, or roll approaching approximately 0°)), one or more images from the plurality of images associated with the subject facing directly to the camera that captured each image, or one or more images from the plurality of images of the subject facing the camera (e.g., perpendicular to the surface of the camera lens of the camera that captured the image). In some embodiments, the reference subset of the plurality of images is associated with a specific ear (left or right) of the subject captured in each of the at least two images from the plurality of images. [EEE15] Estimating the distance from the subject to the camera is Iris size; interpupillary distance; Reference object size; Data from the hardware depth sensor, and Depth from motion algorithms A method described in EEE14, based on at least one of the sets. [EEE16] The method according to EEE14 or 15, wherein estimating the distance from the subject to the camera involves combining the results from two or more distance estimation strategies (e.g., different estimation modalities) using Bayesian inference. [EEE17] The method according to any one of EEE1 to 16, further comprising estimating the head radius for each of the at least two images by projecting 2D facial mesh information onto an image plane. [EEE18] The method according to EEE17, wherein the head radius is estimated based on the distance between the ear plane or axis (e.g., a coronal plane relative to the subject's head containing two tragus; an axis between two tragus) and the face plane or axis (e.g., a plane parallel to the ear plane that minimizes the distance between face mesh landmarks; an axis connecting symmetrical face landmarks). [EEE19] To generate head-related transfer function data corresponding to head-related transfer function data: Extract features (e.g., features related to ears, face, and / or head) from scaled 3D coordinate data corresponding to anthropometric landmarks (e.g., distance, angle, area); This includes applying the extracted features to a generative model to generate data corresponding to one or more PHRTFs fitted to the subject, The method described in any one of EEE2 through EEE18. [EEE20] The stage of calculating the projection error from 2D to 3D; The steps include: initiating a process to capture additional image data of the subject in accordance with a determination that the projection error satisfies a first error criterion (for example, including a determination that the error value is greater than a first threshold); Using the aforementioned additional image data (for example, according to the method of EEE1), the step of generating scaled 3D coordinate data and The method described in any one of EEE1 to 19, further including the method described in any one of EEE1 to 19. [EEE21] The stage of calculating the projection error from 2D to 3D; The process includes the step of generating scaled 3D coordinate data corresponding to a second ear (e.g., right ear, left ear) using image data associated with a first ear (e.g., left ear, right ear), in accordance with a determination that the projection error satisfies a second error criterion (including, for example, a determination that the error value is greater than a second threshold, and / or a determination that the process of capturing additional image data of the subject has been previously initiated), The method described in any one of EEE1 to EEE19. [EEE22] A non-temporary computer-readable storage medium containing instructions, when executed by one or more processors of an electronic device, that cause the electronic device to perform the method described in any one of the EEE1 to 21. [EEE23] An electronic device comprising one or more processors and a memory storing instructions, wherein, when the instructions are executed by one or more processors, the electronic device causes the electronic device to perform the method described in any one of EEE1 to 21. [EEE24] The electronic device according to EEE23, further comprising one or more cameras, wherein the plurality of images are captured by the one or more cameras.
[0092] Several aspects are described below. [Aspect 1] A method for 3D reconstructing anthropometric features from images of a subject: The steps include: obtaining a set of images of the subject from a feed of images captured using an image acquisition device from multiple viewpoints; For each image in the aforementioned set: Detect the 2D positions of multiple human body measurement landmarks, The steps include: estimating the head posture based on the aforementioned 2D position; A step of generating relative 3D coordinates corresponding to at least some of the anthropometric features captured in the set of images, based on the detected 2D position and associated head pose for at least some of the images in the set of images; From the aforementioned relative 3D coordinates The properties associated with the image capture device, and Estimated distance between the subject and the image capture device Based on this, the step of generating scaled 3D coordinates and Methods that include... [Aspect 2] The method according to embodiment 1, wherein the set of images is selected based on a desired uniform angular sampling frequency. [Aspect 3] The method according to embodiment 1 or 2, wherein the anthropometric landmarks include ear landmarks and facial landmarks. [Aspect 4] The method according to embodiment 3, wherein the 2D positions of a subset of anthropometric landmarks are detected only for a subset of images in which the head posture is within a predetermined head posture range. [Aspect 5] The method according to embodiment 4, wherein the 2D position of the ear landmark is detected only for a subset of images in which the head posture is within a predetermined azimuthal angle range. [Aspect 6] The method according to aspect 4 or 5, wherein the detection of an ear landmark includes determining data describing at least one anatomical feature among the auricle, helix, antihelix, tragus, fossa, and concha. [Aspect 7] The method according to any one of embodiments 1 to 5, wherein estimating head posture includes estimating rotational transformations (yaw, pitch, roll) that convert the position of the 2D anthropometric landmark into a general 3D face mesh at the center of the image. [Aspect 8] The method according to any one of embodiments 1 to 7, wherein estimating head posture includes making a request to software in the image acquisition device and obtaining a response including head posture data. [Aspect 9] The method according to embodiment 8, wherein the response also includes a 2D position of an anthropometric landmark. [Aspect 10] To generate relative 3D coordinates: Using the estimated head pose, the 2D position is projected onto a 3D coordinate system; This includes identifying the 3D coordinates of a 2D position as the intersection of projections of that 2D position from multiple images. The method according to any one of the embodiments 1 to 9. [Aspect 11] The method according to embodiment 10, wherein the 3D coordinates of a 2D position are identified as points that minimize errors in the projection of this 2D position from multiple images. [Aspect 12] The method according to any one of embodiments 1 to 11, wherein the estimated distance between the subject and the image capture device is based on the set of images or on other images from the image feed. [Aspect 13] Estimating the distance between the subject and the image capture device is Iris size; Interpupillary distance: Reference object size; Data from hardware depth sensors; and Depth from motion algorithms The method according to any one of embodiments 1 to 12, based on at least one of the above. [Aspect 14] The method according to embodiment 12 or 13, wherein estimating the distance from the subject to the image acquisition device includes combining results from two or more distance estimation strategies using Bayesian inference. [Aspect 15] The method according to any one of embodiments 12 to 14, further comprising projecting a 2D landmark extracted from the image onto the image plane of the image acquisition device, and estimating the head radius r based on the projected 2D landmark and the characteristics of the image acquisition device. [Aspect 16] The aforementioned 2D landmark is related to the tragus of the subject, and the head radius r is The distance between the projections of the tragus in the aforementioned image plane, The estimated ratio between the head radius r and the distance between the face plane to which the estimated distance relates and the ear plane on which the tragus is located. The method according to embodiment 15, determined based on the following: [Aspect 17] The head radius r is r = (Df) / ((2f / x) - α) The method according to embodiment 16, wherein the distance is determined according to the formula, where D is the estimated distance to the face plane, x is the distance between the tragus in the image plane of the image capturing device, f is the focal length of the image capturing device, and α is the ratio of the head radius r to the distance between the face plane and the ear plane. [Aspect 18] The stage of calculating the projection error from 2D to 3D; The step of initiating a process to include additional images in the set of images, based on the determination that the projection error is greater than the error threshold; The step of generating scaled 3D coordinate data using the aforementioned additional image data. The method according to any one of embodiments 1 to 17, further including the following: [Aspect 19] The method according to any one of embodiments 1 to 18, further comprising the step of generating scaled 3D coordinate data corresponding to a second ear using image data associated with a first ear. [Aspect 20] The step of determining the subject's movement metric based on the image feed; The step of including only images in the set of images whose motion metric falls within a predetermined range. The method according to any one of embodiments 1 to 19, further including the following: [Aspect 21] A method for generating a personalized head-related transfer function (pHRTF): The step of obtaining a default set of head-related transfer functions (HRTFs); A step of obtaining a set of scaled 3D coordinates of anthropometric features using the method described in any one of embodiments 1 to 20; The steps include modifying the default set of head-related transfer functions (HRTFs) based on the scaled 3D coordinates, and Methods that include... [Aspect 22] A method for generating a personalized head-related transfer function (pHRTF): A step of obtaining a set of scaled 3D coordinates of anthropometric features using the method described in any one of embodiments 1 to 20; The steps include applying the scaled set of 3D coordinates to the generation model to generate data corresponding to one or more PHRTFs adapted to the subject, and Methods that include... [Aspect 23] An electronic device comprising one or more processors and a memory storing instructions, wherein, when executed by the one or more processors, the instructions cause the electronic device to perform the method described in any one of embodiments 1 to 22. [Aspect 24] The electronic device according to embodiment 23, further comprising one or more image capturing devices, wherein the set of images is captured by the one or more image capturing devices. [Aspect 25] A non-temporary computer-readable storage medium having, when executed by one or more processors of an electronic device, instructions that cause the electronic device to perform the method described in any one of embodiments 1 to 22. [Aspect 26] A method for determining the head radius r of a subject: A step of acquiring at least one image of the subject using an image acquisition device; The steps include: estimating the distance from the image capture device to the subject's face plane; The steps include identifying the 2D coordinates of the subject's tragus in at least one of the images; The steps include: projecting the aforementioned 2D coordinates onto the image plane of the image acquisition device; The distance between projections of the tragus in the aforementioned image plane, The estimated ratio between the head radius r and the distance between the facial plane to which the estimated distance relates and the ear plane where the tragus is located. Based on this, the step of estimating the head radius and Methods that include... [Aspect 27] The head radius r is r = (Df) / ((2f / x) - α) The method according to embodiment 26, wherein the distance is determined according to the formula, where D is the estimated distance to the face plane, x is the distance between the tragus in the image plane of the image capturing device, f is the focal length of the image capturing device, and α is the ratio of the head radius r to the distance between the face plane and the ear plane. [Aspect 28] Estimating distance is: Iris size; interpupillary distance; Reference object size; Data from hardware depth sensors; and Depth from motion algorithms The method according to embodiment 27, based on at least one of the following. [Aspect 29] The method according to any one of embodiments 26 to 28, wherein estimating the distance from the subject to the image acquisition device includes combining results from two or more distance estimation strategies using Bayesian inference. [Aspect 30] An electronic device comprising one or more processors and a memory for storing instructions, wherein, when executed by the one or more processors, the instructions cause the electronic device to perform the method described in any one of embodiments 26 to 29. [Aspect 31] The electronic device according to embodiment 30, further comprising one or more image capturing devices, wherein the set of images is captured by the one or more image capturing devices. [Aspect 32] A non-temporary computer-readable storage medium that, when executed by one or more processors of an electronic device, contains instructions causing the electronic device to perform the method described in any one of embodiments 26 to 29.
Claims
1. A method for 3D reconstructing anthropometric features from images of a subject: The steps include: obtaining a set of images of the subject from a feed of images captured using an image acquisition device from multiple viewpoints; For each image in the aforementioned set: Detect the 2D positions of multiple human body measurement landmarks, The steps include: estimating the head posture based on the aforementioned 2D position; The steps include generating relative 3D coordinates corresponding to at least some of the anthropometric features captured in the set of images, based on the detected 2D position and associated head pose for at least some of the images in the set of images; From the aforementioned relative 3D coordinates The properties associated with the image capture device, and Estimated distance between the subject and the image capture device Based on this, the step of generating scaled 3D coordinates and Methods that include...
2. The method according to claim 1, wherein the set of images is selected based on a desired uniform angular sampling frequency.
3. The method according to claim 1, wherein the anthropometric landmarks include ear landmarks and facial landmarks.
4. The method according to claim 3, wherein the 2D positions of a subset of anthropometric landmarks are detected only for a subset of images in which the head posture is within a predetermined head posture range.
5. The method according to claim 4, wherein the 2D position of the ear landmark is detected only for a subset of images in which the head posture is within a predetermined azimuthal angle range.
6. The method according to claim 4, wherein the detection of an ear landmark includes determining data describing at least one anatomical feature among the auricle, helix, antihelix, tragus, fossa, and concha.
7. The method according to claim 1, wherein estimating head posture includes estimating rotational transformations (yaw, pitch, roll) that convert the positions of the 2D anthropometric landmarks into a general 3D face mesh at the center of the image.
8. The method according to claim 1, wherein estimating head posture includes making a request to software in the image acquisition device and obtaining a response including head posture data.
9. The method according to claim 8, wherein the response also includes a 2D position of an anthropometric landmark.
10. To generate relative 3D coordinates: Using the estimated head pose, the 2D position is projected onto a 3D coordinate system; This includes identifying the 3D coordinates of a 2D position as the intersection of projections of that 2D position from multiple images. The method according to claim 1.
11. The method according to claim 10, wherein the 3D coordinates of the 2D position are identified as points that minimize errors in the projection of this 2D position from multiple images.
12. The method according to claim 1, wherein the estimated distance between the subject and the image capture device is based on the set of images or on other images from the image feed.
13. Estimating the distance between the subject and the image capture device is Iris size; Interpupillary distance: Reference object size; Data from hardware depth sensors; and Depth from motion algorithms The method according to claim 1, based on at least one of the following.
14. The method according to claim 12, wherein estimating the distance from the subject to the image acquisition device includes combining results from two or more distance estimation strategies using Bayesian inference.
15. The method according to claim 12, further comprising projecting 2D landmarks extracted from the image onto the image plane of the image capture device, and estimating the head radius r based on the projected 2D landmarks and the characteristics of the image capture device.
16. The aforementioned 2D landmark is related to the tragus of the subject, and the head radius r is The distance between the projections of the tragus in the aforementioned image plane, The estimated ratio between the head radius r and the distance between the face plane to which the estimated distance relates and the ear plane on which the tragus is located. The method according to claim 15, determined based on the following:
17. The head radius r is r=(D-f) / ((2f / x)-α) The method according to claim 16, wherein the distance is determined according to the formula, where D is the estimated distance to the face plane, x is the distance between the tragus in the image plane of the image capturing device, f is the focal length of the image capturing device, and α is the ratio of the head radius r to the distance between the face plane and the ear plane.
18. The stage of calculating the projection error from 2D to 3D; The step of initiating a process to include additional images in the set of images, based on the determination that the projection error is greater than the error threshold; The step of generating scaled 3D coordinate data using the aforementioned additional image data. The method according to claim 1, further comprising:
19. The method according to claim 1, further comprising the step of generating scaled 3D coordinate data corresponding to a second ear using image data associated with a first ear.
20. The step of determining the subject's movement metric based on the image feed; The step of including only images in the set of images whose motion metric falls within a predetermined range. The method according to claim 1, further comprising:
21. A method for generating a personalized head-related transfer function (pHRTF): The steps include obtaining a default set of head-related transfer functions (HRTFs); A step of obtaining a set of scaled 3D coordinates of an anthropometric feature using the method according to any one of claims 1 to 20; The steps include modifying the default set of head-related transfer functions (HRTFs) based on the scaled 3D coordinates, and Methods that include...
22. A method for generating a personalized head-related transfer function (pHRTF): A step of obtaining a set of scaled 3D coordinates of an anthropometric feature using the method according to any one of claims 1 to 20; The steps include applying the scaled set of 3D coordinates to a generation model to generate data corresponding to one or more PHRTFs fitted to the subject, and Methods that include...
23. An electronic device comprising one or more processors and a memory storing instructions, wherein, when the instructions are executed by the one or more processors, the electronic device causes the electronic device to perform the method according to any one of claims 1 to 20.
24. The electronic device according to claim 23, further comprising one or more image capturing devices, wherein the set of images is captured by the one or more image capturing devices.
25. A non-temporary computer-readable storage medium comprising, when executed by one or more processors of an electronic device, instructions causing the electronic device to perform the method according to any one of claims 1 to 20.
26. A method for determining the head radius r of a subject, which is: A step of acquiring at least one image of the subject using an image acquisition device; The steps include: estimating the distance from the image capture device to the subject's face plane; The steps include: identifying the 2D coordinates of the subject's tragus in at least one of the images; The steps include: projecting the aforementioned 2D coordinates onto the image plane of the image acquisition device; The distance between projections of the tragus in the aforementioned image plane, The estimated ratio between the head radius r and the distance between the facial plane to which the estimated distance relates and the ear plane where the tragus is located. Based on this, the step of estimating the head radius and Methods that include...
27. The head radius r is r=(D-f) / ((2f / x)-α) The method according to claim 26, wherein the distance is determined according to the formula, where D is the estimated distance to the face plane, x is the distance between the tragus in the image plane of the image capturing device, f is the focal length of the image capturing device, and α is the ratio of the head radius r to the distance between the face plane and the ear plane.
28. Estimating distance is: Iris size; Interpupillary distance; Reference object size; Data from hardware depth sensors; and Depth from motion algorithms The method according to claim 27, based on at least one of the following.
29. The method according to claim 26, wherein estimating the distance from the subject to the image acquisition device includes combining results from two or more distance estimation strategies using Bayesian inference.
30. An electronic device comprising one or more processors and a memory for storing instructions, wherein, when the instructions are executed by the one or more processors, the electronic device causes the electronic device to perform the method according to any one of claims 26 to 29.
31. The electronic device according to claim 30, further comprising one or more image capturing devices, wherein the set of images is captured by the one or more image capturing devices.
32. A non-temporary computer-readable storage medium comprising, when executed by one or more processors of an electronic device, instructions causing the electronic device to perform the method according to any one of claims 26 to 29.