Arrangement for generating head-related transfer function filters

By using a camera on a mobile device to acquire and analyze images, and combining machine learning and sparse reconstruction techniques, the high cost and heavy workload of generating personalized head-related transfer function filters in existing technologies are solved, achieving efficient and low-cost filter generation.

CN116912666BActive Publication Date: 2026-03-10APPLE INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-02-20
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies require expensive 3D scanning equipment and a large amount of image or video data to generate personalized head-related transfer function filters, resulting in excessive computational and network connectivity burdens, and making image acquisition difficult.

Method used

By using the camera of a mobile device to capture images, analyzing and selecting appropriate images, and transmitting only the necessary image data to the server for filter generation, the computational and data transmission requirements are reduced by combining machine learning and sparse reconstruction techniques.

Benefits of technology

It enables efficient generation of personalized head-related transfer function filters on low-computing-power devices, reducing data transmission volume and computational requirements, improving the efficiency and accuracy of image acquisition, and shortening filter generation time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116912666B_ABST
    Figure CN116912666B_ABST
Patent Text Reader

Abstract

This invention is entitled "Arrangement for Generating a Head-Related Transfer Function Filter". The invention discloses an arrangement for acquiring images used to generate a head-related transfer function filter. In this arrangement, the camera of a mobile phone or similar portable device is adjusted to take an image. All acquired images are analyzed, and only suitable images are further sent for generating the head-related transfer function filter. The arrangement is further configured to provide instructions to the user to adequately cover the entire head and other relevant body parts.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related application citation

[0002] This application is a divisional application of Chinese national application number 201910126684.2, filed on February 20, 2019, entitled "Arrangement for Generating Head-Related Transfer Function Filters". Technical Field

[0003] This application relates to the arrangement of filters for generating head-related transfer functions. Background Technology

[0004] Audio systems with multiple audio channels are generally known and used in the entertainment industry, such as in movies or computer games. These systems are often called surround sound systems or 3D sound systems. Recently, arrangements have been introduced to achieve even better 3D sound experiences. These arrangements not only have multiple audio channels but also provide object-based audio to improve the listening experience.

[0005] In headphone listening, these setups typically rely on filtering the audio channels using what are known as head-related transfer function (HRTF) filters. A three-dimensional experience is created by manipulating the sound in the two audio channels of the headphones to resemble directional sound reaching the ear canal. A three-dimensional sound experience can be achieved by considering the influence of the auricle, head, and torso on the sound entering the ear canal. These filters are often called HRTF filters. These filters are used to provide an effect similar to the human experience of sound coming from different directions and distances. When the anatomy of a person's body parts (such as the ears, head, and torso) is known, a personalized HRTF filter can be created to make the sound experienced through the headphones as realistic as possible.

[0006] The materials required to generate such a filter include three-dimensional point cloud coordinates describing the surface point cloud, which can be realized by determining the three-dimensional point cloud of the relevant part of the ear. In conventional simulation-based methods, a three-dimensional point cloud of a body part is determined using a 3D scanning device, which generates a 3D model of at least a visible portion of the ear. However, this requires an expensive 3D scanning device that can produce an accurate 3D geometric model of the ear. Because ears can have different geometries, it is possible to generate two filters, so that each ear has its own filter.

[0007] Conventionally, HRTF filters are pre-generated, and for each individual, a filter is selected from a library of HRTF filters derived from acoustic measurements or simulations performed on a small subset of individuals. However, due to technological advancements, personalized filters can be created when the anatomy of the person to whom the filter is designed is known. Anatomical measurements can be performed by acquiring sufficient image or video material to adequately represent the person being measured. However, this is computationally and network-intensive, as long videos and large image sets require significant space. Furthermore, acquiring these images individually is not easy. This increases the number of images or the length of the video required.

[0008] Therefore, an arrangement is needed to acquire the image required to generate the HRTF filter. Summary of the Invention

[0009] An arrangement for acquiring images used to generate a head-related transfer function filter is disclosed. In this arrangement, the camera of a mobile phone or similar portable device is adjusted to take an image. All acquired images are analyzed, and only suitable images are further sent for generating the head-related transfer function filter. The arrangement is further configured to provide instructions to the user to adequately cover the entire head and other relevant body parts.

[0010] In one aspect of the invention, a method is disclosed for acquiring images required to generate geometric data for a head-related transfer function (HRF) filter. The method includes initializing a camera application in a user device for controlling a camera module of the user device; acquiring multiple images using the camera module; selecting an image displaying an anatomical structure, wherein the anatomical structure can be used to generate the HRF filter; determining whether the selected image sufficiently includes the anatomical structure to generate the HRF filter; and if the determination is negative, the method further includes providing an instruction to the user to acquire additional images to obtain images of areas that are not sufficiently covered.

[0011] This aspect facilitates better generation of head-related transfer function filters by providing simple acquisition of the images needed to generate the point clouds required in filter production. Furthermore, it reduces the transmission capacity and computing power required to generate filters at the device or remote service location. Additionally, it improves the geometric accuracy of the point cloud by controlling the quality and angular coverage of the image during image acquisition.

[0012] In its implementation, the method further includes transmitting each selected image, including the anatomical structure used to generate the head-related transfer function filter, to a head-related transfer function filter generation server. It is advantageous to transmit the selected images to an internal or external server or other computing facility with greater computational capacity. The amount of data to be transmitted is reduced when only selected images are sent.

[0013] In its implementation, the method further includes discarding images that do not contain geometry that can be used to generate head-related transfer filters. Discarding unused images to free up memory for other purposes is beneficial.

[0014] In one implementation, the method further includes: preparing a user device for acquiring the image, wherein the preparation includes at least one of the following: selecting a sufficient resolution; turning on the illumination device of the camera user device; adjusting the exposure time; and selecting an appropriate frame rate. It is beneficial to determine suitable settings before acquiring the image. These settings may differ from those preferred by the user for general photography. Therefore, changing the settings will result in a better image for this purpose, and this can reduce the need to acquire graphics for generating point clouds.

[0015] In its implementation, the method, when providing instructions, also includes at least one of the following: displaying visual instructions on the device screen; providing voice instructions to the user; and providing haptic instructions. Providing the user with feedback regarding successful image acquisition is beneficial. This helps to acquire higher-quality images in a shorter time.

[0016] In its implementation, the method further includes detecting and / or marking ear and facial landmarks. Detecting and marking landmarks is beneficial because these landmarks are anatomical features relevant to the production of the filter.

[0017] In its implementation, the method further includes arranging the selected images into at least three datasets, wherein these datasets include: images of the head and upper torso; an image of the left ear; and an image of the right ear. Obtaining images from all body parts that are significant to the filter is beneficial. This will improve the quality of the filter.

[0018] In this implementation, the selection is based on at least one of the following: the visibility of the selected anatomical feature; the quality of the image; and the angular coverage of the image. Advantageously, the image selection can be based on various qualitative measurements, resulting in images that are both high-quality and clearly show the relevant parts.

[0019] In one aspect, a computer program for a server is disclosed, comprising code adapted to induce the method as described above when executed on a data processing system. Advantageously, this arrangement can be provided as a computer program, making it readily available for image acquisition using personal devices.

[0020] In one aspect, an apparatus includes: at least one processor configured to execute a computer program; at least one memory configured to store the computer program and related data; at least one data communication interface configured to communicate with an external data communication network; and at least one imaging device; wherein the apparatus is configured to perform the method described above. Advantageously, this arrangement can be provided as an apparatus so that a user can easily use the apparatus during image acquisition.

[0021] The described arrangement for acquiring images used to generate head-related transfer function (HRT) filters facilitates the generation of personally designed HRT filters without expensive scanning processes. Individuals wishing to obtain their own HRT filters can acquire the desired images using a mobile phone or similar device. The disclosed arrangement is effective because it determines whether the acquired images are usable and only transmits usable images. This not only reduces the need for data transmission but also provides more reliable results. In an alternative example, the images are provided to an application within the same device. In this method, the process reduces the required computational power, making it possible to perform such calculations on devices with lower computational capacity. Furthermore, when less computational power is required, the device's battery will last longer.

[0022] When the person acquiring the necessary images uses the publicly available setup, he / she can immediately acquire all the necessary images. Furthermore, the setup provides immediate feedback indicating whether the acquired images are sufficient. Therefore, the user can rely on the service, eliminating the need for multiple image acquisitions. This reduces the command-to-transmission time for the final head-related transfer filter. Attached Figure Description

[0023] The accompanying drawings, included to provide a further understanding of the arrangement for generating the head-related transfer function filter and forming part of this specification, illustrate an embodiment of the arrangement for generating the head-related transfer function filter, and together with the description, help explain the principles of the arrangement. In the drawings:

[0024] Figure 1 This is an example of a device for generating a head-related transfer function filter, and

[0025] Figure 2 This is an example of a method for generating head-related transfer function filters. Detailed Implementation

[0026] Reference will now be made in detail to the implementation schemes, examples of which are shown in the accompanying drawings.

[0027] In the following description, multiple images have been referenced. In the context of this specification, "multiple images" can refer to a number of still images or images extracted from a video stream, or any combination of both. Multiple images are needed to view the desired features from different angles, allowing for sufficiently accurate determination of the 3D point cloud.

[0028] exist Figure 1 Example of a device 10 for acquiring the image required to generate a head-related transfer function filter is shown. Figure 1 In the example, device 10 is a mobile phone; however, any similar device that follows the principles discussed below may be used. Examples of such devices include tablets, laptops, etc.

[0029] Figure 1 The mobile phone 10 includes a display 11. The display 11 can be a typical mobile display, which is usually touch-sensitive, although it is not required in this example.

[0030] The mobile phone 10 also includes at least one processor 12 configured to execute computer programs and applications. The mobile phone also includes memory 13 for storing computer programs, applications, and related data. Typically, a mobile phone has both volatile and non-volatile memory. This example applies to both types of memory.

[0031] Mobile phone 10 also includes a data communication interface 14. Examples of such interfaces are UMTS (Universal Mobile Telecommunications System) and LTE (Long Term Evolution). Mobile phones can typically access several different network types.

[0032] A common feature of modern mobile phones is a camera 15. This camera includes at least one lens and at least one image sensor. In the case of multiple lenses and acquired sensor images, the acquired sensor images are combined to provide a higher quality image. Typically, cameras (such as camera 15 of mobile phone 10) are capable of acquiring video sequences. In this example, a video sequence at a so-called Full HD 1080p resolution, which is 1920 × 1080 pixels, can be captured. Higher resolutions are also possible. In this example, it is possible to supplement the video sequence by using still images of higher resolution. Modern cameras may also be able to generate three-dimensional images, and other images including depth information of at least some objects in the image. The image may also include additional information, such as lighting conditions, device orientation information, and other similar information providing additional information about the graphics and graphic content. These features can be used in the described embodiments. For example, depth cameras, stereo cameras, or other range imaging devices may be very useful in determining the three-dimensional coordinates of anatomical features considered when generating head-related transfer function filters.

[0033] Mobile phone 10 also includes an audio device 16. This audio device may include a combination of a speaker and a microphone. The speaker can also be used for ordinary calls. Mobile phone 10 also includes a haptic device 17, which can be used to provide feedback to the user of mobile phone 10. This feature is typically used, for example, to notify the user of an incoming call via a vibration alarm.

[0034] exist Figure 2 The image shown is an example of a method for obtaining the image required to generate the head-related transfer function filter. This method can be used in applications such as... Figure 1 The device is a mobile phone 10. However, this is just an example, and any similar device can be used.

[0035] This method is initiated by initializing the mobile phone's camera application (step 20). This initialization typically involves loading and launching the application so that the mobile phone is ready to acquire images. Figure 2 This method also includes setting parameters suitable for the purpose.

[0036] These parameters can be, for example, selecting a video capture mode with the highest possible resolution (such as 1920×1080 or 3840×2160) and an appropriate frame rate. The frame rate doesn't need to be suitable for viewing purposes; however, a higher frame rate provides more material for later use. In addition to the frame rate, an appropriate exposure time can also be selected. If the mobile phone has a lighting device (such as an LED (light-emitting diode) or other light), this lighting device can be turned on to improve capture. Even if several preset options exist, it's not necessary to use all of them. The purpose of the settings is to improve the capture filter to produce the desired features. Therefore, an acceptable image is one that helps extract features from the image, but it may not necessarily be aesthetically pleasing to the human eye. For example, when selecting the optimal exposure time, it's important that important pixels are not overexposed or underexposed.

[0037] Once the settings have been properly configured, multiple images are acquired (step 21). The user of mobile phone 10 uses the camera 15 of mobile phone 10 to acquire multiple images. These images can be acquired in still mode or as a video stream. Instructions may be given to the user, for example, to acquire an image of the left ear first. Imaging stops after images have been acquired, for example, after a video stream for a specific time period or a predetermined number of images has been acquired. The camera stores the acquired images in memory 13. In more advanced implementations, the stopping condition may depend on quality, imaging conditions, etc. For example, it is possible to continue acquiring images until a predetermined angle coverage has been achieved.

[0038] Select the images needed to determine the head-related transfer function from the acquired images (step 22). Process the images in memory 13 by processor 12 to determine if an image is available. Furthermore, some images may be considered unavailable because the area has been sufficiently covered by previous images.

[0039] Several optional steps can be taken when selecting images for further transmission. First, each image can be processed to check its technical quality. This may include, for example, checking if the image is sharpened and properly exposed. During this process, automatic correction algorithms can be used to check if it's possible to improve the image. For example, the variance of a Laplacian filter can be used to evaluate sharpness. A higher variance is generated in the focus box than in the blur box. Frame selection is defined using a dynamic threshold level (the average variance of the video). If the sampling rate is insufficient, the threshold level is lowered until the requested frame rate is achieved.

[0040] Illumination and exposure can be verified by analyzing the highest pixel intensity on the target to confirm that there is no overexposure. This step corresponds to the analysis of selecting the correct exposure.

[0041] Following the technical examination, the desired body parts (such as ears, face, and head) are located on the images that have passed the technical examination, and the technical examination is applied.

[0042] The ear and face are detected using (machine learning) feature detection methods such as CNN (convolutional neural network). The detector is pre-trained using a selected dataset, which typically consists of a large number of image samples of n>1000 images.

[0043] During video capture, feature detection methods may be used to detect ears, and the ROI (Region of Interest) of the ears is drawn on the image. Facial and ear landmarks are detected from the ROI using a pre-trained shape model, and these landmarks are tracked during the capture process. If the ear or facial location and features cannot be detected, the application provides feedback to the user and guides the user to adjust the camera position based on the previously detected features.

[0044] A graphical user interface can guide users to capture multiple images, such as videos, from the correct distance and orientation. This can be done, for example, by displaying the outline of a head or ear on the mobile device's screen. The user is advised to position their head or ear within this outline when shooting video. Additionally, the outline can be rotated to guide the user in changing the shooting direction. Arrows on the screen can indicate the direction the camera needs to move.

[0045] The above feedback applies only when the person acquiring multiple images can see the instructions. This typically occurs only when another person is in charge of acquisition. In cases of unassisted acquisition, tactile and / or audio feedback, rather than visual information, may be provided. Furthermore, all visual, tactile, and audio feedback can be combined or used individually to provide the best possible form of assistance.

[0046] For detected body parts, online visibility detection must be applied. Hair on the ears will affect the final reconstruction, so these will be detected and the user will be notified of the issue. Detection is performed from the ROI detected using the methods described above.

[0047] First, the ear region is segmented using color information. Color-based segmentation can be performed, for example, using a neural network to improve the segmentation results. Edge detection (such as the Canny method) is then applied to the segmented frames to detect fine hairs on the ear. If unwanted hair is detected, the application will notify the user to remove the hair from the ear.

[0048] After selecting an image, processor 12 is configured to determine whether the selected image is sufficient to determine the head-related transfer function filter (step 23). To do this, processor 12 may perform sparse reconstruction of the head / ear.

[0049] Sparse reconstruction refers to point clouds or surface models that are not accurate enough for HRTF processing; however, when the final reconstruction is performed using computing devices capable of providing such reconstruction, the sparse reconstruction is sufficient to provide an estimate of whether the image is accurate enough. Sparse point clouds are generated online using methods such as Fast Simultaneous Localization and Mapping (SLAM). Surface models can be generated using deformable shape models, for example, using Principal Component Analysis (PCA). When performing sparse reconstruction, features from the acquired video stream or image are extracted and tracked. The tracked features are used to improve the estimation of camera position and angle. Information received from other mobile phone sensors, such as gyroscopes and accelerometers, can be used to improve camera localization and absolute zoom.

[0050] At this stage, the user may be instructed to acquire more images if necessary. The quality of the sparse reconstruction can be analyzed, for example, by comparing the original image from the camera with the virtual image generated from the sparse 3D reconstruction. If the features of the sparse reconstruction (such as the outline of the ears) are inconsistent with the original image, the user is instructed to acquire more images. However, it may also be possible to determine if it is possible to create three sufficient sets (step 23). In this example, there are sets for the head and both ears; however, it is possible to include, for example, a separate additional set for the user's body. Accordingly, a lower-quality filter may be created by including a set only for the ears.

[0051] If these sets are insufficient, the method returns to acquiring images via instructions (step 21). If enough images are available, the acquired images are sent to a server, similar to a cloud service used to generate the actual head-related transfer filters. Information obtained from sparse reconstruction may be sent along with the images.

[0052] If these sets are sufficient, the method continues to further transmit the selected images (step 24). Further image transmission can mean transmitting the images to an external device or service, such as a computer, server, or cloud service. However, further transmission to another application is performed within the device used to acquire the images. For example, a mobile phone application can be configured to perform demanding computations in the background, possibly during low-activity periods such as at night, and when the device may be connected to a charger. Therefore, complex processes can be performed even on devices with low computing power.

[0053] In the example above, the method is shown as a sequence of steps; however, the process does not need to be sequential, but can be implemented at least partially in parallel. For example, processing of the first video frame can begin immediately when the user starts acquiring images. Therefore, it is possible to provide the user with information and instructions immediately from the outset.

[0054] As described above, components of exemplary embodiments may include computer-readable media or memory for storing instructions programmed according to the teachings of the present invention, and for storing data structures, tables, records, and / or other data described herein. Computer-readable media may include any suitable medium that participates in providing instructions to a processor for execution. Common forms of computer-readable media may include, for example, floppy disks, floppy disks, hard disks, magnetic tapes, any other suitable magnetic media, CD-ROMs, CD±Rs, CD±RWs, DVDs, DVD-RAMs, DVD±RWs, DVD±Rs, HDDVDs, HDDVD-Rs, HDDVD-RWs, HDDVD-RAMs, Blu-ray discs, any other suitable optical media, RAM, PROMs, EPROMs, FLASH-EPROMs, any other suitable memory chips or cartridges, or any other suitable media from which a computer can read.

[0055] It will be apparent to those skilled in the art that, with advancements in technology, the basic idea of ​​the arrangement for generating the head-related transfer function filter can be implemented in various ways. Therefore, the arrangement for generating the head-related transfer function filter and its implementation are not limited to the examples described above; rather, they can vary within the scope of the claims.

Claims

1. A method for acquiring images needed to generate geometry data for a head-related transfer function filter, the method comprising: initializing a camera application in a user device of a user for controlling a camera module of the user device; acquiring a plurality of images using the camera module; selecting images that display anatomical structures that can be used to generate a head- related transfer function filter; determining whether the selected images include sufficient anatomical structures to generate the head-related transfer function filter by performing a sparse 3D reconstruction and comparing at least one of the plurality of images to at least one virtual image generated by the sparse 3D reconstruction; and if the result of the determination is negative, the method further comprising providing haptic feedback that the determination is negative and that the user adjust the camera position; and acquiring additional images of areas in the selected images that are not sufficiently covered.

2. The method of claim 1, further comprising transmitting each selected image to a head- related transfer function filter generation server.

3. The method of claim 1, further comprising discarding images that do not include geometrical shapes that can be used to generate the head-related transfer filter.

4. The method of claim 1, wherein initializing the camera application comprises at least one of: selecting a sufficient resolution; turning on a camera lighting device of the user device; adjusting an exposure time; and selecting an appropriate frame rate.

5. The method of claim 1, further comprising providing voice instructions to the user to acquire additional images of areas in the selected images that are not sufficiently covered.

6. The method of claim 1, wherein the selecting comprises detecting or marking ears and facial landmarks within the selected images.

7. The method of claim 1, wherein the selecting is based on at least one of: visibility of selected anatomical features within the images, quality of the images, and angular coverage of the images.

8. The method of claim 7, further comprising arranging the selected images into at least three data sets, wherein the sets comprise: images of the head and upper torso; images of the left ear; and images of the right ear.

9. The method of claim 1, wherein selecting images that display anatomical structures comprises using a machine learning feature detector that detects ears and faces in input images.

10. The method of claim 9, wherein the detector is pre-trained with a data set of greater than 1000 images.

11. A non-transitory computer readable medium for holding instructions that can program a server to perform a method comprising: selecting images that display anatomical structures that can be used to generate a head- related transfer function filter from a plurality of images received from a user device of a user, the plurality of images being acquired using a camera module in the user device controlled by a camera application in the user device; from a plurality of images received from a user device of a user, the plurality of images being acquired using a camera module in the user device controlled by a camera application in the user device; determine whether the selected images include sufficient anatomical structure to produce the head-related transfer function filter by performing a sparse 3D reconstruction and comparing at least one of the images to at least one virtual image generated by the sparse 3D reconstruction; if the result of the determination is negative, instruct the user device to provide haptic feedback to the user that the determination is negative and that the user adjust the camera position, and to acquire additional images of areas in the selected images that are not sufficiently covered.

12. The non-transitory computer readable medium of claim 11, further having instructions to program the server to initialize the camera application in the user device for controlling the camera module.

13. The non-transitory computer readable medium of claim 11, further having instructions to program the server to instruct the user device to acquire the plurality of images using the camera module in the user device.

14. A user device for acquiring images needed to produce geometric data for determining a head-related transfer function filter, the user device comprising: a processor configured to execute a computer program; a memory configured to store a computer program; a data communication interface configured to communicate with a data communication network; and an imaging device; wherein the user device is configured to: acquire a plurality of images using the imaging device under control of an application stored in the memory; select an image from the plurality of images that displays anatomical structures that can be used to produce a head-related transfer function filter, and discard images from the plurality of images that do not include geometrical shapes that can be used to produce the head-related transfer filter; and transmit the selected image to a server using the data communication interface, wherein the server will determine whether the selected images include sufficient anatomical structure to produce the head-related transfer function filter by performing a sparse 3D reconstruction and comparing at least one of the images to at least one virtual image generated by the sparse 3D reconstruction, and if the result of the determination is negative, the server instructs the user device to i) provide haptic feedback to the user that the determination is negative, and ii) instruct the user to adjust the camera position and acquire additional images of areas in the selected images that are not sufficiently covered.

15. The user device of claim 14, wherein the memory has stored therein instructions that configure the processor to initialize a camera application by at least one of: selecting a sufficient resolution; turning on a camera lighting device of the user device; adjusting an exposure time; and selecting an appropriate frame rate.

16. The user device of claim 14, wherein the memory has stored therein instructions that configure the processor to provide voice instructions to the user to acquire additional images of areas in the selected images that are not sufficiently covered.

17. The user device of claim 14, wherein the memory has stored therein instructions that configure the processor to select the image based on at least one of: visibility of selected anatomical features within the image, quality of the image, and angular coverage of the image.

18. The user device of claim 14, wherein the memory has stored therein further instructions that configure the processor to arrange the selected image into at least three data sets, wherein the sets include: images of the head and upper torso; images of the left ear; and images of the right ear.

19. The user device of claim 14, wherein the memory has stored therein further instructions that configure the processor to select the image using a machine learning feature detector that detects ears and faces in input images.

20. The user device of claim 19, wherein the detector is pre-trained with a data set of greater than 1000 images.

21. A method for acquiring images needed to determine a head-related transfer function filter, the method comprising: capturing a plurality of images using a range imaging device of a user device; providing haptic feedback to the user using a haptic device in the user device while the range imaging device is capturing the plurality of images; receiving the plurality of images from the range imaging device; determining whether at least one image includes sufficient anatomical structure to produce a head-related transfer function filter by performing a sparse 3D reconstruction and comparing at least one image of the plurality of images to at least one virtual image generated by the sparse 3D reconstruction; and if it is determined that the at least one image does not include sufficient anatomical structure, capturing one or more further images of areas not sufficiently covered in the at least one image using the range imaging device.

22. The method of claim 21, further comprising: selecting at least one image from the received plurality of images that displays anatomical structure for determining whether the at least one image includes sufficient anatomical structure.

23. The method of claim 22, wherein the at least one image is selected by detecting or marking ears and facial landmarks within at least one selected image.

24. The method of claim 22, further comprising: transmitting the at least one selected image having sufficient anatomical structure to a head-related transfer function filter generation server.

25. The method of claim 22, wherein the at least one image is selected based on at least one of: visibility of selected anatomical features within the image, quality of the image, and angular coverage of the image.

26. The method of claim 25, further comprising: arranging the selected at least one image into at least three data sets, wherein the sets include: images of the head and upper torso; images of the left ear; and images of the right ear.

27. The method of claim 22, wherein selecting the at least one image includes using a machine learning feature detector that detects ears and faces in input images.

28. The method of claim 27, wherein the detector is pre-trained with a dataset of greater than 1000 images.

29. The method of claim 22, further comprising discarding one or more images that do not include geometry that will be used to produce the head-related transfer filter.

30. The method of claim 21, further comprising initializing an application in the user device when using the range imaging device by at least one of: selecting a resolution; turning on an illumination device of the user device; adjusting an exposure time; and selecting a frame rate.

31. The method of claim 21, further comprising providing voice instructions to the user for capturing the additional images of areas that are not sufficiently covered in the received images.

32. A method for acquiring images needed to produce geometry data for a head-related transfer function filter, the method comprising: in response to a user input for generating a head-related transfer function, capturing a plurality of images using a range imaging device in a user device; selecting a subset of images from the plurality of images that display anatomical structures of an ear; determining whether the selected subset includes sufficient anatomical structures for producing the head-related transfer function filter by performing a sparse 3D reconstruction and comparing at least one image in the plurality of images to at least one virtual image generated by the sparse 3D reconstruction; in response to determining that the selected subset does not include sufficient anatomical structures for producing a head-related transfer function filter, presenting a notification to acquire additional images of areas in the selected subset that are not sufficiently covered; and in response to a user input for acquiring the additional images, capturing the additional images of areas in the selected subset that are not sufficiently covered using the range imaging device in the user device.

33. The method of claim 32, further comprising instructing the user to first acquire images of a particular one of a left ear and a right ear.

34. The method of claim 32, wherein presenting a notification comprises voice instructions to a user.

35. The method of claim 32, wherein the subset of images is selected based on at least one of: visibility of selected anatomical features within an image, quality of the image, and angular coverage of the image.

36. The method of claim 35, further comprising arranging the selected subset of images into at least three datasets, wherein the sets include: images of a head and upper torso; images of a left ear; and images of a right ear.

37. The method of claim 32, wherein selecting the subset of images comprises using a machine learning feature detector that detects ears and faces in input images.

38. The method of claim 37, wherein the detector is pre-trained with a dataset of greater than 1000 images.

39. The method of claim 32, further comprising presenting a graphical user interface to guide the user to acquire the plurality of images or the additional images, wherein the graphical user interface displays a silhouette of a head or ear on a screen of the user device.

40. The method of claim 39, further comprising the silhouette rotating to guide the user to change a shooting direction of the range imaging device or presenting an arrow on the screen to indicate a direction in which the range imaging device needs to move.

41. The method of claim 32, further comprising providing feedback to the user about a success of image acquisition.

42. An electronic device, comprising: a range imaging device; a haptic device; at least one processor; and memory having instructions that, when executed by the at least one processor, cause the electronic device to: capture a plurality of images using the range imaging device, provide haptic feedback using the haptic device while the range imaging device is capturing the plurality of images, determine whether at least one image includes sufficient anatomical structure to produce a head-related transfer function filter by performing a sparse 3D reconstruction and comparing the at least one image to at least one virtual image generated from the sparse 3D reconstruction, and in response to the at least one image not including sufficient anatomical structure, capture one or more additional images of one or more areas not sufficiently covered in the at least one image using the range imaging device.

43. The electronic device of claim 42, wherein the memory has further instructions to select at least one image from the plurality of images that displays anatomical structure to determine whether the at least one image includes sufficient anatomical structure.

44. The electronic device of claim 43, wherein the at least one image is selected by detecting or marking ears and facial landmarks within the at least one selected image.

45. The electronic device of claim 43, wherein the at least one image is selected based on at least one of: a visibility of selected anatomical features within an image, a quality of the image, and an angular coverage of the image.

46. The electronic device of claim 43, wherein the memory has further instructions to discard one or more images that do not include geometry to be used to produce the head-related transfer filter.

47. A method, comprising: capturing a plurality of images using a range imaging device of a user device; providing haptic feedback using a haptic device in the user device while the range imaging device is capturing the plurality of images; determining whether at least one image of the plurality of images includes sufficient anatomical structure to produce a head-related transfer function filter by performing a sparse 3D reconstruction and comparing the at least one image to at least one virtual image generated from the sparse 3D reconstruction; and ​ ​ if it is determined that the at least one image does not include sufficient anatomical structure, providing audio feedback using a speaker of the user device, the audio feedback indicating to the user to capture one or more further images using the range imaging device of an area not sufficiently covered in the at least one image.

48. The method of claim 47, wherein the audio feedback comprises a verbal instruction to the user for capturing the one or more further images of an area not sufficiently covered in the at least one image.

Citation Information

Patent Citations

  • Quality metrics method and system for biometric authentication

    CN103577801A

  • Scheme for supporting taking picture in apparatus equipped with camera

    KR1020170023494A

  • System and method for exposing video-taking heuristics at point of capture

    US20090244323A1

  • Systems and methods for determining head related transfer functions

    US20130169779A1

  • System and method for providing haptic feedback to assist in capturing images

    US20150334292A1