Arrangement for generating a head related transfer function filter

By capturing ear images with a regular camera and using PCA to generate personalized HRTF filters, the problem of headphones being unable to produce three-dimensional sound has been solved. This enables low-cost and convenient generation of personalized filters, improving the user experience.

CN114998888BActive Publication Date: 2026-02-03APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210310525.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-03-15
Filing Date
2017-03-09
Publication Date
2026-02-03
Estimated Expiration
2037-03-09

AI Technical Summary

Technical Problem

Existing headphones cannot effectively produce three-dimensional sound effects, and traditional three-dimensional scanning devices are expensive and difficult to provide individual head-related transfer function filters for each user.

Method used

By capturing ear images with a regular camera and detecting key feature points, a personalized head-related transfer function filter is generated using a model database and principal component analysis (PCA). This filter is then combined with mobile devices and servers for computation and data processing to produce a unique HRTF filter.

Benefits of technology

It enables the low-cost and convenient generation of unique HRTF filters for each user, improving the three-dimensional sound experience of headphones and making it suitable for home environments and multi-user use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998888B_ABST
    Figure CN114998888B_ABST
Patent Text Reader

Abstract

The present invention discloses an arrangement for producing head related transfer function filters. In producing three-dimensional audio by using headphones, specific HRTF filters are used to modify the sound for the left and right channels of the headphones. Because the morphology of each ear is different, it is advantageous to design the HRTF filters specifically for the user of the headphones. Such filters can be produced by deriving the ear geometry from a plurality of images taken with a common camera, detecting the necessary features from the images, and fitting the features to a model that has been produced from precise scanned ears including representative values for different sizes and shapes. The taken images are sent to a server (52) that performs the necessary calculations and further submits the data or produces the requested filters.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application filed on March 9, 2017, with application number 201780016802.3 and entitled "Arrangement for Generating Head-Related Transfer Function Filter". Technical Field

[0002] This invention generally relates to the arrangement of filters for generating head-related transfer functions. Background Technology

[0003] Audio systems with multiple audio channels are well-known and used in the entertainment industry, such as in movies or computer games. These systems are often referred to as surround sound systems or 3D sound systems. The purpose of such systems is to provide more realistic sound arriving from multiple directions. The most traditional way to achieve this type of audio system is to place multiple audio speakers in a room or space (such as a movie theater). The main purpose of this system is to allow sound to arrive from different directions. Thus, moviegoers will have a more realistic experience when sounds related to a specific action arrive from the inferred direction of the action. In earlier systems, five different full-bandwidth channels and one low-frequency audio channel were typically used; however, this is now even more common and frequently used in home appliances. Therefore, in addition to direction, different channels can be used for different frequency ranges. Low-frequency sounds are typically produced by different types of speakers commonly known as subwoofers.

[0004] Besides the number of audio sources, other factors influence the sound and how realistic it is. The location of the audio sources is essential, and the geometry of the room, furniture, and similar objects affects the sound. For example, in a large movie theater, people in the front row may experience sound differently than those in the back. In a home setting, other furniture, windows, curtains, and similar objects affect the sound. Furthermore, the orientation and physical characteristics of the person watching the movie will influence how they experience the sound. Finally, the geometry of the outer ear (auricle) significantly affects the sound reaching the eardrum from the sound source.

[0005] Because headphones are portable and do not disturb others, they are generally preferred over speaker systems. Traditionally, headphones cannot produce three-dimensional sound, but are limited to two channels (left and right), and the perceived sound is localized within the listener's head. Recently, headphones with three-dimensional sound have been introduced. A three-dimensional experience is produced by manipulating the sound in the two audio channels of the headphones to resemble directional sound reaching the ear canal. A three-dimensional sound experience is possible by taking into account the effects of the auricle, head, and torso on the sound entering the ear canal. These filters are often called HRTF (Head Correlation Transfer Function) filters. These filters are used to provide effects similar to how the human torso, head, and ears filter three-dimensional sound.

[0006] When the geometry of the human ear is known, a personal HRTF filter can be generated so that the sound experienced by means of headphones is as realistic as possible. In a traditional analog-based approach, the geometry is generated by using a three-dimensional scanning device which generates a three-dimensional model of at least a part of the visible part of the ear. This, however, requires an expensive three-dimensional scanning device which can generate an accurate three-dimensional geometric model of the ear. Since ears can have different geometries, two filters can be generated so that both ears have their own filter, respectively. When the three-dimensional geometry is known, the skilled person can use various mathematical methods for generating the HRTF filter from the acquired geometry. Even in the above, only the geometry of the ear is discussed, the skilled person understands that the geometry of the head and / or torso of the person being scanned can be included in the filter generation.

[0007] Reference is made to Figure 1 , Figure 2a and Figure 2b for a better understanding of the framework. Figure 1 is an illustration of an ear as shown in the book "Current Diagnosis & Treatment in Otolaryngology - Head & Neck Surgery" by Lalwani AK. The figures are only referred to for the purpose of disclosing the anatomical construction of the ear. Figure 2a and Figure 2b are two models of an ear generated by using a three-dimensional scanning device or formed by a camera system which involves multiple cameras so that the ear is imaged from multiple directions simultaneously so that three-dimensional coordinates can be calculated.

[0008] In order to improve the quality of a three-dimensional audio system, it would be advantageous to generate an individual HRTF filter for each person using headphones. Thereby, there is a need for an arrangement for providing the geometry required for calculating the filter. SUMMARY

[0009] When generating three-dimensional audio by using headphones, a specific HRTF filter is used for the sound of the left and right channels of the headphones. Since the morphology of each ear is different, it is advantageous to design the HRTF filter specifically for the user of the headphones. Such a filter can be generated by deriving the ear geometry from multiple images taken with a common camera, detecting the necessary features from the images and fitting said features to a model which has been generated from accurate scanned ears including representative values for different sizes and shapes. The taken images are sent to a server which performs the necessary calculations and further submits the data or generates the requested filter.

[0010] In an embodiment, a method for producing geometry data for a head related transfer function filter is disclosed. The method comprises the steps of: receiving at least one image frame comprising an image of an ear; detecting from the image frame a significant point in the ear; producing a point cloud from the points detected from at least one image frame; registering the point cloud with a model database, wherein the model database comprises data derived from an ear model, the ear model comprising geometry data relating to the ear; and processing the model database with the point cloud such that geometry data corresponding to the point cloud is generated by modifying the geometry data in the model database, thereby producing geometry data for a head related transfer function filter.

[0011] In an embodiment, the model database is pre-computed using so-called principal component analysis bases. In a further embodiment, producing geometry data comprises transforming the model database. In another embodiment, the plurality of image frames is a video sequence. In a further embodiment, registering comprises at least one of: translation, rotation and scaling.

[0012] In an embodiment, the above disclosed method is used for producing a head related transfer function filter. In a further embodiment, the above disclosed method is implemented as computer software executed on a computing device.

[0013] In a further embodiment, a three-dimensional camera is used instead of a conventional camera. When using a three-dimensional camera, only one three-dimensional image is sufficient for deriving the above mentioned point cloud, however, two or more three-dimensional images can be combined for producing the point cloud. The point cloud is then processed similarly as in the case of two-dimensional images.

[0014] In an embodiment, an apparatus is disclosed. The apparatus comprises at least one processor configured to execute a computer program, at least one memory configured to store the computer program and related data, and at least one data communication interface configured to communicate with an external data communication network. The apparatus is configured to perform the above described method.

[0015] Benefits of the described embodiments include the possibility to produce individual HRTF filters for a user of a headphone. Furthermore, benefits of the present embodiments include the possibility to produce own filters for each ear. This will for example improve how a user hears the produced audio when watching a movie with three-dimensional sound and will have a more realistic experience.

[0016] A further benefit of the described embodiments includes the possibility of providing an individual filter using only a conventional camera, e.g. a mobile phone. The mobile device is used to acquire image frames, which are then sent to a service that can produce the necessary calculations. This facilitates the production of filters with technically simple tools, such as mobile phones, instead of the conventional need for three-dimensional scanning arrangements. Thereby, low-cost devices can technically be used for producing such filters.

[0017] A further benefit of the described embodiments is the possibility of producing filters at home. This facilitates headphones with individual filters that can be ordered from internet stores. People interested can order a desired headphone set from an internet store or buy from a normal market and produce the filters at home by imaging their own head. This will facilitate easier sales of technically unique headphones.

[0018] A further benefit of the described embodiments is the possibility of producing filters that can be easily changed when multiple people use a headphone set. This will facilitate a better audio experience for all users. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings, which are included to provide a further understanding of the arrangement for producing a head-related transfer function filter and constitute a part of this specification, illustrate embodiments and together with the description help to explain the principles of the arrangement for producing a head-related transfer function filter. In the drawings:

[0020] Figure 1 is an illustration of an ear and parts thereof,

[0021] Figure 2a and Figure 2b is an example of a three-dimensional model of an ear,

[0022] Figure 3 is a flow chart of an embodiment,

[0023] Figure 4 is a flow chart of an embodiment, and

[0024] Figure 5 is a block diagram of an example implementation. DETAILED DESCRIPTION

[0025] The embodiments will now be referred to in detail, examples of which are illustrated in the accompanying drawings. In the following description, the expression "ear" is used to denote the visible part of an ear, which part can be obtained by using a conventional camera that can be arranged to a mobile device.

[0026] In Figure 3In the method of generating the geometry needed for generating HRTF filters, a flow chart of a method for establishing pre-computed PCA bases used in the method is disclosed. The PCA bases are examples of model databases or transformable models. As shown in the flow chart, first, a plurality of samples are scanned, step 30. As explained in the background section of the present invention, the scanning is done by using prior art means for scanning. The number of samples can be, for example, one hundred, but it can be increased later and there is no upper limit. A smaller number can also be used, however, a very small number will decrease the quality of the derived geometry data. Thereby, when later scanning another person with precision, the geometry of their pinna can be included in the same set.

[0027] From the scanning results, an accurate three-dimensional ear canal model is constructed, step 31. In an embodiment, the surface model is constructed from three-dimensional vertices forming triangles. Figure 2a and Figure 2b An example of such a surface model is shown in Fig. 1. The models for different pinnae can have different number of vertices and triangles at this stage. According to conventional methods, these models or the information used to construct these models would be used directly for generating filters.

[0028] The generated models are used as a training data set for generating pre-computed PCA bases used in an arrangement for generating head related transfer function filters, step 32. In the generation, Principal Component Analysis (PCA) is used. PCA is a statistical method where important components are reduced from a data set also including unimportant data, so that the important components can be presented without losing important information.

[0029] For the pre-computed PCA bases, in order to find the physical correspondence between the pinna models, the acquired pinna models are processed so that after the processing they consist of equal number of vertices and the vertices have a physical point-to-point correspondence. Thereby, the same number of vertices corresponds to the same physical location in all processed models. The physical correspondence can be found by using optical flow algorithms or similar.

[0030] Finally, the PCA model is constructed by using the PCA decomposition and using the most important components of the PCA as the geometry basis vectors. Thereby, the acquired geometry has been re-processed into a model that provides a representative value for different pinnae and can be used in an arrangement for generating head related transfer function filters, as will be described below.

[0031] In Figure 4The present invention discloses a flowchart of a method for arranging a mobile device as an imaging device to generate a head-related transfer function filter. The mobile device mentioned above can be any portable camera device capable of executing computer programs and transmitting acquired images to a server via a network connection. Typically, the device is a mobile phone or tablet computer that generally includes a camera and is capable of executing computer programs. The mobile device includes components configured to perform… Figure 4 Software that is part of the disclosed method.

[0032] This method is initiated by launching a computer program to acquire data required to generate the geometry of a human auricle, which is then used to generate a head-related transfer function filter. The program is configured to acquire a sequence of multiple image frames or video image frames, step 40. The program is configured to instruct a person to image their own head, such that at least one ear is imaged from sufficiently different angles. The program may include features configured to determine a sufficient number of image frames; however, the program may immediately send the acquired image frames for further processing and receive an acknowledgment message when a sufficient number of image frames have been acquired. Figure 4 In the method, it is determined that the process is performed at the mobile device, and then the acquired image is sent to the server, step 41.

[0033] The server receives multiple image frames or video sequences, step 42. The received images are processed by determining points representing specific physical locations in the imaged auricle. For example, one or more points may represent the shape of the helix, and one or more points may represent the tragus. There is no upper limit to the number of points determined, and the number can be increased if a more accurate model is desired. Even though the person performing the measurement can decide the points being measured, the concha cymba and the concha cavity (see...) are considered... Figure 1 These are very important points to be measured. Furthermore, as mentioned above, the model does not need to be limited to the auricle but can analyze the entire head. For example, a point cloud including the multiple points mentioned above can be achieved using a motion reconstruction algorithm, step 43.

[0034] In the next step, the PCA model is registered with the point cloud, step 44. Registration here means alignment, and is achieved by translating, rotating, and scaling the point cloud. This changes the size and orientation of the point cloud so that it matches the PCA model, and points corresponding to specific physical locations within the imaged auricle will have corresponding points in the PCA model. At this stage, because each auricle has its own characteristics, it is obviously normal for points not to perfectly match the PCA model.

[0035] After the PCA model has been registered with the point cloud, the point cloud is projected onto a pre-computed PCA basis, step 45. This can be done by solving a minimization problem to obtain the weights for the pre-computed PCA basis functions. In an alternative implementation, steps 44 and 45 can be combined so that the PCA model registration and transformation are performed simultaneously. Thus, the above method derives the geometry of the imaged auricle by generating geometric values ​​from a PCA model that includes geometric values ​​for the sample auricle. Because the resulting set of geometric values ​​is a combination of different geometric structures that best match the imaged auricle, the resulting set can be unique.

[0036] Once the geometry or geometric data values ​​have been generated, an HRTF filter can be generated, such as in the case of a 3D scan of the auricle, step 46. The actual generation of the HRTF filter can be performed as a service at the same place that generated the geometric data, but the geometric data can also be submitted to a customer or another service, such as a headphone manufacturer. An example of a method for generating HRTF filters is described in “Rapid generation of personalized HRTFs Proceedings of AES Conference: 55th International Conference: Spatial Audio, Helsinki, 2014” by T. Huttunen, A. Vanne, S. Harder, R. Reinhold Paulsen, S. King, L. P. R. R. P. R. S ...

[0037] Reference above Figure 4 The described method generates independent and unique geometric values, resulting in a unique HRTF filter for each individual, rather than using an average HRTF filter or selecting an HRTF filter from a pre-computed set of HRTF filters. However, in some cases where two or more auricles are nearly identical, the resulting filter does not necessarily have to be unique. This can occur, for example, in the case of identical twins. However, under normal circumstances, the resulting filter is unique.

[0038] Figure 5 An example of the arrangement used to generate the head-correlation transfer function filter is shown. Figure 5In this context, a mobile device 50, such as a mobile phone including a camera, is used to image the head 51 of a person requiring an HRTF filter. The mobile device 50 includes imaging software configured to instruct the user of the mobile device to capture sufficient images to adequately cover the auricle being imaged. This can be done, for example, by acquiring a video sequence in which the software instructs the mobile device to move around the head so that the auricle is visible from all possible directions.

[0039] The mobile device uses a conventional mobile data communication link to send acquired image frames to server 52. Server 52 can be a physical remote computer, a virtual server, a cloud service, or any other service capable of running computer programs. In this embodiment, server 52 is a server including at least one processor 53, at least one memory 54, and at least one network interface 55. At least one processor 53 is configured to execute a computer program and perform tasks instructed by the program. At least one memory 54 is configured to store the computer program and store data. At least one network interface 55 is configured to communicate with an external computing device such as database 56.

[0040] Server 52 is configured to receive acquired image frames using at least one network interface 55 and process them as described above. For example, when the received image frames are in the form of a video sequence, some image frames can be extracted from the video sequence. The desired physical location of the auricle is then detected from the individual image frames using at least one processor 53, and a point cloud is formed by combining the information extracted from multiple image frames. Those skilled in the art will understand that not all desired locations need to be shown in every image frame, and some of the desired physical locations may not be visible depending on the imaging orientation. The point cloud is then stored in at least one memory 54.

[0041] The point cloud is processed by at least one processor 53. First, the point cloud is registered such that its dimension and orientation correspond to a PCA model stored in at least one memory 54 or database 56. The at least one processor is configured to calculate the geometry of the processed auricle by modifying an existing model, for example, by transforming it using the received point cloud and an existing model containing precise geometric data, to generate a geometry corresponding to the received point cloud. Therefore, a geometry corresponding to the imaged auricle can be generated using imprecise image data (instead of using a precise 3D scanner).

[0042] Server 52 can be configured to calculate the final HRTF filter and send it to the user, allowing the user to install the filter into their audio system. If no filter is generated at this stage, the generated geometry data is sent to the user or to a service used to generate the HRTF filter.

[0043] existFigure 5 The arrangement shows a 3D scanning unit 57. It is used to provide samples for constructing PCA models to the database. The model can be refined upon receiving a new one. However, this has no impact on the user who generates filters using a mobile device. Furthermore, a physical connection to the 3D scanning device is not required, but the necessary data can be sent independently without a continuous connection.

[0044] The methods mentioned above can be implemented as computer software that executes in a computing device capable of communicating with a mobile device. When the software executes in the computing device, it is configured to perform the methods described above. The software is specifically implemented on a computer-readable medium, making it available to computing devices, such as… Figure 5 Server 52.

[0045] The HRTF filters discussed above can be placed in various locations. For example, the filters can be used in a computer or amplifier used to generate sound. The filtered sound is then sent to headphones. Headphones incorporating filters can also be implemented; however, those skilled in the art will understand that this cannot be achieved with conventional two-channel stereo headphones because more sound information must be sent to the filter-incorporated components compared to what can be achieved using ordinary stereo cables. Such headphones are typically connected via a USB connector or the like.

[0046] In the examples discussed above, conventional 2D cameras have been described because they are typically used. Instead of 2D cameras, 3D cameras are used instead of conventional 2D cameras. When using a 3D camera, having only one 3D image for exporting the point cloud discussed above is sufficient; however, two or more 3D images can be combined to generate the point cloud. The point cloud is then processed similarly to the case of 2D images. Furthermore, both 2D and 3D image frames can be used when determining the point cloud.

[0047] In the examples discussed above, the point cloud was generated by a server or computer other than the imaging unit of a mobile phone. It is anticipated that the computing power of mobile devices will increase in the future. Therefore, some steps of the above method can be performed in a mobile device configured to acquire image frames required for further computation. For example, the mobile device can generate the point cloud and send it for further processing. The mobile device may also include software components required for further processing, such that the complete geometric data is computed within the mobile device and then further sent to generate an HRTF filter. In another embodiment, the filter may even be generated within the mobile device.

[0048] As described above, components of exemplary embodiments may include computer-readable media or storage for holding instructions programmed according to the teachings of the present invention and for holding data structures, tables, records, and / or other data described herein. Computer-readable media may include any suitable medium that participates in providing instructions to a processor for execution. Common forms of computer-readable media may include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other suitable magnetic media, CD-ROMs, CD±Rs, CD±RWs, DVDs, DVD-RAMs, DVD±RWs, DVD±Rs, HD DVDs, HD DVD-Rs, HD DVD-RWs, HD DVD-RAMs, Blu-ray discs, any other suitable optical media, RAM, PROMs, EPROMs, flash EPROMs, any other suitable memory chips or cartridges, or any other suitable medium from which a computer can read.

[0049] It will be apparent to those skilled in the art that, with advancements in technology, the basic idea of ​​the arrangement for generating the head-correlation transfer function filter can be implemented in various ways. Therefore, the arrangement and implementation methods for generating the head-correlation transfer function filter are not limited to the examples described above; rather, they can vary within the scope of the claims.

Claims

1. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a device, cause the device to perform a method comprising: receiving a plurality of images of an auricle, wherein the plurality of images are taken at different positions to cover the auricle; determining, as output, a geometric value of the auricle from a statistical model using input based on the plurality of images, the statistical model trained on a dataset comprising a plurality of auricle geometries, wherein the statistical model determines the geometric value by reducing a set of data points representative of the auricle; determining a head-related transfer function (HRTF) filter from the geometric value; and generating three-dimensional sound using the HRTF filter.

2. The non-transitory computer-readable medium of claim 1, wherein the plurality of images of the auricle are taken by a mobile device, and wherein the mobile device is configured to instruct a user to move the mobile device around the user's head to take the plurality of images.

3. The non-transitory computer-readable medium of claim 2, wherein the mobile device is configured to instruct the user to move the mobile device such that the plurality of images are taken at a plurality of different angles to cover the auricle.

4. The non-transitory computer-readable medium of claim 1, wherein the plurality of images are a video sequence.

5. The non-transitory computer-readable medium of claim 1, wherein the statistical model comprises pre-computed principal component analysis bases.

6. The non-transitory computer-readable medium of claim 1, wherein the geometric value corresponds to an important point in the auricle.

7. The non-transitory computer-readable medium of claim 6, wherein the plurality of auricle geometries have the important point.

8. The non-transitory computer-readable medium of claim 1, wherein determining the geometric value comprises solving a minimization problem of basis functions of the statistical model.

9. A method comprising: receiving a plurality of images of an auricle, wherein the plurality of images are taken at different positions to cover the auricle; determining, as output, a geometric value of the auricle from a statistical model using input based on the plurality of images, the statistical model trained on a dataset comprising a plurality of auricle geometries, wherein the statistical model determines the geometric value by reducing a set of data points representative of the auricle; determining a head-related transfer function (HRTF) filter from the geometric value; and generating three-dimensional sound using the HRTF filter.

10. The method of claim 9, wherein the plurality of images of the auricle are taken by a mobile device, and wherein the mobile device is configured to instruct a user to move the mobile device around the user's head to take the plurality of images.

11. The method of claim 10, wherein the mobile device is configured to instruct the user to move the mobile device such that the plurality of images are taken at a plurality of different angles to cover the auricle.

12. The method of claim 9, wherein the geometric value corresponds to an important point in the auricle.

13. The method of claim 12, wherein the plurality of pinna geometries have the important point.

14. The method of claim 9, wherein determining the geometric value comprises solving a minimization problem of a basis function of the statistical model.

15. An apparatus comprising: at least one processor configured to execute a computer program; at least one memory configured to store the computer program and related data; and at least one data communication interface configured to communicate with an external data communication network; wherein the apparatus is configured to: receive a plurality of images of an ear, wherein the plurality of images are taken at different positions to cover the ear; determine a geometric value of the ear from a statistical model as an output using input based on the plurality of images, the statistical model being trained on a dataset comprising a plurality of pinna geometries, wherein the statistical model determines the geometric value by reducing a set of data points representative of the ear; determine a head related transfer function (HRTF) filter from the geometric value; and generate a three-dimensional sound using the HRTF filter.

16. The apparatus of claim 15, wherein the plurality of images of the ear are taken by a mobile device, and wherein the mobile device is configured to instruct a user to move the mobile device around the user’s head to take the plurality of images.

17. The apparatus of claim 16, wherein the mobile device is configured to instruct the user to move the mobile device such that the plurality of images are taken at a plurality of different angles to cover the ear.

18. The apparatus of claim 15, wherein the geometric value corresponds to an important point in the ear.

19. The apparatus of claim 18, wherein the plurality of pinna geometries have the important point.

20. The apparatus of claim 15, wherein determining the geometric value comprises solving a minimization problem of a basis function of the statistical model. ​

Citation Information

Patent Citations

  • Systems and methods for determining head related transfer functions

    CN103455824A

  • Estimation of head-related transfer functions for spatial sound representative

    US6996244B1